EvoPolicyGym
EvoPolicyGym evaluates autonomous policy evolution in interactive RL environments. A harness-model agent iteratively edits an executable policy under a fixed interaction budget. The benchmark records trajectories of programs, submissions, feedback, and selection, and scores agents on held-out episodes.
- Released
- 2026-07-02
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing evaluations often collapse iterative improvement into a final score or confound it with software-engineering progress. EvoPolicyGym isolates the capability to improve policies from bounded feedback, providing trajectory-level diagnostics that distinguish how agents allocate budget and refine policies. This supports comparison of agents on a controlled, reusable protocol.
Motivation
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.