Benchmark Radar
AI BENCHMARK PROFILE

EvoPolicyGym

General AIAgentsEvoPolicyGym Team

EvoPolicyGym evaluates autonomous policy evolution in interactive RL environments. A harness-model agent iteratively edits an executable policy under a fixed interaction budget. The benchmark records trajectories of programs, submissions, feedback, and selection, and scores agents on held-out episodes.

Released
2026-07-02
Readiness
Runnable
Primary field
General AI

Why it matters

Existing evaluations often collapse iterative improvement into a final score or confound it with software-engineering progress. EvoPolicyGym isolates the capability to improve policies from bounded feedback, providing trajectory-level diagnostics that distinguish how agents allocate budget and refine policies. This supports comparison of agents on a controlled, reusable protocol.

Motivation

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.