PACE-Bench
PACE-Bench is a simulator-grounded benchmark with 144 source-to-target adaptation pairs across six physics domains. Each pair presents a code-driven design that succeeds in a source environment but fails in a mutated target environment, and agents must iteratively adapt the design using diagnostic sandbox feedback within a limited attempt budget.
- Released
- 2026-08-14
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing self-evolving agent evaluations assume fixed execution conditions and do not test recovery after environmental shifts. PACE-Bench provides a repeatable protocol to assess an agent's ability to adapt to changing physics, offering insight into the reliability of different self-evolving methods.
Motivation
Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.