Benchmark Radar
AI BENCHMARK PROFILE

PACE-Bench

General AIKnowledge & ReasoningTHUNLP

PACE-Bench is a simulator-grounded benchmark with 144 source-to-target adaptation pairs across six physics domains. Each pair presents a code-driven design that succeeds in a source environment but fails in a mutated target environment, and agents must iteratively adapt the design using diagnostic sandbox feedback within a limited attempt budget.

Released
2026-08-14
Readiness
Runnable
Primary field
General AI

Why it matters

Existing self-evolving agent evaluations assume fixed execution conditions and do not test recovery after environmental shifts. PACE-Bench provides a repeatable protocol to assess an agent's ability to adapt to changing physics, offering insight into the reliability of different self-evolving methods.

Motivation

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.