AgingBench
Evaluates the reliability of deployed AI agents over extended operational lifetimes. The benchmark organizes agent aging into four mechanisms—compression, interference, revision, and maintenance—and uses temporal dependency graphs and paired counterfactual probes to produce diagnostic profiles of the memory pipeline's write, retrieval, and utilization stages. It includes multiple scenarios and memory policies, with scoring based on task performance over sessions.
- Released
- 2026-05-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing agent evaluations are snapshot-based and ignore how agents degrade after deployment. This benchmark measures longevity and provides mechanism-level diagnosis, enabling stage-targeted repair and more dependable agent deployment.
Motivation
Long-lived AI agents are increasingly deployed as persistent operational systems, yet they are still evaluated like freshly initialized models.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.