Benchmark Radar
AI BENCHMARK PROFILE

AgingBench

General AIKnowledge & ReasoningVITA Group

Evaluates the reliability of deployed AI agents over extended operational lifetimes. The benchmark organizes agent aging into four mechanisms—compression, interference, revision, and maintenance—and uses temporal dependency graphs and paired counterfactual probes to produce diagnostic profiles of the memory pipeline's write, retrieval, and utilization stages. It includes multiple scenarios and memory policies, with scoring based on task performance over sessions.

Released
2026-05-25
Readiness
Runnable
Primary field
General AI

Why it matters

Existing agent evaluations are snapshot-based and ignore how agents degrade after deployment. This benchmark measures longevity and provides mechanism-level diagnosis, enabling stage-targeted repair and more dependable agent deployment.

Motivation

Long-lived AI agents are increasingly deployed as persistent operational systems, yet they are still evaluated like freshly initialized models.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.