Benchmark Radar
AI BENCHMARK PROFILE

PAST-Bench

General AIKnowledge & ReasoningGen-Verse

PAST-Bench evaluates recursive self-improvement in personal AI agents by testing whether retained experience improves performance on future tasks. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update capabilities, with matched persistence on/off controls.

Released
2026-08-04
Readiness
Runnable
Primary field
General AI

Why it matters

Whether personal agents actually improve from retained experience has not been systematically tested. PAST-Bench provides a controlled benchmark to measure and attribute cross-session improvement, distinguishing capability gains from pathway evidence.

Motivation

Recursive self-improvement requires agents to turn accumulated experience into better future behavior.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.