PAST-Bench
PAST-Bench evaluates recursive self-improvement in personal AI agents by testing whether retained experience improves performance on future tasks. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update capabilities, with matched persistence on/off controls.
- Released
- 2026-08-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Whether personal agents actually improve from retained experience has not been systematically tested. PAST-Bench provides a controlled benchmark to measure and attribute cross-session improvement, distinguishing capability gains from pathway evidence.
Motivation
Recursive self-improvement requires agents to turn accumulated experience into better future behavior.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.