PM-Bench
PM-Bench evaluates prospective memory in LLM agents through a text-based simulated seven-day week. Agents must maintain user intentions, execute delayed intentions, and monitor latent environment changes while performing ongoing activities.
- Released
- 2026-07-14
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
PM-Bench fills a gap in evaluating agentic AI by measuring an understudied cognitive capability—prospective memory—in a controlled, repeatable setting. It provides a diagnostic tool for comparing LLM agents and guiding interventions to improve reliability in real-world tasks requiring memory for future actions.
Motivation
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.