Benchmark Radar
AI BENCHMARK PROFILE

PM-Bench

General AIKnowledge & Reasoning

PM-Bench evaluates prospective memory in LLM agents through a text-based simulated seven-day week. Agents must maintain user intentions, execute delayed intentions, and monitor latent environment changes while performing ongoing activities.

Released
2026-07-14
Readiness
Paper only
Primary field
General AI

Why it matters

PM-Bench fills a gap in evaluating agentic AI by measuring an understudied cognitive capability—prospective memory—in a controlled, repeatable setting. It provides a diagnostic tool for comparing LLM agents and guiding interventions to improve reliability in real-world tasks requiring memory for future actions.

Motivation

A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.