Benchmark Radar
AI BENCHMARK PROFILE

MEMPROBE

General AIKnowledge & Reasoning

MEMPROBE evaluates long-term memory in LLM agents by reconstructing hidden user-state from agent memory across 50 simulated users with 31 hidden dimensions each.

Released
2026-06-23
Readiness
Paper only
Primary field
General AI

Why it matters

Offers a direct measure of memory fidelity as an auditable artifact, distinct from downstream task success, potentially improving agent memory evaluation.

Motivation

Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.