AI BENCHMARK PROFILE
MEMPROBE
MEMPROBE evaluates long-term memory in LLM agents by reconstructing hidden user-state from agent memory across 50 simulated users with 31 hidden dimensions each.
- Released
- 2026-06-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Offers a direct measure of memory fidelity as an auditable artifact, distinct from downstream task success, potentially improving agent memory evaluation.
Motivation
Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.