AgentMemBench
AgentMemBench evaluates five long-term memory management strategies for conversational AI agents across three public datasets (LoCoMo, MultiDoc2Dial, MSC), covering multi-session dialogue, document grounding, and persona-grounded chat. Scoring uses retrieval metrics (Recall@k, MRR, nDCG@k), Answer F1, LLM-judge faithfulness, memory footprint, and latency over 491 annotated question turns.
- Released
- 2026-06-16
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Long-term memory is a key bottleneck for conversational agents. AgentMemBench provides a controlled, reproducible comparison of memory strategies under identical conditions, enabling practitioners to make informed trade-offs between recall quality, accuracy, and resource cost.
Motivation
Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousands of turns.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.