Benchmark Radar
AI BENCHMARK PROFILE

AgentMemBench

General AIKnowledge & ReasoningAgentMemBench Team

AgentMemBench evaluates five long-term memory management strategies for conversational AI agents across three public datasets (LoCoMo, MultiDoc2Dial, MSC), covering multi-session dialogue, document grounding, and persona-grounded chat. Scoring uses retrieval metrics (Recall@k, MRR, nDCG@k), Answer F1, LLM-judge faithfulness, memory footprint, and latency over 491 annotated question turns.

Released
2026-06-16
Readiness
Paper only
Primary field
General AI

Why it matters

Long-term memory is a key bottleneck for conversational agents. AgentMemBench provides a controlled, reproducible comparison of memory strategies under identical conditions, enabling practitioners to make informed trade-offs between recall quality, accuracy, and resource cost.

Motivation

Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousands of turns.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.