Benchmark Radar
AI BENCHMARK PROFILE

AdaMem-Bench

General AIKnowledge & ReasoningAdaMem Team

AdaMem-Bench simulates weeks of interaction with week-by-week question answering to evaluate memory policies for personalized long-horizon LLM agents. It measures QA accuracy and memory volume across different models.

Released
2026-06-19
Readiness
Paper only
Primary field
General AI

Why it matters

Long-term memory systems often bloat with irrelevant data. AdaMem-Bench provides a controlled environment to assess memory selection strategies, impacting efficiency and accuracy in personalized agents.

Motivation

Long-term memory systems for Large Language Model (LLM) agents typically try to \emph{remember everything}, extracting memories uniformly to retain as many facts as possible.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.