AI BENCHMARK PROFILE
AdaMem-Bench
AdaMem-Bench simulates weeks of interaction with week-by-week question answering to evaluate memory policies for personalized long-horizon LLM agents. It measures QA accuracy and memory volume across different models.
- Released
- 2026-06-19
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Long-term memory systems often bloat with irrelevant data. AdaMem-Bench provides a controlled environment to assess memory selection strategies, impacting efficiency and accuracy in personalized agents.
Motivation
Long-term memory systems for Large Language Model (LLM) agents typically try to \emph{remember everything}, extracting memories uniformly to retain as many facts as possible.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.