DynamicMem
DynamicMem evaluates long-horizon memory in LLM agents through a synthetic benchmark with 15 months of multi-app activity per user across 16 applications, including attributes, habits, and preferences that evolve and must be inferred from scattered evidence. Scoring occurs at quarterly checkpoints, tracking performance as history grows.
- Released
- 2026-06-22
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing memory benchmarks use short, simplified interactions, missing real-world complexity. DynamicMem provides a long-horizon, multi-app evaluation that yields insights into memory failures, such as degradation with history length and retrieval-driven errors, which can guide improvements in memory systems for personal assistants.
Motivation
LLM agents increasingly act as personal assistants that must remember a user's profile over months: who they are (attributes), what they routinely do (habits), and what they prefer (preferences), and keep it updated as jobs, routines, and tastes drift.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.