Benchmark Radar
AI BENCHMARK PROFILE

DynamicMem

General AIKnowledge & Reasoning

DynamicMem evaluates long-horizon memory in LLM agents through a synthetic benchmark with 15 months of multi-app activity per user across 16 applications, including attributes, habits, and preferences that evolve and must be inferred from scattered evidence. Scoring occurs at quarterly checkpoints, tracking performance as history grows.

Released
2026-06-22
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing memory benchmarks use short, simplified interactions, missing real-world complexity. DynamicMem provides a long-horizon, multi-app evaluation that yields insights into memory failures, such as degradation with history length and retrieval-driven errors, which can guide improvements in memory systems for personal assistants.

Motivation

LLM agents increasingly act as personal assistants that must remember a user's profile over months: who they are (attributes), what they routinely do (habits), and what they prefer (preferences), and keep it updated as jobs, routines, and tastes drift.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.