AI BENCHMARK PROFILE
MemTrace
MemTrace evaluates long-term memory in LLM agents at the knowledge point level, probing memory age, question type, and evidence condition. Scoring is based on accuracy across these dimensions.
- Released
- 2026-06-15
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Aggregated accuracy misses important memory behaviors. MemTrace provides finer-grained metrics to reveal bottlenecks in evidence use, guiding improvements in memory systems.
Motivation
LLM agents increasingly maintain long-term memory of user facts across sessions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.