Benchmark Radar
AI BENCHMARK PROFILE

MemTrace

General AIKnowledge & Reasoning

MemTrace evaluates long-term memory in LLM agents at the knowledge point level, probing memory age, question type, and evidence condition. Scoring is based on accuracy across these dimensions.

Released
2026-06-15
Readiness
Paper only
Primary field
General AI

Why it matters

Aggregated accuracy misses important memory behaviors. MemTrace provides finer-grained metrics to reveal bottlenecks in evidence use, guiding improvements in memory systems.

Motivation

LLM agents increasingly maintain long-term memory of user facts across sessions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.