AI BENCHMARK PROFILE
PerMemBench
PerMemBench evaluates personalized memory systems for LLM agents using multi-year, multi-domain interaction histories across 20 user personas, measuring memory retention accuracy.
- Released
- 2026-05-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Universal memory policies waste budget on transient interactions and fail to preserve critical context. PerMemBench enables evaluation of personalization, revealing that accurate gating remains an open challenge.
Motivation
Existing large language model (LLM) based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.