Benchmark Radar
AI BENCHMARK PROFILE

PerMemBench

General AIKnowledge & ReasoningPerMemBench team

PerMemBench evaluates personalized memory systems for LLM agents using multi-year, multi-domain interaction histories across 20 user personas, measuring memory retention accuracy.

Released
2026-05-25
Readiness
Runnable
Primary field
General AI

Why it matters

Universal memory policies waste budget on transient interactions and fail to preserve critical context. PerMemBench enables evaluation of personalization, revealing that accurate gating remains an open challenge.

Motivation

Existing large language model (LLM) based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.