MemOps
MemOps evaluates conversational memory as a sequence of lifecycle operations (remembering, forgetting, updating, reflecting) with structured traces and six categories of operation-level probes, under adjacent-evidence and long-context settings.
- Released
- 2026-07-14
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
This benchmark addresses the gap in memory evaluation by providing operation-level diagnosis rather than final-answer accuracy, revealing specific failure modes in long-context, retrieval-based, parametric, and managed-memory systems, which is valuable for improving memory reliability in LLM agents.
Motivation
Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.