AI BENCHMARK PROFILE
SubtleMemory
SubtleMemory evaluates fine-grained relational memory discrimination in long-horizon AI agents. It contains 1,522 evaluation instances over 10 long histories, grounded in 1,090 relation-controlled memory-variant sets, spanning user-related and non-user-related queries.
- Released
- 2026-06-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing long-term memory benchmarks rarely probe how agents preserve and use relations among memories during downstream tasks. SubtleMemory provides a relation-controlled evaluation to measure this capability.
Motivation
Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.