Benchmark Radar
AI BENCHMARK PROFILE

SubtleMemory

General AIKnowledge & ReasoningSubtleMemory Project

SubtleMemory evaluates fine-grained relational memory discrimination in long-horizon AI agents. It contains 1,522 evaluation instances over 10 long histories, grounded in 1,090 relation-controlled memory-variant sets, spanning user-related and non-user-related queries.

Released
2026-06-04
Readiness
Runnable
Primary field
General AI

Why it matters

Existing long-term memory benchmarks rarely probe how agents preserve and use relations among memories during downstream tasks. SubtleMemory provides a relation-controlled evaluation to measure this capability.

Motivation

Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.