AI BENCHMARK PROFILE
RefMem-Bench
RefMem-Bench is a benchmark for reflective memory in long-horizon dialogue, containing 26K QA instances across eight reflective-memory dimensions and three task formats.
- Released
- 2026-05-31
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It evaluates the ability of models to synthesize fragmented multimodal cues into high-level interpretations, going beyond explicit recall.
Motivation
Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure the reflective memory required to synthesize fragmented, multimodal cues into high-level interpretations.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.