Benchmark Radar
AI BENCHMARK PROFILE

RefMem-Bench

General AIKnowledge & Reasoning

RefMem-Bench is a benchmark for reflective memory in long-horizon dialogue, containing 26K QA instances across eight reflective-memory dimensions and three task formats.

Released
2026-05-31
Readiness
Paper only
Primary field
General AI

Why it matters

It evaluates the ability of models to synthesize fragmented multimodal cues into high-level interpretations, going beyond explicit recall.

Motivation

Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure the reflective memory required to synthesize fragmented, multimodal cues into high-level interpretations.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.