Benchmark Radar
AI BENCHMARK PROFILE

MemoryDocDataSet

General AIKnowledge & Reasoning

MemoryDocDataSet evaluates joint conversational memory and long-document reasoning through 50 synthetic micro-worlds, each with personas, temporal event graphs, real legal documents, multi-session conversations, and QA pairs. Questions are categorized by reasoning type, with Hybrid questions requiring navigation of conversation history to locate the relevant document and extract answers.

Released
2026-06-03
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks evaluate conversational memory or document reasoning separately, leaving a gap in measuring integrated performance. MemoryDocDataSet provides a controlled, synthetic environment to assess systems on tasks that require both capabilities, useful for developing and comparing architectures that unify conversation memory with long-document retrieval.

Motivation

AI systems increasingly need to combine two demanding capabilities: navigating multi-session conversation history and performing deep reading comprehension within long documents.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.