AI BENCHMARK PROFILE
SCALE-QA
SCALE-QA evaluates interleaved conversational memory using 3,000 audited multiple-choice questions across 10 domains, where correct answers depend on causally related evidence from earlier turns in flat unsegmented threads.
- Released
- 2026-08-26
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It tests episode integrity failure in long, mixed-topic conversations, a harder memory regime than benchmarks with explicit topic boundaries, and provides deterministic grading for comparing QA performance.
Motivation
Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.