Benchmark Radar
AI BENCHMARK PROFILE

SCALE-QA

General AIKnowledge & ReasoningLong Context & Memory

SCALE-QA evaluates interleaved conversational memory using 3,000 audited multiple-choice questions across 10 domains, where correct answers depend on causally related evidence from earlier turns in flat unsegmented threads.

Released
2026-08-26
Readiness
Paper only
Primary field
General AI

Why it matters

It tests episode integrity failure in long, mixed-topic conversations, a harder memory regime than benchmarks with explicit topic boundaries, and provides deterministic grading for comparing QA performance.

Motivation

Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.