Benchmark Radar
AI BENCHMARK PROFILE

KVDiagnosis

General AIKnowledge & ReasoningLong Context & MemoryChosenQC

Diagnoses KV-cache compression failures in long-context language models using paired runs and cache, likelihood, attention, and decoding measurements.

Released
2026-08-10
Readiness
Runnable
Primary field
General AI

Why it matters

Provides failure-focused diagnostics that reveal specific causes of compression errors, moving beyond aggregate task scores.

Motivation

KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.