AI BENCHMARK PROFILE
KVDiagnosis
Diagnoses KV-cache compression failures in long-context language models using paired runs and cache, likelihood, attention, and decoding measurements.
- Released
- 2026-08-10
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides failure-focused diagnostics that reveal specific causes of compression errors, moving beyond aggregate task scores.
Motivation
KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.