ContinuityBench
ContinuityBench evaluates stateful failover in multi-provider LLM routing. It measures Continuity Preservation Rate (CPR) and Continuity Latency Overhead (CLO) using synthetic conversation graphs under injected provider failures, with an LLM-as-a-judge scoring context preservation.
- Released
- 2026-07-17
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Production LLM deployments rely on failover mechanisms that often lose conversational context, degrading user experience. This benchmark provides a standardized way to quantify and compare context preservation and latency trade-offs across different routing architectures.
Motivation
In production large language model (LLM) deployments, high API availability guarantees do not equate to conversational continuity.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.