Benchmark Radar
AI BENCHMARK PROFILE

ContinuousBench

General AIKnowledge & ReasoningContinuousBench TeamPeihan Liu

ContinuousBench evaluates differentially private synthetic text by measuring capability gain on QA sets derived from a new quarterly corpus (Geminon procedural data or News articles), with standardized training and evaluation harness.

Released
2026-06-01
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the gap where existing benchmarks are nearly solvable without corpus access, providing a continuously regenerated test to determine whether DP synthesis transmits genuinely new knowledge and capabilities.

Motivation

Differentially private (DP) text synthesis promises to unlock sensitive corpora for model training, but it remains unclear whether DP synthetic data transmits genuinely new knowledge and capabilities present only in those corpora.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.