ContinuousBench
ContinuousBench evaluates differentially private synthetic text by measuring capability gain on QA sets derived from a new quarterly corpus (Geminon procedural data or News articles), with standardized training and evaluation harness.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Addresses the gap where existing benchmarks are nearly solvable without corpus access, providing a continuously regenerated test to determine whether DP synthesis transmits genuinely new knowledge and capabilities.
Motivation
Differentially private (DP) text synthesis promises to unlock sensitive corpora for model training, but it remains unclear whether DP synthetic data transmits genuinely new knowledge and capabilities present only in those corpora.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.