VTC-Bench
VTC-Bench evaluates multiple LLM generations across five domains using Validated Task Coverage (VTC), which measures the number of distinct useful results obtained within k attempts. Tasks are selected from real data where output quality and task-relevant distinctness can be checked automatically without model-based judges.
- Released
- 2026-08-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Traditional evaluations focus on individual outputs or reduce multiple samples to a single score, missing the value of diverse useful results. VTC-Bench provides a reproducible metric for candidate sets, revealing differences in model behavior that single-draw quality metrics do not capture, aiding in model selection for applications requiring multiple outputs.
Motivation
Many LLM applications are most useful when they provide several candidate outputs for comparison, validation, or combination.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.