Benchmark Radar
AI BENCHMARK PROFILE

VTC-Bench

General AIKnowledge & Reasoning

VTC-Bench evaluates multiple LLM generations across five domains using Validated Task Coverage (VTC), which measures the number of distinct useful results obtained within k attempts. Tasks are selected from real data where output quality and task-relevant distinctness can be checked automatically without model-based judges.

Released
2026-08-25
Readiness
Paper only
Primary field
General AI

Why it matters

Traditional evaluations focus on individual outputs or reduce multiple samples to a single score, missing the value of diverse useful results. VTC-Bench provides a reproducible metric for candidate sets, revealing differences in model behavior that single-draw quality metrics do not capture, aiding in model selection for applications requiring multiple outputs.

Motivation

Many LLM applications are most useful when they provide several candidate outputs for comparison, validation, or combination.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.