Benchmark Radar
AI BENCHMARK PROFILE

AVBench

General AIMultimodal Perception

AVBench evaluates audio-video generative models across ten fine-grained dimensions covering visual quality, audio quality, and cross-modal consistency for human-centric scenarios. It uses specialized evaluators trained via preference learning and provides continuous scores from prediction confidence.

Released
2026-05-23
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks for AV generation are coarse and rely on generic multimodal LLMs, leading to inaccurate assessments. AVBench offers automated, human-aligned evaluation with fine-grained metrics, enabling reliable model comparison and serving as a potential reward signal for RLHF.

Motivation

Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.