AVBench
AVBench evaluates audio-video generative models across ten fine-grained dimensions covering visual quality, audio quality, and cross-modal consistency for human-centric scenarios. It uses specialized evaluators trained via preference learning and provides continuous scores from prediction confidence.
- Released
- 2026-05-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks for AV generation are coarse and rely on generic multimodal LLMs, leading to inaccurate assessments. AVBench offers automated, human-aligned evaluation with fine-grained metrics, enabling reliable model comparison and serving as a potential reward signal for RLHF.
Motivation
Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.