AI BENCHMARK PROFILE
HarmVideoBench
Evaluates harmful video understanding in large multimodal models using 1,379 videos and 4,137 multiple-choice questions across three hierarchical dimensions: Observable Evidence, Clip-Internal Meaning, and Beyond-Clip Reasoning.
- Released
- 2026-06-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the gap of shallow binary classification in harmful video benchmarks by testing deep contextual understanding, and the absence of explanatory rationales, providing a diagnostic tool for model evaluation.
Motivation
Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.