Benchmark Radar
AI BENCHMARK PROFILE

HarmVideoBench

General AIMultimodal Perception

Evaluates harmful video understanding in large multimodal models using 1,379 videos and 4,137 multiple-choice questions across three hierarchical dimensions: Observable Evidence, Clip-Internal Meaning, and Beyond-Clip Reasoning.

Released
2026-06-25
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the gap of shallow binary classification in harmful video benchmarks by testing deep contextual understanding, and the absence of explanatory rationales, providing a diagnostic tool for model evaluation.

Motivation

Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.