Benchmark Radar
AI BENCHMARK PROFILE

CoMET-Bench

General AIMultimodal Perception

CoMET-Bench is a benchmark for conditional multi-event temporal grounding in long-form video, with 2,789 queries over 600 videos and a unified evaluation protocol including counting, grounding, and negative-query recognition.

Released
2026-06-13
Readiness
Paper only
Primary field
General AI

Why it matters

Real-world video grounding requires localizing every event satisfying compositional conditions, which existing benchmarks do not jointly handle. This benchmark introduces Rejection-F1 to prevent trivial gaming and exposes gaps in current methods.

Motivation

Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositional temporal and spatial conditions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.