Benchmark Radar
AI BENCHMARK PROFILE

OmniCap-IF

General AIMultimodal PerceptionNJU-LINK

OmniCap-IF evaluates instruction following in omni-modal video captioning with 50 constraint types, 1,920 samples, and checklist-based scoring for format and content correctness.

Released
2026-06-07
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks miss the interplay of audio-visual and user constraints; OmniCap-IF provides a fine-grained evaluation to expose format-content tradeoffs and drive improvements.

Motivation

While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex, multi-faceted user instructions remains largely unexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.