Benchmark Radar
AI BENCHMARK PROFILE

AnyGroundBench

General AIMultimodal PerceptionKeio University

AnyGroundBench evaluates video grounding in vision-language models across five specialized domains (animal, industry, sports, surgery, public security) with spatio-temporal annotations. It provides training and test splits per domain for zero-shot and in-context learning evaluation.

Released
2026-07-02
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the gap between existing benchmark evaluations on general daily-life videos and real-world specialized applications. Provides a structured domain-adaptation protocol to assess model adaptability in specialized fields, enabling comparative evaluation of VLMs in practical scenarios.

Motivation

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.