AnyGroundBench
AnyGroundBench evaluates video grounding in vision-language models across five specialized domains (animal, industry, sports, surgery, public security) with spatio-temporal annotations. It provides training and test splits per domain for zero-shot and in-context learning evaluation.
- Released
- 2026-07-02
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Addresses the gap between existing benchmark evaluations on general daily-life videos and real-world specialized applications. Provides a structured domain-adaptation protocol to assess model adaptability in specialized fields, enabling comparative evaluation of VLMs in practical scenarios.
Motivation
Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG).
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.