AI BENCHMARK PROFILE
VG-GUIBench
VG-GUI-Bench evaluates MLLM-based GUI agents on following video tutorials to complete interactive tasks, with 1,000 long-horizon test cases and four metrics.
- Released
- 2026-06-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
The benchmark addresses video-guided agentic tasks, complementing VideoQA benchmarks for procedural knowledge transfer.
Motivation
Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (VideoQA) benchmarks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.