Benchmark Radar
AI BENCHMARK PROFILE

VG-GUIBench

General AIAgents

VG-GUI-Bench evaluates MLLM-based GUI agents on following video tutorials to complete interactive tasks, with 1,000 long-horizon test cases and four metrics.

Released
2026-06-28
Readiness
Runnable
Primary field
General AI

Why it matters

The benchmark addresses video-guided agentic tasks, complementing VideoQA benchmarks for procedural knowledge transfer.

Motivation

Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (VideoQA) benchmarks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.