GuideMe
GuideMe is a benchmark for evaluating multimodal large language models on streaming video task guidance. It includes 2,458 videos (223.7 hours) with 47,775 interaction samples covering next-step instructions, completion feedback, error detection, and corrective guidance. Assessment uses temporal-semantic matching, behavioral classification, and LLM-as-a-Judge.
- Released
- 2026-07-03
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Existing multimodal models lack closed-loop interactive coaching ability; GuideMe provides a standardized evaluation for real-time procedural guidance, highlighting the gap in error detection and corrective feedback, which is crucial for practical assistants.
Motivation
While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remains unknown.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.