Benchmark Radar
AI BENCHMARK PROFILE

GuideMe

General AIMultimodal PerceptionGuideMe Project

GuideMe is a benchmark for evaluating multimodal large language models on streaming video task guidance. It includes 2,458 videos (223.7 hours) with 47,775 interaction samples covering next-step instructions, completion feedback, error detection, and corrective guidance. Assessment uses temporal-semantic matching, behavioral classification, and LLM-as-a-Judge.

Released
2026-07-03
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing multimodal models lack closed-loop interactive coaching ability; GuideMe provides a standardized evaluation for real-time procedural guidance, highlighting the gap in error detection and corrective feedback, which is crucial for practical assistants.

Motivation

While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remains unknown.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.