AI BENCHMARK PROFILE
MD-VQA
Tests video models on detecting whether a step was executed correctly according to its description, for seen and unseen actions.
- Released
- 2026-08-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Shifts mistake detection to open-set evaluation, requiring models to understand general mistake concepts rather than memorizing steps.
Motivation
Human mistakes are inevitable when following instructions, yet they can lead to severe consequences.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.