MAG
MAG is a benchmark that unifies task execution and guide writing into a single multimodal action and guide task, with grounding over screenshots. It includes a harness for annotation, training, evaluation, and joint metrics.
- Released
- 2026-07-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The evaluation gap is that prior benchmarks separate web agent actions and guide text generation, and often rely on textual DOM rather than screenshots. MAG provides a unified evaluation for multimodal understanding and generation in live environments, which could support development of more capable web agents.
Motivation
Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.