AI BENCHMARK PROFILE
GUI-Primitives
994-item benchmark of contrastive instruction pairs over seven spatial relations in GUIs, isolating whether vision-language models bind relational language to UI elements correctly.
- Released
- 2026-08-22
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides fine-grained diagnostics for spatial reasoning in GUI grounding, a key capability gap in computer-use agents, with public code and predictions.
Motivation
Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.