Benchmark Radar
AI BENCHMARK PROFILE

GUI-Primitives

General AIMultimodal Perception

994-item benchmark of contrastive instruction pairs over seven spatial relations in GUIs, isolating whether vision-language models bind relational language to UI elements correctly.

Released
2026-08-22
Readiness
Paper only
Primary field
General AI

Why it matters

Provides fine-grained diagnostics for spatial reasoning in GUI grounding, a key capability gap in computer-use agents, with public code and predictions.

Motivation

Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.