AI BENCHMARK PROFILE
ParaGUIBench
ParaGUIBench aims to benchmark parallel execution and coordination of multiple GUI agents across separate desktop instances, with a dataset of 233 tasks and efficiency metrics.
- Released
- 2026-07-17
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Could enable evaluation of parallel GUI agent coordination, potentially improving efficiency and success on long-horizon tasks.
Motivation
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs).
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.