Benchmark Radar
AI BENCHMARK PROFILE

ParaGUIBench

General AIKnowledge & Reasoning

ParaGUIBench aims to benchmark parallel execution and coordination of multiple GUI agents across separate desktop instances, with a dataset of 233 tasks and efficiency metrics.

Released
2026-07-17
Readiness
Runnable
Primary field
General AI

Why it matters

Could enable evaluation of parallel GUI agent coordination, potentially improving efficiency and success on long-horizon tasks.

Motivation

Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.