Benchmark Radar
AI BENCHMARK PROFILE

Workflow-GYM

General AIKnowledge & ReasoningWorkflow-GYM Team

Workflow-GYM evaluates AI agents on long-horizon GUI tasks in professional domains, using specialized software environments and economically valuable workflows. Tasks require end-to-end operation of graphical user interfaces, with success rates measured by task completion.

Released
2026-06-09
Readiness
Paper only
Primary field
General AI

Why it matters

Most GUI benchmarks cover simple, short-horizon tasks in general software, leaving a gap for professional, long-horizon workflows. Workflow-GYM provides a fixed protocol for measuring agent performance on such tasks, which is valuable as organizations consider deploying agents for complex professional work.

Motivation

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.