Workflow-GYM
Workflow-GYM evaluates AI agents on long-horizon GUI tasks in professional domains, using specialized software environments and economically valuable workflows. Tasks require end-to-end operation of graphical user interfaces, with success rates measured by task completion.
- Released
- 2026-06-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Most GUI benchmarks cover simple, short-horizon tasks in general software, leaving a gap for professional, long-horizon workflows. Workflow-GYM provides a fixed protocol for measuring agent performance on such tasks, which is valuable as organizations consider deploying agents for complex professional work.
Motivation
Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.