MacArena
MacArena benchmarks computer-use agents on macOS with 421 manually verified tasks spanning 50 applications, running on Apple Virtualization framework on Apple Silicon.
- Released
- 2026-06-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
MacOS GUI challenges are underrepresented in current benchmarks. MacArena provides a harder and more diverse environment, revealing that performance on existing benchmarks may not generalize across platforms, aiding development of robust GUI agents.
Motivation
Computer-use agents (CUAs) operate graphical user interfaces (GUIs) through vision and control primitives, and their capabilities have advanced rapidly, driven in part by standardized online evaluation benchmarks such as OSWorld, which serve both as evaluation tools and as training environments for reinforcement learning.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.