Benchmark Radar
AI BENCHMARK PROFILE

MacArena

General AIAgentsMacPaw

MacArena benchmarks computer-use agents on macOS with 421 manually verified tasks spanning 50 applications, running on Apple Virtualization framework on Apple Silicon.

Released
2026-06-04
Readiness
Runnable
Primary field
General AI

Why it matters

MacOS GUI challenges are underrepresented in current benchmarks. MacArena provides a harder and more diverse environment, revealing that performance on existing benchmarks may not generalize across platforms, aiding development of robust GUI agents.

Motivation

Computer-use agents (CUAs) operate graphical user interfaces (GUIs) through vision and control primitives, and their capabilities have advanced rapidly, driven in part by standardized online evaluation benchmarks such as OSWorld, which serve both as evaluation tools and as training environments for reinforcement learning.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.