Benchmark Radar
AI BENCHMARK PROFILE

MacAgentBench

General AIKnowledge & ReasoningJetAstra

MacAgentBench evaluates computer use agents on macOS desktop automation, with 676 tasks across 25 applications, including GUI and CLI interactions. It uses deterministic rule-based evaluation and fine-grained multi-checkpoint scoring to assess sub-goal completion.

Released
2026-06-21
Readiness
Runnable
Primary field
General AI

Why it matters

MacAgentBench addresses the need for benchmarks that capture framework-augmented agent capabilities and partial progress on long-horizon, multi-application tasks, providing a more granular comparison for real-world desktop automation.

Motivation

Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.