Benchmark Radar
AI BENCHMARK PROFILE

PhoneHarness

Consumer & ProductivityAgents

PhoneHarness Bench evaluates phone-use agents on verifiable mobile workflows with mixed GUI, CLI, and tool actions, scored by observable side effects from auditable execution traces.

Released
2026-06-12
Readiness
Runnable
Primary field
Consumer & Productivity

Why it matters

Mobile agent evaluation often ignores non-GUI actions and side effects; this benchmark measures complete task completion in real device environments, filling a gap for reliable phone automation.

Motivation

Phone agents are increasingly expected to complete real mobile workflows rather than merely predict the next screen action.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.