AI BENCHMARK PROFILE
PhoneHarness
PhoneHarness Bench evaluates phone-use agents on verifiable mobile workflows with mixed GUI, CLI, and tool actions, scored by observable side effects from auditable execution traces.
- Released
- 2026-06-12
- Readiness
- Runnable
- Primary field
- Consumer & Productivity
Why it matters
Mobile agent evaluation often ignores non-GUI actions and side effects; this benchmark measures complete task completion in real device environments, filling a gap for reliable phone automation.
Motivation
Phone agents are increasingly expected to complete real mobile workflows rather than merely predict the next screen action.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.