AI BENCHMARK PROFILE
MobilePA-Bench
Evaluates mobile planning agents on tool-calling and planning across 13 domains and 212 tools in an interactive sandbox, with scoring on three advanced planning dimensions.
- Released
- 2026-08-24
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Closes the gap between GUI-centric and static benchmarks by providing a runtime-grounded evaluation of mobile agents, highlighting reliability issues in frontier LLMs.
Motivation
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.