Benchmark Radar
AI BENCHMARK PROFILE

MobilePA-Bench

Robotics & Autonomous SystemsAgents

Evaluates mobile planning agents on tool-calling and planning across 13 domains and 212 tools in an interactive sandbox, with scoring on three advanced planning dimensions.

Released
2026-08-24
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Closes the gap between GUI-centric and static benchmarks by providing a runtime-grounded evaluation of mobile agents, highlighting reliability issues in frontier LLMs.

Motivation

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.