AndroidDaily
AndroidDaily evaluates mobile GUI agents on 350 daily-use tasks across 94 closed-source Android applications. Automatic scoring is based on a three-tiered guideline system (operational obligations, output quality, negative constraints), with step-level diagnostic judgments.
- Released
- 2026-05-26
- Readiness
- Paper only
- Primary field
- Consumer & Productivity
Why it matters
Fills evaluation gap for real-world closed-source apps where internal states are unavailable, providing a verifiable scoring method based on observable guidelines. Useful for assessing practical deployment of GUI agents.
Motivation
The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments or open-source applications, leaving real-world closed-source applications largely unevaluated.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.