Benchmark Radar
AI BENCHMARK PROFILE

AndroidDaily

Consumer & ProductivityAgents

AndroidDaily evaluates mobile GUI agents on 350 daily-use tasks across 94 closed-source Android applications. Automatic scoring is based on a three-tiered guideline system (operational obligations, output quality, negative constraints), with step-level diagnostic judgments.

Released
2026-05-26
Readiness
Paper only
Primary field
Consumer & Productivity

Why it matters

Fills evaluation gap for real-world closed-source apps where internal states are unavailable, providing a verifiable scoring method based on observable guidelines. Useful for assessing practical deployment of GUI agents.

Motivation

The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments or open-source applications, leaving real-world closed-source applications largely unevaluated.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.