AI BENCHMARK PROFILE
MobileJudgeBench
MobileJudgeBench evaluates LLM-as-judge methods on mobile agent trajectories. It includes 931 human-annotated trajectories from 6 mobile agent benchmarks, covering 4 agent models and 68 apps.
- Released
- 2026-08-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Mobile agent benchmarks increasingly rely on LLM-based judges, yet their reliability is unexamined. MobileJudgeBench fills this gap by providing a standardized evaluation to select reliable judges, improving evaluation fidelity and downstream reinforcement learning.
Motivation
Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.