Benchmark Radar
AI BENCHMARK PROFILE

MobileJudgeBench

General AIAgents

MobileJudgeBench evaluates LLM-as-judge methods on mobile agent trajectories. It includes 931 human-annotated trajectories from 6 mobile agent benchmarks, covering 4 agent models and 68 apps.

Released
2026-08-11
Readiness
Paper only
Primary field
General AI

Why it matters

Mobile agent benchmarks increasingly rely on LLM-based judges, yet their reliability is unexamined. MobileJudgeBench fills this gap by providing a standardized evaluation to select reliable judges, improving evaluation fidelity and downstream reinforcement learning.

Motivation

Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.