AI BENCHMARK PROFILE
MobileWorldSafety
MobileWorldSafety evaluates GUI agents' safety against environmental injection attacks in Android apps. It includes 142 risk tasks on real applications, with programmatically verifiable risk indicators and a two-stage pipeline for verification.
- Released
- 2026-08-18
- Readiness
- Runnable
- Primary field
- Consumer & Productivity
Why it matters
Provides a quantitative measure of GUI agent vulnerability to injection attacks, enabling comparison across agents and supporting development of safer mobile agents.
Motivation
LLM-powered GUI agents that autonomously operate smartphones are rapidly transitioning from research prototypes to early real-world deployment.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.