Benchmark Radar
AI BENCHMARK PROFILE

iOSWorld

Consumer & ProductivityAgentsiOSWorld Team

iOSWorld evaluates phone agents on interactive tasks within a native iOS simulator. It includes 26 custom iOS apps with connected personal data, 133 tasks across single-app, multi-app, and memory/personalization categories. Agents are evaluated with rubric-based scoring.

Released
2026-06-08
Readiness
Runnable
Primary field
Consumer & Productivity

Why it matters

Existing mobile benchmarks lack personalization and interactive evaluation. iOSWorld provides a benchmark with persistent user identity and multi-app tasks, showing significant gaps in multi-app and memory-based performance.

Motivation

A useful phone agent needs to be personally intelligent.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.