Benchmark Radar
AI BENCHMARK PROFILE

SPIEval

General AIKnowledge & Reasoning

SPIEval is a human-curated benchmark for evaluating large language models as mobile assistants that retrieve and reason over personal information scattered across multiple apps. It comprises 250 tasks across five cognitive capabilities, 4,335 fictional personal records in 10 simulated apps, and supports multi-turn interaction through 21 tools. Models receive underspecified user instructions and must search records and invoke tools; final execution calls are compared to human-annotated gold calls at the parameter level.

Released
2026-08-11
Readiness
Inspectable
Primary field
General AI

Why it matters

Existing benchmarks do not target the challenge of leveraging scattered personal information across apps in mobile assistant settings. SPIEval provides a controlled, verifiable evaluation environment grounded in five cognitive capabilities, enabling assessment of model capabilities in realistic mobile contexts. The reported results show substantial room for improvement and highlight fundamental limitations in information localization and search efficiency, offering practical guidance for deploying LLM-based assistants.

Motivation

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.