AI BENCHMARK PROFILE
GUITestScape
Evaluates exploratory GUI testing agents on 61 Android apps with 508 preset defects, using an open-set evaluator that decomposes trajectories into independently diagnosable capabilities.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- Consumer & Productivity
Why it matters
Addresses the lack of open-set evaluation in GUI testing, covering interaction and display defects. Provides a finer-grained assessment of agent capabilities and a verifier integration boost.
Motivation
Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an application and discover defects through its own interaction.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.