Benchmark Radar
AI BENCHMARK PROFILE

GUITestScape

Consumer & ProductivityAgentsGUITestScape Team

Evaluates exploratory GUI testing agents on 61 Android apps with 508 preset defects, using an open-set evaluator that decomposes trajectories into independently diagnosable capabilities.

Released
2026-05-28
Readiness
Paper only
Primary field
Consumer & Productivity

Why it matters

Addresses the lack of open-set evaluation in GUI testing, covering interaction and display defects. Provides a finer-grained assessment of agent capabilities and a verifier integration boost.

Motivation

Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an application and discover defects through its own interaction.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.