Benchmark Radar
AI BENCHMARK PROFILE

UI2App

General AICoding & Software Engineering

Benchmarks visual interaction inference in executable web application generation. Contains 327 screenshots in 45 sets, evaluating executability, navigation reachability, visual fidelity, and interaction inference via the IIS metric.

Released
2026-07-07
Readiness
Runnable
Primary field
General AI

Why it matters

Targets a gap in measuring behavior inference from screenshots, not just visual fidelity. Could help assess models' ability to produce interactive applications.

Motivation

Large language models (LLMs) have demonstrated growing competence in web page generation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.