AI BENCHMARK PROFILE
UI2App
Benchmarks visual interaction inference in executable web application generation. Contains 327 screenshots in 45 sets, evaluating executability, navigation reachability, visual fidelity, and interaction inference via the IIS metric.
- Released
- 2026-07-07
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Targets a gap in measuring behavior inference from screenshots, not just visual fidelity. Could help assess models' ability to produce interactive applications.
Motivation
Large language models (LLMs) have demonstrated growing competence in web page generation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.