Benchmark Radar
AI BENCHMARK PROFILE

MT-Web2Code

General AIMultimodal Perception

MT-Web2Code evaluates coding agents on multi-turn web UI reconstruction and modification tasks across 102 tasks in 16 domains, with a dual-axis protocol measuring target-region fidelity and preservation of unaffected content.

Released
2026-08-04
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks focus on single-turn full-page generation, missing the iterative workflow of real frontend engineering. This benchmark aims to fill that gap and identify weaknesses in multi-turn UI coding agents.

Motivation

Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.