Benchmark Radar
AI BENCHMARK PROFILE

WebIGBench

General AIMultimodal PerceptionCoding & Software Engineering

WebIGBench evaluates code generation for interactive webpages with 103 complex examples and 871 distinct actions, proposing an automated evaluation pipeline for interactive consistency.

Released
2026-05-29
Readiness
Runnable
Primary field
General AI

Why it matters

Fills the gap in benchmarking interactive webpage code generation, offering a public dataset and evaluation method for model comparison.

Motivation

Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyzing a new paradigm for front-end development.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.