Benchmark Radar
AI BENCHMARK PROFILE

LongWebBench

General AIKnowledge & ReasoningLongWebBench team

LongWebBench evaluates structural and functional webpage generation of long webpages. It includes 490 webpages for structural fidelity and 129 for functional interactions, using VLM-based metrics and a DOM-augmented agent pipeline.

Released
2026-06-16
Readiness
Runnable
Primary field
General AI

Why it matters

Current webpage generation benchmarks focus on short static pages, missing long-horizon coherence and interactive functionality. LongWebBench provides a reusable evaluation to assess models on executable multi-step interactions, which is critical for real-world deployment.

Motivation

Recent vision-language models (VLMs) have shown promising progress in generating webpages from visual inputs, yet existing evaluations mainly focus on short, single-screen, and largely static webpages.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.