LongWebBench
LongWebBench evaluates structural and functional webpage generation of long webpages. It includes 490 webpages for structural fidelity and 129 for functional interactions, using VLM-based metrics and a DOM-augmented agent pipeline.
- Released
- 2026-06-16
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Current webpage generation benchmarks focus on short static pages, missing long-horizon coherence and interactive functionality. LongWebBench provides a reusable evaluation to assess models on executable multi-step interactions, which is critical for real-world deployment.
Motivation
Recent vision-language models (VLMs) have shown promising progress in generating webpages from visual inputs, yet existing evaluations mainly focus on short, single-screen, and largely static webpages.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.