Benchmark Radar
AI BENCHMARK PROFILE

Asuka-Bench

General AICoding & Software Engineering

Evaluates code agents on 50 web tasks with underspecified intent and multi-round refinement via browser-rendered behavior.

Released
2026-06-04
Readiness
Paper only
Primary field
General AI

Why it matters

Could address gaps in code-generation benchmarks by testing iterative refinement, but lacks public artifacts for reuse.

Motivation

Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.