AI BENCHMARK PROFILE
Asuka-Bench
Evaluates code agents on 50 web tasks with underspecified intent and multi-round refinement via browser-rendered behavior.
- Released
- 2026-06-04
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Could address gaps in code-generation benchmarks by testing iterative refinement, but lacks public artifacts for reuse.
Motivation
Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.