Benchmark Radar
AI BENCHMARK PROFILE

ClarifyCodeBench

General AICoding & Software Engineering

ClarifyCodeBench evaluates LLMs on clarifying ambiguous requirements for code generation through interactive dialogues. It includes manual annotations of ambiguity types, clarification questions, and ground-truth answers, with metrics for interaction quality.

Released
2026-07-01
Readiness
Paper only
Primary field
General AI

Why it matters

Code generation in practice involves underspecified requirements, but existing benchmarks assume perfect prompts. ClarifyCodeBench addresses this by measuring a critical yet underexplored capability, revealing that strong code generation does not imply effective requirement clarification.

Motivation

Large Language Models have emerged as programming assistants.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.