Benchmark Radar
AI BENCHMARK PROFILE

ICAE-Bench

General AIAgentsCoding & Software EngineeringALEX-nlp

ICAE-Bench evaluates coding agents on interactive project-building tasks. Agents receive a fuzzy product requirement and must clarify missing details via an automated user agent, then implement the project in a container. Scoring uses black-box tests and multi-dimensional diagnostics including functional correctness, semantic/API similarity, structural fidelity, design quality, and interaction quality.

Released
2026-07-23
Readiness
Runnable
Primary field
General AI

Why it matters

Existing coding benchmarks focus on static, fully specified tasks, leaving a gap for interactive, open-ended development. ICAE-Bench provides a reproducible protocol for measuring agent performance in transforming incomplete requirements into working software, which is increasingly relevant for real-world coding workflows.

Motivation

The recent emergence of vibe-coding workflows is changing what coding agents are expected to do.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.