Benchmark Radar
AI BENCHMARK PROFILE

ExplainBench

General AICoding & Software Engineering

Evaluates code explanations from agents by checking whether explanations enable an LLM to correctly answer questions about intended behavior and patch effects, using a question-based suite derived from agent patches.

Released
2026-07-29
Readiness
Runnable
Primary field
General AI

Why it matters

Agent explanations are often untrusted; this benchmark provides a quantitative way to compare explanation quality across agents, which is not captured by existing code-generation benchmarks.

Motivation

Large Language Model (LLM) agents have seen rapid adoption in software engineering.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.