AI BENCHMARK PROFILE
ExplainBench
Evaluates code explanations from agents by checking whether explanations enable an LLM to correctly answer questions about intended behavior and patch effects, using a question-based suite derived from agent patches.
- Released
- 2026-07-29
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Agent explanations are often untrusted; this benchmark provides a quantitative way to compare explanation quality across agents, which is not captured by existing code-generation benchmarks.
Motivation
Large Language Model (LLM) agents have seen rapid adoption in software engineering.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.