Benchmark Radar
AI BENCHMARK PROFILE

RigorBench

General AIAgentsCoding & Software Engineering

RigorBench evaluates autonomous AI coding agents on engineering process discipline across five pillars: Planning Fidelity, Verification Coverage, Recovery Efficiency, Abstention Quality, and Atomic Transition Integrity. It includes 30 tasks in five categories and a composite RigorScore metric.

Released
2026-06-21
Readiness
Paper only
Primary field
General AI

Why it matters

Existing agent benchmarks focus on outcome correctness, ignoring process quality. RigorBench fills this gap by measuring how agents plan, verify, and recover, providing a more comprehensive assessment for reliable deployment in real-world software engineering.

Motivation

Agentic coding harnesses - such as Agent-Skills, Superpowers, and Agent-Rigor - are increasingly deployed to augment underlying LLMs for real-world software engineering tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.