RigorBench
RigorBench evaluates autonomous AI coding agents on engineering process discipline across five pillars: Planning Fidelity, Verification Coverage, Recovery Efficiency, Abstention Quality, and Atomic Transition Integrity. It includes 30 tasks in five categories and a composite RigorScore metric.
- Released
- 2026-06-21
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing agent benchmarks focus on outcome correctness, ignoring process quality. RigorBench fills this gap by measuring how agents plan, verify, and recover, providing a more comprehensive assessment for reliable deployment in real-world software engineering.
Motivation
Agentic coding harnesses - such as Agent-Skills, Superpowers, and Agent-Rigor - are increasingly deployed to augment underlying LLMs for real-world software engineering tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.