AI BENCHMARK PROFILE
GRACE
GRACE evaluates step-level faithfulness of chain-of-thought reasoning in context-grounded tasks, with human annotations and a taxonomy of error categories.
- Released
- 2026-06-15
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Step-level faithfulness assessment addresses the gap where response-level metrics miss localized reasoning failures, providing granular feedback for improving model reliability.
Motivation
Many reasoning tasks require models to reason over input context, from document-grounded question answering to rule-based deduction.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.