AI BENCHMARK PROFILE
LayerRAG-Bench
Evaluates cross-layer reliability of agentic RAG systems on 240 tasks across 8 enterprise domains, with 9 fault scenarios and 2 contract modes; measures success at evidence, tool-contract, authorization, and session-state layers.
- Released
- 2026-07-29
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Groundedness alone misses operational failures. This benchmark isolates which layer a mitigation repairs, supporting targeted reliability improvements rather than blanket fixes.
Motivation
Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.