Benchmark Radar
AI BENCHMARK PROFILE

LayerRAG-Bench

General AIKnowledge & ReasoningSearch & RetrievalMusa Shams

Evaluates cross-layer reliability of agentic RAG systems on 240 tasks across 8 enterprise domains, with 9 fault scenarios and 2 contract modes; measures success at evidence, tool-contract, authorization, and session-state layers.

Released
2026-07-29
Readiness
Runnable
Primary field
General AI

Why it matters

Groundedness alone misses operational failures. This benchmark isolates which layer a mitigation repairs, supporting targeted reliability improvements rather than blanket fixes.

Motivation

Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.