FORCEBENCH
FORCEBENCH is a contrastive stress test for evidence-force calibration in cited RAG. It pairs fixed cited passages with evidence-calibrated claims and force-raised variants across five axes: relation, modality, scope, temporal validity, and numeric specificity. Evaluation is a fixed, locality-filtered set of 198 pairs.
- Released
- 2026-05-27
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Current cited RAG evaluation often treats topical relevance as sufficient grounding, missing cases where a relevant source under-warrants an over-strong claim. FORCEBENCH exposes this citation laundering failure and measures evaluator calibration via monotonicity violation rate and force sensitivity.
Motivation
Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wording.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.