Benchmark Radar
AI BENCHMARK PROFILE

FORCEBENCH

General AIKnowledge & Reasoning

FORCEBENCH is a contrastive stress test for evidence-force calibration in cited RAG. It pairs fixed cited passages with evidence-calibrated claims and force-raised variants across five axes: relation, modality, scope, temporal validity, and numeric specificity. Evaluation is a fixed, locality-filtered set of 198 pairs.

Released
2026-05-27
Readiness
Paper only
Primary field
General AI

Why it matters

Current cited RAG evaluation often treats topical relevance as sufficient grounding, missing cases where a relevant source under-warrants an over-strong claim. FORCEBENCH exposes this citation laundering failure and measures evaluator calibration via monotonicity violation rate and force sensitivity.

Motivation

Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wording.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.