Benchmark Radar
AI BENCHMARK PROFILE

ConflictBench

General AIKnowledge & Reasoning

ConflictBench is a benchmark with ConflictScore metric to quantify how well models acknowledge conflicting evidence in grounding documents, decomposing responses into claims and labeling them against documents.

Released
2026-06-24
Readiness
Paper only
Primary field
General AI

Why it matters

Existing factuality metrics ignore coexisting contradictions; ConflictScore provides a nuanced measure over ConflictBench, offering a corrective feedback mechanism for improving truthfulness.

Motivation

Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting and contradicting evidence coexist.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.