AI BENCHMARK PROFILE
SciConBench
SciConBench evaluates AI agents on open-domain scientific conclusion synthesis using 9,110 questions and expert-written conclusions from systematic reviews. The evaluation decomposes conclusions into atomic facts and measures correctness and comprehensiveness via factual precision and recall.
- Released
- 2026-06-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The benchmark addresses the lack of reliable evaluation for AI agents summarizing scientific evidence in health and other critical fields, providing a way to measure factual accuracy and completeness.
Motivation
Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.