SciR
SciR evaluates LLMs on deduction, induction, and causal abduction in scientific settings, with tasks generated from formal objects and rendered into multi-document scientific discourse. Difficulty is controlled along extraction and inference axes, with verifiable answers.
- Released
- 2026-06-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks either lack mechanistic ground truth or do not resemble real scientific documents. SciR provides a controllable protocol for isolating extraction vs. inference failures, which is valuable for diagnosing model capabilities in scientific reasoning.
Motivation
Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.