AI BENCHMARK PROFILE
SoundnessBench
SoundnessBench evaluates LLMs' ability to judge the methodological soundness of research proposals. It contains 1,099 machine-learning research proposals reconstructed from ICLR submissions, labeled with reviewer soundness sub-scores and audited against source papers.
- Released
- 2026-05-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Autonomous AI research agents need to evaluate research ideas before committing resources, but existing benchmarks do not test this bottleneck. SoundnessBench provides a standardized test for this capability.
Motivation
Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.