Benchmark Radar
AI BENCHMARK PROFILE

SoundnessBench

General AIKnowledge & Reasoninghosytuyen

SoundnessBench evaluates LLMs' ability to judge the methodological soundness of research proposals. It contains 1,099 machine-learning research proposals reconstructed from ICLR submissions, labeled with reviewer soundness sub-scores and audited against source papers.

Released
2026-05-28
Readiness
Runnable
Primary field
General AI

Why it matters

Autonomous AI research agents need to evaluate research ideas before committing resources, but existing benchmarks do not test this bottleneck. SoundnessBench provides a standardized test for this capability.

Motivation

Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.