AI BENCHMARK PROFILE
Lit2Test
A benchmark for falsifiable research ideation, using a six-field contract centered on a falsifying outcome, with pairwise comparisons of proposals from four models.
- Released
- 2026-08-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Introduces a decidable evaluation contract for research proposals, moving beyond free-form judging and enabling reliable comparison of model-generated research ideas.
Motivation
Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.