Benchmark Radar
AI BENCHMARK PROFILE

Lit2Test

General AIKnowledge & Reasoning

A benchmark for falsifiable research ideation, using a six-field contract centered on a falsifying outcome, with pairwise comparisons of proposals from four models.

Released
2026-08-24
Readiness
Paper only
Primary field
General AI

Why it matters

Introduces a decidable evaluation contract for research proposals, moving beyond free-form judging and enabling reliable comparison of model-generated research ideas.

Motivation

Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.