AI BENCHMARK PROFILE
RQ-Bench
RQ-Bench evaluates novelty of research questions generated by LLMs against author-anchored reference questions from recent arXiv papers, using standalone and comparative LLM judging as well as human expert evaluation.
- Released
- 2026-06-10
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the reliability of LLM-based novelty assessment for scientific ideation, indicating that LLM judges may produce a 'novelty mirage' compared to human experts. Useful for researchers evaluating automated scientific review or generation systems.
Motivation
LLMs are increasingly used to generate and judge scientific ideas.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.