Benchmark Radar
AI BENCHMARK PROFILE

RQ-Bench

General AIKnowledge & Reasoning

RQ-Bench evaluates novelty of research questions generated by LLMs against author-anchored reference questions from recent arXiv papers, using standalone and comparative LLM judging as well as human expert evaluation.

Released
2026-06-10
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the reliability of LLM-based novelty assessment for scientific ideation, indicating that LLM judges may produce a 'novelty mirage' compared to human experts. Useful for researchers evaluating automated scientific review or generation systems.

Motivation

LLMs are increasingly used to generate and judge scientific ideas.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.