Benchmark Radar
AI BENCHMARK PROFILE

ResearchQA

General AIKnowledge & Reasoning

ResearchQA is a benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers across eight domains, designed for citation-grounded evaluation with multiple valid supporting passages and grounded refusal.

Released
2026-07-13
Readiness
Paper only
Primary field
General AI

Why it matters

Existing evaluation methods often fail to detect whether answers are supported by verifiable citations. ResearchQA provides a standardized way to assess citation accuracy and groundedness, separating systems more clearly than LLM-evaluator scores.

Motivation

Large language models are increasingly used to assist scientific reading, but existing evaluation methods often fail to detect whether answers are supported by verifiable citations.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.