Benchmark Radar
AI BENCHMARK PROFILE

RWGBench

General AIKnowledge & Reasoning

Evaluates related work generation as citation-centric scholarly positioning. Uses 100 peer-reviewed papers and a 1.09M-document retrieval corpus, with metrics for citation selection, contextual appropriateness, organization, and discourse structure.

Released
2026-05-30
Readiness
Runnable
Primary field
General AI

Why it matters

Fills the gap in RWG evaluation that relies on surface text similarity, offering a citation-centric testbed that aligns with expert judgment and reveals systematic limitations in current systems.

Motivation

Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.