AI BENCHMARK PROFILE
RWGBench
Evaluates related work generation as citation-centric scholarly positioning. Uses 100 peer-reviewed papers and a 1.09M-document retrieval corpus, with metrics for citation selection, contextual appropriateness, organization, and discourse structure.
- Released
- 2026-05-30
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Fills the gap in RWG evaluation that relies on surface text similarity, offering a citation-centric testbed that aligns with expert judgment and reveals systematic limitations in current systems.
Motivation
Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.