AI BENCHMARK PROFILE
SciStyleBench
SciStyleBench diagnoses stylistic bias in LLM-based idea evaluation through controlled stylistic perturbations, metrics, and a mitigation extractor.
- Released
- 2026-08-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Highlights the importance of style invariance in scientific idea evaluation and provides metrics and a mitigation module for improving LLM judges.
Motivation
However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.