Benchmark Radar
AI BENCHMARK PROFILE

SciStyleBench

General AISafety & TrustworthinessSearch & Retrieval

SciStyleBench diagnoses stylistic bias in LLM-based idea evaluation through controlled stylistic perturbations, metrics, and a mitigation extractor.

Released
2026-08-03
Readiness
Paper only
Primary field
General AI

Why it matters

Highlights the importance of style invariance in scientific idea evaluation and provides metrics and a mitigation module for improving LLM judges.

Motivation

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.