Benchmark Radar
AI BENCHMARK PROFILE

SciIntBench

General AIKnowledge & Reasoning

SciIntBench is an adversarial benchmark of 810 prompts across ten responsible-conduct-of-research categories and three scientific domains. Each scenario appears in overt, covert, and benign versions to measure framing-sensitive refusal of misconduct.

Released
2026-05-28
Readiness
Paper only
Primary field
General AI

Why it matters

LLMs are increasingly used in scientific work, but their compliance with research integrity norms is unclear. SciIntBench aims to measure how models handle framing-sensitive ethical scenarios.

Motivation

Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (RCR) norms or help undermine them.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.