AI BENCHMARK PROFILE
SciIntBench
SciIntBench is an adversarial benchmark of 810 prompts across ten responsible-conduct-of-research categories and three scientific domains. Each scenario appears in overt, covert, and benign versions to measure framing-sensitive refusal of misconduct.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
LLMs are increasingly used in scientific work, but their compliance with research integrity norms is unclear. SciIntBench aims to measure how models handle framing-sensitive ethical scenarios.
Motivation
Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (RCR) norms or help undermine them.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.