Benchmark Radar
AI BENCHMARK PROFILE

IntegrityBench

General AIKnowledge & Reasoning

IntegrityBench evaluates language models on research integrity tasks, including misconduct classification, ethical action reasoning, and artifact-grounded decision making, across 36 paired tasks with a 5-level pressure protocol spanning multiple domains and research stages.

Released
2026-06-03
Readiness
Paper only
Primary field
General AI

Why it matters

As language models are used as co-scientists, measuring their integrity under pressure is critical. This benchmark could inform deployment decisions and identify risks of facilitating misconduct or eroding trust, but the evaluation method and reproducibility are not yet specified.

Motivation

Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.