PseudoBench
PseudoBench evaluates agentic auto-research systems on their ability to identify and resist pseudoscientific narratives. It contains 200 curated pseudoscientific claim-evidence pairs across five domains and assesses performance through an end-to-end research pipeline from experiment design to report writing.
- Released
- 2026-06-16
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
As AI agents increasingly participate in scientific research, their susceptibility to generating plausible but misleading studies poses a risk to academic integrity. PseudoBench provides a direct measure of this risk, offering a practical tool for assessing and improving agent safety before deployment.
Motivation
As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.