AI BENCHMARK PROFILE
ASI-Bench
ASI-Bench evaluates AI systems on 60 project-level scientific research tasks across 11 domains, with four guidance levels B1-B4 measuring autonomous execution and innovation.
- Released
- 2026-08-18
- Readiness
- Runnable
- Primary field
- Science & Research
Why it matters
It addresses the gap in evaluating AI's independent scientific exploration and execution, revealing current dependence on human guidance for end-to-end research.
Motivation
Evaluate whether AI agents can independently select methods, execute end-to-end research, and produce verifiable scientific results as human methodological guidance is progressively withdrawn.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.