Benchmark Radar
AI BENCHMARK PROFILE

ASI-Bench

Science & ResearchAgentsApex Intelligence AI

ASI-Bench evaluates AI systems on 60 project-level scientific research tasks across 11 domains, with four guidance levels B1-B4 measuring autonomous execution and innovation.

Released
2026-08-18
Readiness
Runnable
Primary field
Science & Research

Why it matters

It addresses the gap in evaluating AI's independent scientific exploration and execution, revealing current dependence on human guidance for end-to-end research.

Motivation

Evaluate whether AI agents can independently select methods, execute end-to-end research, and produce verifiable scientific results as human methodological guidance is progressively withdrawn.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.