AI BENCHMARK PROFILE
AutoLab
AutoLab evaluates frontier models on long-horizon closed-loop optimization tasks across system optimization, CUDA kernel optimization, model development, and puzzle challenges, with 36 expert-curated tasks and a strict wall-clock budget.
- Released
- 2026-06-03
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Fills the gap in evaluating sustained iterative improvement in agents, moving beyond single-turn or short-horizon benchmarks to test persistence and empirical feedback incorporation.
Motivation
Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and continuously refining artifacts.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.