EdgeBench
EdgeBench evaluates autonomous agents on 134 real-world tasks across scientific discovery, software engineering, optimization, knowledge work, formal mathematics, and games. Each task requires 12+ hours of continuous interaction with multi-level feedback. Scoring tracks agent performance over time (at 2,4,6,8,10,12 hours). A public leaderboard is maintained; 51 tasks and evaluation framework are open-sourced.
- Released
- 2026-07-06
- Readiness
- Runnable
- Primary field
- Science & Research
Why it matters
EdgeBench fills a gap in evaluating agents' ability to learn from real-world environments over extended periods, providing a standardized protocol for comparing long-horizon learning capabilities. It offers a public leaderboard and open-source tasks, enabling reproducible comparisons and tracking of agent improvement over time, which is valuable for model selection and development.
Motivation
Pretraining scaling laws reveal that model capability improves predictably with data and compute.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.