Benchmark Radar
AI BENCHMARK PROFILE

EdgeBench

Science & ResearchAgentsCoding & Software EngineeringMathematics & Formal SciencesByteDance Seed

EdgeBench evaluates autonomous agents on 134 real-world tasks across scientific discovery, software engineering, optimization, knowledge work, formal mathematics, and games. Each task requires 12+ hours of continuous interaction with multi-level feedback. Scoring tracks agent performance over time (at 2,4,6,8,10,12 hours). A public leaderboard is maintained; 51 tasks and evaluation framework are open-sourced.

Released
2026-07-06
Readiness
Runnable
Primary field
Science & Research

Why it matters

EdgeBench fills a gap in evaluating agents' ability to learn from real-world environments over extended periods, providing a standardized protocol for comparing long-horizon learning capabilities. It offers a public leaderboard and open-source tasks, enabling reproducible comparisons and tracking of agent improvement over time, which is valuable for model selection and development.

Motivation

Pretraining scaling laws reveal that model capability improves predictably with data and compute.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.