Benchmark Radar
AI BENCHMARK PROFILE

Long-Horizon-Terminal-Bench

General AIAgentsCoding & Software Engineering

Evaluates long-horizon terminal tasks in a containerized environment with hidden verifiers and dense reward grading across 46 tasks and nine categories.

Released
2026-07-09
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a more demanding evaluation for agentic long-horizon planning and partial credit, addressing gaps in existing terminal benchmarks that only measure final outcomes.

Motivation

AI agents have become capable of autonomously completing short, well-specified tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.