AI BENCHMARK PROFILE
Long-Horizon-Terminal-Bench
Evaluates long-horizon terminal tasks in a containerized environment with hidden verifiers and dense reward grading across 46 tasks and nine categories.
- Released
- 2026-07-09
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a more demanding evaluation for agentic long-horizon planning and partial credit, addressing gaps in existing terminal benchmarks that only measure final outcomes.
Motivation
AI agents have become capable of autonomously completing short, well-specified tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.