AI BENCHMARK PROFILE
Hack-Verifiable Terminal Bench
Adapts Terminal Bench with hack-verifiable environments to automatically detect reward hacking in agent trajectories.
- Released
- 2026-08-22
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides reliable, automatic measurement of reward hacking, enabling comparison of model robustness without human judges.
Motivation
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.