AI BENCHMARK PROFILE
TUA-Bench
Evaluates terminal-use agents on 120 real-world tasks across five families (document editing, email, web info, scientific/engineering workflows) using execution-based scoring.
- Released
- 2026-06-26
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a broad, realistic terminal benchmark beyond coding, showing frontier agents achieve only 65.8% success, highlighting gaps in general-purpose digital work.
Motivation
As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.