Benchmark Radar
AI BENCHMARK PROFILE

TUA-Bench

General AICoding & Software EngineeringFacebook Research

Evaluates terminal-use agents on 120 real-world tasks across five families (document editing, email, web info, scientific/engineering workflows) using execution-based scoring.

Released
2026-06-26
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a broad, realistic terminal benchmark beyond coding, showing frontier agents achieve only 65.8% success, highlighting gaps in general-purpose digital work.

Motivation

As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.