Benchmark Radar
AI BENCHMARK PROFILE

DeskCraft

General AIKnowledge & ReasoningDeskCraft team

DeskCraft evaluates desktop GUI agents on long-horizon professional workflows and human-in-the-loop collaboration across 538 executable tasks in live Ubuntu desktop environments. Covers design, video, audio, and 3D creation software, with execution-based verification.

Released
2026-06-02
Readiness
Runnable
Primary field
General AI

Why it matters

Existing desktop benchmarks simplify tasks and lack human-agent interaction. DeskCraft measures agent performance on realistic workflows and proactive collaboration, identifying gaps in long-horizon delivery and clarification.

Motivation

Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop coordination, where agents proactively seek necessary information and users provide additional instructions, clarifications, feedback, or corrections as the task progresses.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.