Desktop-Delta Bench
Desktop-Delta Bench (DDB) evaluates computer-use models on step-level GUI transition understanding through two tasks: temporal ordering of 3-frame observations and before-after pair classification with five action types, covering 2,013 human-verified instances across ~15 applications.
- Released
- 2026-07-28
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks focus on end-task success or single-frame grounding, missing the ability to reconstruct causal transitions. DDB provides a diagnostic layer for state verification, source tracking, and context-aware control, enabling targeted improvements in desktop CUA reliability and recovery.
Motivation
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.