Can Computer-Use Agents Do Visual Software Engineering?
Software development is more than editing code. Developers run the program, click through its interface, look at what it does, and use those observations to decide what to change and whether a fix worked. Coding agents and computer-use agents (CUAs) have mostly been studied separately — so this integrated loop has gone unmeasured.
What CUA-SWE tests
CUA-SWE is a benchmark, environment, and evaluation pipeline for software engineering with computer use. In a single task an agent may need to:
- modify code and configuration,
- execute commands,
- interact with the running software through its GUI, and
- inspect visual feedback to diagnose and verify the repair.
Tasks span four software-engineering domains. Some go further: the required specification or operational information is available only through the running application's visual interface.
Verified, not judged
Each task ships with deterministic, task-specific tests that check whether the resulting software satisfies the requirements and preserves specified behavior — no LLM judge deciding whether a fix “looks right.”
Code-only vs hybrid computer use
The headline comparison (Figure 1) pits each frontier agent's code-only configuration against a hybrid CUA configuration that can also see and operate the app. Across eight frontier agents, the hybrid setting raises mean task success by +12.8 to +48.6 percentage points; the strongest configuration reaches 59.9% hybrid versus 11.3% code-only. The paper also examines which development behaviors are associated with successful repairs.
For visual software, an agent that can't look at the app is debugging blind.
Frequently asked questions
What is CUA-SWE?
CUA-SWE is an ICLR 2026 benchmark, environment, and evaluation pipeline for software engineering with computer use: agents edit code, run commands, interact with running applications, and inspect screenshots, and deterministic tests verify each repair.
Do computer-use abilities help coding agents fix bugs?
In CUA-SWE, yes: hybrid computer-use configurations improved mean task success over code-only configurations by 12.8 to 48.6 percentage points across eight frontier agents.
How are CUA-SWE tasks graded?
With deterministic, task-specific tests that check whether the software meets the requirements and preserves specified behavior, rather than with an LLM judge.
Based on CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering (ICLR 2026, arXiv:2609.32600) by Prince Zizhuang Wang, Chenhao Liang, Zelong Xu, Aojie Yuan, Xiaolin Zhou, Haiyue Zhang, Yue Zhao, Xiyang Hu, Shuli Jiang. Written by Aojie (Justin) Yuan, USC Fortis Lab.