← Writing

Can Computer-Use Agents Do Visual Software Engineering?

TL;DR — Partly — and seeing the app helps a lot. CUA-SWE (ICLR 2026) is a benchmark, environment, and evaluation pipeline where agents must edit code, run commands, interact with the running application, and inspect screenshots in the same task, across four software-engineering domains. Every task has deterministic, task-specific tests. In the paper's headline comparison, hybrid computer-use agents beat their code-only versions by 12.8 to 48.6 percentage points across eight frontier agents.

Software development is more than editing code. Developers run the program, click through its interface, look at what it does, and use those observations to decide what to change and whether a fix worked. Coding agents and computer-use agents (CUAs) have mostly been studied separately — so this integrated loop has gone unmeasured.

What CUA-SWE tests

CUA-SWE is a benchmark, environment, and evaluation pipeline for software engineering with computer use. In a single task an agent may need to:

Tasks span four software-engineering domains. Some go further: the required specification or operational information is available only through the running application's visual interface.

Verified, not judged

Each task ships with deterministic, task-specific tests that check whether the resulting software satisfies the requirements and preserves specified behavior — no LLM judge deciding whether a fix “looks right.”

Code-only vs hybrid computer use

The headline comparison (Figure 1) pits each frontier agent's code-only configuration against a hybrid CUA configuration that can also see and operate the app. Across eight frontier agents, the hybrid setting raises mean task success by +12.8 to +48.6 percentage points; the strongest configuration reaches 59.9% hybrid versus 11.3% code-only. The paper also examines which development behaviors are associated with successful repairs.

For visual software, an agent that can't look at the app is debugging blind.

Frequently asked questions

What is CUA-SWE?

CUA-SWE is an ICLR 2026 benchmark, environment, and evaluation pipeline for software engineering with computer use: agents edit code, run commands, interact with running applications, and inspect screenshots, and deterministic tests verify each repair.

Do computer-use abilities help coding agents fix bugs?

In CUA-SWE, yes: hybrid computer-use configurations improved mean task success over code-only configurations by 12.8 to 48.6 percentage points across eight frontier agents.

How are CUA-SWE tasks graded?

With deterministic, task-specific tests that check whether the software meets the requirements and preserves specified behavior, rather than with an LLM judge.


Based on CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering (ICLR 2026, arXiv:2609.32600) by Prince Zizhuang Wang, Chenhao Liang, Zelong Xu, Aojie Yuan, Xiaolin Zhou, Haiyue Zhang, Yue Zhao, Xiyang Hu, Shuli Jiang. Written by Aojie (Justin) Yuan, USC Fortis Lab.