Webwright is now CUAWright

CUAWright: A Minimal Unified Interface for All Digital Tasks.

CUAWright is a minimal unified interface for browser and desktop agents. The model gets a terminal, writes code that drives a browser or an Ubuntu desktop, checks the result, and keeps its work on disk.

Yadong Lu, Theodore Lee, Yifei Li, Lawrence Keunho Jang, Tianci Xue, Yu Su, Huan Sun, Ahmed Hassan Awadallah

cuawright desktop
67.9%OSWorld-V2+5.2 over GPT-5.6 Sol alone
88.1%Online-Mind2Web+4.7 over a GUI agent
77.5%Odysseys+44.0 over a GUI agent
−$9.5API cost per taskOSWorld-V2, GPT-5.6 Sol

# idea

Stop predicting clicks. Start writing programs.

Most computer-use agents watch one screen and guess the next click. As models get better at code, that loop becomes the bottleneck. CUAWright keeps the small loop from Webwright and extends it to the desktop.

{ }

Code is the action

A form, a date picker, or a spreadsheet edit becomes a short program with loops and checks, not a long chain of clicks.

~/

The workspace is the state

Scripts, logs, and screenshots stay on disk, so the agent can restart a browser and retry without losing progress.

✓

Done means verified

Text alone never finishes a task. Browser runs rerun a final script; desktop runs end with an explicit submit command.

It's what's not in CUAWright.

Everything else, the agent builds itself as it works.

  • No tools except run_command in a terminal
  • No screenshot on every step by default
  • No multi-agent orchestration
  • No human-engineered RAG or memory system
$ cuawright web

Browser runtime

Writes Playwright scripts, opens and closes browsers freely, and only looks at screenshots when needed.

  • OpenAI, Anthropic, and OpenRouter backends
  • A reusable final_script.py from each solved task
  • Plugin for Claude Code, Codex, and Hermes
$ python explore_results.py   # inspect the page
$ python final_script.py      # rerun in a fresh folder
$ python -m cuawright.webwright.tools.self_reflection
Status: success
$ cuawright desktop

Desktop runtime

Works inside an OSWorld-V2 Ubuntu VM through a guest bridge, with CDP, xdotool, keyboard, mouse, and screenshots.

  • Your own tasks, or scored OSWorld-V2 tasks
  • Pinned release, task setup, evaluator, and provenance
  • Saves manifest.json, trace.jsonl, result.json
$ libreoffice --headless --convert-to xlsx report.csv
$ python check_cells.py       # verify the saved file
$ python /opt/cuawright-tools/submit.py
submitted

# results

Same model, better performance and lower cost.

Each row compares CUAWright with the best existing harness running the same model.

Same model, other harness CUAWright
benchmarkmodelscore (%)gain
OSWorld-V2: partial score. Online-Mind2Web, Odysseys: success rate. WeaveBench: pass rate. BenchCAD: Vision2Code mean IoU. CADGenBench: aggregate score.
Bar charts of CUAWright results on OSWorld-V2, CADGenBench, BenchCAD, Online-Mind2Web, Odysseys, and WeaveBench against the same model in a different harness and other baselines
All six benchmarks, including other baselines.
OSWorld-V2 partial score against API cost per task. CUAWright with GPT-5.6 Sol reaches 67.9% at about $4.5 per task
OSWorld-V2: higher score at lower API cost for both models.

# demo

See it in 78 seconds.

# trajectories

Watch it work, one step at a time.

Ten OSWorld-V2 runs with GPT-5.5 at xhigh reasoning effort. The agent writes scripts, checks its own work, and recovers when something fails. Open any run to step through its reasoning, commands, outputs, and screenshots.

$ python test_simple.py --ckpt_iter 200
✓ Correct predictions:    0 (0.00%)
$ … --ckpt_iter 1000
✓ Correct predictions: 5546 (55.46%)
$ … --ckpt_iter 5000
✓ Correct predictions: 9087 (90.87%)
$ … --ckpt_iter 7000
✓ Correct predictions: 9281 (92.81%)
01 / terminal · pytorch

Train a model until it hits the target

Trains longer and longer runs until test accuracy goes from 0% to 92.8%, then fills in the meeting note.

02 / chrome · xdotool

Solve a jigsaw puzzle without DevTools

Cuts the pieces out of screenshots, finds the layout whose edges fit, then drags each piece with the mouse.

03 / python · pillow

Copy a photo edit from one example

Measures the colour shift in a before/after pair, tests several transfer methods, and applies the best one.

04 / pptx · libreoffice

Fix slide layout by editing the PPTX

Edits the slide XML, takes the title colour from the logo's pixels, and renders after each change.

$ curl …/chart/NVDA
HTTP Error 429: Too Many Requests
$ python fetch.py  # retry, 5 tickers
NVDA 122 days  2025-01-02 → 06-30
AAPL 122 days  …
$ python fill_workbooks.py
Filled daily.xlsx and summary.xlsx
05 / python · openpyxl

Fill stock workbooks from market data

Downloads prices for five stocks, works around a rate limit, and fills two workbooks in 16 steps.

06 / word · pdf · files

Match exam mistakes to past papers

Searches 61 PDFs, matches scans visually with contact sheets, and finds the last paper in the Trash.

07 / freecad · svg

Model a stepped shaft in FreeCAD

Reads exact dimensions from the drawing's vector paths and builds the part with a revolve script.

08 / calc · thunderbird

Build a contract and email generator

Builds macro buttons through UNO and fixes a failing loop with Python helpers: 20 contracts, 20 drafts.

09 / thunderbird · calendar

Schedule thesis defenses from an email

Reads the schedule attached to an email, adds 7 defenses, and removes 2 conflicting events.

10 / 3d slicer · simpleitk

Check and correct a tumor segmentation

Reviews MRI overlay montages slice by slice, then fixes the mask while keeping the BraTS labels.

Scores come from the official OSWorld-V2 evaluator. Nine runs score 0.99 or higher; the FreeCAD run scores 0.90. Open the trajectory viewer →

# quick start

Up and running in three commands.

01 / browser

Run a web task

Writes a trajectory, screenshots, and a reusable script.

pip install -e .
playwright install chromium

cuawright-web main \
  -c base.yaml -c model_openai.yaml \
  -t "<task>" --start-url <url> \
  --task-id demo -o outputs/demo
02 / desktop

Run a desktop task

Needs Python 3.12+, Docker, and the gated OSWorld-V2 snapshots.

pip install -e ".[desktop]"
cuawright setup osworld --root <dir>

cuawright desktop run \
  --setup <dir>/setup.json \
  --credentials <key-file> \
  --results <out> \
  --model <model> --instruction "<task>"
03 / plugin

Add it to your agent

The host agent runs the loop, so no extra API key is needed.

# Claude Code
/plugin marketplace add microsoft/Webwright
/plugin install cuawright@cuawright

# Codex
codex plugin marketplace add \
  microsoft/Webwright
Formerly Webwright. The browser runtime lives on as cuawright.webwright; old commands and imports still work. Webwright page →

# cite

BibTeX

@misc{lu2026cuawright,
  title         = {CUAWright: A Minimal Unified Interface for Digital Agents},
  author        = {Lu, Yadong and Lee, Theodore and Li, Yifei and Jang, Lawrence Keunho and Xue, Tianci and Su, Yu and Sun, Huan and Awadallah, Ahmed Hassan},
  year          = {2026},
  eprint        = {2610.04116},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2610.04116}
}