OSWorld-V2 · GPT-5.5, reasoning effort xhigh · CUAWright desktop runtime

Ten desktop trajectories, step by step

Every step is one model call: the model's reasoning summary, the shell command it ran in the Ubuntu VM, the output, and any screenshot it looked at. Pick a task, then step through with the buttons, the timeline, or the ← / → keys. Yellow marks the key moments.

The agent has one tool, run_command. It gets no screenshot unless it asks for one, and it works without multi-agent orchestration or a hand-built memory system. Every script, helper, and check you see here, it wrote itself.

Why this run is interesting
    Original task instruction

    Key moments

      ← / → steps · K next key moment
      key momentscreenshot or image viewedcommand