Skip to content

What is computer-use training data?

Computer-use training data records an AI agent operating software through the screen: at each step, the screenshot and accessibility tree it saw, the mouse or keyboard action it took and where that action landed, and what changed as a result, closed by a check of the application's final state. Models trained on it learn to work in apps built for people, with no API in between.

One recorded step

Gerra has delivered 96,000 browser and desktop trajectories in this format. This is one step from its Computer-Use Trajectories, as published with the application content and coordinates withheld:

{
  "step": 14,
  "screenshot": "steps/014.png",
  "a11y_tree": "steps/014.a11y.json",
  "action": { "type": "click", "target_role": "button", "grounding": { "x": "<withheld>", "y": "<withheld>" } },
  "observation_delta": "<withheld>",
  "final_state_check": { "deterministic": true, "evaluated_at_end_of_episode": true }
}
  • screenshot and a11y_tree are captured at the same moment, before the action. A capture taken after the click would pair the action with a screen the agent never saw.
  • a11y_tree lists the elements the app exposes, with each one's role, name and state; canvas-drawn and custom controls are often missing, which is why the screenshot rides along.
  • observation_delta keeps the change next to the action that caused it, in the original order.

Grounding: from an intent to a pixel

OSWorld runs 369 tasks on real web and desktop apps. When it was published in April 2024, humans completed over 72.36% of them and the best model 12.24%, and the authors name the main causes: GUI grounding and operational knowledge. Grounding is the step from "click Save" to the exact pixel of the Save button in this window, at this size, in this state.

Data teaches grounding only if every action points at something the agent could see, which is why each step records the target's role next to its coordinates. A click position alone leaves the model to guess which element it belonged to. A role without a position can't teach where to click. Recorded together, each can be checked against the other on replay.

Success is a property of the app

OSWorld gives each task an initial-state setup and an execution-based evaluation script that inspects the result. Gerra's trajectories close the same way: a deterministic check of the application's end state, evaluated once when the episode ends, never a success flag the agent reports about itself. An agent that says the invoice was sent has told you what it believes. The invoice record in the app says whether it was sent.

The check has to be strict to be worth anything. In Gerra's Back-Office Work Tasks, cross-tool computer-use deliverables are graded against a machine oracle where one wrong figure fails the task. Two frontier models were tested on these tasks, with every piece of evidence supplied in context; both scored 0%.

Demonstrations or environments

Reinforcement learning needs the app and its final-state check, not only the record; see what an agent trajectory is. Gerra's back-office environments supply both, recording the same per-step capture and closing each episode with a deterministic check of the environment's end state.