Skip to content

What is an agent trajectory?

An agent trajectory is the step-by-step record of an AI agent doing a task: what it saw, what it meant to do, each tool call with its input and output, and the grade the environment gave the result. Graded successes become training data for supervised fine-tuning and distillation, and failures show where agents break.

What one step holds

Gerra has delivered 96,000 computer-use trajectories, and every ML research environment it builds ships with a run that was executed and graded. In its recorded runs, each step stores the agent's note explaining the call, the tool it called, the input, and the output with a status of ok, error or submitted. Checkpoint labels such as "Implementation" or "Local checks" mark where the run changes phase, which lets long runs be compared phase by phase. A computer-use step adds the screenshot and accessibility tree the agent saw before acting, the grounded action, and the change that followed.

In the replays Gerra publishes, every input and output carries a visibility mark: summarized, excerpted or withheld. Generated source code is withheld while the commands around it stay excerpted, so identifying payloads come out and the sequence of decisions survives.

A graded run, call by call

This is a captured run from Gerra's ML Research Tasks, published as the research reproduction audit replay. The task: audit a reproduction packet for a physics-informed neural network and classify whether its evidence supports the paper's claim.

# Agent's note before the call Tool Result
1 "I'll explore the working directory first." bash Task file, benchmark package and submission directory found
2 "The task is well-specified; check whether a sample packet or precomputed result is available." bash Rubric present; no precomputed result used
3 "I have enough to implement the auditor per the normative spec." write_file Auditor written, source withheld
4 "Run local checks covering the valid case and each named failure mode." bash Seven local cases classified, including valid-reproduction and insufficient-final-accuracy 0.1419
5 "All failure paths classify correctly; submit the completed audit artifact for machine grading." finish Submitted
Grade grader 6 of 6 held-out checks, score 1.0, PASS

The grade says the auditor works. The calls say how it got there: the agent searched for a precomputed result before writing any code, then built an auditor that recomputes the evidence from raw samples instead of trusting a reported score, and exercised every named failure path before submitting. Two runs with the same score can differ on exactly these points, and only the call log separates them.

The grade comes from outside the run

The training diagnostic replay is a 24-call run on a tabular classification task with a held-out accuracy threshold of 0.95. The agent's own cross-validation printed CV accuracy: 0.8100 +/- 0.0470. Its last check confirmed the required columns and the exact 200-row identifier set and order, and its final note read "The submission is valid; hand it off for machine evaluation." The grader returned two checks: submission format PASS, held-out accuracy 0.785, FAIL.

Filter trajectories on the environment's grade, never on the agent's last message. An agent can only validate what it can see, and held-out labels are exactly what it can't see. Keep the per-check breakdown as well: a filter that stores one verdict loses the fact that this run's format passed and only its accuracy fell short of the bar. Gerra's research environments write this rule into each environment's manifest: "trajectory": { "captured": true, "graded": true, "self_reported": false }.

Which runs to keep

  • Graded successes go to rejection sampling, supervised fine-tuning and distillation. Gerra's coding trajectories are graded by the same environment that produced them, so a trace and its finished change arrive with one verdict.
  • Failures are the baseline. Gerra's Real Business Environments ship captured rollouts and failure cases with pass-rate-versus-compute curves, which show where a task family sits before any training.
  • Neither is enough for reinforcement learning. A trajectory is a fixed record, and RL needs the environment itself so the model can make fresh attempts and have them scored. Use trajectories for imitation, environments for RL, and held-out tasks for evaluation.