Skip to content

What is an RL environment?

An RL environment is a task packaged so a model can attempt it thousands of times: a starting state that resets exactly, the tools the agent acts through, and a grader that turns the result into a reward. In reinforcement learning the model learns whatever that grader pays for.

Why the reset has to be exact

For a language-model agent an action is a tool call, and the reward usually arrives once, at the end: the training diagnostic replay makes 24 calls before its only grade. Training then runs every task many times. Group-based methods such as GRPO sample several attempts at the same task and score each one relative to the group: its reward minus the group mean, divided by the group's standard deviation, as described in DeepSeekMath.

That comparison is cleanest when every attempt starts from the identical state; a start that varies adds its own noise to every advantage. Gerra ships each environment with a resettable execution surface: its app-build tasks start every run from identical seeded data, and its coding tasks were accepted against a written requirement that the same submission gets the same score on every run. Why pass/fail rewards waste hard tasks is covered in what RLVR is.

Six families, six reward shapes

Gerra builds environments on real business systems and repositories, and writes others from scratch, such as spec-to-repo coding tasks graded by held-out assertions. Each family below has its anatomy and a recorded run on RL environments for training AI agents.

Delivered The agent starts from Graded by Reward
64,000 graded back-office tasks A real, de-identified ledger with a planted fault An oracle exact to the cent Two scores: the reconciled total, and which voucher was corrupted
96,000 repo tasks and 9,600 spec-to-repo tasks A history-stripped repository, or one specification file A held-out pytest suite, network off Assertions passed over an exact count, 30 to 193 per task in the reference tranche
16,000 exploit-and-patch tasks A running service with an authored flaw One checker per attack stage, or the exploit re-run against the agent's patch Highest stage reached
9,600 app specs and 12,800 browser-graded web tasks A prose spec and a seeded sandbox HTTP calls and a driven browser against the running app Share of a behavior tree satisfied, with dependent checks skipped
96,000 computer-use trajectories A browser or desktop app A deterministic check on the final state Pass or fail on the final state
1,600 ML competition runs, 800 paper replications and 960 post-training runs Data, a paper or base weights, on a stated GPU class The competition's own grader, a paper's rubric, or a benchmark against the base model Score, weighted rubric roll-up, or gain over the base

Repo tasks also keep a full-pass bar, so the same environment doubles as a pass/fail evaluation.

Where the right answer comes from

Every grader needs ground truth, and each source of it fails in its own way.

  • Planted. Inject a known fault into real data and the oracle knows both the exact figure and the record responsible. The risk is a fault that's easy to spot, so back-office instances are run against frontier models before shipping, and any they solve easily are dropped.
  • Mined. A real commit whose tests fail before it and pass after it carries its own answer. The risk is that the answer is still inside the environment. Gerra replays each candidate as a fail-then-pass transition in a history-stripped copy and discards any that don't reproduce.
  • Authored. Write a reference solution and generate the tests from it. The risk is a grader that checks names instead of behavior, so a plausible wrong package must install, pass the structural check, then fail.
  • Borrowed. Use a competition's real grader or a paper's upstream rubric. For a rubric, the risk moves to the automated judge that applies it, which has to be measured before it is trusted.

The tests a grader itself should pass before any model sees it are covered in what a verifier is.

Scorable is not the same as hard

A grader proves a task can be scored. It says nothing about whether the task is hard, and a task every model already solves teaches nothing. Difficulty has to be measured against current models, and measured again when they change. Gerra's exploit-and-patch targets ship a calibration harness that records the highest stage a given model reaches, so difficulty is measured against the model being trained instead of asserted.

Split training from evaluation before the run starts, because a task a model has trained on stops measuring it. Gerra delivers held-out graders apart from what the solver sees, so one family can be split at delivery. What an agent tries when a split or a grader leaks is covered in what reward hacking is.