Skip to content

RL environments for training AI agents

Gerra builds reinforcement learning environments for AI agents: coding, computer use, back-office work, security, full-app builds and ML research. Every environment ships with a grader the agent can’t see, delivered separately from the tasks so a buyer can hold some back for evaluation.

  • 96,000repo tasks delivered
  • 16,000exploit-and-patch tasks delivered
  • 64,000graded back-office tasks delivered
01

Environments and their graders

7 products

Back-Office Work Tasks

  • 64,000graded back-office tasks delivered
Back-Office Work Tasks: what the agent gets and how it is scored
The agent getsA month-end close on a real ledger, with names removed and one fault injected on purpose. The agent has to find the fault, fix it and tie out the close.
ScoreTwo scores kept apart: reconciliation to the exact cent with no tolerance band, and diagnosis of the affected voucher, so a lucky number does not count as a diagnosis.
EvidenceBoth frontier models tested scored 0% on the all-or-nothing deliverables, with every piece of evidence supplied in context.
Task fileinstance.json"oracle": { "expected_cents": "<withheld>", "tolerance": 0 }"scores": ["reconciliation", "diagnosis"]"difficulty_gate": { "frontier_models_tested": 2, "pass_rate": "0%" }

Computer-Use Trajectories

  • 96,000computer-use trajectories delivered
Computer-Use Trajectories: what the agent gets and how it is scored
The agent getsA browser or desktop app. Before each action the agent has a screenshot and the accessibility tree, taken at the same moment.
ScoreA deterministic check of the application’s end state, run when the episode ends. The agent’s own report of success is not used.
EvidenceEach action keeps its target role and screen coordinates, so every step can be checked against what the agent could see.
Task filetrajectory.jsonl"a11y_tree": "steps/014.a11y.json""action": { "type": "click", "target_role": "button" }"final_state_check": { "deterministic": true, "evaluated_at_end_of_episode": true }

Repo Build & Repair Tasks

  • 96,000repo tasks delivered
  • 9,600spec-to-repo tasks delivered
Repo Build & Repair Tasks: what the agent gets and how it is scored
The agent getsA real issue in a real repository, served with the Git history stripped, or a library to write from a spec: invented, or mutated from a well-known package with its public names renamed and its rules inverted. Runs are in Docker with the network off.
ScoreA structural check of the public symbols, then passed held-out assertions over min_passed, the exact count taken from the regenerated suite: 30 to 193 per task in the reference tranche. A task passes only at the full bar, and repository tasks must keep the existing suite green (FAIL_TO_PASS and PASS_TO_PASS).
EvidenceEvery task has to pass its reference install and fail a plausible known-bad package. Where a public original exists, the real upstream package has to fail too, so the public package, installed as is, can’t pass the task: 8 of the 12 reference tasks have one, and the other 4 are invented.
Task filemeta.json"kind": "mutated-domain""min_passed": 193"known_bad_install": "known-bad""upstream_sdist": "<real upstream tarball, plagiarist gate only>"

Exploit & Patch Tasks

  • 16,000exploit-and-patch tasks delivered
Exploit & Patch Tasks: what the agent gets and how it is scored
The agent getsA containerized service that takes a full chain to break, such as injection, then privilege escalation, then data export, across web, memory-corruption and cryptographic targets. The defensive variant asks for a patch.
ScoreOffensive runs score the highest rung reached on the subtask ladder, checked against the live container. A patch scores only if the working exploit fails against it.
EvidenceA target ships only if its exploit succeeds on the original service and fails on the patched one, and each rung is tested with a run that stops at that stage.
Task filetarget.json"chain": ["injection", "privilege escalation", "data export"]{ "rung": 4, "check": "protected records exported" }"exploit_vs_patched": "must fail"

Full-App Build Tasks

  • 12,800browser-graded web tasks delivered
  • 9,600app specs delivered
Full-App Build Tasks: what the agent gets and how it is scored
The agent getsspec.md, a requirement written the way a non-technical customer would write it and the only file the agent receives, plus a seeded sandbox so every run starts from the same data.
ScoreSatisfied substep nodes over the total, checked by HTTP assertions for the API and a driven browser for user-visible workflows. A node whose prerequisite is missing is skipped rather than failed.
EvidenceEvery reference implementation in the shipped tranche scores all-green, and every deliberately partial stub scores partial with skips propagating.
Task filesubsteps.json"grading_tier": "browser"{ "id": "2.1", "check": "created record survives reload", "depends_on": "2" }"stub": "partial required, dependents SKIPped"

ML Research Tasks

  • 1,600ML competition runs delivered
  • 800paper replications delivered
  • 960post-training runs delivered
ML Research Tasks: what the agent gets and how it is scored
The agent getsGPU-backed research: an ML competition with its real grader, a paper to replicate, or real base weights to post-train.
ScoreCompetitions are scored by their own grader and medal-verified. Replications roll weighted leaf criteria into one score. Post-training runs have to beat the base model on a fixed benchmark suite, with a contamination guard on the gain.
EvidenceThe replication judge scores F1 0.946 against the 1,963-leaf upstream rubric, and every shipped trajectory is re-graded by its environment’s own grader.
Task fileenv.json"rubric": { "leaves": 1963, "source": "upstream-authored" }"judge": { "kind": "automated", "measured_f1_against_rubric": 0.946 }"trajectory": { "graded": true, "self_reported": false }

Real Business Environments

Configured to your task

Real Business Environments: what the agent gets and how it is scored
The agent getsA real company’s linked records, with names removed: case, authorization, service record, claim and ledger. The task is expert-written, for example finding and fixing a mismatch between service and billing.
ScoreChecks against expert answers: the affected records reconcile, the real cause is found, and linked records agree afterwards.
EvidenceThe checks stay outside the agent’s view, and baseline runs measure pass rate against compute spent.
02

A business environment, stage by stage

Starting state, task and tools, grading

A service-to-billing workflow

Linked records from one company, names removed.

The case, the authorization, the service record and the accounting entry share a history. The links between them stay intact.

Case record
Request, status and prior history
Authorization
Coverage, approval and limits
Service record
What was delivered and when
Claim and ledger
Billing, posting and reconciliation

Records stay linked across systems and over time.

An example workflow. No customer records or task answers are shown.
03

A graded research run

Recorded step by step

Auditing a reproduction packet

The agent checks a one-dimensional convection PINN reproduction for fabricated metrics, a boundary mismatch or a PDE residual left unsatisfied. Each step shows what it meant to do, the tool it called, what came back, and the grade at the end.

Research reproduction audit

Captured run, de-identified

Pass

Task

Audit a one-dimensional convection PINN reproduction packet. Recompute the evidence from samples and distinguish a valid reproduction from fabricated metrics, a boundary mismatch, an unsatisfied residual, missing curvature gain, or insufficient final accuracy.

01 / 06Task

Toolbash

Intent

Inspect the benchmark packet

Tool call

Inspect the provided repository and the public benchmark specification.

Observed result

The task defines a one-dimensional convection PINN reproduction with a reference PDE, initial condition, periodic boundary, and analytical solution. The working tree contains the model, PDE, reproduction script, requirements, and rubric materials.

repository inspected
benchmark specification located

Recorded state2 records

surfacestatus
repositoryinspected
reference PDElocated

Another run: training diagnosticAll recorded runs

04

Partial credit

Assertion fractions, subtask ladders, substep trees

Why each task scores partial progress

Under pass/fail grading, a run that gets most of a long task right scores zero, the same as a run that did nothing, so on hard tasks the reward carries almost no signal. Partial credit has to be built into each task, because the training harness can’t know what partial progress means.

Three ways to score partial progress

  • Assertion fractions

    An implementation that gets 80 percent of the held-out behavior right scores 0.80 instead of 0. The same task still gives a yes-or-no answer at the full-pass bar for evaluation.

  • Subtask ladders

    A run that reaches rung 2 of 4 on an exploit chain scores 2 of 4. Under pass/fail it scores 0, the same as a run that never authenticated.

  • Substep trees

    When the create workflow fails, the checks that depend on it are marked skipped, so the score separates steps the agent got wrong from steps it never reached.

06RequestReplies usually within one working day

Request a sample

We send real records in the format you would receive, with the field list and where each record came from. Access for evaluation is limited and covered by a mutual NDA.

ProductRL environments

team@gerra.com