Gerra builds reinforcement learning environments for AI agents: coding, computer use, back-office work, security, full-app builds and ML research. Every environment ships with a grader the agent can’t see, delivered separately from the tasks so a buyer can hold some back for evaluation.
Back-Office Work Tasks: what the agent gets and how it is scored
The agent gets
A month-end close on a real ledger, with names removed and one fault injected on purpose. The agent has to find the fault, fix it and tie out the close.
Score
Two scores kept apart: reconciliation to the exact cent with no tolerance band, and diagnosis of the affected voucher, so a lucky number does not count as a diagnosis.
Evidence
Both frontier models tested scored 0% on the all-or-nothing deliverables, with every piece of evidence supplied in context.
Repo Build & Repair Tasks: what the agent gets and how it is scored
The agent gets
A real issue in a real repository, served with the Git history stripped, or a library to write from a spec: invented, or mutated from a well-known package with its public names renamed and its rules inverted. Runs are in Docker with the network off.
Score
A structural check of the public symbols, then passed held-out assertions over min_passed, the exact count taken from the regenerated suite: 30 to 193 per task in the reference tranche. A task passes only at the full bar, and repository tasks must keep the existing suite green (FAIL_TO_PASS and PASS_TO_PASS).
Evidence
Every task has to pass its reference install and fail a plausible known-bad package. Where a public original exists, the real upstream package has to fail too, so the public package, installed as is, can’t pass the task: 8 of the 12 reference tasks have one, and the other 4 are invented.
Exploit & Patch Tasks: what the agent gets and how it is scored
The agent gets
A containerized service that takes a full chain to break, such as injection, then privilege escalation, then data export, across web, memory-corruption and cryptographic targets. The defensive variant asks for a patch.
Score
Offensive runs score the highest rung reached on the subtask ladder, checked against the live container. A patch scores only if the working exploit fails against it.
Evidence
A target ships only if its exploit succeeds on the original service and fails on the patched one, and each rung is tested with a run that stops at that stage.
Full-App Build Tasks: what the agent gets and how it is scored
The agent gets
spec.md, a requirement written the way a non-technical customer would write it and the only file the agent receives, plus a seeded sandbox so every run starts from the same data.
Score
Satisfied substep nodes over the total, checked by HTTP assertions for the API and a driven browser for user-visible workflows. A node whose prerequisite is missing is skipped rather than failed.
Evidence
Every reference implementation in the shipped tranche scores all-green, and every deliberately partial stub scores partial with skips propagating.
ML Research Tasks: what the agent gets and how it is scored
The agent gets
GPU-backed research: an ML competition with its real grader, a paper to replicate, or real base weights to post-train.
Score
Competitions are scored by their own grader and medal-verified. Replications roll weighted leaf criteria into one score. Post-training runs have to beat the base model on a fixed benchmark suite, with a contamination guard on the gain.
Evidence
The replication judge scores F1 0.946 against the 1,963-leaf upstream rubric, and every shipped trajectory is re-graded by its environment’s own grader.
Real Business Environments: what the agent gets and how it is scored
The agent gets
A real company’s linked records, with names removed: case, authorization, service record, claim and ledger. The task is expert-written, for example finding and fixing a mismatch between service and billing.
Score
Checks against expert answers: the affected records reconcile, the real cause is found, and linked records agree afterwards.
Evidence
The checks stay outside the agent’s view, and baseline runs measure pass rate against compute spent.
02
A business environment, stage by stage
Starting state, task and tools, grading
A service-to-billing workflow
Linked records from one company, names removed.
The case, the authorization, the service record and the accounting entry share a history. The links between them stay intact.
Case record
Request, status and prior history
Authorization
Coverage, approval and limits
Service record
What was delivered and when
Claim and ledger
Billing, posting and reconciliation
Records stay linked across systems and over time.
Find and fix the service-billing mismatch.
The agent gets a clear business goal and access to the systems it needs to look into the problem and fix it.
Task
Find and fix the mismatch between service and billing
Evidence
Case history, messages and policy
Actions
Read records, check the mismatch, update the workflow
Limits
The task’s permissions and business rules
Graded on the records, not on the agent’s report.
The final state is checked against expert answers, with checks the agent never sees.
Outcome
The affected records reconcile
Diagnosis
The real cause is found
Consistency
Linked records agree afterwards
Baseline
Pass rate measured against compute spent
Hidden checks and baseline runs show how hard each task is.
An example workflow. No customer records or task answers are shown.
03
A graded research run
Recorded step by step
Auditing a reproduction packet
The agent checks a one-dimensional convection PINN reproduction for fabricated metrics, a boundary mismatch or a PDE residual left unsatisfied. Each step shows what it meant to do, the tool it called, what came back, and the grade at the end.
Research reproduction audit
Captured run, de-identified
Pass
Task
Audit a one-dimensional convection PINN reproduction packet. Recompute the evidence from samples and distinguish a valid reproduction from fabricated metrics, a boundary mismatch, an unsatisfied residual, missing curvature gain, or insufficient final accuracy.
01 / 06Task
Toolbash
Intent
Inspect the benchmark packet
Tool call
Inspect the provided repository and the public benchmark specification.
Observed result
The task defines a one-dimensional convection PINN reproduction with a reference PDE, initial condition, periodic boundary, and analytical solution. The working tree contains the model, PDE, reproduction script, requirements, and rubric materials.
repository inspected
benchmark specification located
Recorded state2 records
surface
status
repository
inspected
reference PDE
located
Intent
Check the packet format
Tool call
Inspect the available reproduction packet and rubric files.
Observed result
The packet format and rubric are present. The auditor can be written against the normative fields without relying on a precomputed result file.
packet format checked
rubric materials located
Recorded state2 records
surface
status
packet fields
checked
rubric
located
Intent
Write the auditor
Tool call
Write a normalized audit artifact that recomputes the packet metrics and returns a classification.
Observed result
The audit artifact is written. Its public role is to recompute evidence from raw packet samples, rather than trust a reported score.
normalized audit artifact written
Recorded state2 records
artifact
status
audit.py
written
output contract
recomputed metrics and classification
Intent
Exercise failure paths
Tool call
Run local checks covering valid, fabricated, boundary, residual, curvature-gain, and final-accuracy cases.
Observed result
The local checks return the expected classifications for the valid reproduction and each named failure mode, including insufficient final accuracy.
failure-path checks completed
outputs summarized
Recorded state2 records
check family
status
valid reproduction
classified
failure paths
classified
Intent
Submit the audit run
Tool call
Submit the completed audit artifact for machine grading.
Observed result
The captured run reports the auditor as complete and ready for the machine grade.
audit artifact submitted
awaiting terminal grade
Recorded state2 records
surface
status
submission
complete
grader
pending
Intent
Terminal grade
Tool call
Evaluate the submitted auditor against six held-out reproduction-audit checks.
Observed result
The machine grader reports a passing result for all six held-out checks.
Assertion fractions, subtask ladders, substep trees
Why each task scores partial progress
Under pass/fail grading, a run that gets most of a long task right scores zero, the same as a run that did nothing, so on hard tasks the reward carries almost no signal. Partial credit has to be built into each task, because the training harness can’t know what partial progress means.
An implementation that gets 80 percent of the held-out behavior right scores 0.80 instead of 0. The same task still gives a yes-or-no answer at the full-pass bar for evaluation.
Subtask ladders
A run that reaches rung 2 of 4 on an exploit chain scores 2 of 4. Under pass/fail it scores 0, the same as a run that never authenticated.
Substep trees
When the create workflow fails, the checks that depend on it are marked skipped, so the score separates steps the agent got wrong from steps it never reached.
We send real records in the format you would receive, with the field list and where each record came from. Access for evaluation is limited and covered by a mutual NDA.