RL environments and grading
What is a verifier in AI training?
A verifier is the check, usually code such as a test suite and sometimes a trained model, that decides whether a model's output is correct and turns that decision into a score. In an evaluation the score is the benchmark number; in reinforcement learning it is the reward, so every flaw in the verifier becomes something the model learns.
Tests, exact matches and learned judges
Most verifiers are code: tests that execute a patch, an exact match on a final answer, an assertion over the end state of an app. SWE-bench grades each task with tests that fail before the real fix and pass after it. A learned verifier is a model that scores outputs. In Training Verifiers to Solve Math Word Problems, one was trained to predict whether a candidate solution is correct, with labels taken from whether the solution reached the right final answer, and the authors note that this marks some solutions with flawed reasoning as correct. A learned verifier inherits the flaws of its labels, so it has to be scored against a trusted reference before its scores become a reward. Gerra measures its paper-replication judge this way: F1 0.946 against a 1,963-leaf rubric.
A grading contract, field by field
Gerra writes each coding task's verifier down as a contract. Abridged from the one published on the Repo Build & Repair Tasks page:
{
"kind": "mutated-domain",
"import_name": "<withheld>",
"public_symbols": ["<5 renamed symbols>"],
"grader": { "pytest_paths": ["."], "min_passed": 193 },
"reference_install": "reference",
"known_bad_install": "known-bad",
"upstream_sdist": "<real upstream tarball, plagiarist gate only>"
}
import_nameandpublic_symbolsdrive a structural check that runs before any behavioral test. A memorized copy of the public library fails there first, because its names differ.min_passedis the exact assertion count, counted from the generated suite. Score is passed divided by 193, so passing 150 scores 0.78, and the task only counts as passed at 193. An estimated denominator makes every reward approximate. The count also sets the reward's resolution: 193 assertions move it in steps of about 0.005, while the six held-out checks in the research reproduction audit replay move it in sixths.- The last three fields exist to test the grader.
Gate results on a reference tranche
Before a task ships it has to clear four gates, three of which run against known inputs. These are the published results for a 12-task reference tranche, from the RL environment supply case study:
| Gate | Input | Required result | Tranche |
|---|---|---|---|
| G1 Reference | Authored correct implementation | Passes at the full ship bar | 12 of 12 |
| G2 Known-bad | Plausible but wrong package | Installs, clears the structural check, then fails | 12 of 12 |
| G3 Plagiarist | Real upstream library | Fails | 8 of 8 gated tasks |
| G4 Isolation | Files the solver can see | Specification only | 12 of 12 |
G2 catches a common grader bug: checking that the right functions exist instead of checking what they return. G3 applies only where a public library exists; the four invented-domain tasks have none. Because the names are changed, the unmodified upstream fails at import; the mutated rules are what stop a recalled copy that has been renamed.
Graders that give partial credit need one more input: an incomplete stub. It must score partial. If it scores zero, the partial credit doesn't work. If it scores near full, the checks can't separate good work from bad, and the reward is noise. Gerra's app-build tasks ship one as stub/, and every stub in the shipped tranche scores partial with skips propagating.
Grader bugs that return plausible numbers
- A test suite copied from upstream asserts the original library's behavior, so it rewards a memorized copy of the thing the task changed. Gerra regenerates every suite from its own reference.
- A right total by the wrong route. Back-office tasks plant a known fault in a real, de-identified ledger. The instance file records
"oracle": { "expected_cents": "<withheld>", "tolerance": 0 }, and the diagnosis of the faulty record is scored separately, so a lucky total can't pass as a diagnosis.