RL environments and grading
What is RLVR (reinforcement learning with verifiable rewards)?
RLVR, short for reinforcement learning with verifiable rewards, trains a model on tasks whose answers a program can check, and uses that check as the reward. There is no human rating and no learned reward model in the loop, so the reward is deterministic and cheap, and the grader decides what the model learns.
Two papers that set the pattern
The Tülu 3 post-training report introduced the name and applied it to math problems and to instructions with checkable constraints. DeepSeek-R1 trained reasoning with rule-based rewards for answer accuracy and output format. Its authors avoided neural reward models for reasoning because they found them susceptible to reward hacking during large-scale RL, and used reward models only for general data such as helpfulness and safety. Agent environments extend the same loop to long tasks: many tool calls, then one graded end state. Gerra has delivered 96,000 repo tasks, 64,000 graded back-office tasks and 16,000 exploit-and-patch tasks of this kind; RL environments for training AI agents lists each family with its grader.
Pass/fail starves long tasks
On a hard task, a pass/fail grader gives almost every attempt a zero. In group-based methods such as GRPO, introduced in DeepSeekMath, each attempt's advantage is its reward relative to the other attempts on the same prompt, so a group that all scored zero has zero advantage and produces no policy gradient; only GRPO's KL penalty still moves the model. DAPO over-samples and filters out prompts whose attempts are all right or all wrong for this reason, which keeps the batch useful but drops exactly the tasks the model can't yet do. Partial credit keeps them in: attempts that pass different fractions of a task have different rewards even when none passes fully. Verifiable partial credit has to be designed into the task, because only the task knows what partial progress means. Three constructions build it in:
- Assertion fractions. Score the share of held-out assertions passed, with the exact count written into the grading contract. In Gerra's coding reference tranche the count runs from 30 to 193 per task. A separate full-pass bar keeps the same task usable as a pass/fail evaluation.
- Subtask ladders. Report the furthest stage reached. A security target that needs injection, then privilege escalation, then data export should record that a run reached stage two.
- Substep trees. Check each user-visible behavior on its own, and mark a behavior whose prerequisite is missing as skipped rather than failed. Attempted-and-wrong and never-reached are different signals.
One graded episode
The training diagnostic replay is a captured agent run on a public tabular classification task with a held-out accuracy bar of 0.95:
| Step | What happened |
|---|---|
| Baseline | Cross-validation near 0.79 |
| Group structure | 117 train/test group overlaps checked as a possible label shortcut; too inconsistent to use |
| Tuning | Best cross-validation about 0.805 |
| Submission | Schema and identifiers valid for all 200 test rows |
| Grade | Submission format PASS, held-out accuracy 0.785, verdict FAIL |
Three details matter for reward design. Against a 0.95 bar a pass/fail reward pays this run zero, and would pay the same for every attempt near 0.8, so the held-out accuracy of 0.785 is the usable signal. The grader reports format and accuracy as separate checks, so a reward can credit the valid submission without crediting the model. And the agent went looking for a label shortcut on its own, so test an environment for one before training on it: had the group overlap predicted the labels, a verifier that only compared predictions would have paid for the leak.
Hard enough to learn from
A task already solved teaches nothing, and neither does a task where every attempt earns the same reward. Gerra's back-office instance files record how unsolved each task is, "difficulty_gate": { "frontier_models_tested": 2, "pass_rate": "0%" }, which is headroom for evaluation. For training, the same file scores "scores": ["reconciliation", "diagnosis"] separately, so an attempt can earn one without the other. Each authored coding task's grader is checked against a correct reference, a known-bad package and, where a public library exists, the real library before it ships.