RL environments and grading
What is benchmark contamination?
Benchmark contamination is when a model has already seen a benchmark's test items or answers during training, so its score reflects what it remembers instead of what it can do. In coding it also covers a less obvious case: if a task asks the model to build something that already exists as public code, that code is the answer, and it is very likely in the training data.
Leaked items and recalled implementations
The familiar kind is direct leakage. Test questions, reference solutions or test files end up in the training corpus, and the model reproduces them.
The kind that matters more for coding agents is recall of the thing being tested. The tasks are new, but they ask for a library or program that exists publicly. A large model reproduces the code it has read, the tests pass, and the score measures retrieval. In reinforcement learning the damage compounds, because the reward attaches to recall and the model learns that recall pays.
Why recency, hidden tests and overlap checks miss it
- Recency. Using problems published after a model's training cutoff helps against direct leakage, and LiveCodeBench collects new problems over time for this reason. But the code a new task asks for can be years older than the task, so a date after the cutoff says nothing about whether the answer was in training.
- Hidden tests. Keeping the test suite private doesn't help when the implementation under test is public. The model has read the implementation.
- Overlap checks. Decontamination that searches training text for matches misses paraphrased copies of test items, as shown in Rethinking Benchmark and Contamination for Language Models with Rephrased Samples. For a task that asks for a public library there is nothing to match: the training data holds the implementation, not the task.
The plagiarist gate
A task can be tested for recall directly. Keep the real upstream library, install it into the task's own environment, run it through the task's own grader and require it to fail. If the genuine public implementation passes, the task can be passed by recall, and it fails the gate. If the implementation fails, the task ships with evidence a buyer can re-run.
Any grader can make the upstream fail by being unpassable, so the gate only counts next to a reference that passes the same grader. The task itself has to sit where recall can't reach. An invented domain has no public implementation to install, such as a calendar with thirteen 28-day months and a leap rule of its own. A mutated domain renames a familiar library's public interface and changes its rules, such as a heap module that pops the largest item where Python's heapq pops the smallest. Renaming alone fails, because a model reads the specification and maps new names onto remembered behavior. The rules underneath have to differ.
The grader around the gate needs its own checks, described in What is a verifier in AI training?.
One tranche: 1,308 held-out assertions, 8 upstream gates
This is the held-out assertion count for each task in one reference tranche of Gerra's coding environments, as published on the Repo Build & Repair Tasks page (task identities withheld):
task,origin,held_out_assertions,plagiarist_gate
01,invented,178,none
02,mutated,121,upstream pinned
03,mutated,74,upstream pinned
04,mutated,63,upstream pinned
05,mutated,193,upstream pinned
06,mutated,149,upstream pinned
07,mutated,51,upstream pinned
08,mutated,61,upstream pinned
09,mutated,30,upstream pinned
10,invented,161,none
11,invented,116,none
12,invented,111,none
,,1308,8 of 12 gated
The eight mutated tasks each pin a real upstream package. The upstream fails the task grader on all eight while the reference passes on all twelve, as reported in the RL environment supply case study. The four invented tasks have no public implementation to gate against. Each count is also the task's exact pass threshold. Installed as is, the upstream fails at import because the task renames the public interface. What stops a model that recalls the library and renames it is the mutated behavior the held-out assertions test.
When the agent trains the model
Contamination turns into a reward hack when the agent builds the training data. In a post-training environment the agent sources data, trains a model and is scored on how far it beats the base model on a fixed benchmark suite. Test items that turn up in the sourced data raise that score with no real gain. Overlap checks miss paraphrased items, so compare the gain on the benchmark with the gain on held-out items the agent had no way to find. A gain that shows up only on the public benchmark is recall.