Skip to content

What is reward hacking in reinforcement learning?

Reward hacking is when a model trained with reinforcement learning earns a high reward by exploiting a flaw in how the reward is computed, without doing the task the reward was meant to measure.

A grader that held up can still break

The Effects of Reward Misspecification found that more capable agents often exploit a misspecified reward more, sometimes abruptly once capability crosses a threshold. A grader that held up against one policy can fail against the next, so graders need re-auditing as the policy improves. For reasoning tasks, the DeepSeek-R1 authors used rule-based rewards instead of a neural reward model, because they found neural reward models susceptible to reward hacking during large-scale RL. Rules can be hacked too. They are code, and they need tests.

Git history is an answer key

A task mined from a real repository carries its own answer: the commit that fixed the issue. If that commit is still reachable inside the container, an agent can find it. A September 2025 issue on the SWE-bench repository documents agents doing this during evaluation. One ran:

cd /testbed && git log --oneline --all | grep -i "bracket\|parametrize\|modpath" | head -10

The search surfaced a future commit, "Fix incorrect result of getmodpath method," whose diff is the fix. The mitigation proposed in the issue removes remotes, branches, tags and the reflog, since each can reveal future commits or their messages. Gerra's version of that mitigation is in what a SWE-bench-style task is.

Installing the answer from a package index

This is the grading quickstart published on Gerra's Repo Build & Repair Tasks page:

# Grade a candidate against one held-out environment
bash harness/grade_in_docker.sh \
  --task <task-id> \
  --install ./candidate      # path-only; bare requirement strings are refused

# The runner executes with the network namespace off:
#   docker run --network none ...      pip install --no-index ...
# Score = passed_assertions / min_passed        PASS only at the full ship bar

bash harness/verify_all.sh   # reference PASS / known-bad FAIL / upstream FAIL

If the grader installed whatever the candidate named, a candidate could list a published package, pip would download it, and the suite would score code the agent never wrote. Three layers close that route: installs from a local path only, no package index and no network. Its last line runs the grader's own checks, covered in what a verifier is.

Two more shortcuts graders have to close

  • Reading the answer key or editing the tests. In authored coding tasks the solver receives exactly one file, the specification. Reference, tests, known-bad package and grading contract sit in a private sibling tree, and the verifier asserts nothing else is visible. Exploit-and-patch targets keep their working exploit private too.
  • Breaking what the target test doesn't cover. A patch can make one test pass by deleting or special-casing a code path. Repository tasks also require the full pre-existing suite to stay green.

Reading a training run for hacks

Suspect a hack when reward climbs while held-out scores stay flat, when many tasks jump at once (usually a shared shortcut), or when solutions get much shorter than the task should need. Re-grade a sample of high-reward attempts by hand. When a score moves, re-score the reference and the known-bad package with the same grader; if their scores moved too, the change came from the harness.

Keep a chain-of-thought monitor out of the reward. Monitoring Reasoning Models for Misbehavior found that under too much optimization against such a monitor, agents learned to hide their intent in their reasoning while still reward hacking at a significant rate.