Agent and coding data
What is a SWE-bench-style task?
A SWE-bench-style task gives a coding agent a real repository checked out just before a known fix, plus a description of the problem, and grades the agent's patch by running tests: the tests the real fix made pass must now pass, and the tests that already passed must still pass. The format comes from SWE-bench, which built 2,294 such tasks from GitHub issues and pull requests in 12 Python repositories.
Base, fix and tests
- The base revision is what the agent sees: the repository at the parent of the fix.
- The reference change is the human patch. It stays hidden and proves the task can be solved. It isn't the answer key, because a different correct patch should also pass.
- The verification is two lists of tests, which the SWE-bench records call
FAIL_TO_PASSandPASS_TO_PASS. The first fail on the base and pass after the fix, which proves the problem was solved. The second pass on both sides, which catches a patch that fixes the issue by breaking something else.
SWE-bench stores the fix and the tests as separate patches and applies the tests only at grading time, so the agent never sees them.
Most commits never become tasks
A commit makes a poor task when it only reformats code, bumps dependencies and fills the diff with lockfile noise, merges unrelated work, or renames things mechanically. It also fails when it needs private services or data, when its tests already pass on the base revision, or when the expected behavior can't be stated without showing the patch. A test that passes before the fix proves nothing about the fix.
A clean history still yields no graded task without a runnable environment. The public sample from Gerra's Full-History Codebases is one complete repository with a 23-commit history, and five of those commits form authentic patch pairs. Its static checks pass: 98 of 98 JavaScript files parse and no relative import fails to resolve. Runtime test success is not claimed, because the exact dependency tree and lockfile were unavailable. Without a runnable environment no pair can show a fail-then-pass transition, and Gerra's rule is to label such a task static rather than claim execution-based grading.
Where the answer leaks in
A task is only fair if the fix is absent from everything the agent can reach. Git history is the obvious leak, since a full clone contains the commit that fixed the issue. The quieter ones:
- a test added later whose name describes the fix;
- a retrieval index built over files from after the base revision;
- documentation updated after the change;
- a dependency pin that resolves to the repaired release;
- commit messages and branch names that name the solution.
Gerra builds repository tasks in history-stripped environments and replays each candidate there as a genuine fail-then-pass transition. Candidates that don't reproduce are discarded. For public repositories one leak survives all of this: the merged fix is on the web, and a model may have read it during pretraining. That problem, and how to test for it, is benchmark contamination. Spec-to-repo tasks reduce it by describing invented domains, or familiar ones whose rules have changed.
What the grader has to accept
Two items on Gerra's acceptance checklist protect the agent's side of the task:
- The problem must be understandable from the task statement alone, without the reference patch.
- The tests must accept any behaviorally correct fix. Matching the human patch is too narrow, because a correct fix can take a different route through the code.
Score observable behavior first, then scope, regressions and unnecessary changes. Gerra's Repo Build & Repair Tasks are graded on behavior: the target tests must pass and the full pre-existing suite must stay green. 96,000 have been delivered.