Repo Build & Repair Tasks
Coding tasks from written specs and real issues, each with a repository that runs and hidden tests that grade the change.
- 96,000repo tasks delivered
- 9,600spec-to-repo tasks delivered
Sample
2 recordsGrading contract for one environment
Real field shape from a shipped task. Identity, symbol names, and mutation rules are withheld here on purpose - publishing them would contaminate a graded environment we sell.
{
"task_id": "05-<withheld>",
"kind": "mutated-domain",
"import_name": "<withheld>",
"public_symbols": ["<5 renamed symbols>"],
"grader": {
"pytest_paths": ["."],
"min_passed": 193,
"source": "Regenerated from the authored reference. Asserts the MUTATED semantics."
},
"reference_install": "reference",
"known_bad_install": "known-bad",
"upstream_sdist": "<real upstream tarball, plagiarist gate only>",
"anti_memorization": "Public surface renamed; every rule inverted from the well-known
analogue. A memorized reproduction fails the structural check first and the
behavioral long tail second - the public package, installed as is, fails the task.",
"provenance": "Authored from scratch. Upstream is retained only to prove it fails."
}Held-out assertion distribution across the reference tranche
Real per-task counts from the 12-task reference tranche. Task identities are withheld.
task,origin,held_out_assertions,plagiarist_gate 01,invented,178,none 02,mutated,121,upstream pinned 03,mutated,74,upstream pinned 04,mutated,63,upstream pinned 05,mutated,193,upstream pinned 06,mutated,149,upstream pinned 07,mutated,51,upstream pinned 08,mutated,61,upstream pinned 09,mutated,30,upstream pinned 10,invented,161,none 11,invented,116,none 12,invented,111,none ,,1308,8 of 12 gated
Specifications
| Formats | Docker, JSONL, Patch, YAML |
|---|---|
| Delivery | Secure download, Docker image, Git bundle |
| Cadence | Continuous build; monthly tranches |
| Spec-to-repo tasks delivered | 9,600 |
| Long-horizon repo tasks delivered | 96,000 |
| Held-out assertions | 30 to 193 per task in the reference tranche |
| Environment | pip-installed in Docker, history-stripped |
| Grading | FAIL_TO_PASS + full PASS_TO_PASS regression |
| Access | Org-scoped credentials with per-customer delivery paths. Held-out graders ship under a separate restricted agreement from the solver-visible specifications so evaluation tasks can be split from training tasks at delivery. |
| Built against | NL2Repo-Bench (ICML 2026), SWE-bench Verified, SWE-bench Pro |
What ships
Spec-to-Repo Tasks
A natural-language library specification plus a held-out assertion suite. The agent writes the library; the suite decides. Every task is validated against a plagiarist gate that proves the answer is not reachable by copying an existing package.
Long-Horizon Repository Tasks
Multi-file, cross-subsystem changes mined from real permissively-licensed repositories - real issue, real fix, real held-out tests - served in a history-stripped environment so the answer cannot be read out of the Git log.
Teacher Trajectories
Full tool-call traces from frontier coding agents solving these environments, with reasoning preserved, graded on the same held-out suites and passed through a Git-history cheat scan.
Fields
11 fields
| Field | Description |
|---|---|
| task_idstring | Stable task identifier within the tranche.Example 07-<withheld> |
| kindenum | invented-domain | mutated-domain | mined-issue.Example mutated-domain |
| import_namestring | Module the structural check imports before any behavioral test runs. |
| public_symbolsstring[] | Symbols the candidate must expose; a memorized upstream package fails here first. |
| grader.min_passedint · assertions | Exact held-out assertion count, counted from the regenerated suite.Example 193 |
| grader.pytest_pathsstring[] | Held-out suite roots. Never shipped inside the solver-visible tree. |
| reference_installpath | Authored correct implementation. Must install and pass at the ship bar. |
| known_bad_installpath | Plausible-but-wrong package. Must install cleanly, then fail the suite. |
| upstream_sdiststring · nullable | Real upstream tarball for the plagiarist gate; null for invented domains. |
| anti_memorizationstring | Written record of every renamed symbol and mutated rule, kept private to the grading tree. |
| provenancestring | Authorship and licensing of the reference implementation. |
Load it
The first commands after an approved delivery.
# Grade a candidate against one held-out environment bash harness/grade_in_docker.sh \ --task <task-id> \ --install ./candidate # path-only; bare requirement strings are refused # The runner executes with the network namespace off: # docker run --network none ... pip install --no-index ... # Score = passed_assertions / min_passed PASS only at the full ship bar bash harness/verify_all.sh # reference PASS / known-bad FAIL / upstream FAIL
Quality checks
4 checksPlagiarist gate
Whether reproducing the real public library passes the task.
The genuine upstream distribution is installed into the task environment and run through the task own grader. The gate requires it to fail.
Real measured: upstream fails on every mutated-domain task in the reference tranche. The four invented-domain tasks have no public analogue to gate against.
Known-bad discrimination
Whether the grader scores behavior rather than mere presence of the right symbols.
A plausible-but-wrong package is required to pass the structural symbol check and then fail the behavioral suite.
Real measured: passes on every task in the reference tranche.
Held-out assertion count
How much graded surface each task actually carries.
Assertions are counted directly from the regenerated suite and recorded as the exact pass threshold in the grading contract.
Real measured: 30 to 193 held-out assertions per task across the 12-task reference tranche.
Copy guard
Whether a submission is byte-near the shipped reference or the real upstream source.
A similarity check runs on every submission with a configurable hard-fail threshold, as defense in depth behind the network isolation.
Methodology-stage. Reported per submission; no aggregate false-positive rate is published.
Ground truth
| What is correct | The authored reference implementation and its regenerated held-out suite. For mined repository tasks, the real upstream fix commit and the tests that commit made pass are the ground truth. |
|---|---|
| How it is established | A structural symbol check runs first, then the held-out suite. Score is the fraction of held-out assertions passed, so an honest 80 percent implementation scores 0.80 rather than zero. A task only counts as passed at the full ship bar. Repository tasks additionally require the complete pre-existing suite to stay green. |
| Agreement | Grading is deterministic and execution-based, so there is no inter-rater statistic. Discrimination is evidenced instead by the known-bad and upstream gates, both of which are required to fail. |
How it is built
- 01
Author or mutate the domain
Each task is either an invented domain with no public analogue, or a familiar domain whose public surface is renamed and whose semantics are deliberately inverted. A model that reproduces the library it remembers is wrong on the graded long tail rather than merely unlucky.
- 02
Regenerate the held-out suite from the reference
Tests are generated from the authored reference implementation and assert the mutated behavior. No upstream test suite is ever vendored in, so the suite cannot agree with a memorized package by construction.
- 03
Structurally isolate the held-out material
The solver receives exactly one file per task: the specification. Reference, regenerated tests, known-bad package, grading contract, and upstream tarballs live in a sibling private tree, and the verifier asserts that nothing but the specification exists in the solver-visible directory.
- 04
Grade offline and path-only
Grading runs in a pinned image with the network namespace off. Installs use a no-index flag and bare requirement strings are refused outright, so the grader can only ever score source the candidate wrote.
- 05
Prove the grader discriminates
Every task is verified four ways before it ships: the reference passes at the ship bar, the known-bad package installs and then fails, the real upstream library run through the task grader fails, and the solver-visible tree contains nothing but specifications.
- 06
Mine and validate long-horizon tasks
Repository tasks are mined by locating paired source and test commits in permissively-licensed projects, then validating each as a genuine fail-then-pass transition inside a history-stripped environment. Candidates that do not reproduce that transition are discarded.
Request a sample
We send real records in the format you would receive, with the field list and where each record came from. Access for evaluation is limited and covered by a mutual NDA.
ProductRepo Build & Repair Tasks