Skip to content

What is a spec-to-repo task?

A spec-to-repo task, the setup NL2Repo-Bench tests, gives a coding agent one written specification and an empty directory, and grades the software it builds with a test suite the agent never sees. The agent decides the package layout, the public interface and every behavior the spec describes, so the task measures whether it can build a working library from prose alone.

How it differs from a repository task

A SWE-bench-style task hands the agent an existing codebase and one problem inside it. A spec-to-repo task hands it nothing to read but the spec. Gerra has delivered 9,600 of these tasks, and in each one the spec is the only file in the agent's workspace; a verifier confirms that before the task ships.

That changes what the grader needs. The hidden tests have to import whatever the agent built, so the spec fixes the module name and the public symbols, and a structural check runs before any behavioral test. The public benchmark NL2Repo-Bench uses the same setup, a single requirements document and an empty workspace, and reports that even the strongest agents stay below a 40% average test pass rate.

It also changes the contamination risk. If the spec describes a library that exists publicly, a model can reproduce the code it read in pretraining. Gerra's tasks describe invented domains, or familiar ones with the interface renamed and the rules inverted. The real upstream library, installed as is, fails all eight mutated tasks' graders in the reference tranche, and the changed rules are what stop a recalled copy that has been renamed. The method is covered under benchmark contamination.

What a wrong build scores

Every task's grading contract names the builds that test its grader: the authored reference, a known-bad build that is plausible but wrong and, for mutated domains, the real upstream library. What a verifier is walks through that contract field by field.

Across Gerra's 12-task reference tranche, known-bad builds pass between 14% and 76% of a task's held-out assertions, 52% overall, and fail every task. The references pass all 1,308 assertions.

Read a partial score against its task's known-bad score, not against zero. Where a wrong build already earns 0.14, a score of 0.80 is far above wrong. Where it earns 0.76, 0.80 is barely better than wrong. Averaging raw scores across tasks lets the easy-to-approximate tasks flatter every model, so measure each score from its task's known-bad score instead. The same gap is the grader's discrimination range: between 24 and 86 percentage points separate a plausible wrong build from a correct one.

The spec is the whole task

  • Every assertion needs a sentence behind it. A hidden test that checks behavior the spec never states measures guessing.
  • Keep the spec private. Gerra doesn't publish task identities, spec text or mutation rules, because a published spec is a contaminated one.

The same construction scales to whole applications. Full-App Build Tasks start from a prose spec written the way a customer would write it, then a driven browser checks each user-visible behavior of the running app.