Skip to content

Full-App Build Tasks

Tasks to build a whole web app, graded by a browser that uses the finished app.

  • 12,800browser-graded web tasks delivered
  • 9,600app specs delivered
01

Sample

1 record

Substep tree for one application specification

Real structure with the specification withheld. The prose requirement is the graded artifact and is not published.

substeps.json
{
  "spec_id": "<withheld>",
  "grading_tier": "browser",
  "substep_tree": [{
    "id": "1",
    "check": "application serves and reaches a usable initial state"
  }, {
    "id": "1.1",
    "check": "seeded records visible",
    "depends_on": "1"
  }, {
    "id": "2",
    "check": "primary create workflow completes end to end"
  }, {
    "id": "2.1",
    "check": "created record survives reload",
    "depends_on": "2"
  }, {
    "id": "3",
    "check": "constraint violation is rejected and surfaced to the user"
  }],
  "harness_selfcheck": {
    "reference": "all-green required",
    "stub": "partial required, dependents SKIPped"
  }
}
02

Specifications

Full-App Build Tasks specifications
FormatsDocker, JSONL, Playwright, YAML
DeliverySecure download, Docker image, Seeded sandbox
CadenceContinuous build; monthly tranches
App specs delivered9,600
Browser-graded web tasks delivered12,800
Reference tranche22 web-app specs with substep trees
Grading tiersHTTP assertions and Playwright browser workflows
Reference checkEvery reference all-green, every stub partial with SKIP propagation
AccessOrg-scoped credentials. Substep trees and reference implementations are delivered separately from the specifications so evaluation and training splits stay clean.
Built againstVibe Code Bench, WebArena, WebBench

What ships

  • Prose-Spec App Builds

    The specification is written the way a non-technical customer would write it. The grader walks a substep tree against the deployed application, so partial progress scores instead of collapsing to zero.

  • Stack x Defect Matrix

    Spec-to-app tasks generated across a matrix of stacks and injected defects, each shipped as a seeded sandbox with deterministic end-to-end assertions.

  • Live-Deployment Grading

    Two tiers: HTTP-level assertions for API behavior and a real browser for user-visible workflows. Grading runs against the running application, not against source inspection.

Fields

8 fields

FieldDescription
spec_idstringIdentifier for the application specification.
spec.mdMarkdownThe requirement written as non-technical prose, the only file the solver receives.
substep_treeobjectOrdered tree of user-visible behaviors, each independently checkable, producing dense partial credit.
grading_tierenumhttp | browser. Determines whether assertions run at the API or through a driven browser.Example browser
sandboxdirectorySeeded starting state so every run begins from identical data.
reference/directoryA complete implementation required to score all-green.
stub/directoryA deliberately partial implementation required to score partial with SKIP propagation.
defectstring · nullableFor spec-to-app tasks, the injected defect drawn from the stack-by-defect matrix.

Load it

The first commands after an approved delivery.

quickstart.sh
# Deploy a candidate and grade it through the browser tier
bash harness/deploy.sh --spec <spec-id> --app ./candidate
bash harness/grade.sh  --spec <spec-id> --tier browser

# Score = satisfied substep nodes / total, with dependent nodes SKIPped.
# Harness self-check: reference must be all-green, stub must be partial.
03

Quality checks

3 checks

Reference all-green

Whether a correct implementation actually scores full marks.

The authored reference is run through the complete substep tree and required to pass every node.

Real measured: every reference in the shipped tranche scores all-green.

Stub partial with SKIP propagation

Whether partial credit and dependent-node skipping behave correctly.

A deliberately incomplete implementation is required to score partial, with nodes that depend on missing behavior marked skipped rather than failed.

Real measured: every stub in the shipped tranche scores partial with skips propagating as designed.

Model difficulty

How far a model gets through the substep tree.

The shipped harness runs any model end to end and reports the fraction of the tree satisfied against the deployed application.

Methodology-stage. The harness ships so difficulty is re-measurable per buyer and per model.

Ground truth

What is correctThe substep tree, anchored by a reference implementation that must satisfy every node and a stub that must not.
How it is establishedTwo tiers. HTTP assertions cover API contracts; a driven browser covers user-visible workflows against the live deployment. Score is the fraction of the substep tree satisfied, with dependent nodes skipped rather than failed when a prerequisite is missing.
AgreementDeterministic against a seeded sandbox. Harness correctness is evidenced by the paired reference-passes and stub-partials checks rather than by inter-rater agreement.

How it is built

  1. 01

    Write the specification as a customer would

    Requirements are authored in non-technical prose rather than as an engineering ticket, so the task measures whether a model can resolve ambiguity into a working product rather than follow a precise order.

  2. 02

    Decompose into a substep tree

    Each specification is broken into an ordered tree of user-visible behaviors. Reward is dense because a run that delivers half the application scores half rather than failing outright.

  3. 03

    Seed the sandbox

    Every task starts from identical seeded data, which is what makes two runs comparable and makes a deterministic assertion possible at all.

  4. 04

    Grade the running application

    Assertions execute against the deployed application - HTTP-level for API behavior, a real driven browser for user-visible workflows. Nothing is scored by reading the source.

  5. 05

    Validate the harness itself

    Before release, the reference implementation must score all-green and the deliberately partial stub must score partial with skips propagating correctly. A scoring change then reflects the model rather than a broken grader.

  6. 06

    Generate the stack-by-defect matrix

    Spec-to-app tasks are produced across a matrix of technology stacks and injected defects, so coverage is systematic rather than whatever happened to be convenient to author.

05RequestReplies usually within one working day

Request a sample

We send real records in the format you would receive, with the field list and where each record came from. Access for evaluation is limited and covered by a mutual NDA.

ProductFull-App Build Tasks

team@gerra.com