Skip to content

ML Research Tasks

Machine-learning competitions, paper replications and post-training tasks, with recorded runs and automatic grades.

  • 1,600ML competition runs delivered
  • 800paper replications delivered
  • 960post-training runs delivered
01

Sample

A recorded run and real records

Watch an agent work through a task

A recorded run from our research tasks, step by step: what the agent meant to do, the tool it called, what came back, and the grade at the end.

Research reproduction audit

Captured run, de-identified

Pass

Task

Audit a one-dimensional convection PINN reproduction packet. Recompute the evidence from samples and distinguish a valid reproduction from fabricated metrics, a boundary mismatch, an unsatisfied residual, missing curvature gain, or insufficient final accuracy.

01 / 06Task

Toolbash

Intent

Inspect the benchmark packet

Tool call

Inspect the provided repository and the public benchmark specification.

Observed result

The task defines a one-dimensional convection PINN reproduction with a reference PDE, initial condition, periodic boundary, and analytical solution. The working tree contains the model, PDE, reproduction script, requirements, and rubric materials.

repository inspected
benchmark specification located

Recorded state2 records

surfacestatus
repositoryinspected
reference PDElocated

Another run: training diagnosticAll recorded runs

Environment manifest for a rubric-graded replication

Real structure. Paper identity, rubric contents, and grader internals are withheld - the rubric is the graded artifact.

env.json
{
  "env_id": "<withheld>",
  "kind": "replication",
  "rubric": {
    "leaves": 1963,
    "source": "upstream-authored",
    "weights": "rubric/weights.json"
  },
  "judge": {
    "kind": "automated",
    "measured_f1_against_rubric": 0.946
  },
  "trajectory": {
    "captured": true,
    "graded": true,
    "self_reported": false
  },
  "gpu_protocol": "A10-class, agent-driven replication",
  "ships": ["grader/", "rubric/", "trajectory/"]
}
02

Specifications

ML Research Tasks specifications
FormatsDocker, JSONL, Checkpoint, Rubric JSON
DeliverySecure download, Docker image, GPU protocol
CadenceContinuous build; monthly tranches
ML competition runs delivered1,600
Paper replications delivered800
Post-training runs delivered960
Replication rubric1,963 leaf criteria, upstream-authored
Judge agreementF1 0.946 against the rubric
AccessOrg-scoped credentials. Rubrics and graders are delivered separately from the environments so evaluation tasks can be withheld from training.
Built againstMLE-bench (OpenAI), PaperBench (OpenAI), PostTrainBench

What ships

  • Competition Environments

    A competition directory, a real grader, and a captured agent run that was medal-verified rather than self-reported. Runs execute on consumer-to-midrange accelerators.

  • Paper Replication Environments

    The real upstream leaf rubric with a weighted grader and an automated judge measured against it, plus a captured replication trajectory that was actually graded.

  • Post-Training Environments

    The agent sources and curates its own data, trains, and has to beat the base model across a fixed suite of benchmarks. Ships with a contamination guard and real weights.

Fields

9 fields

FieldDescription
env_idstringIdentifier for the research environment.
kindenumcompetition | replication | post-training.Example replication
rubric.leavesint · criteria · nullableLeaf criteria in the upstream-authored grading rubric.Example 1963
rubric.weightsJSON · nullablePer-node weighting used to roll leaf scores into a replication score.
grader/directoryThe real grader for the environment, not a proxy metric.
trajectory/directoryA captured agent run that was actually graded, shipped as evidence rather than described.
gpu_protocolstringThe accelerator class and run protocol the environment expects.Example H100 post-training protocol
contamination_guardJSON · nullableCheck that a post-training gain came from training rather than from recall of the benchmark.
base_weightspath · nullableReal starting weights the agent has to improve on.

Load it

The first commands after an approved delivery.

quickstart.sh
# Run one replication environment and score it
bash harness/run.sh --env <env-id> --agent ./candidate --gpu a10

# Weighted leaf rubric -> replication score.
# The automated judge is itself scored against the rubric (F1 0.946).

bash harness/verify_trajectory.sh --env <env-id>   # re-grade the shipped run
03

Quality checks

3 checks

Judge agreement against the rubric

Whether the automated replication judge matches the real rubric.

The judge is scored against the upstream-authored leaf rubric and reported as an F1 rather than assumed correct.

Real measured: F1 0.946 against the 1,963-leaf rubric.

Trajectory verification

Whether the shipped example runs actually achieved what they claim.

Each captured trajectory is re-scored by the environment own grader; competition runs are medal-verified rather than self-reported.

Real measured: every shipped trajectory in the reference tranche is graded evidence.

Base-model improvement

Whether a post-training run genuinely beat its starting point.

The trained model is compared to the base model across a fixed benchmark suite, with a contamination guard applied to the result.

Methodology-stage per run. The comparison and the guard ship with the environment so a buyer can reproduce the check.

Ground truth

What is correctThe real competition grader, the real upstream leaf rubric, and the measured benchmark suite for post-training runs.
How it is establishedCompetition runs are scored by the competition own grader and medal-verified. Replications roll weighted leaf criteria into a score, with the automated judge itself measured against the rubric. Post-training runs are scored by benchmark delta against the base model, filtered through a contamination guard.
AgreementThe replication judge reports F1 0.946 against the upstream rubric. Competition and post-training scoring are execution-based and deterministic.

How it is built

  1. 01

    Use the real grader, not a proxy

    Competition environments ship the actual competition grader and replication environments ship the real upstream leaf rubric. A proxy metric would make the reward cheap to game and the result meaningless.

  2. 02

    Weight the rubric

    Leaf criteria roll up through a weighted tree so a multi-day replication produces a dense scoring surface instead of one subjective judgement at the end.

  3. 03

    Measure the automated judge against the rubric

    The judge that scores replication work is itself evaluated against the rubric rather than trusted, and its agreement is reported as a number.

  4. 04

    Capture and grade a real trajectory

    Every environment ships with an agent run that was actually executed and actually graded. Medal or score claims are verified against the grader, not self-reported by the run.

  5. 05

    Guard post-training against recall

    A post-training environment requires the agent to source data, train, and beat the base model across a fixed suite. A contamination guard separates a real gain from the model having seen the benchmark.

  6. 06

    Run on a stated accelerator protocol

    Each environment declares the accelerator class it expects, from consumer-to-midrange cards for competition runs up to datacenter parts for post-training, so results are reproducible rather than hardware-dependent folklore.

05RequestReplies usually within one working day

Request a sample

We send real records in the format you would receive, with the field list and where each record came from. Access for evaluation is limited and covered by a mutual NDA.

ProductML Research Tasks

team@gerra.com