ML Research Tasks
Machine-learning competitions, paper replications and post-training tasks, with recorded runs and automatic grades.
- 1,600ML competition runs delivered
- 800paper replications delivered
- 960post-training runs delivered
Sample
A recorded run and real recordsWatch an agent work through a task
A recorded run from our research tasks, step by step: what the agent meant to do, the tool it called, what came back, and the grade at the end.
Research reproduction audit
Captured run, de-identified
Pass
Audit a one-dimensional convection PINN reproduction packet. Recompute the evidence from samples and distinguish a valid reproduction from fabricated metrics, a boundary mismatch, an unsatisfied residual, missing curvature gain, or insufficient final accuracy.
01 / 06Task
Toolbash
- Intent
Inspect the benchmark packet
- Tool call
Inspect the provided repository and the public benchmark specification.
- Observed result
The task defines a one-dimensional convection PINN reproduction with a reference PDE, initial condition, periodic boundary, and analytical solution. The working tree contains the model, PDE, reproduction script, requirements, and rubric materials.
repository inspected benchmark specification located
Recorded state2 records
surface status repository inspected reference PDE located
- Intent
Check the packet format
- Tool call
Inspect the available reproduction packet and rubric files.
- Observed result
The packet format and rubric are present. The auditor can be written against the normative fields without relying on a precomputed result file.
packet format checked rubric materials located
Recorded state2 records
surface status packet fields checked rubric located
- Intent
Write the auditor
- Tool call
Write a normalized audit artifact that recomputes the packet metrics and returns a classification.
- Observed result
The audit artifact is written. Its public role is to recompute evidence from raw packet samples, rather than trust a reported score.
normalized audit artifact written
Recorded state2 records
artifact status audit.py written output contract recomputed metrics and classification
- Intent
Exercise failure paths
- Tool call
Run local checks covering valid, fabricated, boundary, residual, curvature-gain, and final-accuracy cases.
- Observed result
The local checks return the expected classifications for the valid reproduction and each named failure mode, including insufficient final accuracy.
failure-path checks completed outputs summarized
Recorded state2 records
check family status valid reproduction classified failure paths classified
- Intent
Submit the audit run
- Tool call
Submit the completed audit artifact for machine grading.
- Observed result
The captured run reports the auditor as complete and ready for the machine grade.
audit artifact submitted awaiting terminal grade
Recorded state2 records
surface status submission complete grader pending
- Intent
Terminal grade
- Tool call
Evaluate the submitted auditor against six held-out reproduction-audit checks.
- Observed result
The machine grader reports a passing result for all six held-out checks.
verdict: PASS held-out checks: 6 / 6 score: 1.0
Recorded state2 records
grade field value held-out checks 6 / 6 verdict Pass
Environment manifest for a rubric-graded replication
Real structure. Paper identity, rubric contents, and grader internals are withheld - the rubric is the graded artifact.
{
"env_id": "<withheld>",
"kind": "replication",
"rubric": {
"leaves": 1963,
"source": "upstream-authored",
"weights": "rubric/weights.json"
},
"judge": {
"kind": "automated",
"measured_f1_against_rubric": 0.946
},
"trajectory": {
"captured": true,
"graded": true,
"self_reported": false
},
"gpu_protocol": "A10-class, agent-driven replication",
"ships": ["grader/", "rubric/", "trajectory/"]
}Specifications
| Formats | Docker, JSONL, Checkpoint, Rubric JSON |
|---|---|
| Delivery | Secure download, Docker image, GPU protocol |
| Cadence | Continuous build; monthly tranches |
| ML competition runs delivered | 1,600 |
| Paper replications delivered | 800 |
| Post-training runs delivered | 960 |
| Replication rubric | 1,963 leaf criteria, upstream-authored |
| Judge agreement | F1 0.946 against the rubric |
| Access | Org-scoped credentials. Rubrics and graders are delivered separately from the environments so evaluation tasks can be withheld from training. |
| Built against | MLE-bench (OpenAI), PaperBench (OpenAI), PostTrainBench |
What ships
Competition Environments
A competition directory, a real grader, and a captured agent run that was medal-verified rather than self-reported. Runs execute on consumer-to-midrange accelerators.
Paper Replication Environments
The real upstream leaf rubric with a weighted grader and an automated judge measured against it, plus a captured replication trajectory that was actually graded.
Post-Training Environments
The agent sources and curates its own data, trains, and has to beat the base model across a fixed suite of benchmarks. Ships with a contamination guard and real weights.
Fields
9 fields
| Field | Description |
|---|---|
| env_idstring | Identifier for the research environment. |
| kindenum | competition | replication | post-training.Example replication |
| rubric.leavesint · criteria · nullable | Leaf criteria in the upstream-authored grading rubric.Example 1963 |
| rubric.weightsJSON · nullable | Per-node weighting used to roll leaf scores into a replication score. |
| grader/directory | The real grader for the environment, not a proxy metric. |
| trajectory/directory | A captured agent run that was actually graded, shipped as evidence rather than described. |
| gpu_protocolstring | The accelerator class and run protocol the environment expects.Example H100 post-training protocol |
| contamination_guardJSON · nullable | Check that a post-training gain came from training rather than from recall of the benchmark. |
| base_weightspath · nullable | Real starting weights the agent has to improve on. |
Load it
The first commands after an approved delivery.
# Run one replication environment and score it bash harness/run.sh --env <env-id> --agent ./candidate --gpu a10 # Weighted leaf rubric -> replication score. # The automated judge is itself scored against the rubric (F1 0.946). bash harness/verify_trajectory.sh --env <env-id> # re-grade the shipped run
Quality checks
3 checksJudge agreement against the rubric
Whether the automated replication judge matches the real rubric.
The judge is scored against the upstream-authored leaf rubric and reported as an F1 rather than assumed correct.
Real measured: F1 0.946 against the 1,963-leaf rubric.
Trajectory verification
Whether the shipped example runs actually achieved what they claim.
Each captured trajectory is re-scored by the environment own grader; competition runs are medal-verified rather than self-reported.
Real measured: every shipped trajectory in the reference tranche is graded evidence.
Base-model improvement
Whether a post-training run genuinely beat its starting point.
The trained model is compared to the base model across a fixed benchmark suite, with a contamination guard applied to the result.
Methodology-stage per run. The comparison and the guard ship with the environment so a buyer can reproduce the check.
Ground truth
| What is correct | The real competition grader, the real upstream leaf rubric, and the measured benchmark suite for post-training runs. |
|---|---|
| How it is established | Competition runs are scored by the competition own grader and medal-verified. Replications roll weighted leaf criteria into a score, with the automated judge itself measured against the rubric. Post-training runs are scored by benchmark delta against the base model, filtered through a contamination guard. |
| Agreement | The replication judge reports F1 0.946 against the upstream rubric. Competition and post-training scoring are execution-based and deterministic. |
How it is built
- 01
Use the real grader, not a proxy
Competition environments ship the actual competition grader and replication environments ship the real upstream leaf rubric. A proxy metric would make the reward cheap to game and the result meaningless.
- 02
Weight the rubric
Leaf criteria roll up through a weighted tree so a multi-day replication produces a dense scoring surface instead of one subjective judgement at the end.
- 03
Measure the automated judge against the rubric
The judge that scores replication work is itself evaluated against the rubric rather than trusted, and its agreement is reported as a number.
- 04
Capture and grade a real trajectory
Every environment ships with an agent run that was actually executed and actually graded. Medal or score claims are verified against the grader, not self-reported by the run.
- 05
Guard post-training against recall
A post-training environment requires the agent to source data, train, and beat the base model across a fixed suite. A contamination guard separates a real gain from the model having seen the benchmark.
- 06
Run on a stated accelerator protocol
Each environment declares the accelerator class it expects, from consumer-to-midrange cards for competition runs up to datacenter parts for post-training, so results are reproducible rather than hardware-dependent folklore.
Request a sample
We send real records in the format you would receive, with the field list and where each record came from. Access for evaluation is limited and covered by a mutual NDA.
ProductML Research Tasks