Back-Office Work Tasks
Accounting and back-office tasks with known correct answers, graded to the cent.
- 64,000graded back-office tasks delivered
- 96,000computer-use trajectories delivered
Sample
A recorded run and real recordsWatch an agent work through a task
A recorded run from our research tasks, step by step: what the agent meant to do, the tool it called, what came back, and the grade at the end.
Research reproduction audit
Captured run, de-identified
Pass
Audit a one-dimensional convection PINN reproduction packet. Recompute the evidence from samples and distinguish a valid reproduction from fabricated metrics, a boundary mismatch, an unsatisfied residual, missing curvature gain, or insufficient final accuracy.
01 / 06Task
Toolbash
- Intent
Inspect the benchmark packet
- Tool call
Inspect the provided repository and the public benchmark specification.
- Observed result
The task defines a one-dimensional convection PINN reproduction with a reference PDE, initial condition, periodic boundary, and analytical solution. The working tree contains the model, PDE, reproduction script, requirements, and rubric materials.
repository inspected benchmark specification located
Recorded state2 records
surface status repository inspected reference PDE located
- Intent
Check the packet format
- Tool call
Inspect the available reproduction packet and rubric files.
- Observed result
The packet format and rubric are present. The auditor can be written against the normative fields without relying on a precomputed result file.
packet format checked rubric materials located
Recorded state2 records
surface status packet fields checked rubric located
- Intent
Write the auditor
- Tool call
Write a normalized audit artifact that recomputes the packet metrics and returns a classification.
- Observed result
The audit artifact is written. Its public role is to recompute evidence from raw packet samples, rather than trust a reported score.
normalized audit artifact written
Recorded state2 records
artifact status audit.py written output contract recomputed metrics and classification
- Intent
Exercise failure paths
- Tool call
Run local checks covering valid, fabricated, boundary, residual, curvature-gain, and final-accuracy cases.
- Observed result
The local checks return the expected classifications for the valid reproduction and each named failure mode, including insufficient final accuracy.
failure-path checks completed outputs summarized
Recorded state2 records
check family status valid reproduction classified failure paths classified
- Intent
Submit the audit run
- Tool call
Submit the completed audit artifact for machine grading.
- Observed result
The captured run reports the auditor as complete and ready for the machine grade.
audit artifact submitted awaiting terminal grade
Recorded state2 records
surface status submission complete grader pending
- Intent
Terminal grade
- Tool call
Evaluate the submitted auditor against six held-out reproduction-audit checks.
- Observed result
The machine grader reports a passing result for all six held-out checks.
verdict: PASS held-out checks: 6 / 6 score: 1.0
Recorded state2 records
grade field value held-out checks 6 / 6 verdict Pass
Instance manifest for a corrupted-close environment
Real structure from a shipped taskset. Ledger contents, fault specifics, and oracle values are withheld - these are graded instances.
{
"instance_id": "<withheld>",
"source": {
"operator": "<opaque token>",
"period": "<withheld>"
},
"fault": {
"kind": "<one of the mixed-fault classes>",
"affected_voucher": "<held out from agent>"
},
"oracle": {
"expected_cents": "<withheld>",
"tolerance": 0
},
"scores": ["reconciliation", "diagnosis"],
"deidentification": {
"identities": "stable tokens",
"business": "opaque token",
"amounts_dates_structure": "preserved",
"leak_audit": "passed"
},
"difficulty_gate": {
"frontier_models_tested": 2,
"pass_rate": "0%"
}
}One step of a computer-use trajectory
Representative shape. Application content and coordinates from the real capture are withheld.
{
"step": 14,
"screenshot": "steps/014.png",
"a11y_tree": "steps/014.a11y.json",
"action": {
"type": "click",
"target_role": "button",
"grounding": {
"x": "<withheld>",
"y": "<withheld>"
}
},
"observation_delta": "<withheld>",
"final_state_check": {
"deterministic": true,
"evaluated_at_end_of_episode": true
}
}Specifications
| Formats | Docker, JSONL, Parquet, Screenshot + a11y tree |
|---|---|
| Delivery | Secure download, Docker image, VM snapshot |
| Cadence | Continuous build; monthly tranches |
| Graded back-office tasks delivered | 64,000 |
| Computer-use trajectories delivered | 96,000 |
| Reference tranche | 20-instance mixed-fault accounting close |
| Oracle | Exact-cent reconciliation plus affected-voucher diagnosis |
| Frontier difficulty | 0% on the all-or-nothing deliverables (both frontier models tested) |
| Access | Org-scoped credentials with per-customer delivery paths. Every artifact ships de-identified and zero-leak verified; source businesses are never identified and never identifiable. |
| Built against | GDPval (OpenAI), HealthBench, FinanceBench, BigFinanceBench, LegalBench, CUAD, OSWorld |
What ships
Controlled-Corruption Environments
A real de-identified ledger is corrupted in a known way. The agent has to find the fault, fix it, and tie out the close. The oracle knows the exact cent and the exact affected voucher, so grading is deterministic.
Computer-Use Trajectories
Real browser and desktop agent runs against real applications, recorded per step with screenshot, accessibility tree, action, and a final-state checker. Coordinate grounding is preserved.
Cross-Tool Reconstruction Tasks
Tasks whose evidence is split across tools, so the model has to reconstruct what happened rather than answer from a single document.
Fields
9 fields
| Field | Description |
|---|---|
| instance_idstring | Identifier for one corrupted-ledger instance. |
| fault.kindenum | The class of corruption introduced into the real de-identified ledger. |
| fault.affected_voucherstring | The record the agent has to identify. Held out from the agent, known to the oracle. |
| oracle.expected_centsint · cents | Exact tie-out value. Grading is to the cent, not to a tolerance band. |
| ledger/directory | The de-identified operating data the instance is built from. Identities are tokenized; amounts, dates, and structure are real. |
| steps[].screenshotimage · nullable | Per-step capture for computer-use trajectories. |
| steps[].a11y_treeJSON · nullable | Accessibility tree at the step, which is what makes an action reconstructable rather than merely visible. |
| steps[].actionJSON · nullable | The action taken, with coordinate grounding preserved. |
| final_state_checkJSON | Deterministic assertion over the end state of the environment. |
Load it
The first commands after an approved delivery.
# Run one corrupted-close instance and score it bash harness/run.sh --instance <id> --agent ./candidate # Two independent scores are returned: # reconciliation : exact cents, no tolerance band # diagnosis : affected voucher identified, scored separately bash harness/calibrate.sh --model <model-id> # re-measure difficulty yourself
Quality checks
4 checksFrontier difficulty calibration
Whether the all-or-nothing professional deliverables are actually hard.
Each eval surface is run against two frontier models in its own environment and graded by the task own machine oracle, in the easiest possible setting where every piece of evidence is supplied in context.
Real measured: 0% on the legal, finance, medical, and cross-tool computer-use deliverables for both frontier models tested. One wrong figure fails the task.
Exact-cent reconciliation
Whether the agent tied the close out correctly.
The submitted close is compared to the oracle value at cent precision, with no tolerance band.
Real measured: deterministic per instance across the reference taskset.
Affected-voucher diagnosis
Whether the agent found the right cause, not just the right number.
The identified record is compared against the known injected fault, scored separately from the arithmetic result.
Real measured: deterministic per instance. Scored apart from reconciliation so a lucky number does not read as a diagnosis.
Leak audit
Whether any identifying material survived de-identification.
Layered automated scans across every shipped artifact, followed by a human read of the delivered bytes rather than of the build own belief about them.
Real measured: zero-leak verified per tranche before release.
Ground truth
| What is correct | The authored corruption. Because the fault is injected deliberately into real data, both the correct end value and the correct causal record are known exactly. |
|---|---|
| How it is established | Two independent scores: an exact-cent reconciliation of the close, and a diagnosis check against the affected record. Computer-use trajectories are closed by a deterministic final-state check on the environment rather than by a self-reported outcome. |
| Agreement | Deterministic against the authored fault, so no inter-rater statistic applies. Difficulty is evidenced by frontier-model calibration rather than asserted, and the harness ships so it can be re-run. |
How it is built
- 01
Originate the operating data
The substrate is real operating data captured from inside businesses actually in motion, not scraped from the web and not synthesized. This is the layer that cannot be obtained any other way.
- 02
De-identify without hollowing out
Identities become stable tokens and the business becomes an opaque token. Amounts, dates, sequence, and structure are preserved, because those are the signal. The rule is to remove who it is, never what happens.
- 03
Verify zero leakage
Every instance passes a leak check before it can enter a tranche. Automated scans run first, and a human read is the catch layer, because the failure mode that matters is the one an automated check already approved.
- 04
Introduce a controlled corruption
A known fault is injected into the de-identified ledger. Because the fault is authored, the oracle knows both the exact tie-out value and the exact affected record, which is what makes the reward deterministic.
- 05
Gate on frontier difficulty
Instances are measured against frontier models before shipping. Ones that are trivially solved do not earn a place in the taskset.
- 06
Capture computer-use trajectories per step
Browser and desktop runs are recorded as a sequence of screenshot, accessibility tree, action, and coordinate grounding, closed by a final-state checker rather than a self-reported success flag.
Request a sample
We send real records in the format you would receive, with the field list and where each record came from. Access for evaluation is limited and covered by a mutual NDA.
ProductBack-Office Work Tasks