Skip to content

Back-Office Work Tasks

Accounting and back-office tasks with known correct answers, graded to the cent.

  • 64,000graded back-office tasks delivered
  • 96,000computer-use trajectories delivered
01

Sample

A recorded run and real records

Watch an agent work through a task

A recorded run from our research tasks, step by step: what the agent meant to do, the tool it called, what came back, and the grade at the end.

Research reproduction audit

Captured run, de-identified

Pass

Task

Audit a one-dimensional convection PINN reproduction packet. Recompute the evidence from samples and distinguish a valid reproduction from fabricated metrics, a boundary mismatch, an unsatisfied residual, missing curvature gain, or insufficient final accuracy.

01 / 06Task

Toolbash

Intent

Inspect the benchmark packet

Tool call

Inspect the provided repository and the public benchmark specification.

Observed result

The task defines a one-dimensional convection PINN reproduction with a reference PDE, initial condition, periodic boundary, and analytical solution. The working tree contains the model, PDE, reproduction script, requirements, and rubric materials.

repository inspected
benchmark specification located

Recorded state2 records

surfacestatus
repositoryinspected
reference PDElocated

Another run: training diagnosticAll recorded runs

Instance manifest for a corrupted-close environment

Real structure from a shipped taskset. Ledger contents, fault specifics, and oracle values are withheld - these are graded instances.

instance.json
{
  "instance_id": "<withheld>",
  "source": {
    "operator": "<opaque token>",
    "period": "<withheld>"
  },
  "fault": {
    "kind": "<one of the mixed-fault classes>",
    "affected_voucher": "<held out from agent>"
  },
  "oracle": {
    "expected_cents": "<withheld>",
    "tolerance": 0
  },
  "scores": ["reconciliation", "diagnosis"],
  "deidentification": {
    "identities": "stable tokens",
    "business": "opaque token",
    "amounts_dates_structure": "preserved",
    "leak_audit": "passed"
  },
  "difficulty_gate": {
    "frontier_models_tested": 2,
    "pass_rate": "0%"
  }
}

One step of a computer-use trajectory

Representative shape. Application content and coordinates from the real capture are withheld.

trajectory.jsonl
{
  "step": 14,
  "screenshot": "steps/014.png",
  "a11y_tree": "steps/014.a11y.json",
  "action": {
    "type": "click",
    "target_role": "button",
    "grounding": {
      "x": "<withheld>",
      "y": "<withheld>"
    }
  },
  "observation_delta": "<withheld>",
  "final_state_check": {
    "deterministic": true,
    "evaluated_at_end_of_episode": true
  }
}
02

Specifications

Back-Office Work Tasks specifications
FormatsDocker, JSONL, Parquet, Screenshot + a11y tree
DeliverySecure download, Docker image, VM snapshot
CadenceContinuous build; monthly tranches
Graded back-office tasks delivered64,000
Computer-use trajectories delivered96,000
Reference tranche20-instance mixed-fault accounting close
OracleExact-cent reconciliation plus affected-voucher diagnosis
Frontier difficulty0% on the all-or-nothing deliverables (both frontier models tested)
AccessOrg-scoped credentials with per-customer delivery paths. Every artifact ships de-identified and zero-leak verified; source businesses are never identified and never identifiable.
Built againstGDPval (OpenAI), HealthBench, FinanceBench, BigFinanceBench, LegalBench, CUAD, OSWorld

What ships

  • Controlled-Corruption Environments

    A real de-identified ledger is corrupted in a known way. The agent has to find the fault, fix it, and tie out the close. The oracle knows the exact cent and the exact affected voucher, so grading is deterministic.

  • Computer-Use Trajectories

    Real browser and desktop agent runs against real applications, recorded per step with screenshot, accessibility tree, action, and a final-state checker. Coordinate grounding is preserved.

  • Cross-Tool Reconstruction Tasks

    Tasks whose evidence is split across tools, so the model has to reconstruct what happened rather than answer from a single document.

Fields

9 fields

FieldDescription
instance_idstringIdentifier for one corrupted-ledger instance.
fault.kindenumThe class of corruption introduced into the real de-identified ledger.
fault.affected_voucherstringThe record the agent has to identify. Held out from the agent, known to the oracle.
oracle.expected_centsint · centsExact tie-out value. Grading is to the cent, not to a tolerance band.
ledger/directoryThe de-identified operating data the instance is built from. Identities are tokenized; amounts, dates, and structure are real.
steps[].screenshotimage · nullablePer-step capture for computer-use trajectories.
steps[].a11y_treeJSON · nullableAccessibility tree at the step, which is what makes an action reconstructable rather than merely visible.
steps[].actionJSON · nullableThe action taken, with coordinate grounding preserved.
final_state_checkJSONDeterministic assertion over the end state of the environment.

Load it

The first commands after an approved delivery.

quickstart.sh
# Run one corrupted-close instance and score it
bash harness/run.sh --instance <id> --agent ./candidate

# Two independent scores are returned:
#   reconciliation : exact cents, no tolerance band
#   diagnosis      : affected voucher identified, scored separately

bash harness/calibrate.sh --model <model-id>   # re-measure difficulty yourself
03

Quality checks

4 checks

Frontier difficulty calibration

Whether the all-or-nothing professional deliverables are actually hard.

Each eval surface is run against two frontier models in its own environment and graded by the task own machine oracle, in the easiest possible setting where every piece of evidence is supplied in context.

Real measured: 0% on the legal, finance, medical, and cross-tool computer-use deliverables for both frontier models tested. One wrong figure fails the task.

Exact-cent reconciliation

Whether the agent tied the close out correctly.

The submitted close is compared to the oracle value at cent precision, with no tolerance band.

Real measured: deterministic per instance across the reference taskset.

Affected-voucher diagnosis

Whether the agent found the right cause, not just the right number.

The identified record is compared against the known injected fault, scored separately from the arithmetic result.

Real measured: deterministic per instance. Scored apart from reconciliation so a lucky number does not read as a diagnosis.

Leak audit

Whether any identifying material survived de-identification.

Layered automated scans across every shipped artifact, followed by a human read of the delivered bytes rather than of the build own belief about them.

Real measured: zero-leak verified per tranche before release.

Ground truth

What is correctThe authored corruption. Because the fault is injected deliberately into real data, both the correct end value and the correct causal record are known exactly.
How it is establishedTwo independent scores: an exact-cent reconciliation of the close, and a diagnosis check against the affected record. Computer-use trajectories are closed by a deterministic final-state check on the environment rather than by a self-reported outcome.
AgreementDeterministic against the authored fault, so no inter-rater statistic applies. Difficulty is evidenced by frontier-model calibration rather than asserted, and the harness ships so it can be re-run.

How it is built

  1. 01

    Originate the operating data

    The substrate is real operating data captured from inside businesses actually in motion, not scraped from the web and not synthesized. This is the layer that cannot be obtained any other way.

  2. 02

    De-identify without hollowing out

    Identities become stable tokens and the business becomes an opaque token. Amounts, dates, sequence, and structure are preserved, because those are the signal. The rule is to remove who it is, never what happens.

  3. 03

    Verify zero leakage

    Every instance passes a leak check before it can enter a tranche. Automated scans run first, and a human read is the catch layer, because the failure mode that matters is the one an automated check already approved.

  4. 04

    Introduce a controlled corruption

    A known fault is injected into the de-identified ledger. Because the fault is authored, the oracle knows both the exact tie-out value and the exact affected record, which is what makes the reward deterministic.

  5. 05

    Gate on frontier difficulty

    Instances are measured against frontier models before shipping. Ones that are trivially solved do not earn a place in the taskset.

  6. 06

    Capture computer-use trajectories per step

    Browser and desktop runs are recorded as a sequence of screenshot, accessibility tree, action, and coordinate grounding, closed by a final-state checker rather than a self-reported success flag.

05RequestReplies usually within one working day

Request a sample

We send real records in the format you would receive, with the field list and where each record came from. Access for evaluation is limited and covered by a mutual NDA.

ProductBack-Office Work Tasks

team@gerra.com