Skip to content

What is enterprise document data?

Enterprise document data is the internal work product of companies, such as engineering, finance, sales and consulting documents, kept as the original files with identities removed so it can be used to train and evaluate AI models. Very little of it is ever published, so a model trained on the web has read far more descriptions of office work than examples of it.

What ships with each document

Gerra has processed 800,000 documents and processes hundreds of thousands more every week; its documents, codebases and business records come from real work at 4,000 companies. Every document ships as the original file, its rendered pages, its text and structure, and a SHA-256 hash of every file. Each layer does a different job:

  • The original file keeps the tables, figures and editable structure that text extraction flattens.
  • The rendered pages are what a person reads, and the layer where logos, stamps and faces survive a text scrub.
  • The hashes bind every later claim to exact bytes. Without them, a sample is whatever was downloaded that day.

For evaluation, documents can come with grounded questions, source bundles, finished human reference outputs and rubrics. They can also come with the records around them. Gerra's Operational Telemetry report documents a cross-tool activity graph from 5 consented companies and 38 tools, including 38.4 million chat messages, 11.2 million emails and 3.6 million files, with each person, team, org, document and task resolved to a single node within its company.

Remove who it is, keep what happened

Gerra's rule is to remove who it is, never what happens. Identities become stable tokens, so the same person is the same token in every file. Amounts, dates, sequence, structure and outcomes stay exactly as they were, because they are the signal. This is the de-identification block from a shipped back-office instance:

"deidentification": { "identities": "stable tokens", "business": "opaque token",
                      "amounts_dates_structure": "preserved", "leak_audit": "passed" }

Scrubbing harder is the common mistake. A corpus scrubbed until nothing identifying could possibly remain has no operational substance left, and blurred figures remove the value along with the risk.

Where automated scrubbing fails

A corpus can pass every automated check and still fail an independent human read. When it does, the checker was usually written with the scrubber's assumptions, so it confirms what the pipeline already believes. The leaks that get through sit outside the text the scrubber was pointed at:

Leak Where it hides What catches it
A real name in a link The URL under renamed link text Extracting every link and reading the target
Faces, signatures, account bars Images and screenshots Looking at every image
Logos, letterhead, watermarked domains The page layer, outside the text Rendering each page and reading it
Per-recipient stamps Repeated on every page Rendering each page and reading it
Health or financial identifiers Inside routine-looking workflows A human read
Counterparties and signatories Anywhere a third party appears A human read

The fix is to verify the delivered file after packaging, the way a recipient will open it, and to have a real sample read by someone who didn't build the pipeline. Every corpus Gerra ships gets that independent read before delivery.

Human-written is only the entry bar

A template filled in by a person is still a template, and a corpus of them teaches format instead of work. Gerra scores quality within each genre against a rubric, with template subtraction as the controlling test, so length can't pass for substance. Suspected model-generated documents are excluded outright rather than scored down, and the exclusions are listed so a buyer can see what was taken out.

Evaluations a model can't have memorized

Documents that were never published can't have been crawled into a model's pretraining data, which makes them hard to contaminate as evaluation material. Gerra's back-office environments are built this way, from de-identified operating data captured inside real businesses rather than scraped from the web.