Gerra / Research indexOpen record
Research, methods, and source material

Evidence has
a paper trail.

Studies, technical notes, field research, and the underlying datasets in one public record. We publish what the data supports, what it does not, and enough of the trail to inspect the difference.

Research works
09
Source datasets
14
Field notes
14
02 / Papers

Research register

Papers, specifications, and technical notes. Ordered by release, with the measured substrate shown beside each claim.

  1. R-08The GPU Rental Market: Price Dispersion and the Cost Curve of ComputeA four-year, cross-provider panel measuring the unusually wide price dispersion of identical accelerators and testing whether GPU prices contain an equity signal.6,216 observations / 38 providersResearch paper/PreprintMay 2026
  2. R-07Retail Financial Sentiment as a Systematic SignalEighteen years of human-labeled bullish and bearish messages tested as a point-in-time systematic signal, including a direct comparison with NLP sentiment from X.426M+ messages / 18 yearsResearch paper/PublishedMay 2026
  3. R-06Mapping Prediction Markets to SecuritiesA cross-platform framework that maps live event probabilities to exposed securities with signed direction, sensitivity, and point-in-time probability history.300+ events / 500+ securitiesResearch paper/PublishedMay 2026
  4. R-05The Falling Price of InferenceA preliminary cross-provider study of hosted-model token prices, separating the clear directional decline in cost from the limits of a mixed-model sample.488 provider-model observationsResearch paper/PreprintMay 2026
  5. R-04From Trained Model to SiliconA reproducible pipeline from a trained logistic decoder to synthesizable Verilog RTL and an auto-generated testbench, measured before and after fixed-point quantization.0.9781 accuracy / simulation passedResearch paper/PublishedMay 2026
  6. R-03Operational Telemetry: A Cross-Tool Activity GraphA data-availability report on consented company activity normalized across communication, documents, CRM, tickets, repositories, people, teams, and tasks.5 companies / 38 toolsTechnical note/Working draftMay 2026
  7. R-02Empirical Characterization of the Sim2Real GapA joint-level comparison of simulated and real trajectories on a full-size bipedal humanoid, isolating where model fidelity breaks down across the body.23 joints / measured transfer gapResearch paper/PublishedDec 2024
  8. R-01Open Robot Training Format v1.0A common representation for robot demonstrations across multimodal observations, action spaces, sensor calibration, coordinate frames, and embodiments.Interoperable training-data schemaSpecification/Working draftDec 2024
03 / Field notes

The archive stays live.

Longer investigations sit beside the formal papers: component histories, market structure, deployment economics, and observations that are useful before they become a specification.

Browse every field note
  1. 01What Code Data Do Codex, Claude Code, and SWE Agents Need?A model-agnostic framework for choosing repository data for terminal coding agents: context, history, environments, tests, and honest evaluation boundaries.Agent data guideJul 19, 2026
  2. 02Reading Federal Software Spending Before the Street DoesUS federal software procurement is public record. A method note on turning USAspending and FPDS-NG awards into a ticker-level revenue signal.Method noteJul 19, 2026
  3. 03How Git History Becomes a Code-Model Evaluation DatasetA practical method for turning authentic commits, patches, and tests into point-in-time coding-agent evaluations without leaking the answer.Evaluation methodJul 19, 2026
  4. 04How to Audit a Codebase Dataset Before You License ItA buyer checklist for repository completeness, Git integrity, provenance, deduplication, privacy, tests, and reproducible release metrics.Buyer guideJul 19, 2026
  5. 05How to Evaluate Retail Sentiment Data Before You Backtest ItA buyer checklist for point-in-time integrity, deleted-message survivorship, human vs NLP labels, spam filtering, and leakage-safe backtests.Buyer guideJul 19, 2026
  6. 06How to Read a Data Vendor's Data RoomWhat a serious data room contains and how to inspect it: dated schemas, delivery-shape samples, checksums, reconcilable manifests, stated limits.Buyer guideJul 19, 2026
  7. 07Agents Fail at Office Work Because the Data Never ExistedModels trained on the public web have never seen how organizations run. What consented, cross-tool company data adds to agent training and evals.Research explainerJul 19, 2026
  8. 08What Prediction Markets Know About Your PortfolioHow live event probabilities map to exposed securities with signed direction and sensitivity, and why point-in-time probability history matters.Research explainerJul 19, 2026
  9. 09Repository-Level Code Data for Coding AgentsWhy complete repositories, cross-file context, tests, and change history are more useful for coding-agent training than disconnected source files.Research explainerJul 19, 2026
  10. 10Why Identical GPUs Rent at Wildly Different PricesThe same H100 rents for $1.20 to $22.02 per hour. What a 38-provider panel shows about GPU price dispersion, and why the equity signal is a null.Research explainerJul 19, 2026
  11. 11The Sim2Real Gap Costs Robotics Billions AnnuallySimulation-to-reality transfer failures create multi-billion dollar losses across the robotics industry, with success rates plummeting from 90% in simulation to under 20% on real hardware, extending development cycles by 6-12 months.Field analysisJun 19, 2025
  12. 12The Complete Robot Activation PlaybookHow top agencies 10X event ROI without doubling budgets - the comprehensive guide to deploying humanoid robots for experiential marketing campaigns.Operating noteJun 8, 2025
  13. 13The QDD Actuator RenaissanceHow a 2016 MIT breakthrough spawned the era of $3k humanoids – and why China now dominates the torque supply chainTechnical historyJun 1, 2025
04 / Source data

The dataset is part of the argument.

Research is linked to the data products it can actually support. Each catalog record carries its collection scope, history, delivery form, and technical detail forward.

A

Markets & signals

Point-in-time data for compute, capital, attention, and event-driven research.

  1. 01Retail Sentiment StreamExclusive retail investor sentiment from the largest social finance platform.426M MESSAGES · SINCE 2009
  2. 02Sports Query StreamThe only sports data source that powers real-time AI tool calls for the largest language models.1BN+ QUERIES · 12 YRS
  3. 03Startup Adoption GraphPrivate company technographics mapped to public market tickers. See what startups adopt before the market prices it in.10M+ COMPANIES
  4. 04Federal Software LedgerTicker-mapped US federal government spending on enterprise software. See which vendors win before the street does.120+ TICKERS MAPPED
  5. 05Event Probability MapCross-platform prediction market events mapped to affected securities with real-time probability streams.500+ SECURITIES
  6. 06GPU Price IndexHourly GPU rental prices across every major cloud, normalized to canonical SKUs. A leading indicator of AI capex before it reaches earnings.22 SKUS · 25+ CLOUDS
  7. 07Inference Cost CurveReal-time inference economics across hosted-model providers - token prices, throughput, and latency as a demand-side read on AI compute.LIVE TOKEN ECONOMICS
B

Software & organizations

Repository and operating histories for agents that need to reason across real work.

  1. 01Workplace Activity GraphCross-tool operational data from collaboration systems. Entity-resolved into a unified schema for AI training and workflow evals.40+ EVENT TYPES
  2. 02Software Delivery GraphEngineering coordination data across source control, issue tracking, and data platforms. Unified SDLC schema for coding agents and SWE evals.50+ SDLC EVENTS
  3. 03Codebase CollectionClone and inspect 178 real-world repositories as complete engineering systems, with 15,238 unique commits preserved for training, reasoning, and evaluation.15,238 COMMITS
  4. 04Company Operating ArchiveComplete operating histories of real companies - code, business data, communications, documents, and databases - as a training corpus for frontier models.WHOLE-COMPANY DEPTH
C

Embodied systems

Demonstration and sensor data grounded in bodies, environments, and time.

  1. 01Robot Demonstration LibrarySuccess-labeled robot manipulation episodes collected via human teleoperation across diverse tasks and embodiments.400K+ EPISODES
  2. 02First-Person Motion LibraryFirst-person human video paired with full-body 3D motion capture - the human-demonstration layer for embodied pretraining.POV + 3D MOCAP
  3. 03Robot Sensor StreamsHigh-frequency proprioception, inertial, and audio streams with sub-millisecond synchronization across embodiments.<1MS SYNC
05 / Evidence standard

Claims should survive contact with the trail.

Our job is not to make every dataset look predictive. It is to establish what was observed, test the strongest interpretation, and leave the boundary visible when the result is weaker than the thesis.

  1. 01

    Start at the source

    Origin, collection context, licensing, and temporal coverage stay attached to the data.

  2. 02

    Preserve the trail

    Manifests, timestamps, mappings, transformations, and checks make the result inspectable.

  3. 03

    Try to break the claim

    Point-in-time tests, holdouts, permutation checks, and leakage review come before the headline.

  4. 04

    Publish the boundary

    Null results, sample limits, exclusions, and unresolved uncertainty belong in the conclusion.

Open collaboration

Have a hard data problem, a result to reproduce, or a corpus that should exist?

research@gerra.com