Gerra
Data that doesn't
exist elsewhere.
We originate, license, and develop proprietary datasets for quant funds, AI labs, and robotics teams. Market signals, software and company records, and physical-AI capture - each from sources only we have.
Trusted by















What we build
Markets & Signals
Proprietary alternative-data signals for systematic funds - from social and sports to the compute markets behind AI.
Retail Investor Sentiment
Exclusive retail investor sentiment from the largest social finance platform.
Sports Query & Answer Logs
The only sports data source that powers real-time AI tool calls for the largest language models.
Startup Technographics
Private company technographics mapped to public market tickers. See what startups adopt before the market prices it in.
Federal Software Contracts
Ticker-mapped US federal government spending on enterprise software. See which vendors win before the street does.
Prediction Market Probabilities
Cross-platform prediction market events mapped to affected securities with real-time probability streams.
GPU Rental Prices
Hourly GPU rental prices across every major cloud, normalized to canonical SKUs. A leading indicator of AI capex before it reaches earnings.
Inference Token Prices
Real-time inference economics across hosted-model providers - token prices, throughput, and latency as a demand-side read on AI compute.
Software & Organizations
Real workflow, codebase, and full-company data for training and evaluating models and agents.
Cross-Tool Workplace Records
Cross-tool operational data from collaboration systems. Entity-resolved into a unified schema for AI training and workflow evals.
Engineering Delivery Records
Engineering coordination data across source control, issue tracking, and data platforms. Unified SDLC schema for coding agents and SWE evals.
Full-History Codebases
Clone and inspect 178 real-world repositories as complete engineering systems, with 15,238 unique commits preserved for training, reasoning, and evaluation.
Whole-Company Operating Archives
Complete operating histories of real companies - code, business data, communications, documents, and databases - as a training corpus for frontier models.
Agent Tasks & RL Environments
Runnable environments with machine graders - repositories, exploit chains, whole applications, real business operations, and GPU research runs. Every task carries a verified reward signal.
Repo Build & Repair Tasks
Natural-language specifications and real issues turned into runnable repositories with held-out test suites, each verified to fail before the fix and pass after it.
Exploit & Patch Tasks
Containerized vulnerable services with working exploits, patches, and a machine-gradeable subtask ladder - multi-stage chains, not single-bug toys.
Full-App Build Tasks
Non-technical prose specifications for whole web applications, graded by a browser driving the running app rather than by reading the diff.
Back-Office Work Tasks
Real de-identified business operations turned into RL instances with exact-value oracles - the accounting close as an agent task, graded to the cent.
ML Research Tasks
GPU-backed environments where the task is to run machine-learning research: win the competition, replicate the paper, or beat the base model.
Physical AI
Multi-modal robot data from our own fleet - teleoperation, human demonstration, and sensor streams for embodied AI.
Robot Teleoperation Demonstrations
Success-labeled robot manipulation episodes collected via human teleoperation across diverse tasks and embodiments.
First-Person Human Motion
First-person human video paired with full-body 3D motion capture - the human-demonstration layer for embodied pretraining.
Robot Sensor Recordings
High-frequency proprioception, inertial, and audio streams with sub-millisecond synchronization across embodiments.
How we work
Originate.
License.
Develop.
Originate
We operate our own robot fleet and build collection infrastructure from scratch. The data exists because we created it.
License
We hold exclusive, multi-year licenses to private data platforms. Nobody else can sell you this data.
Develop
We clean, structure, and deliver in the format you need. Backtest-ready for quant funds. ML-ready for AI labs.
From the research desk
Evidence,
published.
Papers, buyer guides, and field notes behind the catalog - including the results that did not survive validation.
Ready to see the data?
Ask for a sample of anything in the catalog and we will send real records in the delivery format, with the schema and provenance that come with them.
Request a sample
Tell us what you are building and which lane matters. We reply personally, usually within a working day.