Gerra
Data that doesn't
exist elsewhere.
We originate, license, and develop proprietary datasets for quant funds, AI labs, and robotics teams. Market signals, software and company records, and physical-AI capture - each from sources only we have.
Trusted by















What we built
Markets & Signals
Proprietary alternative-data signals for systematic funds - from social and sports to the compute markets behind AI.
Retail Sentiment Stream
Exclusive retail investor sentiment from the largest social finance platform.
Sports Query Stream
The only sports data source that powers real-time AI tool calls for the largest language models.
Startup Adoption Graph
Private company technographics mapped to public market tickers. See what startups adopt before the market prices it in.
Federal Software Ledger
Ticker-mapped US federal government spending on enterprise software. See which vendors win before the street does.
Event Probability Map
Cross-platform prediction market events mapped to affected securities with real-time probability streams.
GPU Price Index
Hourly GPU rental prices across every major cloud, normalized to canonical SKUs. A leading indicator of AI capex before it reaches earnings.
Inference Cost Curve
Real-time inference economics across hosted-model providers - token prices, throughput, and latency as a demand-side read on AI compute.
Software & Organizations
Real workflow, codebase, and full-company data for training and evaluating models and agents.
Workplace Activity Graph
Cross-tool operational data from collaboration systems. Entity-resolved into a unified schema for AI training and workflow evals.
Software Delivery Graph
Engineering coordination data across source control, issue tracking, and data platforms. Unified SDLC schema for coding agents and SWE evals.
Codebase Collection
Clone and inspect 178 real-world repositories as complete engineering systems, with 15,238 unique commits preserved for training, reasoning, and evaluation.
Company Operating Archive
Complete operating histories of real companies - code, business data, communications, documents, and databases - as a training corpus for frontier models.
Physical AI
Multi-modal robot data from our own fleet - teleoperation, human demonstration, and sensor streams for embodied AI.
Robot Demonstration Library
Success-labeled robot manipulation episodes collected via human teleoperation across diverse tasks and embodiments.
First-Person Motion Library
First-person human video paired with full-body 3D motion capture - the human-demonstration layer for embodied pretraining.
Robot Sensor Streams
High-frequency proprioception, inertial, and audio streams with sub-millisecond synchronization across embodiments.
How we work
Originate.
License.
Develop.
Originate
We operate our own robot fleet and build collection infrastructure from scratch. The data exists because we created it.
License
We hold exclusive, multi-year licenses to private data platforms. Nobody else can sell you this data.
Develop
We clean, structure, and deliver in the format you need. Backtest-ready for quant funds. ML-ready for AI labs.
From the research desk
Evidence,
published.
Papers, buyer guides, and field notes behind the catalog - including the results that did not survive validation.
Ready to see the data?
Browse our catalog, request trial access, or tell us what you need.