Requirement
Identical accelerators rent at very different prices. The same card, the same week, quoted across dozens of providers, can differ by a multiple rather than a few percent.
A desk we work with asked the follow-on question: if compute is the primary input cost for a whole sector, do input prices lead the equities of the companies buying them?
| # | Requirement | Acceptance condition |
|---|---|---|
| R1 | Comparable units | A price difference must reflect the market, not heterogeneous bundling |
| R2 | Point-in-time | Every observation carries what was knowable at that timestamp |
| R3 | Leakage-safe test | No model may see information dated after the return it predicts |
| R4 | Published limits | Whatever the result, the boundary is stated rather than implied |
R1 is where this class of study usually fails, and it fails silently.
Study panel construction
This is the panel built for the equity question. The productized feed is a separate and much denser dataset, covered at the end.
| Property | Value |
|---|---|
| Observations | 6,216 |
| Providers | 38 |
| Window | Four years |
| Matching | By accelerator, like for like |
| Normalization | Committed term vs on-demand, bundled storage and egress, card variant, quoted availability |
A listed hourly rate is not a price until the bundle is normalized. Committed-term and on-demand quotes are different instruments. Storage and egress sit inside the rate sometimes and outside it other times. The quoted card is not always the variant the label implies. Availability is occasionally quoted for capacity the provider does not hold.
A panel that skips this produces dispersion that is really heterogeneous units, after which the researcher discovers a signal that is an artifact of their own cleaning. That failure looks identical to a real finding until someone tries to trade it.
Primary finding
Dispersion survived normalization.
This is the durable result and it is genuinely unusual. In most commodity markets an identical good converges to a narrow band. This one does not, and the width is not explained by the bundle differences R1 removed.
Signal test
| Element | Method |
|---|---|
| Features | Cross-provider price level and dispersion, derived from the normalized panel |
| Target | Forward equity returns for the exposed names |
| Fitting | Walk-forward |
| Constraint | At each step the model sees only information available at that step |
Walk-forward is not a refinement here, it is the entire test. A model fit on the full history and evaluated on the full history will find a relationship in a panel this size whether or not one exists. The only version worth running is the one where the model is asked about a future it has not seen.
Result
| Question | Result |
|---|---|
| Standalone directional alpha, GPU prices to exposed equities | None found |
| Scope of the null | This panel size, this period, this feature set |
| Published | Yes, as a research note with the methodology attached |
| Effect on the product | Dataset is not sold as a trading signal |
Nothing usable came back on the directional question. The result was published rather than shelved, and it is why the catalog entry describes a compute-market panel instead of an alpha source.
What the panel is for
The null closed one use and clarified another.
| Use | Supported |
|---|---|
| Standalone directional equity signal | No |
| Procurement benchmarking and provider selection | Yes |
| Competitive analysis of the providers themselves | Yes |
| Capacity-tightness and cost-curve measurement | Yes |
| Component input inside a broader model | Yes, where it is not asked to carry a directional call alone |
The panel measures the compute market: who is quoting what, where capacity is tight, how fast the cost curve is moving, and how far a given provider sits from the floor. That is infrastructure intelligence, and it is what the licence covers.
Worth separating two things that are easy to conflate. The study above ran on a deliberately sparse four-year panel, built that way because reach mattered more than density for a question about multi-year equity returns. The productized GPU Price Index is a different dataset built for a different job: 262,146 price observations across 40 providers and 33 canonical SKUs, collected hourly, with the source URL, collection method, and raw payload pointer retained on every row. Companion tables carry marketplace utilization and inference token pricing across 381 models.
Standard applied
A data vendor's positive claims are worth roughly what its negative claims cost it. Publishing a null on your own dataset is inexpensive to do and rarely done, which is what makes it informative.
Every quantitative dataset in the catalog is held to the same three conditions: leakage-safe validation, stated baselines, and published limits sitting next to the published capability.