observed_at
requiredWhen the price was collected from the provider pricing surface.
- Type
- timestamp
- Unit
- UTC
2026-04-02T00:00:00ZINFERENCE DEMAND DATA
Real-time inference economics across hosted-model providers - token prices, throughput, and latency as a demand-side read on AI compute.
Representative records in the delivery format, ready to inspect before licensing the full dataset.
Current collected record (output price)
Representative. The study measures output price per million tokens on cheapest-available models; tiers are not pooled.
{
"observed_at": "2026-04-02T00:00:00Z",
"provider": "provider-a",
"model_family": "llama",
"model_listing": "llama-3.1-8b-instruct",
"output_price_per_mtok_usd": 0.02,
"input_price_per_mtok_usd": 0.01,
"model_tier": "small_open"
}Marketplace utilization snapshot
Real row shape from the reference sample. 1,366 snapshots, every one carrying a computed ratio.
provider,sku_canonical,instance_total,instance_available,instance_rented,utilization_ratio,snapshot_ts marketplace-a,RTX_3090_24GB,2368,1120,1248,0.527,2025-12-03T09:36:00Z
Every field, its type, whether it can be null, and a representative value.
When the price was collected from the provider pricing surface.
2026-04-02T00:00:00ZOne of four hosted-inference providers.
provider-aOpen-weight lineage the listing is grouped under (Llama, Qwen, Mixtral, DeepSeek).
llamaProvider raw model product name before family grouping.
llama-3.1-8b-instructNormalized output-token price - the quantity the study actually measures.
0.02Input-token price where listed.
0.01Coarse tier flag; the study does not pool tiers.
small_openContext window listed for the model at that snapshot.
131072Marketplace inventory: instances listed for the provider and SKU.
2368Marketplace inventory: instances currently rented.
1248Rented over total for the provider, SKU, and region at that snapshot. The supply-tightness read.
0.527| Field | Type | Constraint | Description |
|---|---|---|---|
| observed_at | timestamp · UTC | required | When the price was collected from the provider pricing surface. e.g. 2026-04-02T00:00:00Z |
| provider | string | required | One of four hosted-inference providers. e.g. provider-a |
| model_family | string | required | Open-weight lineage the listing is grouped under (Llama, Qwen, Mixtral, DeepSeek). e.g. llama |
| model_listing | string | nullable | Provider raw model product name before family grouping. e.g. llama-3.1-8b-instruct |
| output_price_per_mtok_usd | float · USD / 1M output tokens | required | Normalized output-token price - the quantity the study actually measures. e.g. 0.02 |
| input_price_per_mtok_usd | float · USD / 1M input tokens | nullable | Input-token price where listed. e.g. 0.01 |
| model_tier | string | nullable | Coarse tier flag; the study does not pool tiers. e.g. small_open |
| context_window | int · tokens | nullable | Context window listed for the model at that snapshot. e.g. 131072 |
| instance_total | int | nullable | Marketplace inventory: instances listed for the provider and SKU. e.g. 2368 |
| instance_rented | int | nullable | Marketplace inventory: instances currently rented. e.g. 1248 |
| utilization_ratio | float · 0..1 | nullable | Rented over total for the provider, SKU, and region at that snapshot. The supply-tightness read. e.g. 0.527 |
Input and output token rates per model per provider over time - the unit economics of inference, tracked continuously.
Time-to-first-token, generation throughput, and latency percentiles - congestion signals that move before capacity announcements.
Request success and error rates across providers - a real-time read on where inference demand is outrunning supply.
Collect from the public pricing surfaces of four hosted-inference providers.
Normalize to a per-million-token basis for output tokens.
Group heterogeneous model listings by the underlying open-weight family so comparisons stay within recognizable lineages.
Observed models span small open-weight checkpoints to frontier-scale hosted deployments priced very differently; tiers are not pooled into a single level.
Report the direction - a steep decline at the cheap end - as robust, and treat absolute price levels as preliminary because the cheapest quote in any period may reflect a different tier than in another.
What each evaluation measures and how it is run. Where no benchmark is published, we show the methodology and say so.
Measures
How the price of the cheapest available output tokens moved over time.
Method
Track the cheapest available output price per million tokens across the four providers over the window.
Result
Descriptive, published as preliminary: fell from about $0.13 per million tokens in mid-2024 into the $0.01 to $0.03 range in 2025-2026. Direction robust; absolute level preliminary due to mixed tiers.
Measures
Whether rented-versus-available instance counts identify capacity tightening on a marketplace.
Method
Poll marketplace inventory per provider, SKU, and region, recording instances total, available, and rented, and derive a utilization ratio per snapshot.
Result
Real measured: 1,366 utilization snapshots across the reference window, every row carrying a computed ratio. Latency and success-rate telemetry is a separate collection surface and is quoted only where it is actually collected.
What correct means for this data, and how it is established.
Ground truth
For the descriptive layer, the observed posted output-token prices across the four providers on a cheapest-available basis. For the proposed demand signal, there is no ground truth yet because the data is uncollected.
How it is established
Descriptive aggregation of the cheapest-available output price over time per provider and family. No predictive grader exists; the proposed demand-side validation is future work gated on API access.
Throughput collapse and rising latency across providers signal surging model demand days to weeks before it shows up in chip orders.
Track the falling price of intelligence per token and model the margin structure of hosted-inference and model-API businesses.
Compare price, speed, and reliability across providers for the same model - who is winning the inference market in real time.
Delivery
S3, REST API, Parquet
Formats
Parquet, JSON, CSV
Auth
Licensed for internal research and model development. Sourced from public pricing surfaces; no PII or MNPI.
Cadence
Continuous polling with daily aggregates.
Request a sample
Real records in the delivery format, with the schema and provenance that come with them. Evaluation access is restricted-scope and moves under a mutual NDA, so tell us what you are building and we will send the slice that fits.