A daily index over GPU rental pricing from roughly twenty clouds, built to the standard a cash-settled contract would demand: reproducible years later, defensible to a counterparty who lost money on it, and resistant to a participant who wants the number somewhere else.
A struck-through value is what the estimator produced and the gates refused to publish. Open the dashboard to see every provider behind a number.
A PCIe card on Ethernet in a spare-capacity marketplace and an SXM card on an NVLink fabric under a datacentre SLA differ in price by a factor of three, and they are not substitutes for the buyer. An index over a heterogeneous good is meaningless without a standard unit.
So every index defines exactly one good, and every input is restated as that good or discarded — the same pattern as a Platts standard cargo or a Baltic route. The standard good is the benchmark; everything else is a spread to it.
…
Benchmark administration has a standard answer to inputs of unequal quality: prefer transactions to quotes, and quotes to judgement. A value resting only on rate cards is published with a flag, because “what vendors say they would charge” is not the same claim as “capacity traded here.”
The estimator is built around that question rather than around statistical elegance. Each defence below answers one specific attack.
Collapse to a per-provider median first. A venue publishing forty variants of one box gets one vote, not forty.
Screen on median absolute deviation, which has a 50% breakdown point, rather than standard deviation, which does not.
Cap any single provider at 35% of total weight, applied iteratively, and refuse to publish below a provider floor.
Gates are evaluated every day. A thin day withholds rather than printing a number nobody could defend.
MAD is used precisely because you cannot defeat it with the outlier it exists to catch. Fed deliberately hostile inputs, it stopped working entirely.
The trap is that the failure requires agreement. As long as honest providers differ even slightly, MAD is positive and the screen works — which is why a test suite built on realistic-looking prices passed happily. It only breaks when the market agrees exactly, which is also the moment one extreme quote does the most damage.
The fix falls back to a symmetric ratio band. A percentage test would have been wrong twice over: it flags a provider 50% below consensus, which is ordinary market structure, and it caps at 100% on the downside, so it could never catch one cent against a three-dollar consensus.
four providers at $3.00, one at $1,000,000 median = $3.00 MAD = median(0, 0, 0, 0, 999997) = 0 sigma test undefined -> screen returns published value: $200,002
Now covered by property-based tests that generate markets rather than choose them, asserting ten invariants that must hold for any market at all.
A gap in the series is a fact about the market; an interpolated value is a fiction about it. Carrying yesterday's number forward on a thin day is exactly the behaviour that makes a benchmark manipulable — it converts “nobody traded” into “the price was unchanged,” and it rewards whoever stayed quiet.
An index that printed a number for a GPU only one venue quotes would be reporting one vendor's rate card dressed up as a market.
A methodology document that only lists strengths is a marketing document. These are the limits that matter most, and the full list runs to ten.
Every input is an offer or a rate card. Nobody here observes a completed rental at a price. That is the largest gap between this and a benchmark that could responsibly settle a contract, and no estimator sophistication closes it.
The economically dominant volume is private multi-year bilateral capacity that never appears on a rate card and prices differently. Anything settling against this carries that basis risk.
Where one venue sells the same box two ways, the ratio between the prices is a direct observation of what it charges for the difference. The community factor survives that test at +1.2% over 41 pairs. The spot factor turns out to be uncalibratable from public data, because the only venue quoting both tiers applies an identical ratio to every SKU — a discount policy, not a market spread.
This card used to end “the rest have no observable check at all.” That was false, and false only because nobody looked. The same test applies to form factor wherever a venue quotes one GPU as both PCIe and SXM: 21 pairs across 6 venues imply 1.085 against an asserted 1.18, so the schedule is roughly 8% too aggressive on the factor touching the largest share of inputs. Reported, not enforced — moving a factor would break the reproduction of every historical value. Interconnect and node size still have no check, and that sentence should now be read with suspicion.
Two runs an hour apart on an unmoved market produced dispersion of 0.560 and 0.365 — withheld, then published. The gate is correct in intent and too coarse at these provider counts. Recorded as an open defect rather than patched with a guessed constant.
Every provider behind every number, with weights, screens, and the gate results that published or withheld it.
The full methodology, including the section on everything it gets wrong.
What running it against live pricing actually turned up — including a forward curve that cannot be derived.