Skip to content

Cheaper tokens, richer labs

4 min read
#ai#economics#inference

OpenAI cut the price of Luna, its cheapest model, by 80% on July 30: input tokens went from $1 to $0.20 per million, output from $6 to $1.20. Terra, the mid-tier model, got a smaller 20% cut. Sol, the flagship, didn’t move. OpenAI’s own explanation was efficiency, not desperation. The cut came from “improvements in serving efficiency” across its training and inference stack, not from matching a competitor at a loss. That phrase is worth sitting with, because it tells you the cut wasn’t the whole gain. If serving a token got cheaper and the price dropped by less than the cost did, the difference didn’t vanish. It became margin.

That’s the pattern underneath the headline, and it shows up more clearly at Anthropic. Over the same stretch, Anthropic’s gross margin on inference infrastructure went from 38% to over 70%, per SemiAnalysis’s cost modeling, while the company was also cutting list prices. Opus dropped from $15/$75 per million tokens to $5/$25 across two generations, a two-thirds cut. Falling sticker price and rising margin, at the same time, is not the story a simple price war tells. A price war compresses margin. This compressed cost faster than it compressed price.

Some of that gap comes from how the price is actually realized rather than how it’s listed. SemiAnalysis puts the effective blended price on Opus running agentic work at around $0.99 per million tokens, well under the $5/$25 sticker. Cache hit rates on repeated context now clear 90% and get billed at a steep discount. The sticker price is what you’d pay cold; almost nobody pays it cold. The number that matters — cost to the lab per unit of real task completed — has been falling faster than the number a vendor puts on a pricing page. The lab keeps the difference.

Where did the efficiency actually come from, and who’s declining to charge for it? Hardware, mostly. Nvidia’s GB300 systems run roughly 17 to 32 times the throughput of the prior generation on the same workload, depending on precision. Nvidia hasn’t repriced accordingly — list rates for the new chips have moved down, not up, even as the value delivered per chip climbed by an order of magnitude. Nvidia left obvious pricing power unclaimed, plausibly to stay out of antitrust crosshairs. That value sat on the table for whoever was willing to pick it up. The labs picked it up. Compute got radically cheaper to provide. The price of a frontier token fell by a smaller fraction than the cost did, and the spread went to gross margin, not back to the chip layer and not fully back to the customer either.

This is the opposite of what happened to cloud bandwidth. When CDN pricing collapsed through the 2010s, it collapsed because bandwidth was fungible — dozens of vendors sold the same bit-for-bit product, so efficiency gains had nowhere to go but straight through to price. There was no differentiation to protect and no supply constraint to lean on, so the market did what commodity markets do. Frontier inference doesn’t behave that way, at least not yet, because two conditions bandwidth never had are both true here: open models are still a real notch behind the frontier for hard work, and demand for frontier tokens outstrips what labs can serve. Pricing power survives exactly as long as both hold.

Both conditions are shakier at the bottom of the model stack than at the top, which is why this looks like two markets instead of one. Chinese open-weight providers already account for something like 46% of token volume by some estimates, and that’s precisely where undifferentiated capacity behaves like bandwidth: falling price, thinning margin, the CDN pattern intact. Luna’s 80% cut reads less like an efficiency dividend and more like a defensive move to stay ahead of that commodity layer, while Sol sits untouched because nothing at that tier is commoditized yet.

So the per-token chart is measuring the wrong layer. Watch gross margin, not sticker price, and expect the two to keep diverging at the top of the model stack for a while yet — the constraint holding it up is supply and a quality gap, not goodwill, and neither is closing fast. The builders sitting on top of these models are the ones who end up looking like the commodity layer. Recent survey data already puts average AI product gross margins around 52%, well under the 80-90% that defined mature SaaS. The margin didn’t disappear from the stack. It moved up it.