AI Demand Could Push Compute Prices Up Tenfold
Dwarkesh Patel argues that if frontier AI revenue grows far faster than compute capacity, the gap will have to emerge in higher margins, higher compute prices or a shift of hardware toward inference. He expects physical supply constraints—from fabrication capacity to wafer allocation—to limit compute growth even as more capable models raise the value of each unit. That dynamic, he says, would favor labs able to extract more useful work from scarce capacity and could deepen concentration in AI infrastructure.

The gap between AI revenue and compute supply has to show up somewhere
Dwarkesh Patel starts from a conditional mismatch: if frontier-lab revenue continues growing around 10x annually while available compute grows only around 3x, the economic surplus must appear through higher lab margins, higher compute prices, a larger share of hardware devoted to inference rather than training, or some combination.
Epoch AI’s displayed capacity chart estimates cumulative compute-capacity growth at 3.4x annually, with a 90% confidence interval of 2.8x to 4.1x.
Patel treats the revenue scenario as deliberately extreme rather than inevitable. He says Anthropic ended the prior year at roughly $9 billion in revenue and could finish the current year at $100 billion to $150 billion; another year of comparable growth would imply $1 trillion. Whether that happens depends, he says, on whether AI becomes useful enough to earn it.
He argues that all three adjustments are already underway. He cites reported Anthropic inference margins rising from about 40% in the middle of the prior year to above 80%. Semianalysis’s displayed H100 rental-price chart shows prices rising after an early-2026 low; Patel says spot compute prices were more than 40% above their February trough.
Inference is taking a larger share of the hardware base, too. Epoch AI’s displayed accounting of OpenAI’s 2024 cloud-compute expense shows roughly $2 billion for inference and about $5 billion for R&D compute, much of it experiments and other work rather than final training runs.
Patel says inference represented only about a quarter of OpenAI’s compute in 2024 and may now be closer to half or higher. But labs want inference revenue to fund the training and experiments for their next models, not to consume the bulk of their capacity. That leaves higher margins or higher compute prices as the more consequential ways to close the gap while preserving training investment.
Expensive capacity makes efficient models more valuable
A lab that gets more useful work from each unit of compute can both outbid rivals for scarce capacity and charge more for its own model. Dwarkesh Patel frames the mechanism through the Alchian–Allen effect: once an H100 costs $20 an hour, using a weaker model that burns more tokens to reach the same result becomes much less attractive.
The consequence is more than a generic preference for better models. Higher compute prices amplify the value of efficiency at inference. A model that delivers an equivalent result with less compute has, in Patel’s formulation, “created more compute” from a fixed physical supply. Its provider can command a larger premium because it economizes on the input everyone is competing for.
This gives leading labs an advantage on both sides of the market. As they become better at monetizing compute, Patel says, they can bid against competitors for capacity using a resource more productively than those competitors can. And customers facing expensive inference have a stronger reason to pay for the model that uses fewer tokens.
Patel expects that dynamic to price out some applications that are popular while AI remains comparatively cheap. Once models can do work performed by top humans, he says, Google, Anthropic, or OpenAI may be willing to pay more for tokens used to automate AI research than ordinary users will pay for lower-value consumer uses.
If a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over 250K a year.
Patel says that is more than 15 times the current H100 spot price, before considering that an AI system can work outside normal hours. He acknowledges that a sudden abundance of software-engineer-equivalents might reduce their marginal value. But, invoking the objection to a fixed “lump of labor,” he suggests innovation and specialization could keep the marginal value of labor—and therefore compute—astonishingly high.
Frontier procurement is already above spot prices
Frontier labs cannot simply buy isolated spot instances when demand rises. They need clusters large enough to operate efficiently, flexibility over the resource, and security for model weights and customer information. The relevant price is therefore the cost of securing a large, specialized tranche of capacity, not merely the headline rate available to an ordinary cloud customer.
Patel doubts that margins alone will absorb the revenue-versus-capacity mismatch. Margins above 90%, he says, would require leading models to stay sufficiently better than alternatives that competition does not erode the premium. He finds it wild to assume that margins on intelligence would persist at that level.
As a case study, he points to the Wall Street Journal report displayed on screen that Google would pay SpaceX $920 million a month, from October 2024 through June 2025, for data-center capacity including at least 110,000 Nvidia chips. Patel describes the fleet as a blend of GB200s and GB300s and says Google’s implied price was about twice the spot hourly price for those GPUs—while spot prices were themselves already elevated from February.
The supply response has hard physical limits
The historical weakness of scarcity arguments is that market signals can induce substitution and innovation. Patel invokes the Simon–Ehrlich bet, commonly used to illustrate that point, while noting that Ehrlich might have won had the wager been made in a different decade.
His distinction is that compute supply is less elastic than commodity extraction and has fewer ready substitutes. Dwarkesh Patel divides the roughly 3x annual expansion into three contributors: about 1.4x from Moore’s Law, 1.2x from new fabrication capacity, and 1.8x from AI taking wafer allocation that previously served smartphones and PCs.
He does not see easy ways to accelerate any of them. Maintaining Moore’s Law would itself be a “miracle,” he says. New fab capacity is constrained through 2030 and possibly beyond by the need to build additional ASML EUV machines. Reallocating leading-edge wafers has a built-in ceiling.
The Semianalysis chart displayed in the source projects accelerators at 90% of TSMC N3 wafer demand in 2026. Separately, Patel describes AI’s share of that leading-edge capacity as moving from 60% to 86% by the end of the following year; the figures are related but are not presented as identical measures.
Once AI cannot keep taking wafer share from phones and PCs, Patel argues, even sustaining 3x annual capacity growth becomes questionable. His thesis is not permanent chip scarcity, but a period in which supply cannot expand as fast as smarter systems raise the economic value of each unit of compute.
Scale economies turn scarce capacity into a concentration risk
Patel expects compute eventually to become cheap in a far more automated world, where robots can turn raw materials into chips and costs approach those of inputs and processing. His concern is the intervening “pre-singularity” regime.
Model businesses have unusually strong scale economies, he says: training acquires skills once, and those capabilities can be shared across many users. Human labor requires each additional worker to be trained separately. Revenue growing faster than compute is consistent with that asymmetry.
I wish we didn't live in a world with such strong economies of scale for intelligence because I'm worried about power concentration, but it seems we do.
If scarce compute becomes more valuable, the labs best able to monetize frontier capacity gain leverage in two directions: they can pay more for hardware than rivals, and they can charge more for models that economize on that hardware.


