AI Selloff Tests Whether GPU Scarcity Can Fund the Buildout
Gavin Baker, an investor focused on AI infrastructure, argues that July’s selloff in AI and semiconductor stocks reflected fears about financing and model-layer disruption rather than deterioration in compute demand. He says rising GPU rental rates, accelerating hyperscaler operating cash flow and growing private-lab and open-source workloads support a different reading: capacity contracted at older prices can reset higher and fund more of the buildout internally. The thesis fails, he says, if GPU prices remain depressed, demand weakens, debt becomes essential or regulation blocks new data-center power.

The thesis turns on whether AI capacity can reprice before credit becomes indispensable
Gavin Baker describes July’s AI-equity decline as “2022 in a month”: many AI-linked names fell 40% to 60% from their highs in a straight line, even as the quantitative signals he encountered in Silicon Valley appeared, to him, to be accelerating rather than decelerating.
His thesis requires a specific sequence to hold. Demand for compute must remain strong; installed GPU capacity committed under older contracts must reprice higher as those contracts expire; and the resulting operating cash flow must fund much of the next phase of the buildout. If that happens, Baker believes the market is mistaking a temporary free-cash-flow squeeze for a deterioration in the underlying economics of AI infrastructure.
| What Baker says must hold | Evidence he cites | What would challenge it |
|---|---|---|
| Compute scarcity persists | GPU rental prices have risen, including for older hardware | A sustained, sharp fall in GPU pricing or widespread excess capacity |
| Contracted capacity resets at higher prices | Current spot pricing is above rates embedded in older long-term GPU contracts | Renewals at lower rates or an inability to monetize new capacity |
| Demand is broader than public data captures | Private labs, inference clouds, and open-source use appear to be growing | Aggregate weakness or plateau across labs and open-source providers |
| Buildout can be funded internally | Operating cash flow at major hyperscalers is accelerating, by Baker’s calculation | Operating cash flow stalls while debt becomes necessary to sustain capex |
| Capacity can be energized | Developers and equipment suppliers are responding to power constraints | Moratoria or other restrictions that prevent new capacity from coming online |
The contrary case is clear in his formulation. AI demand could flatten, GPU rental prices could fall sharply and remain low, operating cash flow could fail to accelerate, and hyperscalers could become materially dependent on debt just as credit conditions worsen. Political opposition could also prevent data centers from being built or energized regardless of demand for their output.
Baker says he spent the week trying to disprove his own position. The question he repeatedly posed was whether anyone had seen a negative quantitative AI metric or a clear deceleration. His answer was largely no, with one qualified exception: third-party indicators suggesting Anthropic’s growth curve may have slipped below its earlier trajectory. He says that interpretation was disputed by Anthropic shareholders, while OpenAI and open-source inference appeared to be accelerating.
The difficult-to-observe part of the market, in his view, is private demand: Anthropic, OpenAI, and U.S. inference-cloud providers including Fireworks, Baseten, and Modal. Public investors can see semiconductor-company cash flow and hyperscaler free cash flow, but have less visibility into private labs and workloads served through open-source systems. That missing layer can make the standard picture—semiconductor cash flow rising while cloud companies consume cash building AI infrastructure—look more one-sided than it is.
The most consequential observation for Baker is GPU pricing. In 2024 and 2025, even bullish observers generally expected the cost of renting GPUs to decline gradually as supply arrived. He instead says that older GPUs and current-generation systems have become more expensive to rent. He describes speaking with a startup that had rented several thousand B200 GPUs at a rate in the mid-$2 range per GPU hour and expected to pay just under $4 for an effectively identical cluster seven months later.
That pricing gap matters because much of the installed compute base was financed or sold under long-term agreements at rates below current spot pricing. As those agreements expire, Baker expects capacity to reset to higher prices. Spot prices could decline from today’s levels and the reset could still be upward, provided spot remained above the rates embedded in the older contracts.
He treats operating cash flow, rather than free cash flow, as the appropriate measure of whether that reset is starting to work. Baker says aggregate operating cash flow reported by Microsoft, Meta, and Amazon accelerated from 28 to 32; after adjustments for what he characterizes as unusually large one-time expenses, including legal costs, he puts the change at 28 to 35. Heavy capital expenditure is visibly depressing free cash flow. The relevant question, he argues, is whether newly installed capacity generates enough operating cash flow to make the investment self-funding.
The underlying fundamentals are improving, and stocks—Nvidia is actually, as we record this, at its lowest forward P/E of the last 10 years. The market 100% thinks they are significantly overearning.
Baker is explicit that the market could be right. Nvidia’s valuation may reflect a justified belief that current earnings are temporary and GPU economics will normalize down sharply. He notes that, by his account, semiconductor equities were cheaper only around Christmas 2018 and the onset of COVID, both of which were subsequently V-shaped recoveries. He offers that comparison as a measure of how extreme he considers current valuation, not as a forecast that the same pattern will recur.
Open source may shift margins away from models without reducing compute demand
Patrick O'Shaughnessy identifies stronger open models as an important catalyst for the selloff. Baker says investors treated releases including GLM 5.2 and Kimi K3 as evidence that frontier labs would lose pricing power and that AI-infrastructure demand would weaken.
Baker’s counter-thesis depends on separating the price paid by an end user from the cost of producing a token. Open-source systems can reduce the model layer’s margin and make AI cheaper to use without reducing the compute, memory, electricity, cooling, and data-center capacity needed to serve the workload. If lower prices increase usage, he says, open source transfers economic surplus away from frontier-model providers while expanding demand for infrastructure.
A token is a token, and it takes the same amount of FLOPS, watts, space, cooling to make.
Baker acknowledges that tokens are not literally identical. But he argues that investors were treating a shift from expensive frontier tokens to lower-priced open-source tokens as a volume decline, when it may instead have been a mix shift. The Silicon Data token index, which he says does not capture every token, appeared to dip and flatten during the selloff. His explanation is that activity moved toward open-source output following releases such as GLM 5.2 and Kimi, affecting what the index could show without necessarily indicating less underlying compute use.
The economic mechanism is margin compression at the model layer. Baker describes frontier-model inference as potentially carrying margins of 80%, 90%, or even 95%, while open-source systems may operate at far lower margins. The customer sees lower prices. But the cloud providers that run Anthropic and the open-source systems still charge for the underlying compute. In that account, lower model-layer margins and more elastic demand can direct more dollars toward GPUs, data centers, and memory rather than less.
He expects a multi-model architecture rather than a single winner. AI-native companies can fine-tune open models using proprietary data, put them behind a router, and send difficult tasks to frontier systems such as Claude or Grok for planning, checking, or orchestration. In some cases, Baker says, users can get slightly better outcomes at about half the price.
That lower price does not necessarily imply fewer GPU hours. A company may stabilize or reduce its AI budget after installing a router, but generate more tokens because cheaper models make more use cases economical. An end-user spending reduction may reflect reduced model-layer margins rather than a reduction in the compute consumed behind those model layers.
This architecture could also change the durability of AI-native software businesses. A company that began as a frontier-model “wrapper” can gather domain-specific data, tailor an open model with supervised fine-tuning or reinforcement learning, and reduce dependence on a single provider’s pricing and terms. Baker cites Fireworks’ Nexus as an example of a product that ingests customer data, tailors a model, and routes requests between proprietary and frontier systems.
He does not dismiss the opposing possibility. Some people he knows expect frontier labs to distill capabilities so effectively that they can serve every intelligence level at structurally lower cost, leaving limited room for independent open-source providers. Baker calls that possible, but not his base case, partly because AI-native companies are accumulating specialized data and inference-cloud providers are improving at customization and routing.
Efficiency improvements remain a meaningful uncertainty. Baker says people in the field appear to believe continual learning and sample-efficient learning may be nearing a solution. If models could reach useful capability with fewer training tokens and then learn efficiently after deployment, that could create a discontinuity in training demand. He and O’Shaughnessy regard such an outcome as beneficial, but Baker doubts it would necessarily reduce total infrastructure demand over time: he expects training to become a relatively small share of semiconductor demand relative to inference.
Credit matters because a debt-funded buildout can unwind quickly
The bearish case Baker takes most seriously is credit. His own thesis works only if a large share of the AI buildout can be financed from operating cash flow. If the industry instead needs substantial debt to keep adding capacity, it becomes vulnerable to the classic capital-cycle problem: fixed repayment demands arrive even when supply and demand temporarily move out of balance.
Baker points to the signals that made him take that possibility seriously: real yields had risen, credit spreads had widened, and Meta’s recent bond issue priced worse than he expected. O’Shaughnessy also notes widening Nvidia credit-default-swap spreads. Baker calls those undeniable facts, even if some private-capital investors interpret the move as banks hedging commitments rather than as a judgment that AI infrastructure is fundamentally impaired.
A debt-fueled buildout can unwind quickly because supply does not need to exceed demand by much before highly leveraged owners face pressure. He compares that failure mode to the internet buildout. That is why he regards July’s individual narratives about open source, token data, or Meta potentially renting compute as secondary to the financing question.
Baker argues that consensus estimates may be understating the revenue productivity of incoming Blackwell and Rubin capacity. In his view, the market is effectively modeling those future gigawatts at monetization rates closer to Ampere, an older Nvidia generation, rather than current-generation economics. If hyperscalers monetize at that assumed level, he estimates they could produce about $1.3 trillion to $1.4 trillion in operating cash flow. If they monetize at a discount to current Blackwell economics but above Ampere-era levels, he estimates something closer to $2 trillion.
The difference could remove roughly $700 billion of projected credit demand, improve credit ratios, and make external funding easier if companies still chose to use it.
The risk is a potential “Blackwell air pocket.” Hyperscalers have committed hundreds of billions of dollars to equipment that is initially used heavily for training, which Baker says does not itself generate a direct return. Earlier in the year, he says, the market was willing to look through that gap. July brought several concerns together just as operating cash flow began to improve, and investors stopped looking through it.
His answer is not that credit will never be needed. It is that constrained credit could preserve the value of existing capacity: if fewer new GPUs can be financed, the GPUs already energized become scarcer and potentially more valuable. But the argument remains conditional on sustained demand and an ability to monetize capacity above the rates embedded in older contracts.
The adoption case matters only insofar as it supports that cash-flow durability. Baker estimates that perhaps 250,000 to 500,000 people are using agentic AI today, while the industry already faces an acute compute shortage, and asks what demand could look like at 100 million or 500 million users. He sees adoption arriving in uneven waves: leading companies are already using model routers to manage cost, AI-native companies are treating tokens as a core input to growth, and many businesses have barely adopted AI.
At the most AI-intensive companies O’Shaughnessy sees, token spending can reach 20% to 25% of compensation; Baker says he has heard of a company at 50%. Both frame the more constructive outcome as growth rather than mass labor replacement: AI-native firms may avoid hiring as many people, while spending on models rises alongside revenue. They cite charts from Cognition, Ramp, and Stripe that they say associate higher AI spending with faster growth, though Baker acknowledges those comparisons may not adequately control for industry.
Memory supply and financing have become part of the competitive product
Gavin Baker argues that the market is shifting, especially in memory, from pursuing every increment of short-term pricing upside toward long-term supply agreements, or LTAs. These agreements can take different forms, including customer prepayments and contractual price floors and ceilings. Investors may read them as suppliers trading away immediate upside for stability. Baker’s interpretation is that the agreements reflect the cost of being unallocated in a future shortage.
More memory per unit of compute can produce more tokens per FLOP, Baker says, making memory capacity one of the most important levers for raising output and reducing cost. In that environment, an allocation affects competitive position. A buyer that breaks an LTA during an oversupply period may secure a lower price in the short term but damage its future standing with suppliers.
His hypothetical is Google breaking an agreement during an oversupply. If the industry later cuts capacity and moves back into shortage, suppliers could favor buyers that honored their commitments. The penalty would not merely be a higher price. A company could lose the supply it needs while direct competitors gain it.
The argument rests on the fact that the relevant buyers compete directly. Baker names Amazon with Trainium, Google with TPUs, AMD, and Nvidia, along with startups seeking allocations. He contrasts this with Apple’s historical position in consumer electronics, where its purchasing scale made suppliers highly dependent on its volume. In AI compute, a supplier can redirect scarce allocation to a buyer’s competitor.
Nvidia, in Baker’s account, has recognized that the competitive product now includes financing. He describes a “credit wrapper” under which an outside lender finances a GPU buyer while Nvidia participates in a revenue share if GPU prices remain above a floor. He distinguishes this from conventional vendor financing because Nvidia is not itself lending the money.
If you need to be able to finance the chips, and you do, nothing’s more financeable than an Nvidia GPU.
Baker thinks the structure can give Nvidia equity upside in AI companies, a royalty-like participation in cloud revenue, and a way to bridge the period between heavy capital spending and later operating cash flow. It can also deepen Nvidia’s competitive position because a challenger may pay more at TSMC, pay more for HBM DRAM, and have less access to affordable financing.
That is why he doubts leading AI labs will readily reduce their compute commitments. He argues that Anthropic’s earlier restraint on compute helped OpenAI regain ground, and that the recent competitive gains of Grok reinforce the same lesson: in a capability race, underbuying compute can be as dangerous as overbuying it.
SpaceX is relevant to this argument chiefly as an example of the value of bringing clusters online quickly and selling capacity into a shortage. Baker says the company has demonstrated an ability to add compute faster and at lower cost than most peers, and that the market absorbed a large amount of its capacity without an evident slowdown in demand. But he is skeptical of a public report suggesting that SpaceX could add 8 gigawatts of compute in 18 months, calling that scale extraordinary and implausible enough that he does not fully believe the report.
The more general point is that power, equipment, memory, and financing have become interlocking inputs to competitive position. Companies that secure all four may be able to participate in the next model cycle; those that lose an allocation or cannot finance a cluster may not simply pay more—they may cede share.
Political permission is a constraint on the entire underwriting case
Gavin Baker calls regulation the largest AI risk. The failure mode is not necessarily an engineering shortfall or a model plateau, but local and political opposition that prevents data centers from being built or energized. He points to a New York data-center moratorium as a warning that this could become a material constraint.
Baker’s position is that the industry has not made an adequate public case for data-center development. He describes a public narrative in which data centers raise household power bills, consume local water, and eliminate jobs. He argues that some current development agreements pair behind-the-meter power arrangements with local investment and trade employment.
Under those arrangements, he says, local power prices can fall and construction, maintenance, upgrades, and operations can support ongoing work for electricians, plumbers, HVAC contractors, and other trades. He also argues that developers can place facilities away from dense residential areas and make community commitments such as funding local infrastructure.
He says an erroneous water-use claim has been especially influential: an author, in his account, overstated data-center water use by 10,000 times, later acknowledged the error, and yet the claim continued circulating. O’Shaughnessy compares that dynamic to the persistent Popeye-and-spinach story, in which a misplaced decimal allegedly created an enduring misconception about spinach’s iron content.
Baker’s prescription is a communications and political effort. The industry should explain concrete local terms rather than assume the benefits are self-evident: what power arrangements have been made, what infrastructure is being funded, what work persists after construction, and what broader benefits AI may provide. He also says the sector should connect the compute buildout to scientific and medical applications, including rare-disease research.
The investment implication is direct. If companies cannot make that case, regulatory constraints may limit new power capacity regardless of demand, GPU economics, or available financing. In Baker’s framework, that would constrain the supply side of the market—and prevent the operating-cash-flow expansion needed to validate the industry’s capital spending.



