Orply.

Underbuilding, Not Overbuilding, Is AI’s Near-Term Infrastructure Risk

Gavin BakerDavid Georgea16zMonday, August 31, 202615 min read

a16z investor David George and Gavin Baker argue that AI’s central infrastructure risk is not overbuilding but failing to add compute fast enough. They base that view on what they describe as rapid data-center paybacks, bottlenecks in power and equipment, and a demand base still limited to a small group of heavy users that could expand to hundreds of millions. In their account, that scarcity could allow frontier labs, open models, cloud providers, chipmakers and applications to grow together rather than divide a fixed pool of value.

Fast paybacks and constrained supply make underbuilding the central risk

Gavin Baker and David George see an unusual combination in AI infrastructure: estimated paybacks measured in months, physical constraints on new capacity, and demand still concentrated among a relatively small population of intensive users. If those conditions persist, they argue, the greater risk is not an ordinary overbuild but a shortage of compute.

Baker points to disclosures from Nebius and CoreWeave and estimates that Nebius can achieve a nine- to ten-month payback under a particular set of assumptions. In his illustration, one gigawatt costs $50 billion to bring online, while customer prepayments cover 50% to 60% of that cost. The remaining investment can be monetized in the spot market, where he believes paybacks may be faster. George says that even using a smoothed-out pricing level rather than the highest prevailing rate produces compelling economics.

9–10 months
Baker’s estimated Nebius payback under a customer-prepayment scenario

Baker believes SpaceX could earn back compute investment faster because it brings large clusters online quickly and may monetize capacity at higher rates. The broader point is that he has rarely seen businesses able to deploy tens or hundreds of billions of dollars while potentially achieving a sub-one-year payback.

Financing reinforces that case, in their telling. Baker says Nvidia-based systems can be financed, reducing the equity required from the buyer. George describes the financing as relatively low-cost. Baker attributes its availability to private-credit firms and banks underwriting the assets, to useful lives being extended as models improve, and to the prospect that monetization per gigawatt rises as token spending becomes more valuable.

The implication is not that every layer of AI must compete for a fixed pool of value. George rejects the familiar framing that asks which part of the stack has to lose: frontier labs or open models, applications or clouds, infrastructure providers or model companies.

This is not an “or” thing, it’s an “and” thing.

David George

Frontier models can improve, older and cheaper models can become useful, open models can spread, well-executed applications can grow, and cloud and inference providers can benefit from the same expansion in demand, George says. Baker adds that not every application will win, but sees room for labs, open models, clouds, inference providers, chip companies, and strong applications to expand together.

The conventional bear case remains intact. Baker notes that railroads, steel mills, automobiles, radio, television, PCs, and the internet all produced periods of overvaluation and overbuilding. Debt-funded construction is especially vulnerable, he says, because it requires prompt returns rather than returns that may arrive years later. But this buildout also faces constraints in power, wafers, memory, networking equipment, labor, land, copper, and permits. Those bottlenecks, in Baker’s view, make an unconstrained overbuild less likely than the historical analogy suggests.

The demand base is still narrow, even as usage compounds

David George argues that current AI revenue gives a misleading picture of diffusion. He puts roughly $180 billion of revenue on the back of something like 30 million heavy paying users, then immediately allows that the real number may be materially lower. Gavin Baker says he would take the under.

Within companies, George sees a pronounced power law. The highest-spending engineers can consume 10 or 100 times as many tokens as the median engineer. Older-economy companies that are using AI effectively may spend around 1% of human compensation on tokens, he says, while AI-native companies may spend high single digits or more than 10%.

Against that concentrated use, George places a potential base of roughly 1.5 billion knowledge workers.

1.5B
Knowledge workers George identifies as a potential demand base

Baker offers Atreides as an anecdotal example of how rapidly consumption can move. He says its internal token spending rose 100-fold from March through August. After two people gained access to Grokbot Enterprise, he expected spending could rise another 10- to 20-fold in a month.

George’s point is that heavy token usage is not necessarily wasteful if it produces valuable work. Baker describes building a podcast summarizer, Substack summarizer, X summarizer, and sentiment tracker through Grok bots in seconds—tools he says would have taken hours to build through Claude Code. For him, the shift was not simply faster coding, but a reduction in the effort required to turn an idea into a working tool.

George distinguishes those uses from the next stage. Summarizers and preparation systems remain reactive: they improve a user’s information and judgment but do not execute work. He is testing systems that examine his workflow, recommend automations, and take action after approval. If users routinely authorize agents to do work rather than merely answer questions, he expects token consumption to become much more open-ended.

Diffusion could disappoint, George acknowledges. But their shortage thesis rests on a specific possibility: demand broadens from a small group of highly intensive users to hundreds of millions of people faster than power, equipment, and infrastructure can be built to serve them. High utilization alone would not establish that conclusion.

Frontier labs can choose capability over current revenue

Gavin Baker expects major frontier labs to generate substantial operating cash flow while continuing to forgo free cash flow, recycling proceeds into GPUs and other accelerators in pursuit of capability gains. The economics differ from those of mature internet companies, he argues, because a lab’s current revenue partly reflects a decision about how to allocate scarce compute.

Revenue depends not only on customer demand but on which checkpoint a lab releases, where it prices along the capability-cost curve, and how it divides capacity among inference, internal research, and training. Baker’s illustration is deliberately simple: a company with 10 gigawatts of power might put eight into inference and two into research and training, monetize inference at $60 billion per gigawatt, and report $480 billion in annual revenue. If a research breakthrough persuaded management to redirect eight gigawatts into training, annual revenue could fall to $120 billion even if the decision improved long-run competitiveness.

8 GW
Inference allocation in Baker’s illustrative 10-gigawatt lab model

David George says this trade-off was less acute for mature internet and cloud companies, where infrastructure serving current revenue and long-term investment were more separate. A frontier lab can sacrifice present monetization to pursue a future capability advantage. Baker expects public markets to become more accustomed to that dynamic, though he notes that stock volatility after an IPO can affect morale, recruiting, and retention—and may limit how sharply management is willing to cut revenue.

The pair see strategic caution as costly when capability improvements can rapidly alter competitive position. Baker says Microsoft slowed capital expenditure when it should have continued investing, contrasting that decision with what he characterizes as more aggressive choices by OpenAI and SpaceX. He recalls Dario Amodei’s formulation of the dilemma: spend too little and lose share; spend too much and risk bankruptcy. Baker’s conclusion is that in a rapidly advancing market, underinvestment can be nearly as consequential as overspending.

A supply shortage could turn intelligence into a scarce luxury

David George expects capacity to remain tight through 2028 even with forecast buildouts, which he believes could be delayed by political resistance. Gavin Baker is therefore more concerned about massive undersupply than oversupply.

In that scenario, the price of intelligence could rise rather than fall. George recalls an argument that token costs could increase by as much as 10 times—not as a forecast, but as a conceivable supply-and-demand outcome. Users already select frontier tokens even when cheaper models can handle many tasks, he says. His explanation is that frontier output creates enough surplus for users to pay up.

Baker worries about the distributional outcome. Restrictions on data-center construction could create “compute inequality,” he says, in which large companies and wealthy people can obtain the best systems while smaller businesses and ordinary users cannot. Advertising may eventually support low-cost consumer products, George says, but advertising businesses take time to build. There could be an interim period in which advanced intelligence cannot be broadly delivered at very low prices.

Open models do not eliminate the physical constraint. Baker disputes the idea that open-model tokens are free: for comparably sized models, he says, producing an open token takes essentially the same compute as producing a frontier token. The difference is largely the margin added on top. He also distinguishes open weights from open source, pointing to the Kimi license’s stipulated 30% share of generated revenue. George adds that Kimi can be token-hungry enough that its cost per completed task is high.

Open models may lower margins, diversify suppliers, and give organizations greater control over deployment. In Baker’s account, they do not abolish the cost of compute.

Data centers need a public case grounded in everyday gains

Gavin Baker and David George argue that the AI industry has not made its infrastructure case in terms that matter outside technology and finance. Staying ahead of China may be strategically important, George says, but it is too abstract for people focused on affordability and on how AI will change their lives.

A Brookings commentary shown on screen—“Competing AI strategies for the US and China,” by Kyle Chan—captures the China-competition framing that George considers correct but insufficiently concrete for voters.

Baker wants companies to make a more direct case: identify businesses that improved, workers who gained opportunities, and towns that gained revenue. He praises Meta’s practice of highlighting small businesses that benefited from advertising and argues that AI companies, chip providers, and infrastructure builders should do the same.

The truth shall set you free, but only if you tell it.

Gavin Baker

Baker says data centers have created demand for electricians, plumbers, and HVAC technicians, and may offer a more attractive economic path than college for some workers. He says data centers paired with behind-the-meter generation can dramatically increase local tax revenue and revitalize small towns.

A CBS MoneyWatch article shown on screen, “Data center frenzy is spurring a jobs boomlet for blue-collar workers,” illustrates the employment claim Baker wants the industry to foreground. Baker and George also point to Loudoun County, Virginia, as an example of a high-income area with dense data-center development and substantial tax revenue.

George calls the water-consumption concern “totally debunked.” Baker says the industry is becoming better at addressing environmental concerns and describes natural gas as a relatively clean fuel. He argues that lower U.S. natural-gas prices than those in Europe and Asia give American manufacturing a cost advantage because electricity is a central industrial input. Combined with the data-center boom, he sees that advantage as part of a broader reindustrialization of America.

Baker also alleges that an organized Chinese Communist Party-funded campaign against U.S. data centers is being amplified through TikTok. George’s related concern is that the industry cannot rely on promises of future medical breakthroughs. It must show tangible, everyday benefits beyond chat interfaces and search substitutes. Baker agrees that companies should stop merely talking about curing cancer and produce real breakthroughs, while maintaining that the infrastructure case already includes more immediate claims about wages, construction, tax bases, and productive capacity.

Orbital compute is a potential release valve for terrestrial constraints

Gavin Baker and David George treat orbital compute as a possible response to terrestrial scarcity, not as a replacement for Earth-based data centers. George rejects the image of a giant building in space. The relevant system, he says, is more like an airplane-sized arrangement of racks with solar wings. Baker describes a sun-synchronous orbit, solar power, and a radiator kept in shadow for cooling.

Baker says SpaceX considers the engineering problem simpler than building a Starlink satellite, whose phased arrays and movement create additional complexity. George takes a more qualified view: he sees no physics reason the concept cannot work, while acknowledging that its costs appear imposing. His confidence rests on what he sees as a repeated pattern at Musk-led companies, where initially difficult launch, satellite, and vehicle economics improved sharply over time.

The hinge, Baker argues, is Starship reusability. His illustrative terrestrial data-center cost is $50 billion per gigawatt, including $35 billion of IT equipment and $15 billion for power, cooling, labor, and related infrastructure. He expects the latter category to become more inflationary on Earth as materials and skilled labor tighten. In orbit, solar power and radiative cooling could remove much of it, leaving launch cost as the decisive variable. If reusable Starship launch costs fall below $1 billion, he argues, the economics could change quickly.

Earth-based training remains important in Baker’s account. Closely colocated GPUs retain advantages, and latency and speed-of-light limits remain real. But he expects a growing fraction of global compute to move into orbit. He says Elon Musk has stated that SpaceX and Nvidia co-designed a Rubin rack intended for a fourth-quarter 2027 launch. Even if that date slips by two quarters, Baker regards 2028 as close enough to matter for capacity planning.

The broader SpaceX story, they say, does not need orbital compute to work immediately. But for the supply question, its relevance is narrower: cheaper, reusable launch could eventually create swing capacity if power, cooling, labor, and materials on Earth become increasingly constrained.

The enterprise prize is control over a portfolio of models

Gavin Baker expects enterprises to use an ensemble of models rather than rely on one dominant system. No model will be best across every point on the capability-cost curve, he argues. Companies may fine-tune capable open base models on proprietary data, retain control of that customized intelligence, and call frontier models for planning or verification.

That structure protects enterprise context, which Baker treats as a company’s core intellectual property. Sharing all internal data with a frontier lab may be hazardous to a company’s financial health, he says. Fine-tuning a model on data the company controls offers an alternative: it owns the model, its capabilities, and its cost structure.

A router would conceal the complexity from users. Baker says his understanding is that Grokbot already draws on Gemini, Sonnet, Llama, and Opus behind such a layer, though he expects its owner to seek a more fully first-party system over time. David George describes the intended division of labor: use a frontier model for planning and lower-cost models for execution.

This world is friendlier to Microsoft than one controlled by just two frontier labs, Baker argues. He says Microsoft failed to build a competitive frontier model but may benefit from a plural market, including American open models that he expects Nvidia to help fund. Chip companies can use their cash flow to finance major training runs, he says.

The strategic prize is what George calls the abstraction layer between intelligence and an organization’s users. It must connect sensitive enterprise context, choose and route models, manage cost, establish privacy assurances, continuously update the system, and make the result feel seamless.

The speakers emphasize that this is not ordinary middleware. Baker compares it to operating a national retailer: the proposition sounds simple until a company must stock, price, staff, and maintain a thousand stores across 50 states. The enterprise AI layer faces an analogous operational challenge, with data access, model updates, routing, and user trust all needing to work together.

They see Cursor as a product-led example. Baker says that while many labs presented grand visions of AGI or superintelligence, Cursor focused on making a useful product. George says it met customers and technology where they were and has worked its way toward autonomy. Coding has a special advantage because it is documented and verifiable; much of enterprise work is neither.

Harvey’s progress in legal work is relevant for the same reason: legal tasks are unusually documented and somewhat verifiable. Tax and compliance may follow. But the broader market will draw competitors from every direction: Microsoft, Databricks, Snowflake, Palantir, inference providers, specialized applications, Salesforce, Workday, and products such as Fireworks Nexus.

Baker says vertical integration may matter over the long run because a company that does not own compute will struggle to become the lowest-cost provider. The winner, in this framing, will be the company that can make heterogeneous intelligence useful, trusted, and operationally invisible.

Nvidia’s advantage is a system of supply, finance, and adoption

Gavin Baker sees Nvidia’s advantage not merely in chip performance but in a system that joins supply-chain access, financing, and customer adoption. Nvidia-based data centers are easier to finance than alternatives, he argues, because lenders understand the assets and Nvidia can offer residual-value guarantees.

Baker gives a $50 billion data-center illustration: a buyer might contribute $15 billion of equity and finance the other $35 billion. If a residual-value guarantee is below Nvidia’s gross profit from selling into the facility, he argues, Nvidia can offer it with limited risk and potentially gain additional revenue-share upside. He contrasts this with TPUs, which he believes may require at least twice the equity contribution and carry higher rates on the debt portion.

That cost-of-capital gap can determine which customers can build at scale, particularly when leading labs are willing to pay heavily for compute. Baker believes Nvidia’s guarantees can help other buyers compete rather than allowing an Anthropic-and-OpenAI duopoly to dominate capacity.

He characterizes Nvidia as vertically integrated but horizontally open: a company with accelerators, CPUs, Ethernet switches, DPUs, and systems spanning several forms of networking. A rival may produce a strong accelerator, but system-level competition requires more than one chip.

Baker estimates that Nvidia has locked up 70% to 80% of relevant supply, including fab capacity, DRAM, NAND, lasers, capacitors, and components required to build racks. His larger point is that Nvidia anticipated a full bottlenecked system earlier than rivals and now makes multihundred-billion-dollar commitments alongside suppliers and financiers.

That is why he advises accelerator startups not to challenge Nvidia directly. His rule of thumb is that each 1% of accelerator market share might be worth $100 billion. The rational strategy, in his view, is to find a defensible niche and integrate with Nvidia’s ecosystem.

Open models fit that system rather than threatening it, Baker argues. Lower margins on tokens can lead to greater token consumption, which in a supply-constrained market means more demand for compute. Nvidia’s commercial interest in a fragmented model landscape therefore aligns, he says, with wider access to multiple models.

Hardware competition remains punishing. A chip can look promising through design, tape-out, emulation, and simulation, then fail when it arrives from the lab. If it fails completely, the company may need hundreds of millions or billions more and wait years for another attempt. Even a functioning chip can lack product-market fit.

In a market where almost every credible allocation sells out, Baker says shipments alone do not reveal true customer preference. Deal structures are more informative. A chip company investing in a customer can create a favorable arrangement when the investment remains below expected gross profit. A residual-value guarantee below gross profit, coupled with revenue share, can also be attractive. Warrants tied to a fixed token price may work if the chip’s performance improves faster than the issuer’s stock price. Unbounded warrants can become negative present value for the issuer because the better its stock performs, the more expensive the deal becomes.

For Baker, these arrangements indicate which suppliers have leverage, which customers require subsidies, and which infrastructure assets sophisticated financiers are willing to underwrite.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free