Orply.

Better AI Models Are Expanding, Not Reducing, Compute Demand

John CooganJordi HaysTae KimTBPNWednesday, July 29, 202610 min read

Tae Kim argues that the AI selloff rests on a static view of demand: investors assume more efficient models mean cheaper compute and, ultimately, excess data-center capacity. He contends that more capable reasoning, open-weight and agentic models instead make new workloads commercially viable, expanding total demand for compute, memory, networking and power even as the cost of individual tasks falls. Kim’s more speculative extension is that recursive self-improvement could add another major source of demand within months.

Better models can expand compute demand rather than erase it

Tae Kim sees the pressure on AI-linked equities as a sentiment unwind built around a static view of demand. The bearish story, as he describes it, runs from greater model efficiency to cheaper compute, then from cheaper compute to excess data-center capacity. Kim’s argument is that this misses what happens when models become more capable: new workloads become economically useful, usage rises, and total demand can grow even if the cost per unit of intelligence falls.

Kimi is the immediate example. Kim says fears around the open-weight model resemble the DeepSeek selloff: investors see an apparently more efficient system and infer a coming compute glut. But he describes Kimi as a 2.8-trillion-parameter model that requires substantial infrastructure to serve. Its launch, he says, overwhelmed its servers, and its published materials describe a recommended deployment on a server with 64 GPUs.

2.8T
parameters Kim says Kimi has

Kim’s reading of DeepSeek is that efficiency did not reduce the need for infrastructure. Reasoning models created demand because they made more tasks viable. He expects agentic systems to have the same effect: as models can carry out multi-step work, companies have more reason to use them in software development, product design, customer operations, and other internal workflows.

When you have more capable models that come out people find uses for them.
Tae Kim · Source

The question, in Kim’s framing, is not simply whether inference gets cheaper, but whether the volume of work expands faster than the per-task cost falls. He expects the next six to nine months of agentic deployment to be materially larger than the market anticipates.

John Coogan agrees that adoption remains shallow outside a concentrated group of AI-native companies. In many ordinary businesses, he says, AI use is still limited in both the number of employees using it and the amount of time they spend with it. Kim points to a claim attributed to Aravind Srinivas and relayed by Jared Sleeper of On Deck: usage would be 100 times higher if every company adopted AI at the level of the most advanced firms.

Jordi Hays draws a line between nominal adoption and operational adoption. Giving an employee a ChatGPT Pro account may be useful, he says, but it is not the same as putting coding agents into production or using AI to automate repetitive work. The more consequential diffusion is the integration of models into the workflows that determine how quickly a company can operate and improve.

Kim places annual corporate spending on IT and knowledge management at roughly $6 trillion. Against that, he estimates OpenAI and Anthropic’s combined annual recurring revenue at about $120 billion. His point is that the addressable corporate budget is still vastly larger than the current frontier-model revenue base. If model providers keep expanding into that budget, he sees $200 billion, $300 billion, or $400 billion of combined revenue in the next few years as plausible.

$6T
annual corporate IT and knowledge-management market Kim cites

The most aggressive extension of Kim’s thesis is recursive self-improvement: his expectation that AI models may soon use compute to help develop and improve their own capabilities. He believes this could arrive within three, six, or nine months, based on what he reads as signals from OpenAI personnel, Anthropic’s public statements, and frontier researchers’ responses to his own posts. This is a forecast rather than an operating result. If it happens on the timetable he expects, Kim argues, it would add a new source of demand on top of reasoning and agentic workloads.

We had this exponential ramp for reasoning, exponential ramp for agentic, and then if RSI actually happens ... that's going to soak up an unbelievable amount of compute.
Tae Kim

The capex case depends on demand arriving before the buildout is complete

Tae Kim answers concerns about runaway capital expenditure by emphasizing the lag between construction and utilization. Data centers and fabs take time to build, he says. Companies that expect demand in the next year or two need to secure power, shells, chips, memory, and networking capacity before that demand appears in current financial results.

Kim rejects the premise that hyperscalers are necessarily financing an ever-larger buildout from a fixed revenue base. Azure, he says, is growing 40%; Google Cloud, 80%; and Amazon at high double digits. In his view, if cloud revenue continues growing at those rates, the operating cash flow available for investment grows with it.

If revenue is growing 40 to 80% this year, next year, and the year after, that's more revenue you have, that's more operating cash flow you have to invest.
Tae Kim · Source

Kim calls this the difference between return on investment and “return on revenue.” A company can assess whether a particular AI expenditure produces direct returns, but it also has to consider what happens if competitors use AI to improve their product, customer service, and development speed first. He points to AT&T’s use of generative AI, described at AMD’s Advancing AI event as putting 100 models into production and consuming a trillion tokens per month. In his framing, AI deployment can protect existing revenue as much as it creates a standalone AI revenue line.

The inputs Kim emphasizes are operational signals, not proof that every data-center investment will pay out: rising market estimates, constrained component supply, and cloud businesses that are already growing. He treats them as evidence that demand is expanding faster than investors assume.

SignalWhat Kim highlightsHow Kim interprets it
AMD market forecastLisa Su raised AMD's 2030 CPU and agentic-compute forecast from $120B to $220B in three monthsKim reads the revision as a sign that AMD is seeing stronger demand than it had previously expected.
SK Hynix capacityKim says executives described customers requesting five to six times more capacity than the company could serveHe takes the mismatch as an indication that memory and component demand exceeds present supply.
Hyperscaler cloud growthKim cites Azure at 40%, Google Cloud at 80%, and Amazon at high double-digit growthHe argues that continued growth would expand the cash flow available to fund infrastructure.
The operating signals Kim uses in support of his demand thesis

Kim treats the AMD revision as particularly meaningful. Public-company executives who have run established businesses for decades, he argues, do not sharply increase their estimates of the total addressable market without seeing a major change in customer demand. He also cites SK Hynix executives as saying customers were asking for five to six times more capacity than the company could provide, while the company planned to double capacity over five years.

The market’s immediate concern is that this spending resembles a dot-com-style excess. Kim’s counterargument is conditional: if GPU cloud and inference are highly profitable services, capital deployed to add supply need not signal weak underlying demand. He cites a Morgan Stanley estimate that inference can carry 60% to 80% profit margins. That estimate, and the demand indicators he cites, are central to his case for the buildout; they do not settle whether particular projects, financing arrangements, or capacity forecasts will prove justified.

That is also how he reads AI spending at Meta. A leaked remark from Mark Zuckerberg about slower-than-expected agentic progress was interpreted as evidence that Meta might retreat from investment. Kim says the statement was taken from an internal town hall and that subsequent reporting pointed instead to higher capital expenditure. Meta has had to reset its frontier-model operation, he says, and clearer progress may take six to 12 months; he does not take that reset to mean the company has become less willing to spend.

Jordi Hays puts the commercial challenge more sharply. Meta may gain from AI-enhanced ads, but investors still need to understand where its broader combination of frontier models, agents, internal tools, and open-weight releases produces significant revenue. Kim’s response is that Meta is already receiving returns through its existing business, while the contentious part of the spending is its renewed bid to compete in frontier models.

The Texas lease is smaller when read as a long-term obligation

Tae Kim argues that data-center financing stories should be read in terms of their actual obligations, not their largest possible headline number. His example is the Financial Times report describing NVIDIA as tenant for a $50 billion Texas data center.

The number initially looked alarming to him. But, as the terms were described in the discussion, NVIDIA’s 15-year lease commitment was worth about $20 billion; renewal options would bring the figure to $50 billion over 30 years. Kim’s narrower point is that the reported lease should be annualized and evaluated against those disclosed terms. On that basis, he characterized the commitment as roughly $1 billion to $2 billion a year rather than an immediate $50 billion exposure.

John Coogan adds that being a tenant does not make the lease an annual loss. NVIDIA could operate or rent out the resulting capacity in pursuit of a profit. Kim calls the commitment a rounding error relative to what he describes as NVIDIA’s roughly $320 billion annualized revenue run rate, which he expects to rise to $400 billion or $500 billion in the following year.

The lease arithmetic does not, by itself, establish that the facility will be profitable or that the arrangement is strategically sound. Kim’s point is more limited: the headline aggregate value can obscure the length and conditional nature of the commitment, and should not be treated as though NVIDIA had incurred a $50 billion cost immediately.

Kim takes a more reserved position on separate reports that NVIDIA was in talks involving OpenAI, SoftBank, and a possible backstop. He says the terms had not been disclosed and declines to speculate, arguing that the market should wait for the actual structure, metrics, and numbers before drawing conclusions.

He offers one interpretation of NVIDIA’s investments in the AI supply chain. The company has previously invested in CoreWeave and, Kim says, had recently bought stakes in optical companies Lumentum and Coherent. Critics may view such arrangements through the lens of vendor financing; Kim sees NVIDIA financing capacity among companies that will be needed to serve an expected surge in GPU-cloud demand.

In that view, investment is part of NVIDIA’s operating advantage rather than a separate financial maneuver. It helps build suppliers and cloud capacity required to deliver systems at scale when components are scarce. Whether those bets validate that interpretation depends, in Kim’s own account, on the demand growth he expects to materialize.

CUDA’s durability depends on production reliability and system scale

Tae Kim sees the prospect of AI-generated kernels as a reason to revisit, but not dismiss, NVIDIA’s CUDA advantage. The claim advanced around AMD’s AI event is that agentic coding could make it easier to write software for different chips and reduce NVIDIA’s pricing power.

Kim says AMD has a clear incentive to argue that CUDA is no longer a major obstacle. He notes that Kimi’s blog post described creating a GPU kernel, but argues that a technical demonstration does not establish that an AI-generated stack will be reliable under real production workloads. CUDA’s value, in his account, comes from software repeatedly optimized and debugged through enormous use.

CUDA is not NVIDIA’s only advantage, he argues. Kim also points to the company’s co-design across GPUs, CPUs, networking, and the systems that connect them. That integration matters for the large clusters required to train and serve frontier models, where a component-level comparison can miss bottlenecks elsewhere in the system.

Kim further emphasizes NVIDIA’s balance sheet and purchasing power. He says the company has secured constrained optical components, TSMC wafer capacity, and HBM memory, allowing it to build systems while rivals may face supply limits. In Kim’s view, NVIDIA’s scale allows it to prepay and make supply commitments that smaller competitors cannot match.

On power and data-center constraints, Kim similarly distinguishes between the existence of bottlenecks and an inability to grow through them. Jensen Huang had cited bottlenecks in power, data-center shells, components, and energy, Kim says, but also said the chip industry had enough supply to double revenue each year. Kim reads those comments as an indication that NVIDIA has planned around the constraints; he also argues that consensus revenue estimates for the following year remain well below what doubled revenue would imply.

Open-weight progress is part of the same demand argument

Tae Kim describes NVIDIA’s open-weight letter as an unusually broad industry response to Anthropic’s position. He says companies representing roughly $18 trillion in aggregate market capitalization had signed on, with Apple the notable holdout in his account. Kim sees the mobilization as a response to the possibility that the White House or Congress could impose restrictions on open weights or model distillation.

The policy question matters to his investment thesis because Kimi and DeepSeek are being interpreted as signals of cheaper, less hardware-intensive AI. Kim disputes that reading. Kimi’s scale, serving requirements, and initial demand are the concrete facts he emphasizes; in his view, capable open-weight models widen access to useful AI rather than eliminate the need for the infrastructure behind it.

John Coogan suggests that concerns about open weights may also turn on who releases them and who bears the consequences of misuse. He distinguishes American hyperscalers, which have large existing businesses and reputational exposure, from foreign publishers that may have less direct exposure to downstream incidents in the United States. Kim’s account is that the stated safety and distillation concerns have become intertwined with possible regulatory action.

That dispute does not change his broader conclusion. For Kim, the relevant economic question is whether open models make more reasoning and agentic applications viable. If they do, he argues, greater model availability becomes another reason to expect demand for compute, memory, networking, and power to rise.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free