Orply.

Faster AI Inference Could Expand Infrastructure Spending

Andrew FeldmanEd LudlowBloomberg TechnologyFriday, July 24, 20265 min read

Cerebras CEO Andrew Feldman argues that faster AI inference can expand infrastructure demand by making AI output more productive, rather than simply reallocating a fixed pool of spending among chip suppliers. Under Cerebras’s partnership with AMD, AMD’s Helios system would process prompts while Cerebras hardware generates answers, a division Feldman says combines GPUs’ strength in parallel processing with Cerebras’s high-speed token generation. He contends that the commercial question is not token cost alone, but the productivity customers can extract from faster responses.

Faster inference can expand demand rather than divide it

Andrew Feldman frames the AMD-Cerebras partnership around a commercial claim: faster AI responses can increase the value customers get from AI, making them willing to spend more rather than merely shifting a fixed infrastructure budget between suppliers.

Fast inference is productive inference, and where AI is productive people are willing to spend more and more and more.
Andrew Feldman · Source

Asked how the companies would divide revenue from a combined offering, Feldman rejected the premise that the arrangement is principally about allocating an existing market. “Speed makes markets bigger,” he said. Both companies have substantial backlogs, in his account, while customers around the world are demanding fast inference at extraordinary scale.

The first deployment of the combined system is planned for the Cerebras cloud later this year, Feldman said, followed shortly afterward by broader availability. His argument is that lower latency is not simply a benchmark result: making AI faster can make its output more productive, which in turn can support greater spending.

That distinction informs Feldman’s response to the industry’s emphasis on token costs. Cerebras aims to provide more tokens per unit of power, more tokens per dollar, and faster token delivery. But Feldman said the central difficulty is not necessarily that tokens are expensive; it is measuring the productivity a buyer gets from them.

The system assigns prompt processing and answer generation to different machines

The proposed architecture rests on a division within inference itself. As Andrew Feldman describes it, processing a prompt—analyzing the query before an answer begins—is a computationally parallelizable task. Generating the answer has different requirements, particularly when the aim is to produce tokens at very high speed.

AMD’s Helios system is intended to process the prompt, while Cerebras generates the answer. The result is a disaggregated inference flow in which two machines perform separate parts of one request. Bloomberg Technology’s on-screen partnership graphic described the split plainly: AMD products “decipher queries” and Cerebras “handles answers.”

Feldman says GPUs are particularly well suited to the parallelized prompt-processing stage, and describes Helios as the system for that work. He positions Cerebras as the system for response generation, claiming that the combined flow is the fastest available and delivers unusually high throughput.

The problem isn't that tokens are expensive, the problem is that it's hard to measure how much productivity you're getting from tokens.
Andrew Feldman

Cerebras does not use HBM, Feldman said, and therefore can generate tokens for less without depending on that constrained component. The broader point in his account is not merely cheaper output: faster response can make an AI application more usable and its productivity easier to capture.

Open standards let Cerebras pursue a wider ecosystem strategy

Ed Ludlow asked why Cerebras and AMD had chosen a disaggregated design rather than a more tightly integrated server approach, comparing it with Nvidia’s announced relationship with Groq. Feldman said he could not assess Nvidia’s approach because it had not yet been delivered to market. Cerebras’s own route, he said, depends on open, standards-based I/O.

Using standards-based high-speed Ethernet rather than proprietary I/O allowed Cerebras and AMD to build their joint solution quickly, Feldman said. In his view, that openness also gives Cerebras a way to work with other vendors whose components could be part of a high-speed inference system, instead of operating as a closed stack or “walled garden.”

The company says this design follows its earlier work with AWS Trainium. Feldman cited that integration as evidence for the approach, saying AMD and AWS were using Cerebras in disaggregated configurations for high-speed inference delivery. He said the same standards-based model could let Cerebras engage with Google or other component makers as well, though those potential integrations were presented as an opportunity rather than as announced deployments.

The AMD relationship itself predates the new server arrangement. Feldman said AMD had invested in Cerebras during mid- and later-stage financing rounds, and that the companies had been discussing ways to collaborate for a long time. He has known AMD CEO Lisa Su since AMD acquired his prior company, he said, and described the Helios-Cerebras design as an idea that emerged from ongoing discussions with Su, CTO Mark Papermaster, and other AMD leaders.

Cerebras focuses on what it can control amid uncertainty over model costs

On the question of whether Chinese AI developers are pursuing lower costs per token than U.S. developers, Andrew Feldman was explicitly uncertain. He said Cerebras can produce tokens for less because it does not depend on HBM, but said he did not know what Chinese model makers spend to develop their models.

Feldman characterized the issue as unresolved: lower apparent costs could reflect borrowed or distilled technology, he said, or they could reflect technical inventions that permit genuinely cheaper model development. He said Cerebras has visibility into the compute requirements of some U.S. frontier labs, including OpenAI, which he identified as a large customer, but lacks a comparable understanding of Chinese model-development costs.

He was similarly noncommittal on the appropriate policy toward open-weight models. OpenAI and Anthropic, he said, made the market through enormous investment, while Chinese and other followers are attempting fast-follow strategies. Rather than prescribe an industry-wide answer, Feldman returned to the variables he said Cerebras can affect: tokens per watt, tokens per dollar, and the speed with which those tokens arrive.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free