Goodput, Not FLOPS, Measures What Frontier AI Systems Actually Deliver
Amin Vahdat, Google’s chief technologist for AI infrastructure, argues that theoretical chip speed is a poor measure of AI capacity: at large scale, failures and recovery determine how much useful work a system delivers. He calls that measure goodput, and says it should guide decisions across the infrastructure stack, from chip design and networking to power and storage. As AI workloads shift toward long-running agents, Vahdat argues, the systems supporting them must evolve beyond accelerators alone.

Delivered work matters more than theoretical compute
Amin Vahdat argues that FLOPS—the theoretical computing capacity of a chip—is a poor measure of what an AI system actually delivers. A workload rarely runs on one chip alone. Its performance depends on how accelerators are composed with one another, how CPUs feed them data, how the network connects them, and how the software uses the whole system.
The more useful measure, he says, is goodput: the useful work a workload completes under real operating conditions. FLOPS utilization—the fraction of theoretical capacity used by a particular workload—is one way to assess how effectively a chip is being used. But goodput also accounts for workload slowdown, failures, and recovery. The question is not only how much computation a system can perform in theory, but how long it takes to produce the answer the workload is meant to deliver.
The difference becomes clear when a large synchronous job fails. Training, serving, and agentic workloads may require thousands or tens of thousands of components to coordinate at microsecond or millisecond intervals. If one component stops, the others may have to wait. The system must detect the problem, work out what failed, restore a checkpoint, and restart. In some cases, it may have to redo computation or start again from the beginning.
Vahdat compares that wasted effort to solving a problem on paper: if you make a mistake at step four and have to return to step one, you have still done work, but that work has not brought you closer to the answer. The relevant measure is the time it took to solve the problem, including interruptions and recovery—not how many steps were attempted.
At a scale of 100,000 accelerators, he says, something may fail multiple times a day and, depending on the configuration, multiple times an hour. These are complex systems made up of chips, chiplets, memory, and networking; any of those elements can fail. Nor is the failure problem limited to hardware. Compiler or runtime bugs, operating-system issues, and model problems can all reduce end-to-end performance even if the chips themselves are working properly. A nominally capable accelerator cannot deliver its theoretical performance if the surrounding software or system prevents the workload from using it effectively.
There is no single common cause to eliminate. Vahdat describes failures as a “long tail of constant discovery”: new products and configurations keep exposing issues, sometimes in high-speed networking, sometimes in hardware, and sometimes in software. If a cause were common and well understood, he says, it would likely have been found and fixed. The operational challenge is to find a fault in a huge volume of telemetry, recover quickly, and keep doing so across jobs that may run for hours, days, or weeks.
That is why he frames accountability around the delivered performance of workloads that matter, under actual failure conditions. Goodput is a Google term, he says, though he sees more people across the industry adopting it. It makes reliability part of the performance question rather than treating it as a separate property of the hardware.
Goodput also changes how Vahdat thinks about capacity. Google’s target is to roughly double effective serving capacity every six months, measured as token-generation capability. That does not mean doubling the number of FLOPS. The aim is to increase the number of tokens the system can generate with its available infrastructure. Vahdat says the gains may come as much or more from software as from hardware: model improvements, runtime changes, and dozens or hundreds of smaller optimizations can accumulate.
Asked how gains divide among silicon, models, and software, Vahdat says he does not have an exact breakdown. In his experience, most of the gains in intelligence per watt often come from the model side. Software also matters because it helps ensure the hardware is used effectively. He describes intelligence per watt as a useful measure, and also frames the system goal as goodput per watt. Hardware remains an important multiplier: year-over-year performance improvements of 2× or more are possible, he says, and improvements above the hardware layer can build on that.
The six-month target is therefore an outcome to pursue across the stack, not a forecast that any one component will double. Hardware improvements can raise the ceiling; model and software changes determine how much of that potential becomes useful serving capacity. Failures and recovery determine how much of the work actually reaches the user.
Purpose-built infrastructure depends on a durable workload
An AI data center still contains familiar ingredients: buildings, electrical and mechanical infrastructure, cooling, networking, storage, and computing equipment. The difference, Vahdat says, is how closely the site may be designed around a particular generation of AI hardware.
Traditional data centers are planned as long-lived buildings, perhaps over a 20- to 30-year horizon, while the hardware inside may last about six years. A general-purpose facility must therefore accommodate several generations of servers, networks, storage, and accelerators. AI sites can be more purpose-built: the building, cooling, and power distribution may be designed alongside the equipment that will occupy it.
The physical difference between racks helps explain why. A storage rack might draw 10 to 40 kilowatts; a TPU or GPU rack can draw hundreds of kilowatts today, and Vahdat says people are discussing megawatt-scale racks in the coming years. A row designed for many storage racks will have different power-distribution and networking requirements from one built around a small number of dense accelerator racks. Storage racks need comparatively little networking, especially when built around hard drives. A facility made fully fungible across all those uses could end up too large or overbuilt for its intended purpose.
Google’s TPU program began in 2013, when, Vahdat says, conventional wisdom held that custom accelerators for a narrow set of workloads were a bad bet. General-purpose processors, improvements in performance, and familiar programming models seemed likely to win. The first TPU focused on inference, initially for applications such as language translation and voice recognition. A later generation extended the approach to training. Transformers and recommender systems then broadened the range of work the chips could support.
The decision to build separate TPU 8i and 8t chips for inference and training followed a calculation about how large and durable each workload was likely to become. If inference were expected to account for only 2% or 5% of a chip’s lifetime use, Vahdat says, a specialized chip might not justify the cost, even if it were twice as fast. But if serving were projected to account for 30%, 40%, 50%, or 60% of that chip’s lifetime use, specialization could make sense. Those percentages describe a possible forecast of the chip’s use over its life, not a claim about inference’s share of the overall market.
That calculation is complicated by how long chips remain in service. A new architecture commits resources to a workload years before its full demand is known, and specialization trades flexibility for speed and power efficiency. A workload that is dominant for a month or two may not be durable enough to design around: the opportunity to benefit from the chip could be too brief. The work is partly technical and partly a projection about what will persist, and what performance gain specialization can deliver.
The two TPU 8 variants retain some flexibility. Both can run the other’s main workload, even though each is better at the work it was designed for. That matters over a hardware lifetime of roughly six years: if one type of capacity is left over, it can still be used for the other task. A chip that could only perform its specialized workload would require a much more exact forecast of demand across that period.
Specializing to transformers alone would still leave choices. Vahdat describes transformers as relying on linear-algebra operations, including vector and matrix multiplication and softmax, that can be supported in hardware. Further specialization could target a particular model architecture—its layers and the shapes of the matrices and vectors it uses—but that would narrow the design further. He notes that some companies are exploring specialization at that model level.
Customers do not face a simple choice between interchangeable TPUs and GPUs. Vahdat says GPUs are more general-purpose, and that different workloads are better suited to different options. Google uses GPUs internally and sells them through Google Cloud, as well as offering TPUs. The reference stacks provided for both types of accelerator give customers a starting point, but customers may adapt or specialize those stacks for their own use cases. Vahdat’s stated aim is to provide the option that best meets a customer’s needs, rather than treat one kind of accelerator as the answer for every workload.
Co-design works across a pipeline, but late changes have a high bar
Amin Vahdat says co-design can produce gains across the layers of a system, from chip architecture and software to networking, storage, cooling, and power delivery. If each layer is optimized in isolation for broad compatibility, mismatches between them can leave performance on the table. Improvements at several layers can compound into a larger end-to-end gain.
The alternative is a system designed to run across different clouds, hardware, networks, and software stacks. Such a system is easier to move when new capacity becomes available and reduces dependence on a particular infrastructure. Its cost is that it may be less efficient than a system tuned to a narrower set of components. Vahdat frames this as a trade-off between flexibility and potential performance, not as a reason to eliminate portability altogether.
He declines to speculate about whether differences in other companies’ model architectures result from differences in their compute stacks. For Google, he describes the work with DeepMind as a close partnership. Researchers and engineers can work together when a model improvement might benefit from hardware support, including when a chip is already partway through development.
The practical opportunity depends on where the chip sits in a five- or six-stage, multiyear pipeline. Some chips are in production. Others have returned from manufacturing and are being debugged and brought up. Some are in implementation and close to tape-out; others are still in design or concept. The teams can work on all of these stages, but the options are not the same. Production hardware can be measured against current models. Designs still in concept can be compared with a wider range of possible future workloads. A chip close to tape-out leaves less room to change.
When DeepMind researchers identify a model improvement that might benefit substantially from hardware support, the teams can examine what changes are feasible. They may adjust the hardware, adjust the model to fit a different hardware option, or decide that a requested change cannot be made in time. Vahdat says that, in some cases, Google may delay tape-out by a week or two if the expected benefit warrants it. He distinguishes that from a small gain: a tweak worth 1% or 0.5% on a chip about to tape out is unlikely to justify stopping the program. A large opportunity may.
This is a consequential advantage of working within the same company, in his account. The teams can get together for days or weeks to examine a proposed change and its effects on both sides. A hardware change that would be hard or impossible to make near tape-out can sometimes be intercepted; even then, changing a chip in flight is a major decision, not a routine response to every new model idea.
Further out in the pipeline, the teams can compare possible hardware designs with projections of where model architectures may go. Google’s simulation infrastructure helps estimate how workloads would map onto different designs, supporting iteration before a chip reaches implementation. Vahdat says the teams work in the same buildings and often the same rooms, with frequent interaction. He also says hardware is planned two to five years ahead, while model researchers are not necessarily working on that horizon. Researchers who have grown up alongside the hardware program have learned that a late change to a chip is not like a software change that can ship in a couple of weeks.
The basic TPU design has helped make that long-range work possible. Vahdat says its medium-level architectural primitives have remained in place since v1, although they have been extended. Those include large matrix-multiply units, a sparse core for vector and scatter-gather operations, and the ability to access remote memory through the interconnect. The point is not that every detail has stayed fixed, but that a stable base has carried across generations of models and neural-network algorithms. That base gives the teams something to adapt without requiring them to predict every future model in detail.
Long-horizon agents increase demand beyond accelerators
Long-horizon agents change the rhythm of data-center workloads, in Vahdat’s account. In a conventional human-to-model interaction, a person may take seconds or tens of seconds to read an answer and decide what to ask next. An agent can process a response and issue another request in milliseconds, without a person setting the pace.
That faster loop increases accelerator demand, but it also raises demand for CPUs, networking, and storage. The agent’s next step may require reasoning about a response, gathering context from local or remote DRAM, or retrieving information from an SSD or hard drive. The work is not just generating tokens; it is coordinating the sequence of actions and data the model needs.
That creates a design question: put CPU racks beside accelerator racks, or keep the systems in separate buildings? Accelerator racks are denser and generally need more networking than CPU racks, so placing them together makes it harder to specialize each building for its requirements. Separate buildings preserve that specialization but require substantial networking between them. Vahdat notes that once traffic leaves a building, reliability, cost, and latency become more difficult; latency could reach hundreds of microseconds or more, depending on queuing.
The need to serve agents and other workloads also affects where capacity is built. Training benefits from concentration: a large cluster with short network distances can be advantageous. Serving, by contrast, needs to be distributed around the world, close to users and across different model endpoints. Google cannot rely only on older training clusters as they are freed up. Training capacity may be concentrated in a small number of large sites, even within the same region or part of a continent, while serving needs to reach places on other continents as well.
Nor are serving clusters necessarily cheaper per megawatt. Serving can require accelerators, compute, networking, and storage to be colocated, while training can use a denser and more uniform deployment. Serving clusters may also be smaller and less vertically integrated because individual endpoints and models have to be distributed geographically. The appropriate mix depends on where the workload must run and how its components need to communicate.
Optical switching can reroute connections without moving fiber
Amin Vahdat describes optical circuit switching as one way to change how data-center networks are configured. Google introduced wave division multiplexing—carrying multiple signals on a single fiber—and optical circuit switching for communication between racks, he says, around 15 or 16 years ago.
Unlike an electrical packet switch, which reads packet headers and forwards packets through a network, an optical circuit switch directs light from an input fiber to a selected output fiber without converting the traffic into the electrical domain. In Google’s system, programmable micro-mirrors steer the light between ports. The switch can be configured to map each incoming fiber to an output port by changing the orientation of the mirrors.
That makes it possible to reconfigure connections without physically moving fiber. The technology can create direct optical connections between groups of racks that communicate frequently—for example, compute and storage clusters serving the same workload. It can also help expand or contract the network spine under software control.
In TPU systems, Vahdat says, the same capability can redirect connections from a failed TPU rack to a spare rack, without moving fiber. The system can keep a spare rack available and redirect the light to it when another rack fails. The change can happen in milliseconds.
The light still travels through fiber for most of its path. Vahdat says free-space transmission across a large building would face attenuation and bandwidth loss, as well as the difficulty of aiming connections across a three-dimensional space. At the optical circuit switch, the light is directed by mirrors. He describes networking more broadly as an area of rapidly growing capability and demand in data centers.
Power is a long-term constraint, but the grid can share capacity
Amin Vahdat says the constraints on data-center development shift over time, and no single one is always the most difficult. But if he had to name the most fundamental long-term constraint, he would choose power. Abundant clean energy could ease many problems, he says, but when it will be available at scale remains uncertain.
Google’s preferred approach is to work with utilities and remain connected to the grid, rather than build all its own generation. A gigawatt-scale requirement has to be planned years ahead. The utility may be able to provide, for example, 700 megawatts in a year when the data center needs a gigawatt. The remaining capacity could mean waiting, arranging some local generation, or using other sources such as solar and batteries. Local generation could also potentially supply power back to the grid during periods of high demand.
Vahdat’s argument for this approach is partly about sharing capacity. If a company tried to supply a gigawatt at very high reliability on its own, he says, it might need to build two gigawatts of generation to provide redundancy. He presents that as a consequence of pursuing very high reliability, not as a general requirement for every data center. Working with the grid allows capacity to be shared across a wider base. In his account, the arrangement can benefit the company, the utility, and residents.
That arrangement also has a cost-allocation condition. Vahdat says Google works to cover the infrastructure costs created by its demand, such as transmission lines that need to be built or upgraded and additional utility substations. Without that, utility investment for a large new load could in theory contribute to higher rates for other customers. The planning is therefore not just a question of whether power can be delivered, but also of coordinating with the utility over several years and accounting for who pays for the required upgrades.
How large to make a data center is another trade-off. A larger training cluster can reduce network distances, but concentrating too much capacity in one place creates a single point of failure and depends on power being available there. Vahdat recalls debate at Google more than a decade ago about whether to put everything into one data center. A single site might simplify some aspects of operation, but it would concentrate risk and, at today’s scale, could not house all of Google’s requirements.
Sites may range from a rack at the edge of a network or at an internet service provider, to tens or hundreds of megawatts, to training clusters approaching a gigawatt. The right scale depends on the purpose and location. Google uses models and simulators to inform those choices, but Vahdat describes sizing as an art as well as an engineering decision.
The lifecycle of the hardware adds another layer. Vahdat says seven- and eight-year-old TPUs are still fully utilized, but that does not mean old systems never get replaced. Once hardware is fully depreciated, the power efficiency of newer generations can make an upgrade worthwhile. Google thinks in terms of replacing whole TPU pods—groups of thousands of chips and many racks—not simply swapping individual chips. The new pod may not fit the old footprint exactly, so decommissioning and retrofitting require planning even when the next generation’s dimensions are not yet known.
Open interfaces preserve customer choice
Amin Vahdat says co-design need not require customers to accept a closed, end-to-end stack. Google supports specialized systems while also trying to keep the interfaces between components open.
He gives the example of JAX, Google’s model-development framework, and PyTorch, which many customers prefer. Google could require customers using TPUs to adopt JAX, but instead offers PyTorch/XLA so that unmodified models can run on the platform. The aim is to let customers use the framework they choose while still making specialization possible where they want it.
Vahdat compares open interfaces to the role of IP in the development of the internet. In his account, IP became a common layer that different software and hardware could connect through, rather than requiring a single tightly controlled stack. Google’s position, he says, is to support open standards and, ideally, open source around them, while allowing customers to plug in more specialized components when they choose.
That principle also informs Google’s support for both TPUs and GPUs, along with other accelerators. GPUs are more general-purpose, Vahdat says, and the right choice depends on the workload. Google uses and sells GPUs as well as TPUs; its stated aim is to offer customers an option that fits their needs, rather than insist on one design.
The next system may be a denser rack—or an orbital data center
Amin Vahdat says AI is changing his team’s work beyond software engineering. Engineers use AI for software development, testing, rollouts, and aspects of design. He says hardware engineers are using AI at a level comparable to software engineers, and that the time from design kickoff to tape-out and the time required for chip bring-up are shrinking. AI is also helping with data-center planning by bringing relevant information together for human decision-makers; he does not describe it as replacing their judgment.
Google is also pursuing orbital data centers, which Vahdat calls a “moonshot.” He connects the idea to the challenge of energy production. In his account, space offers about 40% more solar power because the atmosphere does not attenuate it. A sun-synchronous orbit could expose solar cells to sunlight 98% to 100% of the time, compared with roughly 28% to 35% on land, he says. Together, those conditions could reduce or remove the need for batteries and provide a carbon-free power source.
The difficulties are substantial. Cooling in space is harder than it may seem, and failures are harder to repair. Networking would also change: instead of fiber between components, orbital systems might use free-space optics, with lasers aimed at receivers and calibrated in real time. Vahdat says he sees no fundamental showstoppers, while emphasizing that there are many challenges.
His picture of a frontier supercomputer in 2036 is similarly tentative. The “cone of uncertainty” is wide, he says, but he expects greater integration. A rack may contain hundreds or more accelerators, be manufactured centrally, and need only a small fiber bundle rather than the elaborate cabling visible in today’s systems. It could draw multiple megawatts. The installation might be as simple as wheeling the rack into place and connecting water, power, and fiber.
Vahdat leaves open the possibility that such a rack could be launched directly into space and connected to a module by a robotic arm. For now, both the ground-based rack and the orbital version remain projections. The practical choices he describes are shaped by power, networking, reliability, and how much specialization a workload can support.

