AI Buyers Are Choosing Models on Cost, Governance, and Deployment
John Coogan argues that the latest model releases show why benchmark leadership is becoming an inadequate guide to AI capability and adoption: OpenAI’s GPT-6 Astra may set a new mark on ARC-AGI-3, but ARC’s Mike Knoop says the result still falls short of evidence for general intelligence. Coogan and Jordi Hays contend that buyers will increasingly judge models by practical demonstrations, cost, speed and data-governance terms rather than leaderboard gains alone. Coogan makes the same distribution argument about Nvidia’s reported $13bn Hugging Face acquisition, framing it as a bid to connect open-source developers to the compute infrastructure needed to deploy their work.

Benchmark saturation is not the same as general intelligence
John Coogan framed a crowded release cycle—new models from Anthropic, Google, Meta, and OpenAI—as a contest that is becoming harder to settle with a leaderboard alone. The most consequential question was not which model led a benchmark, but what benchmark leadership establishes when scores begin to approach saturation.
According to Coogan’s account of OpenAI’s launch materials, GPT-6 Astra posted a 99.9% result on ARC-AGI-3. Before the launch, an on-screen post from Lisan al Gaib had circulated lower but similarly striking figures: 98.6% on ARC-AGI-3, 97.6% on FrontierMath Tier 4 v2, 74.1% on DeepSWE v1.1, and 100% on ExploitBench.
Coogan described ARC-AGI-3 as a set of video-game-like tasks that require a model to learn how an unfamiliar environment works, think creatively, and operate efficiently toward an objective. In the displayed response to Astra’s result, Mike Knoop of ARC described the benchmark more formally as testing whether models can make sense of unfamiliar environments and operate autonomously toward goals within them.
Coogan also cited a 64.6% Astra score on TerminalBench Science at maximum reasoning effort, with a stated cost of $26, and a 0% score on the ExploitGym honeypot, where lower is better. He called Astra the world’s best computer-use model, while also saying he wanted to test that claim in practical use.
Knoop’s response supplied a more qualified interpretation than the headline score alone. He called Astra the new state of the art on ARC-AGI-3 and said it represented “a qualitatively large leap towards AGI” whose pace was “frankly surprising.” But he also wrote that ARC lacked evidence to call it AGI yet.
The stated reason was that open-ended invention remains unsolved. ARC was still studying the gaps between current systems and human capabilities, Knoop said, and intended to make that problem the basis of ARC-AGI-4. In that framing, Astra’s performance suggests an ability to navigate difficult bounded environments; it does not yet demonstrate that the system can originate new conceptual terrain.
That qualification clarifies Coogan’s broader concern with the release cycle. He said the industry may be reaching “the end of the benchmark era,” because benchmarks have become difficult to interpret and have drawn accusations of bench hacking. The Artificial Analysis Intelligence Index can compress many evaluations into a single meta-benchmark, he said, but a score does not by itself settle whether a model will be dependable for a company’s actual work.
Jordi Hays said users have now spent enough time with competing models to develop their own internal benchmarks. In practice, he said, people increasingly rely on demonstrations, reports from trusted users who work across multiple systems, and specific examples of what models do well or badly. A convincing demonstration—building a game, solving a novel problem, or operating a computer—may matter more to prospective users than a small gain on an aggregate index. It still does not resolve the question ARC raised about general intelligence.
A strong model still has to be cheap, governable, and deployable
The commercial contest is not reducible to a frontier score. Coogan used Anthropic’s Claude Fable 5.1, Google’s Gemini 3.8 Flash, and Meta’s Muse Spark 1.3 to show how price, throughput, data controls, and deployment requirements can pull buyers in different directions.
Coogan said Fable 5.1 scored 66 on the Artificial Analysis Intelligence Index, the highest result yet on that aggregate measure. He put Opus 5 at 63 and an earlier Fable release at 62. Anthropic also said an improved caching system would make ordinary workloads 25% cheaper and long-horizon agentic work 45% cheaper.
Gemini 3.8 Flash was presented as a performance-per-dollar proposition rather than an unqualified frontier leader. Coogan said it scored 73.7% on DeepSWE, just behind Opus 5, and had received an independent Intelligence Score of 59. At roughly 300 tokens per second, he said, the model combined high speed, low cost, and strong coding performance on that benchmark.
Muse Spark 1.3, Coogan said, scored 75.4% on DeepSWE, ahead of Gemini 3.8 Flash and above Opus 5 and GPT-5.6 Soul on that measure. It did not lead across the board: Opus still beat it on several professional-work and computer-use evaluations, he said. The releases therefore did not produce a single clean winner. The relevant trade-offs differed by workload, with benchmark performance, speed, and cost all shaping the case for adoption.
Data policy may be equally material for enterprise buyers. Coogan said one explanation for weaker-than-expected enterprise adoption of Fable was customer demand for zero data retention. Companies, he said, do not want closed-model labs collecting their private information.
He said Anthropic’s no-ZDR approach would be replaced with Enterprise Frontier Safeguards. Under the arrangement Coogan described, some data would be retained and use would still be monitored for hostile activity, but data would not go directly into Anthropic’s databases. Hays said Anthropic was testing functionality that would permit data retention on infrastructure owned by the customer.
The commercial implication is straightforward: model quality does not remove data-governance requirements. A technically strong system can remain difficult to deploy when its terms for retaining and monitoring information do not fit how a prospective customer handles sensitive internal data.
Enterprise demand is concentrated—and its meaning is unsettled
The concentration of AI spending raises a separate question: whether current enterprise revenue reflects broad adoption or dependence on a narrow group of buyers.
Coogan highlighted an on-screen post by Ara Kharazian of Ramp Economics Lab stating that 80% of OpenAI’s and Anthropic’s enterprise revenue comes from 1% of their business customers. The displayed chart showed the top 1% of businesses accounting for an increasing share of enterprise spending from late 2024 through mid-2026, with the next 9% and all other customers representing much smaller shares.
Kharazian’s post characterized the pattern as a concentration risk unseen in other software categories tracked by Ramp. It said the top 1% of buyers skew heavily toward technology businesses and AI products and services. The concern, as presented in the post, is correlation: if these companies face a market correction together, the revenue base supporting the leading AI providers could be exposed to the same pressure at once.
Coogan offered a different reading. He said the top 1% of American companies by sales generate 80% of total revenue, a comparison he considered more relevant than employment concentration. He separately said the top 1% of businesses employ 65% of the workforce.
His interpretation was that large enterprise AI purchases are likely to be consumption-based rather than collections of $20 or $200 subscriptions. A large company may treat AI spending more like a variable operating expense, potentially scaling with its revenue and activity. Coogan estimated annual AI spending at roughly $150 billion, or around a quarter of a percent of total U.S. business revenue.
He summarized the ambiguity by saying that the chart could be a “Rorschach test” for views on AI: some readers may see concentrated spend as a healthy reflection of where business revenue already sits, while others may see an unstable customer base. The source leaves that dispute open. Kharazian’s chart identifies a concentrated revenue structure; Coogan’s comparison proposes that concentration may be consistent with the distribution of sales among American firms.
Nvidia is buying a front door to open-source AI
Nvidia’s reported acquisition of Hugging Face, which Coogan put at roughly $13 billion, was presented as a strategic bet on distribution and infrastructure demand rather than ownership of the most capable proprietary model.
Coogan called Hugging Face “the GitHub of AI”: a place where developers can upload, discover, test, modify, version, and demonstrate models. Unlike a frontier-model lab, he said, Hugging Face does not attempt to own the smartest models. Its role is to host and organize the ecosystem around them.
The scale Coogan cited is the core of the strategic case: more than 18 million developers, 200,000 companies using the product, 3 million models, and more than 500,000 datasets.
| Hugging Face measure | Reported scale |
|---|---|
| Developers | Over 18 million |
| Companies using the product | 200,000 |
| Models | 3 million |
| Datasets | Over 500,000 |
Hugging Face did not begin as model infrastructure. Coogan said its founders initially built an AI companion aimed at teenagers: a cute, emotional digital friend that users could name, text, send selfies to, and trade emojis with. The product had reached a million messages a day and more than 100 million messages in total by 2018, he said, but did not become a durable consumer business.
Its developer flywheel began after Google released BERT in 2018. Coogan said BERT arrived as a paper and an implementation in TensorFlow, while many developers preferred PyTorch. Hugging Face converted BERT from TensorFlow to PyTorch and released the conversion free of charge. Developers adopted it, and the company expanded from a code library into a platform for publishing models, attaching datasets, discussing changes, maintaining private repositories, and building demonstrations through its Spaces product.
The company’s earlier strategic position was neutrality. Coogan described Hugging Face as “Switzerland of AI,” supporting competing clouds, chips, frameworks, and models. He said its 2023 financing included Salesforce, Google, Amazon, Nvidia, AMD, Intel, Qualcomm, and IBM, reinforcing its presentation as a platform that could serve multiple infrastructure ecosystems.
Nvidia ownership creates a consequence to watch: whether that neutral positioning remains persuasive to users and partners that compete with Nvidia. Coogan did not suggest Nvidia had already made Hugging Face closed or exclusive. His argument was that Nvidia has a clear incentive to make open-source AI more powerful and more commercially attractive.
Nvidia already sells chips to closed-model labs, Coogan said, but those labs are also building ASICs and participating in an ongoing infrastructure tug of war. Open source offers another route to sustained compute demand. If developers find a model and framework through Hugging Face, then move on to buy inference, compute, or GPUs, Nvidia has an opportunity to be present at that next step.
Jordi Hays raised a more specific version of the question. Hugging Face already offers model routing and ranks inference providers, he noted. Could it compete more directly with products such as OpenRouter and Ramp’s router?
Coogan did not make a firm prediction. He described the acquisition as serving two goals: preserving Nvidia’s strength in the closed-model ecosystem while becoming a front door for growing open-source development. He also argued that the deal sends a signal that developers can build successful companies and careers around open source.
Coogan said Nvidia’s reported $12.9303 billion purchase price corresponded to the decimal code for the hugging-face emoji and could also be read as a hexadecimal green associated with Nvidia. He treated that as an entertaining piece of symbolism. The more durable rationale in his account was that Hugging Face’s developer network can direct open-source experimentation toward the infrastructure required to run it.

