Orply.

Kimi K3 Shows Open Models Closing the Frontier Gap

Jordi HaysJohn CooganTBPNTuesday, July 21, 202611 min read

John Coogan argues that Moonshot’s Kimi K3 challenges both the assumption that Chinese open-weight models will remain far behind proprietary frontier systems and the view that greater model efficiency weakens demand for AI infrastructure. In his account, efficiency shifts compute bottlenecks rather than removing them, while capable downloadable weights make cyber controls and the protected returns needed to finance frontier training runs harder to sustain. The article presents a related constraint in media: Netflix is putting generative AI into production workflows, but *The Odyssey*’s debut points to the continuing value of audience trust in a director’s name.

Kimi K3 challenges both the open-model gap and the compute bear case

John Coogan treated Moonshot’s Kimi K3 as a challenge to two assumptions that had begun to harden around AI: that Chinese open-weight models would fall materially behind proprietary frontier systems, and that more efficient model architectures would necessarily weaken demand for chips and data-center infrastructure.

Kimi’s weights had not yet been released at the time of the discussion; the model was gated behind an app, with a planned release expected if plans held. But its benchmarks and early tests had already produced a familiar cycle of excitement and suspicion. Coogan’s view was that the performance looked impressive without being inexplicable. Many lab leaders, he said, have expected open models to remain a few months behind the frontier rather than to vanish from contention altogether. Kimi appeared to fit that catch-up pattern.

The headlines for Kimi are impressive, but not wildly surprising based on the current trend lines and model progress.

John Coogan · Source

An unidentified panelist contrasted Kimi with GLM-4.2, which had appeared perhaps close to a year behind leading systems. Kimi seemed materially closer. Coogan was particularly interested in what that closeness would look like beyond broad benchmarks: where Moonshot had focused its training, which domains it had emphasized, where it might still lag, and how much practical performance depended on the surrounding harness rather than the base model.

He suggested that increasingly visible differences among leading models may reflect those choices. Some labs visibly prioritize mathematics, some code, some writing, some chess. Kimi’s strong performance in front-end evaluation raised the question of whether Moonshot had deliberately optimized for that kind of work, and what other capabilities might emerge as users test it in different environments.

Coogan laid out two broad explanations for the model’s emergence. The ordinary explanation is that Moonshot assembled strong researchers and enough compute, then chose open weights as a commercial strategy: gain attention, attract talent, create usage, and eventually sell hosted APIs and services. He compared the logic to Red Hat, whose software could be run independently but still supported a major business around consulting, hosting, and enterprise services.

The darker account circulating online included possible evasion of chip controls, extraction of frontier-lab data through intermediary companies, distillation from restricted systems, or direct theft of intellectual property. Coogan did not resolve those claims. He said the eventual explanation could contain elements of both genuine technical achievement and less transparent practices. Reports that Kimi identified itself as Claude or GPT were, in his view, similarly hard to assess and potentially straightforward to correct through training.

The cybersecurity evidence made the stakes of that uncertainty more concrete. A visual shown from Guillermo Rauch described internal “stealth evals” in which Kimi K3 was “top-tier at cybersecurity,” with “raw IQ,” rather than simply optimized against visible benchmarks. Rauch’s evaluation put Sol ahead at substantially higher cost, while saying Fable refused to complete the relevant run. Coogan took that result as a partial counterweight to a simple distillation account: the public models from which Kimi might supposedly have been copied are, in theory, heavily constrained in cybersecurity assistance.

That same result sharpens the tension around open access. A model that can be distributed as weights may widen the availability of strong defensive capability, but it also lowers the barrier to powerful cyber assistance outside the controlled interfaces where frontier providers can impose restrictions.

Efficiency shifts the bottleneck rather than ending the infrastructure race

The immediate hardware bear case is straightforward: if a model uses less memory or requires less compute at inference, demand for NVIDIA hardware, high-bandwidth memory, DRAM, and networking should decline. The SemiAnalysis thread shown on screen argued that Kimi K3 points in the other direction.

It said Kimi’s linear-attention approach, described as KDA, can reduce some networking requirements associated with KV-cache transfers. But the model’s more than 2.8 trillion parameters still require a large scale-up domain to store its weights. SemiAnalysis argued that this is precisely the kind of large-model inference for which NVIDIA’s NVL72 configuration is well suited. Lower KV-cache networking demands, it added, can be outweighed by the bandwidth required for “wide EP,” an optimization that distributes weights across GPUs.

2.8T+
Parameters cited for Kimi K3

Coogan’s broader point was that efficiency gains can move constraints rather than erase them. Kimi’s own capacity notice reinforced that argument. The company said that demand in the preceding 48 hours had pushed its GPUs close to existing limits, causing it to pause new subscriptions temporarily while prioritizing current users and adding capacity.

Jordi Hays added that access to compute does not necessarily require moving physical hardware into China. A Chinese lab, he said, could establish U.S. entities and obtain remote compute from several domestic sources. Coogan contrasted that possibility with TPU-based infrastructure, which could require closer coordination with Google, while GPUs are distributed through clouds, resellers, and smaller providers.

Nicolas Bustamante’s post expanded the meaning of “compute constrained.” The bottlenecks, he wrote, include land for data centers, permits, electricity, grid connections, chips, construction, and cooling. Demand is growing faster than supply, in his telling, even before AI is seriously used at broad scale. His conclusion was that companies with large cash flows from non-AI businesses are best able to fund the hundreds of billions of dollars in capital expenditure required to remain at the frontier.

Coogan did not see Kimi as a reason to abandon the AI infrastructure trade. Demand, he argued, exists at multiple levels: frontier intelligence, near-frontier intelligence, and older systems made cheap enough to become nearly invisible infrastructure. The task for companies is to find useful applications, then reduce the cost of delivering that intelligence until it can be embedded widely.

That is not a guarantee that model providers win. Hays pointed to companies that seem unremarkable or tenth in a category, only to reappear months later growing faster than software businesses from a few years ago. Coogan’s explanation was deployment. Consumer companies may have better distribution or subscription mechanics. Enterprise companies may win by changing workflows inside an organization rather than merely offering an API. A legal-AI company that can make a product work within a law firm has done more than hand the firm a model endpoint.

Open weights make capability harder to contain—and returns harder to protect

John Coogan described open-source AI as culturally aligned with the internet, crypto, downloadable media, and the appeal of retaining something useful outside a corporate service. The attraction, in his formulation, is a machine with broad knowledge and reasoning ability that can remain available even if a provider changes its policies or disappears.

There really is no debating that open source is just cool.

John Coogan

That appeal sits alongside a concrete security concern. Coogan said capable open models could make phishing and fraud more personalized and convincing. Rather than producing generic spam, an attacker could research a person’s insurance coverage, imitate an insurance agent, and tailor the message with relevant details. Reasoning, research, and adaptation can make attacks harder to recognize as automated.

He did not present defenders as permanently ahead. Drawing on prior conversations with Palo Alto Networks’ Nikesh Arora and CrowdStrike’s George Kurtz, Coogan said security firms had been anticipating more capable systems and that, in his account, a period of frontier-model exclusivity had given defenders time to patch systems. But the advantage is temporary: the contest becomes one between increasingly customized attacks and increasingly capable defensive layers.

Coogan expected discussion over how the United States should respond to Kimi’s planned release, and listed possibilities ranging from a direct or soft ban to public messaging encouraging people not to use a foreign product. He identified no policy decision. His point was instead that a broadly downloadable model creates a different problem from a hosted service: conventional restrictions become more difficult once capable weights are distributed.

Dean Ball’s comments brought the economics into focus. He called Kimi “a very good model,” said it seemed roughly on par with the best public models from the first quarter of 2026 in agentic coding, and noted that it appeared token-hungry. That qualification matters because a model that looks efficient in one dimension may still be expensive to operate at scale.

Ball was also surprised that the Chinese state would permit open sourcing a model of this capability. His explanation assigned roughly 75% of the decision to what he called “strategic blindness” or insufficient AGI-focused concern, and 25% to China’s limited capacity for serving customers through closed inference systems under U.S. export controls. In that account, open weights are partly ideological and partly a practical response to the difficulty of selling Chinese closed-source AI services abroad.

Coogan supplied the enterprise logic behind that difficulty. American companies already debate whether they should send code and internal data to domestic frontier-model providers. Convincing them to upload an entire codebase to a Chinese closed-source API would be still harder. Open release can therefore distribute a model without requiring users to place proprietary data inside a foreign hosted service.

The deeper dispute is whether open weights accelerate or slow the frontier race. Ball’s position was that open-weight models are “inherently decelerationist.” They may create an effectively ungovernable AI environment that some accelerationists welcome, but they can also reduce the expected return on a major training run. If a lab’s model can be copied, distilled, or matched within months, investors may be less willing to finance progressively larger runs.

Coogan accepted the tension. A frontier lab must fund researchers, experiments, and compute even if a single core training run has not yet reached $50 billion. A period of monopoly or oligopoly makes it easier to earn back that investment. Open models can expand the ecosystem of users and deployers while weakening the protected window that helps finance the next frontier system.

A director can be a franchise, but not an owned one

Jordi Hays cited a $264 million worldwide opening for The Odyssey, calling it the year’s largest non-animated opening at that point and, according to Lucas Shaw’s post, the biggest opening of Christopher Nolan’s career. For Coogan, the more strategic implication was that Nolan’s name now functions as an asset audiences will follow across subject matter.

$264M
Worldwide opening cited for The Odyssey

Coogan said the film carried a $375 million total budget: roughly $250 million in production costs and $125 million in global marketing. In his estimate, a run of slightly more than $1 billion worldwide could produce around $250 million in pretax lifetime profit for Universal and Comcast. That would be a major outcome for a film and its studio, but only about 0.2% of Comcast’s market value.

The more consequential figure came from the Wall Street Journal article shown on screen: 53% of Odyssey attendees said the director was their number-one reason for seeing it. Coogan contrasted that with attendance driven by actors, source material, genre, or premium formats. The audience, in this view, was not simply buying an adaptation of Homer; it was choosing the next Christopher Nolan film.

The film’s $124 million domestic debut was third-largest of the year, behind Toy Story 5 and Super Mario Galaxy. Coogan nevertheless expected repeat viewing and IMAX ticket premiums to give it a longer theatrical run. He described it as a breakout for an R-rated historical film, a category that usually does not open like major family animation or established blockbuster IP. It also outperformed newer franchise entries including The Mandalorian and Grogu, Supergirl, and Scream 7.

Universal’s bet was unusually large: an R-rated adaptation of an ancient Greek poem, shot across Europe with mostly practical effects and a 35-foot Trojan horse. Hays, who had not yet seen it, said the adaptation omitted some “hijinks and gaffs” from the poem that he would have liked included, while regarding the film as broadly solid.

The distinction Coogan drew is between an owned franchise and earned audience trust. Disney can own Marvel or Star Wars without relying on one filmmaker’s continued participation. A studio can benefit from Nolan’s following, but it cannot own his reputation in the same way; the commercial asset travels with the director. Coogan wondered how firmly Nolan was tied to Universal precisely because the relationship looks valuable but carries that key-person risk.

He tied the result to a wider shift in studio strategy. As Marvel, Fast & Furious, and Transformers have stumbled, the Wall Street Journal’s account suggested that audiences—particularly younger ones—are responding more to filmmakers with recognizable voices. Coogan cited director-driven successes and deals involving Curry Barker, Ryan Coogler, Zach Cregger, and Greta Gerwig as signs that studios are again competing for people whose audiences may follow them across genres rather than only for libraries of familiar IP.

Netflix is putting generative AI into the production workflow

Jordi Hays highlighted Netflix’s disclosure that generative-AI workflows had been used in roughly 300 productions during 2026, with the largest concentration in post-production. Netflix’s framing was operational: AI is becoming part of the production toolkit across concept development, previsualization, post-production, and delivery.

The company said creative partners were using the tools to produce higher-quality output more quickly and at lower cost. In some cases, Netflix said, productions would have left out key shots or sequences without generative AI. It also maintained that the technology was not being introduced to replace writers, directors, actors, or other creative professionals.

The American Experiment, a docuseries about the American Revolution, was the concrete proof point Hays offered. According to his account of comments by co-CEO Ted Sarandos, the series included 17 minutes of AI-enhanced footage that expanded the scale of the project beyond what would have been financially feasible through traditional production methods.

Netflix also described AI uses beyond production itself. It said it is applying large language models to title discovery and understanding member preferences, while developing voice search and AI-powered natural-language search. Its position is that generative AI is becoming embedded in the systems used to make, package, surface, and deliver entertainment—not a separate category of content and not, in its account, a substitute for creative labor.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free