Orply.

Open-Weight Models Are Eroding Frontier Labs’ Pricing Power

Alex KantrowitzRanjan RoyAlex KantrowitzMonday, July 20, 202614 min read

Alex Kantrowitz and Ranjan Roy argue that Moonshot’s Kimi K3, by approaching frontier-model performance at a lower price and with planned open weights, weakens the case for paying a large premium to OpenAI or Anthropic. As capable models proliferate, they say, advantage will depend less on benchmark leadership than on products, infrastructure, trusted data practices and partnerships—areas where Google’s execution problems and OpenAI’s conflicts with allies expose different vulnerabilities.

The model layer is becoming a price war

Alex Kantrowitz sees Moonshot’s Kimi K3 as a challenge to the business logic of frontier AI labs—not because it is unequivocally the world’s best model, but because it has arrived close enough to the leading closed systems to narrow the reason for paying a substantial premium.

Kimi K3 is a 2.8 trillion-parameter model that Moonshot says rivals leading offerings from OpenAI and Anthropic. Across the benchmark dashboards displayed during the discussion, it traded wins with leading models in coding, general-agent, and visual-agent tasks. It surpassed Fable 5 on Program Bench and SWE Marathon in one coding comparison, and led on SpreadsheetBench 2, Automation Bench, and BrowseComp in the general-agent display. But it was not a clean sweep: Fable 5 led on several measures, including GPQA-AA v2 Elo, while GPT-5.6 Sol led on Terminal Bench 2.1.

BenchmarkLeading model shownKimi K3 result
Program BenchKimi K377.8, ahead of GPT-5.6 Sol at 77.6 and Fable 5 at 77.0
SWE MarathonKimi K342.0, ahead of Fable 5 at 40.0 and GPT-5.6 Sol at 39.0
GPQA-AA v2 EloFable 5Kimi K3 scored 1660, behind Fable 5 at 1780 and GPT-5.6 Sol at 1745
Terminal Bench 2.1GPT-5.6 SolKimi K3 scored 86.3, behind GPT-5.6 Sol at 88.0
BrowseCompKimi K391.2, ahead of Fable 5 at 90.4
Selected results from the benchmark comparisons displayed for Kimi K3 and competing models

That distinction matters. Kimi is not markedly better than the best closed models, Kantrowitz says, nor is it radically cheaper than every alternative. Its significance is its proximity to the frontier. A Chinese model expected to be available with open weights, and roughly in the range of GPT-5.5, GPT-5.6, and Opus 4.8, changes the calculation for companies deciding whether they need the most expensive proprietary intelligence for a given task.

$3 / $15
Kimi K3’s stated price per million input / output tokens

Ranjan Roy argues that the final benchmark delta matters less than crossing a threshold of practical quality. If a model is in the “general quality range” of a leading system, it is viable for an expanding share of work. K3’s pricing, Roy notes, is not a bargain-basement offer: at $3 per million input tokens and $15 per million output tokens, he describes it as 40% cheaper than GPT-5.6 and 70% cheaper than Fable. Its proposition is competitive quality, lower cost, and greater control—not merely a cheaper approximation of the frontier.

That is different from the earlier DeepSeek shock in Roy’s telling. DeepSeek established that useful models could be much cheaper. Kimi K3 suggests that a model can be close to the top tier and still weaken the premium frontier labs expected to command. Meta’s MuseSpark 1.1 and SpaceX’s Grok 4.5 add to the pressure: each offers competitive performance in some areas at lower prices.

The likely buyer behavior, both speakers suggest, is greater model interoperability. Enterprises can route lower-intensity or specialized work to cheaper systems while reserving the hardest tasks for a frontier model when necessary. The central shift is not that K3 replaces every leading model. It is that an open-weight alternative close enough to the frontier makes a single-provider strategy harder to justify.

In Kantrowitz’s framing, OpenAI and Anthropic built their API businesses around a straightforward proposition: build intelligence materially better than the alternatives, then sell access at a premium. Comparable proprietary systems from vertically integrated companies, combined with capable open-weight systems, make that premium harder to defend.

Open weights could move value away from the model provider

Moonshot’s planned open-weight release is nearly as consequential as K3’s benchmark results, Alex Kantrowitz argues. A company with sufficient infrastructure and engineering talent could download the weights and build on the model itself rather than buy a fully managed implementation from a frontier lab.

That does not mean an individual can run K3 on a laptop. Its scale makes self-hosting a proposition for governments, very large companies, and cloud providers. But those are precisely the institutions that spend heavily on AI. A large bank, a government, or a company such as Apple can weigh several options: hire OpenAI or Anthropic and use their models and forward-deployed engineers; buy a hosted model through a cloud provider; or work with an open-weight model that it can customize and operate with more direct control over the stack.

Kantrowitz raises a further possibility: once a model is available through cloud platforms, the listed Moonshot API price may no longer be the only relevant price. A cloud provider might operate the model more efficiently, bundle it with infrastructure, or use the model to draw customers toward other cloud spending. Those are possible commercial outcomes, not settled ones. But they point to the same risk for the labs: the model becomes an input to a broader infrastructure and software business instead of a high-margin endpoint controlled exclusively by its creator.

Investor Gavin Baker, quoted by Kantrowitz, frames K3 as potentially negative for OpenAI and Anthropic while benefiting much of the rest of the AI economy. His concern is not that inference gets cheaper because it requires less hardware. Baker’s point is the opposite: an open-source model of similar size and architecture requires the same amount of compute to run as a closed one. K3’s per-token price, he says, is roughly comparable to GPT-4 Turbo’s, which may suggest lower computational efficiency rather than a fundamentally lower infrastructure requirement.

The economic question is where margins accrue. Baker argues that if two or three frontier labs held 90% inference margins, they could become dominant buyers of power, data-center capacity, semiconductors, and hyperscaler services. They could also expand into the application layer. More competition at the model level would reduce that concentration and leave more value for infrastructure providers, neoclouds, and software companies.

Competition from Google, Meta, and SpaceX matters alongside Kimi K3 for the same reason. Those companies can benefit from AI even if a model itself is not their primary profit center. A vertically integrated company can accept lower model-layer margins if it earns elsewhere—in cloud usage, advertising, devices, subscriptions, or other services. OpenAI and Anthropic have less room for that tradeoff if inference remains central to their businesses.

Ranjan Roy sees this as especially acute for Anthropic, which shifted aggressively toward the enterprise. Its financial narrative, he says, depended on investing heavily in training, producing a leading model, and then earning high inference margins. Open-weight models and lower-cost rivals challenge the premise that model leadership will be sufficiently scarce or durable.

Kantrowitz does not think that necessarily means less infrastructure investment overall. He offers a counter-scenario: if OpenAI or Anthropic lost their central positions while highly capable models remained available, Amazon, Google, and Microsoft could still operate the data centers and sell the compute. Customers would use broadly accessible models, while cloud platforms would earn from delivering the infrastructure rather than from one lab’s monopoly-like hold on intelligence.

Roy notes that the current buildout is complicated by circular financing: hyperscalers are funders, customers, and infrastructure suppliers to leading labs. But a more competitive model market need not mean empty data centers. It could mean the data-center business is owned more directly by the companies already operating it.

A product moat depends on trust as much as capability

The possibility of commoditized intelligence forces a harder question for OpenAI and Anthropic: what remains distinctive if they cannot preserve a large capability gap?

Alex Kantrowitz cites two reasons Baker does not regard Kimi K3 as unambiguously negative for the labs. The first is that ChatGPT and Claude—their interfaces, workflows, and “harnesses”—may matter more than the underlying models. The second is the hypothesis that the labs possess substantially more advanced internal checkpoints and are using them for recursive self-improvement. If one reached that point even months before others, Baker suggests, it might establish a lasting lead.

Kantrowitz is skeptical of the second proposition. The speed with which open-weight and competing models have approached the frontier makes him doubt that a sizable intelligence advantage can be hoarded for long. He points to researcher Ryan Greenblatt’s expectation that, if Kimi and similar companies retain open-weight policies, an open-weight system straightforwardly at “Mythos-level” cyber capability could arrive within about five months.

I think we should assume, and we’ve talked about this on the show, that intelligence is going to commoditize.

Alex Kantrowitz

Ranjan Roy says the standard AGI thesis still has adherents: if a model becomes sufficiently generally intelligent, it could remove the need to choose the right model for the right task. But he sees less confidence in that story than before. The practical trend is toward orchestration, specialized deployment, and cost-performance optimization—not toward a single system so dominant that all such choices disappear.

That does not erase the possibility of a strong business. It changes the nature of the bet. If intelligence is abundant, Kantrowitz says, OpenAI and Anthropic must win because they build better products than everyone else who can access similar models. That is a much broader competitive field than a race among two or three labs for the best model.

Roy argues that OpenAI once demonstrated the importance of product clearly. The original ChatGPT experience—the presentation of a model “thinking,” streaming text, and a conversational interface—felt meaningfully different from a raw API response. But he believes both OpenAI and Anthropic have recently placed less emphasis on product interface and more on raw model capability. He points to backlash when ChatGPT’s Mac app appeared to remove parts of the chat experience, though Kantrowitz says OpenAI began restoring some of those features.

Kantrowitz’s view is more qualified. Better intelligence has enabled better products, he says, and labs retain an advantage from being close to their own model roadmaps. They know what capabilities are arriving and can design products around them rather than waiting to infer changes through an external API.

Anthropic’s Claude Code illustrates that possibility. Kantrowitz says Anthropic’s revenue has grown roughly tenfold since he interviewed CEO Dario Amodei the prior year. At that time, Amodei confirmed that more than half of revenue came through the API. More recently, Claude Code head Boris Cherny would not confirm that figure and said the Claude products had contributed meaningfully to revenue. Anthropic, in Kantrowitz’s view, can build a substantial business by turning its own capabilities into products rather than merely selling underlying model access.

Roy agrees that Claude Code is a strong product even though it operates through a command-line interface. The product was not merely the model; it was the harness and the experience around it.

But the product strategy creates a conflict with customers. If a lab uses enterprise data to improve its products, then expands into adjacent application categories, customers may worry that they are helping to train a future competitor. Roy raises Figma’s Claude integration, the overlap between Cursor and Claude Code, and Anthropic’s ability to move upmarket. A company that plugs deeply into a model provider may eventually conclude it is feeding proprietary context to a firm that can replicate part of its business.

Kantrowitz’s answer is that avoiding that risk becomes easier when companies have credible alternatives, including models such as Kimi K3 that they may be able to customize themselves.

Microsoft CEO Satya Nadella supplied a broader statement of the concern. In what he calls the “Reverse Information Paradox,” the buyer pays twice for intelligence: first with money, then with the proprietary knowledge needed to make the system useful. “The better you want the model to perform,” Nadella wrote, “the more of that knowledge you have to feed it.”

Nadella’s argument is that the imbalance grows over time: the seller learns more about the customer through use of the AI system, while the customer has limited visibility into what the seller is learning in return. Kantrowitz reads the message as a clear appeal from Microsoft to enterprises: build with Microsoft because it will protect their information rather than turn their usage into a competing-product advantage.

Roy says the data issue changes the direction of model development. At Writer, where he works, the company has foundation models trained on synthetic data, he says. He contrasts that with OpenAI and Anthropic, which he says have been open that non-enterprise data is, at least in theory, used for training. That can create a flywheel: the provider gets more data, improves its model, attracts more use, and gathers more data.

For a period, Roy says, companies appeared less concerned about whether providers trained on their data. The issue has returned quickly as providers increasingly launch applications and vertical products. The compressed pace of the AI cycle has made the shift unusually abrupt: dynamics that might unfold over years in a conventional software market are changing within weeks.

The strategic response is being tested at Google and OpenAI

The narrowing model moat does not affect every company equally. Google can afford to treat model quality as one part of a larger business spanning Search, Cloud, YouTube, Maps, advertising, and devices. OpenAI, by contrast, is trying to preserve a central position while its relationships with major partners become increasingly strained.

Google’s immediate problem is execution. A Bloomberg report discussed by Alex Kantrowitz said Gemini 3.5 Pro was months behind schedule because it had failed to meet internal goals, particularly in coding. The delay had frustrated engineers, researchers, and managers concerned that OpenAI and Anthropic were moving ahead. Google’s broad product portfolio compounds the problem: models need to be prepared for integration with Search, Maps, YouTube, and other products, creating multiple layers of stakeholders and release friction.

Ranjan Roy emphasizes that Google was not always behind. Gemini 3, launched in November 2025, briefly made the company appear competitive again. But over the following seven or eight months, Roy says, Google’s consumer experience did not seem to improve at the pace of the field. The models may not have become worse, but competitors became noticeably better.

Google still has a meaningful strength in multimodality, Roy says, particularly image and video capabilities. He also finds himself using AI Overviews more often within Google Search. That may become a major business even if Google is no longer winning the frontier-model contest. But Google’s model momentum, especially in coding, looks weaker.

Kantrowitz offers a theory rather than a reported explanation for the slowdown. After DeepSeek, he says, Google may have leaned too heavily into the idea that smaller, lower-cost Flash models would drive greater usage and increase demand for cloud and AI products. The market then shifted as larger models became capable of more autonomous coding work. If Google had committed organizationally to smaller, purpose-built models, he suggests, it may have taken time to reorient around the big-model race.

Roy sees clearer evidence of an old Google failure mode in Bloomberg’s account of overlapping coding efforts. Google Cloud, Google DeepMind, and the Android team were all building AI coding tools. That resembles Google’s history of overlapping messaging products and fragmented ownership: capable teams pursuing related goals without a unified product strategy.

The report also described engineers with a more purist view that important code should remain human-written to satisfy Google standards. Kantrowitz allows that core software deserves caution, but says a company gets into trouble if it treats hand-written code as a general principle while the rest of the industry is automating development.

The compute question sharpens the management critique. Google licenses substantial capacity through its cloud business while also presenting AI as one of the most important technologies it has ever faced. Kantrowitz cannot reconcile a shortage of internal compute with that priority. Roy counters that Google Cloud is itself a massive, fast-growing business operating with separate incentives; reallocating resources is not a simple executive instruction. Kantrowitz’s response is that resolving exactly that sort of conflict is management’s job.

Both expect Google to recover eventually. Its talent and resources remain formidable. But the current delay is evidence that scale alone does not create strategic coherence.

OpenAI faces a different constraint: retaining partners when alternatives are becoming more credible. Apple’s lawsuit against OpenAI brings that risk into focus. Apple alleges that former employees, using Apple-issued laptops, discussed plans to exfiltrate Apple data, bring Apple components to interviews, and exploit access to Apple’s roadmaps. Kantrowitz characterizes the alleged conduct as an unusually clumsy case of corporate espionage.

Roy agrees that the conduct Apple describes appears extraordinary. He compares it with Uber’s dispute with Waymo, but says even that episode had more “cloak and dagger” than what Apple alleges here.

Apple has sent legal preservation letters to dozens of OpenAI employees, according to Kantrowitz. Apple has also said that what it identified may be only the “tip of the iceberg.” OpenAI’s stated position is that it takes the allegations seriously but is not aware of evidence that the complaint has merit.

Kantrowitz expects Apple to pursue discovery rather than settle quickly. His larger point is that OpenAI has turned important former allies and partners into opponents at moments when their cooperation could matter. Elon Musk, an OpenAI cofounder, is now an opponent. Dario Amodei, an early OpenAI employee, leads Anthropic. Microsoft, OpenAI’s central financial backer and infrastructure partner for years, is publicly warning customers about the data risks of buying intelligence directly from model providers. Apple, which partnered with OpenAI to integrate ChatGPT into Apple Intelligence, is now suing it.

Roy notes the scale of Apple-to-OpenAI hiring discussed by the two: OpenAI has roughly 400 employees who came from Apple, while about 40 employees had reportedly received legal communications related to the matter. If a meaningful share of that group were implicated in improper conduct, he says, the consequences would be substantial.

The issue is not simply reputational. Large technology companies need cloud providers, hardware partners, distribution channels, data partners, and customers willing to trust them with sensitive systems. Apple and Google can compete while still cooperating when both benefit; Google’s Gemini role in Siri is an example of that pragmatic interdependence.

As model differentiation narrows, those relationships become harder to treat as secondary. A frontier lab may continue to build strong products and retain considerable market power. But Kantrowitz argues that there may come a time when OpenAI needs Microsoft, Apple, or another former ally—and finds that the relationship has already been spent.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free