Orply.

Kimi K3 Releases Open Weights at Frontier Coding Scale

Károly Zsolnai-FehérTwo Minute PapersWednesday, July 29, 20265 min read

Károly argues that Kimi K3 makes competitive frontier-level coding capability available under open weights, even if its 2.8-trillion-parameter scale puts local deployment beyond most users. The model’s reported benchmark results place it close to leading proprietary systems and first on the displayed SWE Marathon test, while Kimi’s technical report says it still trails the strongest proprietary models overall. His case is that releasing the weights and methods matters as much as the demos, enabling cheaper access, distillation and follow-on open-model work.

A frontier-scale model whose weights can be kept

Károly Zsolnai-Fehér calls Kimi K3 remarkable not simply for its coding demos, but for the terms of its release: the 2.8-trillion-parameter model is open weights. In his framing, people can download and retain the weights “for free, forever,” rather than depend on access to a proprietary frontier system that could be restricted or withdrawn.

The clips attributed to K3 show a broad range of generated interactive software. A macOS-style desktop, credited on screen to @mweinbach, includes Finder, widgets, FaceTime history, Mail, and an App Store receipt; Károly describes it as a “mostly working copy” of macOS. Other clips show a cowboy sketch becoming the basis for a playable fantasy world, an agent debugging an invisible collision mesh in a western game, and iterative adjustments to a black-hole render. An Animal Crossing-style game includes tasks, dialogue, fishing, and item collection.

Kimi’s technical report describes K3 as a Mixture-of-Experts model with 2.8 trillion total parameters but 104 billion activated parameters, native vision capability, and a one-million-token context window. Its abstract says K3 still trails the two proprietary systems it names as most powerful—Claude 3.5 and GPT-4.0 Sol—while consistently outperforming other open and proprietary models in its evaluation suite.

This is an open-weights model. Yes, we can download and own the weights for free, forever. And nobody can take this from us.

Károly Zsolnai-Fehér · Source

Near the leaders on coding, and first on SWE Marathon

The displayed coding dashboard places Kimi K3 close to leading systems across several programming and software-engineering evaluations, with all models shown at maximum thinking effort. That is what makes the open-weights release consequential in Károly’s account: it is not merely a freely available model with eye-catching demos, but one positioned among the strongest displayed systems on coding and long-horizon software work.

K3 is particularly close on Terminal Bench 2.1, where it scores 88.3 against GPT-4.0 Sol’s 88.8. It ties GPT-4.0 Sol for the highest displayed Program Bench score, at 77.8. The clearest lead comes on SWE Marathon: K3 is shown at 77.8, ahead of Fable 5 at 70.8, while GPT-4.0 Sol and Opus-4.8 are listed at 14.0 and 13.7.

BenchmarkKimi K3Highest displayed scoreLeading displayed model
DeepSWE67.572.0GPT-4.0 Sol
Terminal Bench 2.188.388.8GPT-4.0 Sol
FrontierSWE81.288.6Fable 5
Program Bench77.877.8Kimi K3 / GPT-4.0 Sol
SWE Marathon77.877.8Kimi K3
Scores displayed in the coding-benchmark dashboard, with maximum thinking effort.

The report’s own framing is more qualified than a blanket claim of superiority: it says K3 trails the proprietary models it identifies as most powerful, while outperforming the other open and proprietary models in its evaluation suite. The displayed results nevertheless support Károly’s larger point that competitive coding capability is no longer confined to systems whose weights remain unavailable.

The efficiency claim is about training progress, not inference cost

K3’s central technical claim is an approximately 2.5× improvement in scaling efficiency over Kimi K2. Károly Zsolnai-Fehér is explicit about the intended meaning: it is not a claim that K3 is 2.5 times cheaper or 2.5 times faster to operate. It means roughly 2.5 times more learning progress from the same amount of training computation.

2.5×
claimed improvement in overall scaling efficiency over Kimi K2

He explains that claim through two architectural ideas named in Kimi’s report: Kimi Delta Attention, or KDA, and attention residuals.

KDA is presented as an alternative to repeatedly rereading every earlier token in a long discussion. Károly’s analogy is an institute where every researcher must reread every document written by everyone else. KDA instead maintains a carefully updated notebook: the model reads and updates that compact state, while older information can gradually fade. In his account, this lets the system handle a very long discussion while continuing to contribute meaningfully.

Attention residuals address what happens as information passes through many layers. Károly compares it to a document revised in a sequence of departments. A later department normally receives only the latest version; with attention residuals, it also retains access to useful earlier drafts, like a version history. His shorthand is that KDA maintains and corrects memory, while attention residuals retrieve earlier information across model depth.

The report attributes the gain to a combined system: KDA and attention residuals, Stable LatenMoE—which it says activates 16 of 896 routed experts per token—refined training and data recipes, reinforcement learning across general, agentic, and coding domains, and infrastructure work for training and deployment at this scale.

Open weights matter even when the original model is too large to run at home

Károly Zsolnai-Fehér does not suggest that K3’s full release makes frontier capability immediately local for most users. At 2.8 trillion parameters, he says, it is “stupendously” large, and most people cannot afford to run it at home. The practical routes he identifies are trying it on the web, subject to availability, or using an API that he characterizes as far cheaper than current frontier models. Even people who never use K3 directly, he argues, could benefit if that competition pushes token prices down.

His longer-term accessibility argument rests on what happens after the original release. Huge models are often distilled into smaller systems intended to retain similar capabilities, he says. Károly expects a distilled version of K3 to become much easier to run, while the disclosed techniques and weights give other open-model developers material to build on.

With every paper like this, we make all the other open models work better. We are building this together.

Károly Zsolnai-Fehér

That is the basis for his broader claim about open science: not that K3 itself is trivial to deploy today, but that a competitive model, released with its weights and technical methods, can improve the capabilities available to later open systems and their users.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free