Orply.

Open Weights Let DeepSeek V4 Pro Escape Single-Provider Pricing

Károly Zsolnai-FehérTwo Minute PapersWednesday, August 19, 20265 min read

Two Minute Papers host Károly Zsolnai-Fehér argues that DeepSeek V4 Pro 0813 makes open model weights a material alternative to dependence on a single AI provider: although DeepSeek has raised API prices, developers can retain the MIT-licensed model and run it through competing hosts or their own infrastructure. He says the release delivers its largest benchmark gains in software engineering and harder data-science tasks, while still trailing the cited closed-model baseline in full-stack and tool-use tests. Zsolnai-Fehér attributes the improvement to post-training specialist models, multi-teacher distillation and DSpark inference techniques rather than a new base architecture.

Open weights turn a price increase into a market choice

Károly Zsolnai-Fehér argues that DeepSeek V4 Pro 0813 changes the practical meaning of a hosted-model price increase. DeepSeek has raised its own API prices by roughly 2.5× to 5×, he says, but the model’s MIT-licensed weights remain available. That means a developer is not locked into DeepSeek’s own endpoint: the same model can be run elsewhere, at a price determined by the host—or on self-managed infrastructure for those with sufficient hardware.

The release is not a small model for casual local deployment. It is distributed across 90 safetensor shards, with the shown individual files ranging from roughly 13.6 GB to 14.2 GB each. Zsolnai-Fehér says he would like to host it himself but does not have the necessary hardware. Open weights therefore do not mean that every user can practically host the model at home. They mean that access is contestable: users can choose another host, deploy on infrastructure they control, or retain the checkpoint rather than depend on a single API provider.

He points to cloud providers such as Lambda and to a competitive market of hosted options. A provider comparison in the source lists a 366% spread across 17 providers for the same model. The lowest shown input price is $0.43 per million tokens through merge-gateway; other listed prices include $1.10 through nano-gpt and roughly $1.30–$1.32 through DeepInfra, Hugging Face, DigitalOcean, and Together AI.

366%
Shown price spread across 17 providers for the same model

The point is not that inference is free in every setting, but that an open-weight release separates the model from any one seller’s pricing and service decisions. Zsolnai-Fehér sees that competition as pressure on frontier labs to improve their own offerings quickly.

He also treats ownership of the weights as a product property in its own right. Users can keep running the checkpoint they selected rather than depend on a provider to maintain a particular version. An on-screen illustration contrasts that with a proprietary-model interface that rejects a prompt after a particular word is entered; it serves as an example of the control he says open weights provide, not evidence that every closed model behaves that way.

Pro gains are substantial, but uneven across the tasks shown

The release is positioned as a substantial improvement over DeepSeek V4 Flash and over V4 Preview, despite retaining the same underlying architecture as Preview. Zsolnai-Fehér calls the advance striking precisely because the structure is unchanged while capability has improved markedly in less than four months.

On the shown benchmarks, V4 Pro scores 62.7 on DeepSWE, a software-engineering measure, versus Flash’s 54.4—a gap of 8.3 points. On DSBench-Hard, a data-science measure, Pro scores 67.2 versus 59.6 for Flash, a 7.6-point lead. The same comparison places Fable 5 at 70.0 on DeepSWE and 68.3 on DSBench-Hard. On those two measures, Pro is much closer to Fable 5 than Flash is.

BenchmarkV4 Pro 0813V4 Flash 0731Fable 5
DeepSWE (software engineering)62.754.470.0
DSBench-Hard (data science)67.259.668.3
DSBench-FullStack (data science)71.168.777.2
Toolathlon-Verified (agentic tool use)74.170.377.9
Shown comparisons between DeepSeek V4 Pro, V4 Flash, and Fable 5

But the results qualify any blanket claim of near-parity. On DSBench-FullStack, Pro leads Flash by 2.4 points, 71.1 to 68.7, while Fable 5 remains at 77.2. On Toolathlon-Verified, the Pro–Flash gap is 3.8 points, 74.1 to 70.3, and Fable 5 leads both at 77.9. The largest visible gains are concentrated in the software-engineering and harder data-science comparisons; the FullStack and tool-use results still leave a larger gap to the shown Fable baseline.

For a buyer, that makes the release less a general declaration of superiority than a task-dependent option. The open-weight and host-choice case is strongest where the remaining benchmark deficit is acceptable for the workload, rather than where a closed baseline’s lead is decisive. The benchmark story is strongest where Pro’s gains over Flash are widest; elsewhere, the case for switching depends more heavily on the value of open weights, host choice, and cost.

The demonstrations are intended to make the difference concrete. In a Rubik’s Cube construction task, Flash produces a partial object with missing geometry and black areas; Pro produces a complete cube, scrambles it, and then solves it. Zsolnai-Fehér describes this as a much stronger understanding of the object’s three-dimensional structure.

The claimed improvement comes after pre-training

Károly Zsolnai-Fehér says that “a lot of the magic happens after pre-training.” His account begins with separately post-trained specialist checkpoints: one for mathematics, one for coding, and one for agentic tool use. He stresses that these specialists should not be confused with the experts in a mixture-of-experts model. Mixture-of-experts components are smaller pieces within one neural network; the math, coding, and agentic specialists he describes are separately trained model checkpoints.

Zsolnai-Fehér describes DeepSeek’s process as using more than 10 specialist teachers to train a final student model through distillation. In his simplified account, a student produces what it would do, a teacher supplies what it would have done, and the student is adjusted to behave more like its teacher. Repeating that process with many teachers, he says, allows a final model to absorb specialist capabilities without changing the base architecture.

The second component is DSpark. Its paper, DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation, describes speculative decoding as accelerating large-language-model inference by decoupling draft generation from target verification. Zsolnai-Fehér characterizes DSpark as drafting several tokens ahead rather than predicting only one token at a time, and says it improves on earlier techniques.

He says DeepSeek reports up to a 78% increase in generation speed for V4 Pro—a practical inference gain available in current use, not merely a research result.

Up to 78%
Reported faster generation for DeepSeek V4 Pro with DSpark

Zsolnai-Fehér emphasizes how quickly the method moved from paper to deployment: he says the DSpark research appeared about six weeks earlier and is now being used broadly. The source associates DSpark with DeepSeek and with models from Nemotron, Qwen3, Mistral, and Gemma.

He also flags DeepSeek Harness, presented as an in-browser developer tool in developer preview, with source code included and a plugin-based design. The source offers no detailed account of its mechanism. Zsolnai-Fehér calls it a novel, powerful “no-agent harness” and says he plans to examine it separately.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free