Orply.

Post-Training Lifts DeepSeek Flash Above Its Larger Pro Model

Károly Zsolnai-FehérTwo Minute PapersMonday, August 3, 20264 min read

Two Minute Papers’ Karoly Zsolnai-Fehér argues that DeepSeek’s V4-Flash-0731 shows how much capability can be added through post-training rather than a larger base model. DeepSeek reports that the 384B-parameter model retains the prior Flash architecture and size but substantially improves on agentic and software benchmarks, including a 54.4 DeepSWE score versus 7.3 for Flash-Preview. Zsolnai-Fehér’s case is that improved planning, verification and error recovery—not new underlying capacity—account for the shift.

A re-post-training run appears to have remade DeepSeek’s Flash model

Károly Zsolnai-Fehér calls DeepSeek-V4-Flash-0731 a striking example of what post-training can do. DeepSeek says the model retains the architecture and size of DeepSeek-V4-Flash-Preview and was “only re-post-trained.” Yet its benchmark comparison reports large gains across agentic and software-oriented evaluations.

The reported improvements are not marginal. On DeepSWE, Flash-0731 scores 54.4, against 7.3 for the prior Flash preview—a 7.45-fold increase. CyberGym rises from 38.7 to 76.7; DSBench-Hard from 6.2 to 39.0; and AutomationBench (Public) from 10.8 to 25.1. Zsolnai-Fehér notes that benchmarks are not everything, nor do these tables cover every possible test. But he argues that a change at this scale cannot simply be dismissed.

BenchmarkFlash-0731Flash-PreviewPro-Preview
Terminal Bench 2.182.761.872.1
CyberGym76.738.752.7
DeepSWE54.47.312.8
Toolathlon-Verified70.349.755.9
DSBench-FullStack68.137.041.8
DSBench-Hard39.06.221.6
Selected results from DeepSeek’s comparison; Flash-0731 leads Pro-Preview on all nine listed tests.

The 9/9 result against Pro-Preview is central to Zsolnai-Fehér’s interpretation. He describes Pro as roughly five times larger, yet Flash-0731 leads it on every listed test. In the terms of the release, the implication is not that DeepSeek added more base-model capacity to Flash; it is that a model with the same stated size and architecture was made materially more capable through what happened after base training. Flash-0731 is also competitive with, though not uniformly ahead of, the listed GLM-5.2 and Opus-4.8 scores where those comparisons are available.

7.45×
reported DeepSWE gain for Flash-0731 over Flash-Preview

Post-training changes the model’s sequence of actions

Károly Zsolnai-Fehér describes post-training as giving the same underlying “brain” a new playbook: it teaches the model when to use particular abilities, how to plan, how to check its work, and how to recover after errors.

Károly Zsolnai-Fehér · Source

His builder analogy makes the distinction practical. Before post-training, the builder can see the relevant pieces but uses them poorly and creates a mess. Afterwards, the same builder learns a strategy: establish a base, test each component, correct mistakes early, and think ahead. The builder’s toolbox and raw knowledge have not changed; the order and discipline of its actions have.

That is why Zsolnai-Fehér treats the reported result as an unusually clear demonstration of post-training’s leverage. The claim is not simply that a model can be made better after base training, but that planning, verification, and recovery behavior can produce a large performance shift without a larger model or a different architecture. The Hugging Face page shown in the source similarly describes Flash-0731 as an open-weights instruction-tuned model with enhanced agentic capabilities, the same DeepSeek-V4-Flash structure, and an attached speculative-decoding module.

The practical question is how users can actually run it

Károly Zsolnai-Fehér treats the ability to download and retain the weights as at least as important as the benchmark scores. The Hugging Face page shown in the source lists DeepSeek-V4-Flash-0731 as an MIT-licensed open-weights model with 384 billion parameters. His practical point is that users can obtain the weights rather than rely solely on a hosted interface.

Károly Zsolnai-Fehér

That does not make Flash-0731 a routine local deployment. Zsolnai-Fehér says it requires a “beefy” machine to run locally, consistent with the 384B-parameter figure displayed on the model page. For users without that hardware, he identifies cloud compute and DeepSeek’s API as routes to use it, and characterizes the API as cheap compared with frontier companies.

The contrast he draws is with hosted products that impose session or weekly caps. Downloadable weights remove dependence on those particular access constraints, while the hardware requirement remains. The operational choices are therefore distinct: run the model on sufficiently capable local equipment, use rented compute, or access it through an API.

His forecast remains conditional. If this pace of progress continues, he says, a free open model could come close within less than a year to what he calls today’s “billion dollar fable level intelligence,” compressed enough to run on a beefy laptop. He says that possibility had sounded impossible to him only weeks earlier, but does not present it as a certainty.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free