Orply.

Qwen 3.8 Max Could Pressure Closed AI Providers on Price and Openness

Károly Zsolnai-FehérTwo Minute PapersWednesday, August 5, 20265 min read

Qwen 3.8 Max could put pressure on OpenAI and Anthropic not by leading every benchmark, but by combining frontier-style multimodal agents with lower claimed API prices and planned open weights, argues Two Minute Papers’ Károly Zsolnai-Fehér. He presents the model as capable of carrying out and revising work over days, while acknowledging that its displayed software-engineering scores trail leading closed rivals. The more consequential prospect, he says, is whether smaller Qwen releases can bring comparable capability to local users.

A capable open model at a lower displayed price could pressure frontier labs

Károly Zsolnai-Fehér presents Qwen 3.8 Max as a competitive challenge to OpenAI and Anthropic on three linked terms: frontier-style multimodal and agent demonstrations, materially lower claimed API costs, and an upcoming release of model weights. It is not shown as the leader in every category. The argument is that it may be capable enough, cheap enough, and open enough to alter what users expect from the closed-model providers.

Zsolnai-Fehér says Qwen 3.8 Max may be “five to 10 times cheaper depending” on the task. That is a broader comparative claim, distinct from the task-specific costs in one displayed Flappy Bird-style agent comparison: Qwen3.8-Max is listed at $0.0248, against $0.150 for GPT-5.6 Sol and $0.253 for Opus 5. He argues that the combination of capability and low API pricing could force other providers to reduce prices.

The broader open-model market already contains very cheap options. A VALS.AI screen labels DeepSeek V4 Flash “the cheapest model above 60%” on its index: 64.0% at $0.06 per test, compared with $2.08 for GLM 5.2 at 65.0% and $2.34 for Kimi K3 at 74.7%. Zsolnai-Fehér calls DeepSeek Flash “super quick” and “super cheap,” then positions Qwen as the higher-end open offering.

$0.06
Cost per test shown for DeepSeek V4 Flash on the VALS index

Qwen’s materials show application-generation interfaces across several domains: an architectural workspace with a floor plan and rotatable building model, a rehabilitation guide with an interactive leg model, and a Blender scene in which the system identifies that a television faces the wrong direction and says it will rotate it. The demonstrations show application-generation interfaces with visual inspection and correction.

The product pitch is work that continues while the user is away

Qwen presents 3.8 Max as multimodal—“it has eyes and ears,” in Károly Zsolnai-Fehér’s description—with a one-million-token context window and a focus on agentic workflows: systems intended to pursue multi-step tasks using tools, intermediate decisions, and checks.

Its demonstrations lean heavily on persistent task execution. One interface takes an 8-hour, 14-minute series comprising 10 episodes and bonus material, claims to index every frame, divides the footage into scenes of varying length, and constructs a visual scene graph. Another describes a 100-hour video memory, then answers a request for every moment when “Tim” is eating by narrowing 2,191 macro-scene nodes into clips for a highlight reel.

The central appeal is endurance. Qwen’s videos show the model designing chips for 12 hours while an engineer fishes, verifying protein sources while a biology professor plays tennis, reconciling finances while a CFO climbs, and generating contracts while a lawyer relaxes. Zsolnai-Fehér characterizes the pitch as productivity and tranquility, rather than a system marketed through intrusion into other people’s systems.

It sat there thinking for 16 days. Starting from an empty folder, writing, testing, and repairing its own code.
Károly Zsolnai-Fehér

Zsolnai-Fehér says that run could reproduce research papers and “even improve them meaningfully.” The accompanying Qwen screen describes an “oh-my-cli” project built from scratch and continuously iterated by the model. It reports self-testing and self-fixing, 127 pull requests closed by day three, and a dynamic workflow-engine release by day 10.

He contrasts that claim with systems that “tap out within minutes to hours.” The model is being presented as able to sustain implementation, testing, debugging, and revision over days while the user does something else.

Open weights matter most if the smaller models stay capable

The full Qwen 3.8 Max model is described as too large for most people to run at home, even if its weights are released. Károly Zsolnai-Fehér says Qwen has committed to releasing them soon, but treats the anticipated smaller models as the development with wider practical relevance.

He points to Qwen 3.6’s 27-billion- and 35-billion-parameter models as the precedent. He calls them “the Toyota Corolla of the AI world”: dependable and accessible rather than exotic or hardware-bound. A montage of Reddit and forum posts calls Qwen 3.6 27B “a BEAST” and includes one user’s claim that it was the first local model to hold up against Claude Code.

For users with modest resources, Zsolnai-Fehér says the smaller releases are “the real news.” His prospect is a capable local daily driver that people can own and run for free.

The flagship is shown producing an educational fossil-gallery site, building a Three.js Hogwarts castle against an acceptance checklist, and generating a financial-analysis interface with sub-agent tracking. Those flagship demonstrations do not establish how future smaller releases will perform; the expectation of useful lower-scale models rests instead on the earlier Qwen 3.6 releases.

The benchmark record is uneven, and that is part of the case

Károly Zsolnai-Fehér cautions that benchmarks vary: “some of them are gamed, some of them less so.” The displayed results show Qwen 3.8 Max leading on some visual measures and trailing on the software-engineering measures presented.

On Qwen’s charts, the model leads the best rival by 7.8 points on ERQA embodied reasoning, 77.8 to 70.0, and by 3.8 points on PerceptionBench visual perception, 63.5 to 59.7. The same materials show it behind on software engineering: 73.5 on FrontierSWE, 15.3 points below the displayed leader at 88.8, and 67.7 on SWE-Pro, 12.3 points below the leader at 80.0.

The benchmark Zsolnai-Fehér emphasizes is Humanity’s Last Exam, which he calls a “devilishly difficult academic benchmark.” A CAIS dashboard shown in the source displays 2.7% HLE accuracy for an earlier model-era result. Qwen’s own chart reports 56.2 for Qwen 3.8 Max with tools—above 50%, but below the displayed best-rival score of 64.5.

56.2
Qwen 3.8 Max score shown on Humanity’s Last Exam with tools

For Zsolnai-Fehér, the movement from roughly 2% for the strongest closed systems when the benchmark began to an open model above 50% a little more than a year later is evidence of rapid open-model progress. He regards HLE as particularly indicative of real-world performance for scholarly users, while still distinguishing it from benchmarks he considers easier to game.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free