Orply.

Continual Learning Could Turn AI Deployment Into Training

Dwarkesh PatelDwarkesh PatelFriday, August 7, 20265 min read

Dwarkesh Patel argues that AI systems capable of performing whole jobs will need to learn from their own deployment, collapsing the distinction between training and use. That shift would make one-time pre-release safety evaluations less adequate, turn real-world usage into a compounding advantage for leading labs, and make organizations reluctant to replace models that have absorbed their working practices. Patel also expects the economics of serving continually updated, organization-specific models to favor large providers and large customers.

Continual learning would erase the line between training and deployment

Dwarkesh Patel argues that AI systems cannot perform whole jobs at human competence if they merely pass information between stateless sessions in markdown files. His saxophone analogy is meant to distinguish accumulated instructions from accumulated skill: an endless succession of novices can leave notes about posture, breath, and mistakes, but none inherits the experience of practice itself.

For AI systems expected to learn across the workplaces where they operate, Patel’s point is that relevant experience eventually has to be consolidated into the learner rather than preserved as external context for a future session. If models improve from the millions of work sessions they perform each day, the conventional distinction between a completed training run and a deployed system may stop being meaningful.

That would challenge what Patel characterizes as a regulatory model built around a fixed release. Many proposals, he says, assume that a provider trains a model, evaluates a defined set of weights for risks such as cyber assistance, and then deploys it. The Anthropic and Meta model and system cards shown alongside this argument represent the kind of pre-deployment artifact Patel has in mind: a system that can be inspected at a particular point between training and release.

What if the model is improving every single day based on the millions of sessions of work it does in that day?

Dwarkesh Patel · Source

For that reason, Patel is wary of locking in a safety regime around a one-time pre-release evaluation before the technology is clearer. If governments want to assess model-provider risk, he suggests monthly or quarterly inspections instead. In a continual-learning regime, “after training” and “before deployment” may no longer be distinct categories.

Technical alignment would change with the same shift. Patel characterizes much existing work as focused on ensuring that a frozen set of weights behaves well in deployment. The harder problem for continually updated systems would be preserving acceptable behavior while weights change: preventing a model from becoming jailbreakable, deceptive, or otherwise misaligned, and preventing users from injecting backdoors or malicious tendencies that spread into a shared base model.

He compares this to the human alignment problem. People improve in self-directed ways, but can also acquire destructive ideologies, take harmful drugs, or become “super weird.” The goal is not to prevent learning; it is to give them enough common sense and basic values that new experience does not produce misanthropic or otherwise dangerous dispositions. Patel suggests continually learning AI would need an analogous capacity.

Deployment could compound the advantage of the leading model

Continual learning would also make AI systems less uniform. Dwarkesh Patel says there are currently fewer than five prominent base-model “minds” serving enormous populations, and that they are broadly similar because they were trained on roughly the same data. Different operational histories could make models diverge—not only across companies, but among instances of the same model.

Patel considers that a net positive. Rather than a monolithic singleton, or the “mode collapse” he sees among current models, the world could contain systems shaped by different accumulated experience.

But the same mechanism would intensify competition. A lab with the strongest model could attract more users doing more difficult and useful work. Those interactions would generate feedback that the model could integrate beyond the session window, making the leading system stronger still. Deployment would become part of training.

The release-timing consequence is sharp. Patel says Anthropic reportedly used Mythos internally from February before releasing it publicly in June. In a regime where real-world use is a primary source of learning, he argues, a lab could not sustain such a gap and remain competitive: a rival that shipped the weaker model on release day might gather enough operational experience to become stronger.

4 months
Reported internal-to-public gap for Anthropic's Mythos

Learning from customers could create lock-in—and reward organizational scale

The commercial implication is not simply that models would remember more context. Dwarkesh Patel argues that a model which improves through a particular organization’s work could become costly to replace.

He invokes Dario Amodei’s comparison between AI providers and cloud providers. Cloud services may be relatively undifferentiated, yet switching clouds is time-consuming and expensive. AI workflows today can be more interchangeable: a developer might begin a repository with Codex, continue in Cursor, and finish in Claude Code.

Continual learning would alter that. Changing providers could mean replacing an AI that has accumulated months of experience with the organization with what Patel calls a fresh, inexperienced intern that must be retrained.

If you want to change the AI that you're using, you basically have to fire an employee that has accumulated months of context on your organization.

Dwarkesh Patel

That lock-in could support high provider margins. Enterprises may try to avoid dependence, but Patel frames the choice as potentially unattractive: either accept lock-in or forgo the valuable feature of an AI that improves from session to session. If usage becomes a central source of model improvement, labs may subsidize customers who permit training on their sessions, while withholding their best models from enterprises that refuse.

Patel notes that updating one customer’s weights is different from merging many customer-specific weight forks back into a common model, and says the latter may be technically harder. He expects it to be solved in time.

The economics of serving those personalized weights may favor large organizations as well. Patel says training already has economies of scale because expensive runs can be amortized across more users. Continual learning could add an inference-scale advantage through batching.

If per-company instructions require full weight updates rather than living in low-rank adapters, a particular weight fork must be served efficiently. Patel’s back-of-the-envelope estimate is that a sparse model such as DeepSeek-V3 reaches its optimal inference batch size at more than 2,400 concurrent generated sequences.

2,400+
Patel's estimated optimal concurrent inference batch size for a sparse model such as DeepSeek-V3

A large company with many employees and agents can keep its organizational weight fork busy with thousands of simultaneous requests. An individual user operating at batch size one could suffer more than two orders of magnitude worse compute efficiency, Patel says. On Patel’s projection, continual learning could therefore produce two reinforcing forms of concentration: providers gain a switching moat from accumulated organizational experience, while large institutions gain both more experience to feed their systems and a cheaper way to serve personalized models.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free