AI’s Next Training Paradigm Depends on Learning From Deployment
Dwarkesh Patel argues that frontier AI labs are betting too much on reinforcement learning from verifiable rewards: training models across vast numbers of checkable, replayable tasks in the hope that this produces general agents. In his account, verifiability is not enough; the domains that matter most are often too slow, messy, and non-repeatable to be “grindable” training environments. The next paradigm, Patel suggests, will depend on whether models can turn scarce deployment experience into durable updates to their weights, through continual learning methods such as on-policy self-distillation, “dreaming,” or something not yet invented.
Dwarkesh Patel·Jun 26, 2026·14 min read