Orply.

On-Policy Distillation Trains LLMs on Their Own Generated Responses

Shervine AmidiStanford OnlineThursday, September 3, 20263 min read

Stanford adjunct professors Shervine and Afshine Amidi argue that pretraining gives large language models broad text-prediction ability, but not necessarily the capacity to follow instructions, reason reliably, or work in specialized domains. They position fine-tuning, reinforcement learning, preference optimization, verifier-guided training and distillation as post-training methods for shaping that capability, highlighting on-policy distillation because it lets a stronger teacher assess answers the student model generates itself.

Post-training shifts the question from scale to use

Shervine Amidi says much of AI’s progress came from pretraining: training larger models on more data. But broad text-prediction capability is not the same thing as reliably following an instruction or performing well on a demanding task.

For ? afshine-amidi, the question is therefore no longer only how to make a model bigger. It is also how to make it reason better, follow instructions better, solve harder problems, and adapt to specialized domains.

The question is no longer only how do we make a model BIGGER. It is also, how do we make it reason better? Follow instructions better? Solve harder problems? And adapt to specialized domains?

? afshine-amidi · Source

The Amidis place fine-tuning, reinforcement learning, preference optimization, verifier-guided training, and distillation on the post-training side of that divide. These techniques do not replace pretraining in their account; they address what happens after a model has acquired broad capability, when the objective becomes shaping its performance on instructions, reasoning work, difficult problems, or domain-specific tasks.

On-policy distillation changes what the teacher evaluates

Afshine Amidi highlights on-policy distillation as a particularly interesting variation on a familiar student–teacher arrangement.

In standard distillation, a stronger teacher produces examples and a smaller student learns to match the teacher’s outputs. On-policy distillation instead has the student generate responses during training. The teacher evaluates those responses, making the student’s attempted solution—not a teacher-produced example—the object around which learning is organized.

ApproachWho generates the responseWhat the student learns from
Standard distillationThe stronger teacherTeacher-generated outputs
On-policy distillationThe studentThe teacher’s evaluation of the attempt
The source distinguishes standard and on-policy distillation by the origin of the training response.

That distinction changes the training setting in a way Shervine Amidi says better matches practical use. A deployed model generates its own answers; on-policy distillation trains around that same basic condition, with a stronger model assessing what was produced.

This aligns the learning process with how the model behaves in practice.

Shervine Amidi

Reasoning makes the path through an answer matter

For Shervine Amidi, the alignment is especially important in reasoning tasks. The relevant material is not limited to a final answer: intermediate steps, mistakes, and corrections matter as well.

The source does not claim that every evaluation explicitly identifies or repairs each error. Its point is narrower: when training is organized around a model’s generated attempts, those features of its reasoning are present in the setting being evaluated. That differs from a setup in which the learning target is solely an output already supplied by the stronger teacher.

Afshine Amidi describes the broader aim as making recent post-training advances intelligible from first principles and connecting them to practical systems. On-policy distillation supplies a compact example of that approach. The technical change is simple to state—who generates the response—but it bears on a practical concern: whether training reflects the way a model will actually be asked to reason and respond.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free