Orply.

LLM Design Trades Model Capacity Against Compute Per Token

Shervine AmidiStanford OnlineFriday, October 9, 202622 min read

Large language model design is a question of how to spend computation, not simply how many parameters a model has. In Stanford’s CME295 lecture, adjunct professor Afshine Amidi uses training budgets, sparse expert routing and attention mechanisms to show how designers allocate work across a model and its tokens; decoding choices, in turn, shape how the model generates answers. The distinction is between a model’s total capacity and the computation it uses for each token.

The central design problem is spending compute where it changes the answer

A large language model’s size is not the same thing as the amount of computation it must use for every input. That distinction connects several of the design choices discussed by ? afshine-amidi: how much data to train on, which parameters to activate for a token, and how much memory to spend keeping earlier computations available during generation.

The starting point is the decoder-only transformer. Its basic operation is next-token prediction: given the tokens already in a sequence, assign probabilities to possible next tokens. In the simplified model described in class, repeated blocks of causal multi-head attention and feed-forward computation process the sequence. Attention lets tokens use information from earlier tokens; the feed-forward component is where Afshine located much of the computation.

That backbone is a result of trying different ways to adapt the original transformer, not a requirement built into the idea of language modeling. The original design split the work between an encoder and a decoder. The encoder formed representations of the input; the decoder generated output autoregressively, attending both to its earlier generated tokens and to the encoder’s representations. That division suits a task such as machine translation, where source and target text play different roles. Later variants kept only one side of the architecture.

BERT kept the encoder. Its bidirectional self-attention lets each input token attend to tokens before and after it. A special [CLS] token at the start of the sequence participates in that attention; its resulting embedding can be used as a representation of the input for a downstream task. For example, a projection from that representation can classify a sentence or tweet as positive or negative. An individual token’s embedding can instead support a token-level prediction, such as whether that token is a noun or a verb.

BERT’s training has two stages. Pretraining uses proxy tasks—tasks intended to help the model learn about text, rather than the application itself. In masked language modeling, selected tokens are corrupted and the model tries to recover the originals from context. In the example shown, “A cute teddy bear is reading” becomes “A [MASK] teddy bear is [MASK].” The model uses its representation of the corrupted positions to predict “cute” and “reading.” Afshine cautioned that the actual corruption procedure is more involved than simply replacing selected tokens, but the explanatory point was the same: predict missing material using surrounding text.

BERT also used next-sentence prediction in the approach described in class. The model receives two sentences separated by [SEP] tokens and predicts whether the second follows the first. Afshine presented masked language modeling as a way of learning how tokens fit into text and next-sentence prediction as an attempt to learn how sentences are sequenced across a document.

After pretraining, BERT is adapted to a particular task. For sentiment classification, a projection maps the [CLS] representation to the relevant classes. This stage requires labeled examples, but Afshine characterized it as needing relatively little data—on the order of a hundred examples. The stated advantages were good performance and a modest data requirement for fine-tuning. The costs were that fine-tuning remains necessary and that an encoder built to produce representations does not offer an obvious general route to text generation. Afshine noted that encoder-only generation is possible but deferred the explanation to a later lecture.

GPT takes the other route: retain the decoder and express tasks as text-to-text problems. A sentiment request can be written as a prompt—“Is this sentence positive or negative?” followed by the sentence—and answered with generated text such as “positive.” The underlying operation remains next-token prediction. The model predicts a token, adds it to the sequence, and predicts another using the expanded context. Afshine described this as autoregressive decoding.

This framing makes fine-tuning unnecessary for every new task, though the point was not that it is never useful. A model can be prompted to answer a classification question in text; fine-tuning may still improve its performance. The key architectural difference is that the input is not assigned a separate encoder. In the decoder-only design, input and generated tokens are handled under causal attention: each token can attend to itself and earlier tokens, not future ones. In the original encoder-decoder design, input tokens can attend bidirectionally to one another, while the decoder generates causally.

Afshine’s explanation for the popularity of decoder-only models was partly practical. An encoder-decoder system treats input specially, a sensible arrangement for machine translation, where source and target have distinct roles. For general text generation, it may be less clear what should be separated as input. Decoder-only models treat everything alike and offer a simpler text-to-text framing. That is a design advantage in the settings discussed, not a claim that encoder-decoder models cannot generate text.

The same question—what work must be done for each input?—reappears in scaling, expert routing, and inference. A model can have many parameters without necessarily using every one for every token. It can reuse past calculations rather than recompute them. It can limit attention to nearby tokens in an individual layer. These choices do not remove computation; they determine where it is spent and what costs are accepted in exchange.

Training scale is constrained by how the compute budget is divided

Scaling-law research frames the first resource tradeoff: increasing model size and training data can improve performance, but both consume a finite training budget. The slide attributed to Kaplan and colleagues plotted test loss against compute, dataset size, and parameter count. Afshine pointed to the curves as evidence that loss fell as those quantities increased in the experiments presented.

The same work was described as showing greater token efficiency for larger models. In the plotted comparison, larger-model curves reached a given test loss after processing fewer tokens than smaller-model curves. Afshine’s interpretation was that a larger model could perform better earlier in training, needing fewer tokens to reach a particular level of performance. The relationship helps explain why model sizes and training-token counts grew in the early 2020s. It does not imply that tokens cease to matter or that increasing parameters alone is always the best use of a fixed budget.

A fixed budget creates a different question from whether scale helps in general. Training a larger model costs more per token, so a given compute allowance permits fewer training tokens. The practical choice is how to allocate compute between parameters and data.

Afshine distinguished two similar-looking measures. FLOPs, with a lowercase “s,” counts floating-point operations and measures total computation. FLOPS, or FLOP/s, measures operations per second and describes hardware speed. The class gave an order-of-magnitude figure of around 10²⁵ FLOPs for training an LLM. The distinction is between the total work required and the rate at which a machine can do it.

The 2022 compute-optimal training paper by Hoffmann and colleagues addressed how to divide a training budget between model size and token count. Afshine summarized its conclusion as finding a “sweet spot” and said that many models at the time were undertrained: too large relative to the number of tokens they had seen, given the compute spent. In this use of “undertrained,” the issue is not that the models had too few parameters. It is that their parameter counts were high relative to the training data processed.

They were too big for the amount of tokens that they were trained on. They were under trained.

? afshine-amidi · Source

The classroom comparison illustrates the tradeoff. Afshine recalled a model of about 70 billion parameters trained on roughly 1.4 trillion tokens and a larger model of more than 200 billion parameters trained on fewer tokens under the same compute budget. He described the smaller model as more performant. The point of the example is the controlled budget: spending more on parameters leaves less available for tokens.

The rule of thumb presented was roughly one parameter for every 20 training tokens. The slide’s table included a 67-billion-parameter model paired with 1.5 trillion tokens, consistent with that approximate ratio. The ratio is a useful account of the paper’s compute-optimal setup, not a universal prescription for every architecture, dataset, or objective.

ExampleParametersTraining tokensTokens per parameter
Smaller model in the classroom comparisonAbout 70 billionAbout 1.4 trillionAbout 20
Larger model in the classroom comparisonMore than 200 billionFewer than the smaller modelNot specified
Model listed in the paper’s table67 billion1.5 trillionAbout 22
The examples illustrate the approximate tokens-per-parameter ratio and the allocation tradeoff under a fixed compute budget.

There is a second budget beyond training. A model also has to be served, and a very large model can be expensive to run when it receives many requests. Afshine raised this as a separate consideration: the allocation that is compute-optimal for training need not be the most economical choice for inference. The class did not quantify the serving tradeoff, but it makes the distinction important. Training asks how to spend a bounded amount of work to produce a model; inference asks what resources the model consumes while answering requests.

Afshine also qualified the scaling-law discussion when asked about mixture-of-experts models. The paper he had cited studied dense models, not MoEs. If researchers believe their architecture differs materially from the setup studied, he said, they may run smaller experiments using their own setup to estimate a scaling relationship. The dense-model result does not settle how to train MoEs.

A language model, in the definition shown in class, assigns probabilities to sequences of tokens. In the decoder-only case, the model’s immediate task is to assign probabilities to possible next tokens. “Large” refers to more than one dimension: billions or more parameters, hundreds of billions or more training tokens, and substantial compute. Those dimensions interact, but they are not interchangeable. A larger parameter count raises model capacity; more training tokens provide more training examples; and compute limits how much of each can be used.

Sparse routing separates total capacity from work per token

If every parameter in a large model were used for every input, increasing total size would also increase the computation required for each forward pass. The class introduced mixture of experts as a response to that problem. Afshine’s analogy was a room of specialists: if a question is about mathematics, it may not need to be sent to every specialist present.

A mixture-of-experts model has experts and a gate, or router. The router uses the input to decide which experts should contribute and to what extent. A dense mixture assigns different weights to experts but still activates them all. That preserves the computation cost the design is meant to avoid. A sparse mixture selects a subset, often the top k experts, and combines the selected experts’ outputs according to their routing weights.

In the transformer design discussed in class, the MoE mechanism typically replaces the feed-forward network. The experts are feed-forward neural networks, and the router makes its selection for each token. Afshine distinguished this from attention: attention is where tokens communicate, while the feed-forward network is where much of the computation happens. The motivation is to increase the model’s total number of parameters without activating all of them for each token.

Mixture typeRouter outputExperts activated
Dense MoEWeights assigned across expertsAll experts
Sparse MoEWeights for a selected top-k subsetOnly selected experts
Dense routing changes the experts’ relative contributions; sparse routing also limits which experts run.

Sparse routing introduces a problem during training. If certain experts are selected frequently early on, they receive more updates and can improve; experts selected less often receive fewer updates and can remain weak. That imbalance can reinforce itself. Afshine called this routing collapse: the model relies on a small set of experts instead of using the full collection. He described the consequence as an effective model size smaller than the total number of experts and parameters.

The class showed a load-balancing auxiliary loss as one response. The formula uses the fraction of tokens routed to each expert and the average routing probability assigned to that expert. Its purpose, as Afshine explained it, is to encourage use of the experts more evenly. In response to a question about the discrete selection, he said the fraction of tokens is treated as a constant for the relevant gradient calculation, while the routing probability is where the gradient passes through. The routing can therefore still be adjusted, while the auxiliary term discourages the same experts from always being selected.

Afshine emphasized that routing collapse is one of the challenges of training an MoE-based model, not the only one. He showed code from a Mixtral example in which colors marked which expert received each token at a particular layer. The displayed token assignments were diverse. That illustration showed routing choices in one example; it was not presented as a general test of expert quality.

The class also discussed two design directions associated with a 2024 paper. One is to include shared experts that activate regardless of the router’s selection. Afshine’s explanation was that some computation is useful for every token, and shared experts can reduce the pressure for routed experts to learn the same general foundation before specializing. He presented this as a design people use, not as a guarantee that shared experts improve every model.

The other direction is to make experts more fine-grained. Rather than use only eight or 16 experts, the class described designs with dozens or hundreds available and a smaller number selected for each token. Afshine’s intuition was that a larger pool gives experts more opportunity to specialize. The performance chart shown in class compared several configurations and, in the lecturer’s account, indicated better performance for the finer-grained arrangements. That is the evidence presented in the class, not a universal rule about the best number of experts.

The displayed model examples had 256 or 384 total experts, with six to eight routed experts active at a time. Some examples included one shared expert and one did not. Their listed model sizes ranged from hundreds of billions of parameters to just over one trillion. The figures demonstrate the distinction between total and active experts in the examples, rather than establish a standard configuration for all MoEs.

These models make the training-versus-inference distinction concrete. A model can contain many expert parameters while routing each token to only a subset. That changes the amount of computation used per token relative to activating every expert, but it does not make the system costless: the model still has to store its parameters, learn useful routing, and manage the computation associated with selected experts. Sparse routing offers a way to scale capacity without requiring every expert to run for every input.

Position and attention changes target different costs

Self-attention gives tokens direct connections to other tokens, but a model also needs information about their order. Without position information, the connections do not by themselves indicate where tokens occurred in the sequence. The original transformer added positional information to token embeddings.

One option is to learn a separate embedding for each position and add it to the token embedding. The difficulty identified in class is that if training uses sequences of a given length, learned embeddings for positions beyond that length have not been encountered. The original transformer also tried fixed sinusoidal embeddings, which can be computed at positions beyond those seen in training.

The class motivated the sinusoidal construction with a clock. Hour, minute, and second hands move at different rates; considered together, their positions give a more precise sense of time than any one hand alone. In the positional vector, different dimensions vary at different frequencies. Afshine described lower-index dimensions as changing more quickly and higher-index dimensions as changing more slowly. Taken together, these patterns represent a position.

The construction also relates similarity to relative distance. For two positions, m and n, the dot product of the corresponding vectors can be expressed as a sum of cosine terms involving m − n. In Afshine’s explanation, this makes the dot product a function of the difference between the positions. Nearby positions tend to have higher similarity than distant positions. He said the original paper found learned and fixed approaches to have roughly similar performance; the reason he emphasized for the fixed approach was that it could be evaluated at positions beyond those seen in training.

The class then questioned where positional information should enter the model. Adding it to token embeddings affects all later layers, including the feed-forward network. Afshine argued that relative position is especially relevant in attention, where queries and keys determine which tokens interact. This motivated methods that act directly in the attention calculation.

One approach adds a bias to the query-key scores. The class described T5’s bias as learned per attention head and dependent on a bucketed distance between positions. Afshine identified two concerns: the bias has to be learned, and it may not generalize to every possible distance. ALiBi—attention with linear biases—was presented as a deterministic alternative in which the bias varies linearly with the distance between positions.

The method Afshine described as commonly used was RoPE, or Rotary Position Embeddings. Instead of adding a position vector to token embeddings or adding a distance bias to attention scores, RoPE rotates the query and key vectors according to their positions. In two dimensions, multiplying a vector by a rotation matrix changes its angle. When the query at position m and key at position n receive position-dependent rotations, their interaction depends on the difference between those positions.

Afshine worked through the geometry on the board: write a vector using its cosine and sine coordinates, multiply it by a rotation matrix, and the result has an angle shifted by the matrix’s rotation. The point is that applying position-specific rotations to queries and keys produces an attention interaction that reflects relative position. For dimensions greater than two, the rotation is applied in two-dimensional blocks, with angles that vary by block.

Afshine also described the broad behavior of query-key interaction with distance as tending to decay as positions grow farther apart, while warning that it is not strictly monotonic. He presented that pattern as a way for attention to reflect relative distance, not as a simple rule that every farther token receives less weight. In response to a concern that this might prevent distant tokens from interacting, he said the method does not remove those connections. Closer tokens may be more likely to matter, while farther tokens remain available to the model.

Position handling is one part of attention efficiency. Another is how many tokens each token can attend to. Full attention lets a token attend to every eligible token, but the class noted that doing this at every layer involves substantial computation. Sliding-window attention limits a token to a local range. The slide also showed designs that interleave local and global attention layers.

Afshine addressed the concern that local attention might recreate the difficulty of carrying information across long distances. A transformer has multiple stacked blocks. A token may attend locally in one layer to representations that already incorporate information from earlier layers. Across several layers, information can travel beyond a single window. He compared the effect to a convolutional network’s receptive field: a sequence of local operations can let a unit depend on a broader region than any one operation directly covers.

This explanation is a qualification, not a claim that local attention is equivalent to unrestricted attention. The direct connections in a given layer are still limited to the window. The argument is that stacking layers can expand the range of information available indirectly, while reducing the work done by each local attention operation.

Generation creates a separate memory cost. In a decoder-only model, each new token attends to itself and earlier tokens. At each step, the model compares the new query with keys for previous tokens and uses their values. Recomputing those earlier keys and values every time would repeat work, so a system can save them during decoding and reuse them for subsequent predictions.

The saved keys and values consume memory, and the amount grows as the generated sequence gets longer. Multi-head attention produces multiple key and value representations for each token, which increases the material that may need to be kept. Multi-query attention shares one key and value across query heads; group-query attention shares keys and values within groups of query heads. Standard multi-head attention instead uses separate projections for each head.

Afshine described group-query attention as common and gave its memory motivation: reduce the number of distinct keys and values that must be retained during autoregressive generation. Queries are used for the current step; they do not need to be saved in the same way. These arrangements do not eliminate the need to retain information from prior tokens. They reduce the number of separate key/value representations that must be maintained.

Model design also determines how much variation and reliability generation can offer

The decoder’s final representation is projected from the model dimension into the vocabulary, and a softmax converts the resulting scores into probabilities for the next token. Decoding is the choice of how to turn that distribution into a sequence.

Greedy decoding selects the highest-probability token at each step. Its simplicity comes with a limitation: choosing the best next token at each step does not guarantee the completed sequence will have the highest joint probability. Beam search keeps k candidate paths, extends them, and retains the most likely candidates at each step. It looks beyond a single immediate choice, but costs more computation. As Afshine stressed, it is not a full search over every possible sequence, which would be too expensive.

Both methods are deterministic under the same conditions: they return the same path rather than offering alternative responses. Afshine illustrated why that might be undesirable in a chatbot. If different people ask the same greeting and receive exactly “Good, and you?” each time, the interaction can feel robotic. He explicitly treated this as a thought experiment, noting that real context can change between requests and thereby change the output even under a deterministic scheme.

Sampling offers a way to produce varied plausible continuations. Instead of always selecting the highest-probability token, the system draws from the distribution. The class described two common restrictions on the candidate set. Top-k sampling draws among the k most probable tokens. Top-p sampling draws among the smallest set of tokens whose cumulative probability reaches a threshold, such as 90%.

Temperature changes the distribution before sampling. At a low temperature, the distribution becomes sharper: high-scoring tokens receive more probability and the others less. At a high temperature, the distribution becomes flatter. Afshine described temperature as a continuous control rather than a binary setting for determinism. The appropriate choice depends on whether the task benefits more from repeatability or variation.

The class named summarization and evaluation as cases where consistency can be useful. For example, a repeatable summary or consistent ratings from an LLM judge may matter. In many other settings, Afshine said, a nondeterministic path may be more attractive. This is a tradeoff, not a claim that one decoding strategy is best for every task.

Even a fixed decoding rule may not produce identical output in every real inference setup. The class described inference engines that batch requests, processing a user’s request alongside others. Depending on implementation, the order of operations within tensor calculations can change. The model’s weights and intended operations may be fixed, but batch invariance is not respected in every case.

The explanation offered was floating-point non-associativity. In exact arithmetic, regrouping additions does not change their result. With floating-point values, which are approximations, the order of additions can matter: (a + b) + c may differ from a + (b + c). A small change in intermediate calculations can affect later scores and, potentially, the next-token choice. Afshine noted that the details depend on implementation and recommended a technical blog for readers who wanted to examine the issue more closely. The point was that theoretical determinism does not necessarily guarantee identical results across all execution arrangements.

Reliability can also mean controlling the form of an output. Asking a model to produce JSON may work, the class said, but it is not guaranteed; malformed output can cause downstream parsing to fail. Guided decoding addresses this by restricting which next tokens are allowed according to the output syntax. If the model is generating JSON, the allowed next-token set can be limited to tokens that keep the sequence valid. This is a stronger constraint than prompting the model to follow the format and hoping it does.

The same focus on control appears in prompting. The context window includes both the prompt and the tokens generated so far. It is not just the text supplied by the user: generated tokens are fed back into the model as generation proceeds. Consequently, a long answer uses part of the context budget that might otherwise be available for input.

The class gave approximate context and output sizes based on examples shown, while emphasizing how quickly the specifications were changing. The examples included context limits around one million tokens and output limits often in the order of 100,000 tokens, alongside a newly announced model with an output limit of one million. These were examples from a particular moment, not a guarantee that every model has those limits. A model’s input limit and its maximum output length may differ.

For scale, the class offered rough conversions: a token is about three-quarters of a word, and a page is about 500 tokens. At that scale, 10,000 tokens is comparable to a technical report, 100,000 to a book, and a million to an encyclopedia. These are approximations; the number of tokens depends on the text. The key accounting point is that output tokens also consume the context window.

More context does not necessarily mean better recall. Afshine described “context rot” through a Needle in a Haystack test: place a fact somewhere in a long context and ask the model to retrieve it. The chart shown in class indicated lower retrieval accuracy at larger context lengths in the example, with degradation when the fact was placed between 10% and 50% of document depth. This is a risk of burying relevant information among increasing amounts of other material, not proof that every model fails in the same way at every long context length.

That risk is one reason Afshine said context compaction and starting a new chat can remain useful practices. A large context limit makes it possible to supply more material; it does not guarantee that every detail will remain equally accessible. The distinction between capacity and reliable use of capacity runs through the other examples too: total parameters versus active experts, theoretical decoding rules versus inference behavior, and a large context window versus successful retrieval.

Prompting trades model changes for longer and more structured inputs

In-context learning changes the prompt rather than the model’s weights. A zero-shot prompt asks the model to do a task directly. A few-shot prompt includes examples of questions and answers, showing the format or pattern desired. In the class’s bear example, a direct question about the bear’s age is zero-shot; examples about other bears provide demonstrations before the new question.

Afshine said examples can help the model learn the expected format and may implicitly convey a reasoning pattern. He also described a tradeoff: examples take effort to prepare and add input tokens, increasing cost and latency. The class’s slide listed those drawbacks alongside the potential performance benefit.

He further said that few-shot examples were becoming less common as reasoning models improve, and that demonstrations can make the model overfit to the particular cases supplied. In some cases, people instead craft instructions that express what the examples were meant to teach. This is Afshine’s account of a changing practice, not a claim that examples are no longer useful.

Chain-of-thought prompting extends demonstrations by including explanations of how an answer is reached. The class cited a 2022 paper in which the authors found that examples containing reasoning paths improved performance. In the bear example, the answer explains that the bear was born in 2020 and is four, then uses that result to answer how old it will be next year. The demonstration contains an intermediate explanation, not only the final answer.

Afshine cautioned against treating chain of thought as merely an in-context learning technique. He said the same idea would arise in a later discussion of reasoning models and thinking tokens. Explanations also increase token use, cost, and latency. The class presented them as a possible performance aid with a resource tradeoff.

Self-consistency takes the idea further by generating multiple reasoning paths and aggregating their answers. In the example shown, several paths reach five while one gives a different result; majority voting selects five. Afshine described this as a way to make answers more stable on tasks such as arithmetic, while noting that multiple generations add cost. It is useful when the benefit of agreement among paths justifies that additional computation.

Across these techniques, prompting does not bypass resource constraints. Examples and explanations add tokens to the input; generating multiple paths adds inference work; long prompts compete with outputs for context; and more context does not guarantee better retrieval. The practical question is what the task needs—a demonstration, an explanation, a constrained format, or several candidate solutions—and whether the resulting cost is worthwhile.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free