The Scheduler Can Bottleneck Inference Without Doing the Heavy Compute
Charles Frye, who works on inference at Modal, argues that an inference engine’s central challenge is not generating tokens but serving them efficiently for the workload at hand. The scheduler can limit performance even though it does far less computation than the GPU, because it decides which requests reach the accelerator and how they are batched. Frye traces how workload demands, cache management and host-side overhead shape latency and throughput, and why operators need evidence from the deployed system to diagnose problems.

The hard part is not generating the next token
An inference engine takes tokens in and returns tokens out. A basic version can be assembled quickly with PyTorch and Transformers, Charles Frye says. The engineering challenge is making it produce those tokens efficiently: keeping latency and cost under control while serving the workload people actually have.
That challenge spans several layers. A server handles requests and responses; tokenizers and detokenizers translate between external inputs and model tokens; a scheduler decides what work reaches the accelerator; and model code runs on the GPU. Most of the expensive computation happens on the GPU, but the scheduler can still limit performance. It is the gatekeeper that chooses what the GPU does and manages the resources needed to do it.
Frye, who works on inference at Modal, frames inference engineering as work across the stack, from application design to linear algebra and hardware. The economic distinction matters too: training is a cost center, while inference is a revenue center. Companies may pay to use a model through a service even when they do not buy the model weights themselves. And while a relatively small number of organizations produce foundation models, many more deploy and customize them. Inference also supports post-training, which requires generating samples from models.
In the end, training is a cost center, and inference is a revenue center.
The engine’s design follows from its central job: run model work at high speed without letting the machinery around that work get in the way.
The scheduler is the sort of gatekeeper to the GPU.
A single application can contain two different workloads
The same API can serve applications with very different performance requirements. Frye groups them into three broad types.
A chatbot or coding assistant has a person waiting on the other side. The “plus” in Chatbot+ means the system may also call tools or interact with external services. These applications often reuse a substantial prompt prefix, produce relatively short outputs—tens to hundreds of tokens—and have tight latency budgets. A user expects a response quickly; Frye says a delay of a few hundred milliseconds can already matter.
A background agent may do similar work, such as writing a code change, but without responding directly to a person waiting in the interface. It may have minutes to produce a pull request, or hours if the expected output is comparable to work by an engineer. The application can tolerate more delay, though it may still benefit from prefix reuse and may generate long outputs.
A data processor turns unstructured material, such as a PDF or video, into structured information. It may use a relatively small model and produce a short result, such as fields to put in a database. Because it processes different documents, it generally has little prefix reuse beyond a system prompt. Its main requirement may be aggregate throughput rather than low latency for any one request.
These distinctions become clearer when a request is split into prefill and decode. During prefill, the model processes the input tokens. During decode, it generates output tokens, usually one or a few at a time. Prefill and decode are not just successive stages from the user’s perspective; they are different workloads inside the engine. Prefill does more arithmetic per request. Decode does less arithmetic per token, but repeats the work sequentially as output grows.
That split changes what “fast” means. Applications and operators need to consider query volume, input and output length, how much prompt content can be reused, and the latency budget. Two particularly important latency measures are time to first token and time per output token, also called inter-token latency. Query volume is difficult to forecast because it depends on users and can vary over time. Token counts are also variable: input length comes from users, while output length depends partly on when the model decides to stop. Those quantities need to be measured and benchmarked, not simply assumed.
Follow the request to see where the engine’s work begins
At the outermost level, an inference server accepts an HTTP or RPC request and returns a response. The inference engine is the part that prepares and runs the model work inside that interface. The distinction matters because HTTP serving is usually not the hard part. Frye describes tens or hundreds of requests and responses per replica, perhaps thousands at most; general-purpose hardware can handle the server I/O. The performance-sensitive work sits further inside.
Inputs can include text, images, video, or other formats. Pre-processing turns them into tokens, and model execution maps those tokens to tensors on the GPU. In the simplest next-token view, the model takes a tensor containing the current sequence and produces the next entry. Post-processing turns the generated output back into tokens and then into a response that the client can use.
In the SGLang architecture Frye used as an example, this work is split across processes: server I/O, tokenizer processes, a scheduler, model runners, and detokenization. Components communicate with one another rather than running as a single undifferentiated program. There are implementation differences across engines; for example, Frye said tokenizer and detokenizer live in the same process in vLLM’s engine manager, while they are separate in the SGLang architecture he described.
A request moves from the server to tokenization, then to the scheduler. The scheduler considers what resources are available and organizes requests into batches for the model runners. Accelerators benefit from processing multiple requests in parallel, so the engine collects work and runs it together. Frye compares this to collecting several database queries that will each scan the same table: if loading the table is the hard part, it can make sense to do the queries together. The model runners return tokens or logits; as tokens are ready, they can pass to detokenization and back to the server for the response.
The tokenizer does not have to communicate directly with the detokenizer to keep their behavior consistent, Frye said. They use the same underlying tokenization libraries, and tokenization is the more complicated side of the mapping. He qualified his answer rather than promising that every implementation behaves identically, but said direct communication between the two components is not generally needed.
Nor does a long time to first token necessarily mean tokenization is slow. Tokenizers can run in parallel, often across multiple processes, but long inputs still take longer to process in the model’s prefill work. At sufficiently large sequence lengths, that work grows roughly linearly with the input size. If the input must be split into multiple batches, queuing can add delay more than once. Frye noted that the first token in “time to first token” means the first output token, not completion of tokenization.
Frye also recalled a presentation from Crusoe that, as he remembered it, reported that existing tokenizers generally do not show up as a bottleneck for models around 20 billion parameters or larger, but can matter for smaller models processing very long contexts. He offered this as a rough guide, not a universal threshold. He said he had not often needed to focus on tokenizer latency himself: in his experience, prefill and especially decode more often dominate. The tokenizer is typically on a millisecond scale or faster and runs once per request, while decode latency recurs for every output token. He also pointed out a contrast: people commonly discuss KV caching, but rarely tokenizer caching, even though caching tokenization is possible.
The scheduler can bottleneck a GPU it barely computes on
The model forward pass is familiar territory to many machine-learning engineers: it contains much of the PyTorch code and runs on the accelerator. GPUs are expensive and capable of enormous amounts of computation, so it is natural to focus performance work there. But the scheduler can constrain how much useful work reaches them.
Someone has to decide which work runs on the GPU and manage the associated resources. In implementations Frye discussed, the scheduler may be literally single-threaded; more broadly, it is semantically a single point of control. It does far less computation than the GPU, yet can become a bottleneck because it sits between requests and the accelerator.
Asked whether parallel schedulers might solve this, Frye argued that coordination can cost more than it saves in the current setup. The host-side process needs to be faster than the GPU work it is arranging. Adding locks or coordinating concurrent access to GPU resources increases complexity. As an illustration, he said that even reducing scheduling to a microsecond would not make GPU work start sooner if the preceding work still takes tens or hundreds of milliseconds. He was not claiming that scheduler performance can never become limiting: he said that, in most cases, the host process can run fast enough without parallelizing it, while the trade-off depends on the workload. He also suggested the pressure may increase as GPUs get faster and the amount of work on the GPU falls.
This is also why engines use separate processes for different components. The server, tokenizers, detokenizer, and model runners need their own threads of control so they can operate independently. Python’s global interpreter lock has historically limited how threads in one process can execute Python code concurrently, making processes a practical way to separate compute-intensive work.
Python remains common because it has strong model and GPU support, including deep C and C++ interoperability. The model worker’s role is often to tell the GPU what to do; the GPU operations themselves are implemented in compiled, heavily optimized code. Rewriting everything in Rust would not automatically accelerate those operations. The scheduler may be a more plausible candidate for a lower-level rewrite: it is host-side, manages GPU resources, and need not contain all the model-specific PyTorch logic. Frye suggested that pressure on this component may grow as GPUs get faster and the amount of work per operation falls.
Concurrency brings a related question: how much KV cache can each request occupy, and what happens when demand exceeds available memory? Frye said the scheduler generally favors requests that already have KV cache allocated, so their previous computation need not be repeated. If cache space must be cleared for another request, material may be moved to CPU memory and then disk. He compared this to swapping in an operating system: it can be useful to have, but he considers cache thrashing or data being flushed to disk a sign that the engine is in a failure state. The fallback may help avoid going hard down; it is not a way to make repeated cache movement cheap.
Engines organize model work; kernel libraries do much of the computation
Inside the model runner, the engine’s forward-pass code controls which operations run and in what order. Adding support for a model architecture generally requires implementing that control flow in each engine. Reference implementations in libraries such as Transformers can help, and SGLang and vLLM have paths that use them with some automatic optimization, but Frye said production use commonly relies on more specifically implemented model code.
The underlying kernels are often supplied by external libraries. These libraries perform operations such as attention and matrix multiplication; an engine selects and organizes them rather than implementing every operation itself. That is one reason Frye sees limited differentiation between engines in raw GPU performance. Much depends on whether an engine gets out of the GPU’s way and constructs useful batches.
The important operations include cross-token computation, such as attention, and per-token computation, such as the model’s multilayer perceptron. Dense models use the same MLP for every token. Mixture-of-experts models route tokens to different smaller networks. That routing can require communication as well as computation, and the work may be handled in separate kernels or combined into a larger kernel. Frye pointed to DeepSeek’s DeepGEMM and DeepEP among the libraries used for matrix multiplication and routing, and mentioned attention backends including FlashAttention, FlashInfer, CUTLASS, and Triton. Which options are available and how they are configured vary by engine and change over time.
Batch construction is a more distinctive part of an engine’s design. The engine must decide which requests to run together, how to queue them, and whether to mix prefill and decode work or run them separately. Different choices have different performance implications. Frye said an engine may need to support pure prefill batches, pure decode batches, or mixtures, and to switch between them dynamically. He regards a queue that keeps growing as a problem: ideally, queued work is flushed as quickly as possible.
Frye also described several ways work can be distributed across accelerators. Tensor parallelism splits matrix multiplications across GPUs; data parallelism splits the batch; pipeline parallelism divides forward-pass steps across GPUs; and expert parallelism distributes mixture-of-experts components. Context parallelism splits a sequence across GPUs, which he described as still new in open-source engines. Prefill/decode disaggregation, in his view, is a special case of pipeline parallelism. The kernels need to support the chosen parallelism, while the model-forward code in the engine handles the associated communications. Frye noted that communication is where many problems with these techniques arise.
Tokenization is comparatively well served by existing libraries for text, but multimodal inputs complicate the path. An image or video requires more than mapping Unicode bytes to tokens; video processing, for example, may involve extracting frames before turning them into tensors. Frye described this as an area where traces still reveal opportunities for optimization.
KV cache saves computation, but its capacity sets limits
Attention requires comparing tokens across a sequence, and Frye described its load-bearing computation as quadratic. KV caching stores information after it has been computed, trading storage for less repeated computation. The constraint is GPU memory: storing data somewhere else and loading it back can be slower than recomputing it, because GPUs are fast.
Prefix reuse adds another opportunity. If requests share the same initial text—such as a long system prompt used by an agent—the engine can reuse the computation for that shared part. Frye illustrated this with several continuations of “Thou shalt not”: the shared prefix can be reused up to the point where the text branches, but not beyond it. A collection of such shared prefixes has a tree-like structure. Reuse can matter when an agent repeatedly sends a system prompt tens of thousands of tokens long.
Frye described KV caching in pages or radix-like structures. In a page-based layout, tokens are grouped into blocks; a page-size-one layout instead maps individual tokens to blocks. The engine stores reusable information in a cache and reconstructs the tensors needed by attention kernels. He said some of this work, once a defining engine concern, has moved into the attention kernels. The engine’s remaining problem is especially about managing cache capacity and layout.
CUDA graphs and speculation address different sources of decode cost
Even when the GPU has work to do, CPU overhead can interrupt progress. Each kernel launch may require host-side work to decide what to run and where. A host-device synchronization can block that process. The goal is not to eliminate CPU work—the CPU still has to make decisions and return results—but to make sure it does not keep the GPU waiting.
CUDA graph capture addresses repeated launch overhead. The engine records a sequence of GPU operations and their dependencies as a graph: which kernel runs, how it changes data at a pointer, and what later operations depend on that result. Instead of doing separate host-side work to launch each kernel, the engine can launch the graph with a single CPU-side operation. Frye showed a profiler trace in which many operations are launched through a graph and said the technique lets engineers relax some of their concern about what is happening inside the model forward pass.
Decode has a different constraint: it is sequential, and each token can require loading the model’s parameters—or the active parameters of a mixture-of-experts model—again. Frye described the resulting memory-bandwidth demand as enormous: producing tokens every few milliseconds while moving tens or hundreds of gigabytes of model data per token is difficult.
Speculative decoding tries to make those sequential steps more parallel. A separate speculator model proposes several tokens; the target model checks them together. With the appropriate sampling procedure, including rejection sampling, the target’s output can match what sequential decoding would have produced, apart from numerical differences. In effect, the method turns some decode work into a small prefill.
The speedup depends on how many proposed tokens the target accepts. Frye argued that improving the speculator can yield a roughly linear increase in decode throughput as acceptance length rises. Many engine optimizations produce incremental gains that require difficult performance work. A better speculator, by contrast, may offer a route to much larger gains. He gave two-, four-, and eight-fold speedups as examples of what might be obtained by investing in a better speculator, not as guaranteed results for every deployment. He said the technique had emerged as important for reaching high token rates on Nvidia hardware with large models, without requiring a different accelerator.
Correctness and performance need evidence from the deployed system
Inference services can fail in several ways. An application may behave badly even when the engine is working correctly. The engine can also produce incorrect model behavior, including through a tokenizer bug. Frye warned that model releases can include problems in tokenization or chat templates, which are often beaten down within a couple of weeks but still require attention. Performance problems include regressions that emerge only after a replica has run for a long time, and differences between replicas in a heterogeneous deployment.
His advice is to evaluate the deployment that will actually run, not just an abstract model or local setup. Run evaluation and benchmarking scripts against the deployed engine before release; record the results and traces so they can be investigated later. In production, retain traces and connect user feedback about correctness back to the request that produced the output. For tokenizer problems in particular, log token IDs as well as the rendered text. For performance, collect more metrics than seem necessary: a correlation that is invisible in a sparse dashboard may make a regression easy to find.
Useful measures include time to first token, time per output token or inter-token latency, end-to-end time, request rate, input and output tokens, queueing, and prefill, cached prefill, and decode activity. Frye also recommended looking at percentiles and averages, per-replica and aggregate measurements, and GPU temperature, power, kernel utilization, and memory utilization. The point is not that every metric will be useful every day. It is that production regressions can have causes that are hard to infer from one headline latency number.
A dashboard from Frye’s museum-placard application made the relationship visible. The application took what a camera saw, passed it through a vision-language model, and returned a short description in the style of a museum placard. Frye ran three replicas and showed a traffic spike. As requests accumulated, time to first token rose: requests were waiting before entering prefill. Inter-token latency rose too, first modestly and then more sharply as decode slowed amid contention on the underlying GPU.
The dashboard broke latency out further, including per-container P50, P95, and P99 measurements. That makes it possible to compare replicas rather than relying only on a single aggregate number. In Frye’s example, new replicas came online and began taking requests; the amount queued then fell back toward baseline. He presented adding replicas as the primary way to relieve congestion in many cases.
When dashboard metrics are not enough, system profilers can expose the whole engine: its processes, GPU work, and delays between them. Frye described using Nsight Systems or the Torch profiler to investigate a replica that ran more slowly than others. A trace can show work across processes and GPUs, making a delay in one replica visible rather than reducing the symptom to an aggregate latency measure. In one case, he said, the trace revealed that one GPU was consistently slower because of a NUMA-awareness problem. Power measurements can also help identify a slow replica; he showed a chart where the slow replica drew less GPU power.
Small reference engines make the architecture easier to inspect
Frye’s suggested starting points for studying implementations are mini-sglang and nano-vllm, simpler versions of engine architectures intended to be easier to read. He described them as useful for people and for coding agents, since their smaller scope makes it easier to fit the relevant code into an agent’s context. He also recommended Aleksa Gordic’s walkthrough of vLLM and tools such as DeepWiki for repository diagrams and queries, while noting that repository documentation can be frozen and is best paired with access to the current code.
The practical implication is that understanding an inference engine does not require beginning with every optimized kernel. A reader can first follow the request across the server, tokenizer, scheduler, model runner, and detokenizer, then ask how batching, caching, and model execution fit together. That architecture-level view helps explain why a deployment can be limited by a scheduler or queue even when its most visible component—the GPU—is not the whole story.


