← All transcripts

What Is an Inference Engine, Anyway? — Charles Frye, Modal Transcript, AI Summary & Key Points

AI Engineer · 3 days ago · Science & Technology · 59:59 · EN

Watch on YouTube

Answer

An inference engine is the performance-sensitive machinery between an inference server's request and response: it tokenizes inputs, schedules and batches work, runs model operations on accelerators, manages state such as the KV cache, detokenizes outputs, and coordinates the surrounding processes.

AI Summary

An inference engine turns external requests into generated tokens through server I/O, preprocessing and tokenization, scheduling, GPU model execution, detokenization, and response handling. The difficult engineering work is achieving high performance and efficient token economics rather than implementing the basic token-in/token-out loop. Workloads differ: interactive chatbots need short latency, background agents allow longer processing times, and document processors prioritize aggregate throughput. Prefill processes input tokens in a high-arithmetic-intensity forward pass, while decode generates output tokens through repeated lower-arithmetic-intensity runs. Scheduling, batching, KV-cache management, GPU execution, host overhead, and observability determine production performance.

Key Points

  • An inference engine can be implemented quickly with PyTorch and Hugging Face Transformers, but making it highly performant and cost-efficient is the challenging part.
  • Inference engines generally use communicating processes for server I/O, tokenization and detokenization, scheduling, and GPU-bound model execution.
  • Inference is a revenue center because people pay for systems that turn tokens and context into new tokens, while training is characterized as a cost center.
  • A museum placard application sends visual input through a vision-language model and produces an art-description placard; its workload sits between the lowest-latency and highest-throughput applications.
  • Chatbot-plus workloads involve an interactive person or agent, tool calls, high prefix reuse, short outputs of tens to hundreds of tokens, and latency tolerance of a few hundred milliseconds at most.
  • Background agents perform similar tasks without direct user interaction; producing a pull request may take a few minutes, while higher-quality software-engineering work may take a few hours.
  • Data processors usually handle relatively small models, transform unstructured inputs such as PDFs or videos into structured data, have low prefix reuse, and prioritize aggregate throughput over latency.
  • Prefill processes input tokens in one forward pass, while decode performs multiple runs that produce one or a few output tokens at a time.

AI in practice

Agents

  • Deep Wiki — Create an explorable architectural reference for inference-engine codebases and answer questions about them. 1 held 58:47

Business ideas

Help AI labs and companies deploy, operate, benchmark, debug, and optimize inference engines that serve language, multimodal, agent, and document-processing workloads. The value comes from improving latency, throughput, correctness, GPU utilization, and queue behavior rather than simply exposing a model API.

For
AI labs and companies that want to own or rethink their inference infrastructure, including teams deploying chatbots, background agents, and document processors.
Solves
A model deployment can become slow or expensive because of queueing, scheduler bottlenecks, GPU underutilization, KV-cache pressure, tokenizer problems, host overhead, correctness regressions, or inconsistent performance across replicas.
  • A museum placard application used a vision-language model to describe what a camera saw as an art piece; under increased traffic, first-token latency and inter-token latency rose because requests queued, and adding replicas reduced congestion.
🔒  Build steps and tools for 1 idea. Unlock

Tools & resources

8 items

HNo. 5105
AIAINotes.us AI product

Hugging Face Tokenizers

In the AINotes directory

A Hugging Face tokenization tool used to convert text into tokens and to detokenize token sequences back into text, including as part of an inference request's processing pipeline.

Mentioned in
1 video
Kind
AI
HNo. 5104
AIAINotes.us AI product

Hugging Face Transformers

In the AINotes directory

A machine-learning library used with PyTorch for basic inference and as a reference implementation of model forward passes.

Mentioned in
1 video
Kind
AI
MNo. 5109
AIAINotes.us AI product

Mini-SGLang

Open source · sgl-project/mini-sglang

Mini-SGLang is a compact implementation of the SGLang large-language-model inference and serving architecture. Its roughly 5,000-line, type-annotated Python codebase is intended as both a usable inference engine and a readable reference for understanding LLM serving systems. The framework serves models through an OpenAI-compatible API or an interactive shell. Its implementation includes radix caching for reusing KV cache across shared prefixes, chunked prefill, overlap scheduling to run CPU scheduling alongside GPU computation, tensor parallelism across multiple GPUs, and optimized kernels using FlashAttention and FlashInfer. Mini-SGLang runs on Linux with NVIDIA CUDA dependencies and supports installation from source or via Docker. The repository documents deployment on one or more GPUs, including models such as Qwen and Llama, and provides offline and online benchmark configurations.

Mentioned in
1 video
Kind
AI
NNo. 5106
AIAINotes.us Tool

Nsight Systems

nvidia.com

Nsight Systems is an NVIDIA performance-analysis tool used to inspect engine processes and GPU work when investigating performance issues. In the cited use, it supports tracing and examining the interaction between inference-engine processes and GPU activity.

Mentioned in
1 video
Kind
Other
PNo. 0240
AIAINotes.us AI product

PyTorch

Open source · pytorch/pytorch

PyTorch is a Python machine-learning library for tensor computation and deep neural networks, with CPU and GPU execution. Its tensor library provides NumPy-like operations, while its tape-based autograd system uses reverse-mode automatic differentiation for differentiable tensor operations and dynamically defined networks. The project also includes neural-network, compilation, multiprocessing, and data-loading components, and supports extensions through Python packages such as NumPy, SciPy, and Cython.

Mentioned in
6 videos
Kind
AI
PNo. 5107
AIAINotes.us AI product

PyTorch Profiler

pytorch.org

PyTorch Profiler is a performance-analysis tool in the PyTorch ecosystem used to inspect inference-engine traces and investigate problems such as NUMA-awareness issues. It is associated with PyTorch, the open-source deep learning framework maintained within the PyTorch Foundation ecosystem.

Mentioned in
1 video
Kind
AI
SNo. 0237
AIAINotes.us AI product

SGLang

Open source · sgl-project/sglang

SGLang is an open-source serving framework and inference engine for large language models and multimodal models, developed by the SGLang project. It provides an alternative serving stack for deployed models, with runtime and kernel optimizations for inference workloads. Its documented capabilities include RadixAttention, a zero-overhead batch scheduler, cache-aware load balancing, structured-output decoding, speculative decoding, disaggregated prefill and decode, model parallelism, and support for GPU and TPU backends. The project also publishes integrations and optimizations for current open models and multimodal, image, and video-generation workloads.

Mentioned in
4 videos
Kind
AI
VNo. 0236
AIAINotes.us AI product

vLLM

Open source · vllm-project/vllm

vLLM is an open-source library and inference-serving engine for deploying large language models on GPUs and other hardware accelerators. Originally developed in UC Berkeley's Sky Computing Lab, it is maintained by a broad community and provides an OpenAI-compatible API server, as well as Anthropic Messages API and gRPC support. Its serving stack manages attention key-value memory with PagedAttention, continuous batching, chunked prefill, prefix caching, streaming generation, and disaggregated prefill, decode, and encode. It also supports speculative decoding, quantization, optimized attention and GEMM/MoE kernels, automatic kernel generation and graph-level transformations, structured outputs, tool calling, multiple decoding algorithms, and distributed inference through tensor, pipeline, data, expert, and context parallelism. The project integrates with Hugging Face models and supports decoder-only, mixture-of-experts, hybrid state-space, multimodal, embedding, retrieval, reward, and classification models. It can run across NVIDIA, AMD, Intel, and other supported accelerators, as well as x86, ARM, and PowerPC CPUs, and can be installed with uv or pip or built from source.

Mentioned in
10 videos
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of What Is an Inference Engine, Anyway? — Charles Frye, Modal — AI Engineer (59:59). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:12 All right. Um, so I'm going to go ahead and get started while people get seated. Um, we got a pretty tight turnaround. A lot of really wonderful workshops at this event. Um, so I want to make sure we have plenty of time. Uh so I prepared for 10 or 15 minutes of questions and I think it'd be best especially given this is a workshop if people just interrupted me as I was going raise your hand and I'll call on you and answer your question.

00:41 Um and so just like let's keep this interactive. Um I think it's it's better that way. Um, and it seems like talked to a couple of people, it seems like we got a lot of folks here who know their who know their inference, uh, know their engines, and so we can get a a sophisticated conversation going. Um, so the title of this talk is what is an inference engine anyway?

01:04 And the goal is to sort of peel back some of the layers that are behind APIs and uh what is go actually going on inside of the code that runs your inference uh when you run it yourself. So roughly you know there tokens go in tokens come out um and a lot of the time you get to avoid thinking about what happens in between. Um there's this really high quality open source software out there for running your uh running your language or multimodal model inference.

01:38 You've got SGA, VLM, Tensor RTLM. Many people write similar things to these in-house. Uh they design these these uh engines that like at a high level just taking in some tokens putting out some tokens is not that hard. Um you can write something like this in using PyTorch and hugging face transformers in like an hour or two especially if you're good with agents.

02:04 Um but making it uh extremely high performance so that you can produce tokens with good tokconomics is the challenging part. Um and that d like sort of drives the design and architecture of these systems. Um so just as like a very quick sketch um before we jump back to the high level and talk about why why we're here. Um the uh these are generally architected as a set of communicating processes.

02:31 One that does IO does HTTP or RPC serving and responds to the outside world. Um and then uh there's a little bit of pre and postprocessing with tokenizers and detokenizers. uh and this wraps uh the sort of core of the system where the real work gets done um in sort of two pieces. One a sort of scheduler that like defines and schedules work to run on the accelerator, usually a GPU.

03:01 Um and then the actual code that runs on the GPU. Um and so this these two components in the bottom right here are where the sort of majority of the performance sensitive work gets done and where the complicated problems are. [snorts] Um so that's where we'll spend the majority of our time uh discussing. Um, all right. So, with that stage set, um, I think, uh, I've been to AI engineer since it was the summit and talked about, uh, running, uh, running inference, uh, every time.

03:39 And the answer to this question gets better and better each time. Uh, the question of like why inference and why you should care about it, why you should think about being uh, an inference engineer who operates inference engines. um our uh fearless leader and uh uh one-man conference army Sean Wang uh tweeted recently uh that everybody in AI infrastructure the boring stuff uh like inference engines not the sexy research stuff um is finally getting filthy rich um and so what uh I think he's pointing to here is there's

04:16 tremendous demand uh for these uh to run inference uh there's demand at the labs there's demand within companies who want to own their own stacks. There's demand to set up and rethink the infrastructure stack and all of this wraps this core inference workload. Um but Sean is also pointing out that a lot of what people have been thinking about up to this point um is sort of the research side, the production of models.

04:44 So um why not, you know, why not think about training? Why not go to the training engines uh talk with your precious time? I think the fundamental reason why this is such a critical place uh for engineering work to get done and why it's such an like cool place to work as an engineer is that in the end training is a cost center and inference is a revenue center.

05:05 Um so when you train models um it turns out um people have not generally been able to make businesses work where they sell model weights to people directly. Um and so you can't really turn a training setup into like a proper self-supporting system. Uh but uh inference once you have a model um is something that people will pay for. People will pay for a system to you know take in their tokens and their context to produce new tokens, produce software for them, whatever.

05:34 Um there's a lot of opportunities to do inference. we have set up where there are a small number of essentially model foundaries who produce foundation models that are then deployed and customized by inference engineers. Um and then even for training inference has uh snuck its way in. A lot of the post-training adaptation of these foundation models requires that you generate samples from the model uh in inference and a lot of the same technologies come to play.

06:04 A lot of the same concerns are present. Um, and then finally, sort of more selfishly as an engineer, there's like right now inference is still very new, even though we're years into this new technology. Um, and so you end up kind of straddling the whole stack from application architecture uh through linear algebra down to like electrons and hardware.

06:24 So it's a an engineer's playground. Um, so we'll see a lot of that as we go. Um I'll maybe in the interest of time since we kind of had to get started very quickly I'll skip over this but I've been working on inference for a couple of years mostly uh at modal um since coming from the sort of training world before that um and helping a lot of teams sort of deploy and optimize their inference on hundreds or thousands of GPUs on the modal platform.

06:48 Um I've also been making a lot of uh inference applications myself. This is a little art piece that I made. Um, we uh uh had it in the uh Legion of Honor just this past week. Um, this art piece takes what it sees, passes it through a vision language model, one of the Quenes running on uh uh served via the SGA inference engine and running on modal. Um, it takes that input and it describes it as though it were an art piece.

07:16 So, it produces one of these little museum placard uh type things. And as everybody knows, you're not art until somebody writes one of these little placards to describe what you've done. Um, so I pointed it I pointed it at the line on the way in. Um, and it came up with a cute little description. Anyway, so this is the kind of, you know, uh, sort of in the middle of the set of inference applications between the lowest latency and the uh, highest throughput applications that get deployed.

07:43 Um so uh I'll start by talking in more detail about those workloads like what workloads land on these inference engines and then jump quickly jump move from that quickly into architecture how these engines are designed some of the key techniques uh that fit within that architecture to make their uh improve their performance. Um and then lastly, a key piece for making these things work in production, observability and uh like debugability of these engines for correctness and for performance.

08:14 All right, so let's start with workloads really quick. Um uh I like to break things down into kind of three pieces. There's chatbot plus, which is like the original killer app of inference, something like chat gypt or clawed code. there's a person waiting on the other side interacting with this inference system and providing additional context. Um the plus uh indicates that not only does it talk to a person, it also probably interacts with external systems via tool calls.

08:38 Um then you have background agents. They often do similar tasks to something like clawed code. Um but they do it in the background, not directly in response to a person. So something like Devon, Ramp inspect, uh OpenClaw or Herbies for a non-coding case of a background agent. Um so uh this is uh characterized by a similar sort of affordances or workloads but a different performance profile.

09:03 And then lastly data processors. Um these generally run relatively small models. They take unstructured data like a PDF and they turn it into structured data. Um they extract some information from it. Um usually the sort of thing you might insert in a database. This is something like the uh reductto document processing platform or something like Fathom that takes a a video and extracts the transcript.

09:30 Um okay so under the hood what do these uh workloads look like? Um you send in some tokens from a client uh that goes to some server somewhere. Um for now we're going to focus on cloud deployment of inference because that's where the majority of it happens right now. Obviously in the future a lot of these inference engines are going to end up running on people's uh machines like in a couple of years.

09:51 Um yeah, go ahead. >> The size of the screen >> brightness of the screen. Um yeah, so the question was about adjusting the brightness. I'm not sure if we can uh fix that. That's what I get for making dark mode slides, but yeah. Um yeah. All right. So a client submits um an input uh and then it's processed by our engine running on a server somewhere.

10:18 Um first we process the input tokens in the prefill phase and then we generate the output tokens in the decode phase. Um, and so we have uh one uh one prefill uh run uh does a forward pass through the model and then multiple decode runs to produce one or a few tokens at a time and then once we're done return them to the user or generate a tool call and interact with an external system.

10:48 Um so with these work these both the external interface and this sort of internal pre-filled decode split defines these uh the sort of metrics that you care about for a workload that sort of tells you how to deploy your inference engine and and how to measure its success and performance. So you have something like queries per second in and queries per second produced.

11:09 Um this one pretty tricky to think about uh under control of your users. um uh you probably have a better time thinking of it uh at the engine level thinking of how many queries per second you can handle per replica of your inference engine. Um and then you have aggregate demand coming from users. Uh then with those queries you want to look at how many tokens there are per query um input and output aka prefill and decode.

11:34 Um so this one is tricky because it's not only under control of users but it's also under the control of a model. Uh so the determination of when to finish processing a query is not under your deterministic control. It's under the control of the model inside the inference engine. Um so this is something not both of those sources of nondeterminism means you need to kind of like measure this benchmark it to get it.

11:58 Um and then uh when we do our processing we uh we produce some intermediate artifacts that we want to cache in the KV cache. Uh and then the degree to which that cache is reusable is a key component of the definition of our workload. Um this is also mostly under control of users but you can get like depending on the workload there's actually like pretty uh relatively fixed patterns of of prefix reuse.

12:23 Um uh and then finally coming from the application layer um you have some uh latency budget some SLO of like how quickly you want to return tokens. That could be how quickly you finish something like the prefill phase and have your first tokens ready to send to users. Or it can be how quickly you can generate each set of tokens during your decodes and get them back to users.

12:47 So that's your time to first token and your time per output token or inter token latency. Uh so this is going to be the like primary constraint that you operate under. Okay. Um, and we can map those back onto our application archetypes. Uh, like chatbot plus background agent, really high prefix reuse. There's some interactivity going on whether that's a person or an agent.

13:10 Um, the chatbot plus generally has relatively short decodes, relatively short outputs of tens to hundreds of tokens. Um, and a very tight latency tolerance on that. Like a person uh wants to see something in a few hundred milliseconds at the very most, or they get bored, uh, they get angry, they turn away. Um, with a background agent, you maybe have a few minutes to produce a PR, right?

13:30 Um, if it's as high quality as an engineer, maybe as like what would be produced by a good software engineer, maybe you have a few hours. Um, and so there's not as tight of a of a budget on your milliseconds of network overhead or whatever. Um, and then with a data processor, you generally have very low prefix reuse. Um, you are seeing different documents each time.

13:49 There's maybe a system prompt, but it's generally small relative to the other inputs. Um, and then you have a very short decode. you're producing some structured object that's generally short relative to the input. Um the structured version is shorter than the unstructured version. Um and then the demand here is less on latency generally and more on aggregate throughput because you're generally like ripping through a database backfill type job in a lot of these.

14:17 All right. Um and so maybe just want to call out uh like now that we've looked through this you could see that like even though at the application layer this feels like kind of one it's one single workload from the perspective of the user once you start drilling down into what the engine is doing it really feels like there's two subworkloads going on.

14:37 there's the like workload to do your prefills to do the processing of input prompts. Um, and then there's work to do the decodes to produce the outputs. Um, and these uh they sort of operate semi-independently inside the engine. Um, and so need like a sort of like different perspective and they follow often like different code paths. Um, and they're sort of defined by the fact that the per user in the decode phase, you do you're doing far less arithmetic, far less math per uh per decode per user than you would do in

15:11 the prefill phase, which is a lot like much higher arithmetic intensity, a lot more uh uh a lot more math per request. Um, and so that's your fundamental constraint, but also the sort of space in which you operate in configuring an engine and then like adding features or um, adjusting the engine behavior to serve these two things. Okay, so I wanted to make sure we're all on the same page about some of the core stuff about what these engines are actually doing before we peel open uh, the insides and start looking at

15:42 them. So what is the architecture of an inference engine? Um so um this is a little little diagram showing at a high level what an inference engine does. So the goal of an inference server is to take an HTTP request or a gRPC request and turn it into an HTTP response to the client. Um and so it does that by passing through all these things down here.

16:06 Um so the idea in this diagram um is that if you follow any set of arrows you get the same thing. So uh you can sort of peel back a layer of the arrows and look inside and say all right so what what is the the inference server does this high level thing what's the next level down the inference engine is the sort of like interesting part of it is the is the part that is not doing the HTTP serving because we aren't in one of those situations where HTTP serving is the hard part is your bottleneck it's not like a proxy um

16:35 it's uh it's a this is a case where you're handling tens hundreds maybe thousands at most of requests and responses per replica and so general purpose hardware just you know kind of eats this for lunch. Um so the real interesting stuff is all at the layer below this. Um so at the top is your like your server IO for HTTP requests. At the bottom is the actual engine and it's like collection of processes.

17:00 Um and you can kind of break this out into three pieces. The request comes in expressed using like sort of human friendly um or external representations. So you have Unicode strings, you have PNG images, you have MP4 videos or WebRTC streams. Um, and these need to be pre-processed before they can get passed to the actual inference. So they get turned into collections of tokens.

17:24 um which then get mapped into tensors on the GPU and then we can actually do the inference part of the inference engine which takes you know in the simplest case it takes a tensor with K entries in it and gives you a tensor with K plus one entries in it and these are machine learning models that predict the next token they guess what would the next entry in this big tensor look like um based off of the data I've seen in training whether that's data from the internet or data from what a helpful agent would do or RL data

17:58 um whatever that you learn to predict this next tensor entry. So that's the core right there. Um and that is what is wrapped by the rest of the inference engine. Um so then we have post-processing on the side here. So post-processing turns that new tensor into a response. Uh and that's generally done by this detokenizer that goes back to the sort of like token space first and then uh and then back to our um back to the human and external system friendly representation in the response so it can be returned to the user.

18:31 question. >> Could you explain the difference between the inference server and the inference engine and where does >> Great question. Yeah. So the question was what is the difference between an inference server and an inference engine and where does something like VLM lie? Um so inference server uh means that you've wrapped some like input output around the core operations of like taking requests and turning them into responses.

18:59 Uh so VLM has both of these like VLM operates as a an HTTP or gRPC server. Um and then you the engine layer is sort of like part of the VLM or suing SDK as like VLM. Um so most deployments are look like are deployed as a server with like a single request single response kind of interface um as depicted here. Um yeah, does that answer your question? Yeah.

19:30 Um cool. Um I said this earlier but some people have come in since like please interrupt raise your hand with questions if I'm like if there's anything confusing or if you'd want to dive deeper. Yeah. >> Rom. >> Yeah. So the question from Rustom is is this like a stateful or stateless uh workload? So the answer is it is stateless [snorts] if you don't care about performance.

20:02 Um it's stateful once you do care about performance which is pretty much every case. So and the primary state that you accumulate over time for requests is the KV cache of like previously computed things for for a previous request attached to some notion of session. um and that handling things like sessions actually lives outside of the inference server/inference engine in general.

20:25 So the an individual engine replica can track this KV cache but is up to the wrapping layer which might be something like Nvidia Dynamo or LMD to handle like defining what a session means. So it's like the clients in that architecture sort of define the the sessions. Um and that is where a lot of the trickiness comes in. Um, so we won't talk too much about that because it's kind of outside the there's enough to talk about with engines um already, but yeah, uh, come find me afterwards um, if you want to talk about that

20:55 layer. Um, right. So this diagram, we also talked about it previously, but I just want to return to it now that we've seen a little bit more. So this outer layer is our server IO process. This um, is this is specifically SGA's architecture. There's some slight differences with VLM, but this server IO process communicates with the outside world and then with our pre-processing and post-processing.

21:19 Those things communicate with auler process um which is the one that defines and schedules work onto the GPU, handles things like queuing. Um and then the not ready for that yet. Um, and then the one or more processes that are sort of closely tied to the GPU are the ones that do the real work of doing the inference. Um, so the yeah, so the key things to pull out here is that if you are an old school ML person who's been in this world for a long time, the model forward passes layer is the one that you're familiar with.

21:53 This is where the vast majority of the sort of like PyTorch uh and similar code shows up. Um and this is the pl like the GPU or the accelerator in general is the component that costs the most um and is capable of the most work like paflop per second scale operations. So this is where the like hardcore shave a shave a microcond performance engineering is focused.

22:19 Uh but critically thisuler process out here or it's equivalent in another architecture is your is the bottleneck to that component. Um and so this is it's generally operates um it is literally singlethreaded in implementations we'll talk about but is also like semantically it's single threaded. It's this one thing that is defining somebody has to define the work that happens on the on the GPU.

22:44 somebody has to manage a bunch of resources there and there's some some fiddly like tricky things to do concurrently. And so this becomes this uh like semantically singlethreaded uh like point of control that can potentially despite the fact that it's not doing much it's not doing as many operations uh as the work on the GPU it actually becomes your bottleneck.

23:09 It's the sort of gatekeeper to the GPU. And so this is the other place where the majority of like performance engineering and care needs to go in the design and operation of an inference engine. Um so we'll see more on that as we go but but that's like the high level takeaway about these things. >> So is the solution to that bottleneck problem? Uh yeah so the question was is the solution to that bottleneck to you know operate the schedulers in parallel or um uh or to improve some of the external components that consume

23:47 from it. I think the current state of affairs is that this is not reached the like that you can operate these server processes in Python and you just have to be faster on the host side than the GPU side and there's not the like complexity increase that you get from trying to manage like locks or coordination on the GPU resources is much that complex lexity increase is much greater than the gain um because even if you were able to get it down to a microscond the work wouldn't start happening until the previous work on

24:28 the GPU had finished and that's you know depends on the particular workload but you're talking generally tens of milliseconds to hundreds of milliseconds and so the like um yeah there are general most of the time you can get around it we we'll talk a little bit more about this so maybe ask again once we get to if if that wasn't satisfying. Um but yeah, great question.

24:53 More good questions, please. Okay. Um so then just looking at that was like the architecture diagram that you think of as your like engineering. Um let's look at the like life of a request in this architecture. So the um the server like receives an input from the outside world. It says okay, we've got a request. um it sends it to the tokenizer to prepare or pre-process this request that then goes into the scheduler that decides like okay how many resources is this thing going to consume what resources are available um

25:22 and then that runs this core loop of creating um batches of requests to go in the model runners. So generally most of the accelerators like benefit from operating on multiple requests in parallel. Um so especially GPUs but also TPUs and and many others. So you want to collect a bunch of requests together and run them. Um if you're a database person you might imagine like I'm g I'm about to run a sequential scan so I might as well grab like three or four queries that are all going to like read this entire table and do

25:57 all of their um like you know do all their queries at once. Um because the hard part is that like sequential scan of loading the entire table. Um so there's a similar thing going on here. Um uh so this core loop operates over and over again start a batch and outputs the like uh the tokens and the probability distribution over tokens the the logits. Um that comes from the the model runner or runners.

26:20 Um so that operates um sort of continually. Uh and then as tokens are ready uh they can be passed to the detokenizer to be like turned back into a proper response to the outside world and then sent back to the server. So yeah, this is what's interesting about this stuff is that it is simple relative to what's going on inside databases, file systems, uh like web servers.

26:49 Um, partly because it's fairly new, but partly because there's just this core bit at the middle that you just want to operate at maximum speed. Yeah. Question. >> The question was, is there communication between the tokenizer and the detokenizer? Um, to my knowledge, there is no direct communication between the tokenizer and the detokenizer in SG lang.

27:12 Um, in VLM, they actually live inside the same process. the um manager, the engine manager. And so like it'll be a little harder to determine if there's any communication there. Um but in general, they actually don't need to communicate with each other. They just need to operate essentially the same um system for mapping from like inputs to uh tokens or tokens to outputs.

27:34 Yeah. Did that answer your question? >> Yeah. I'm not sure >> Right. So the the question was about basically maintaining coherence between the tokenizer and the detokenizer and how would you handle maybe like ambiguous um like ambiguous inputs. Um the answer is that yeah they they run the same like underlying li like underlying libraries like the usually like the tokenizers library from hugging face.

28:25 Um and the I would say the tokenization part is the one that is harder than the detokenization part. the detoization part is like a little bit more entirely under under the control of the person implementing the the detokenizer. Um yeah. Um >> software >> yes. So the question was does the software allow the tokenizer and detokenizer to always make the same decisions?

28:55 I would say yes. I'm trying I'm like trying to make sure that I'm not overpromising here. Um but I don't think you need communication between them to ensure that the same decisions are made. Yeah. Question. Sorry, the person. Yeah. >> Yes. >> Right. So first pointing out that the tokenizer can be um parallelized and generally is like multiple processes um and yet you still see longer time to first token uh if you have longer input sequences.

29:35 So why is that? Um so the answer is that processing a longer sequence of tokens takes longer inside of this core loop. Um so it's the like primary driver. Um if you have a 100,000 tokens that will take let's call it 10 times 10,000 tokens. Um once you get to a certain uh size it's basically linear in the input sequence size. Um then in a like running inference engine there's also effects from like you break up the input tokens into multiple pieces and run them as separate batches.

30:05 So you can also get additional if there's like queuing and congestion that adds more to the um uh processing of a uh of a long batch um than to a short batch. You experience the queueing delay multiple times possibly. Um but maybe the one underlying um point to clarify like producing the first token means producing the first output token not the tokenization process.

30:30 Um and then finally not a part of the question but something I wanted to point out. It is interesting that you generally have multiple deto tokenizer processes um and one dtokenizer in the sgang architecture at least. Um and the answer there is that input tokens are come in at like a much higher rate than output tokens are produced. Um and that's related to the fact that prefill can operate at much like higher rates of total tokens per second than decode can in general.

30:57 Um, and so you there isn't as much pressure on the detokenizer component of the system to operate like at the as quickly. Um, and it's also in some it's bottlenecking your response to the user, but it's not bottlenecking the core model forward pass resource. So it's also like off the hot path a little bit. Um, yeah. >> Yeah. So the question was how much of the latency comes from this part in the tokenizer versus this part in the core loop.

31:32 This is going to depend on your workload. Um I would say there was a nice thing from uh who was that Cruso I think um that pointed out that the existing tokenizers generally don't show up as a bottleneck if you have models that are let's call it like 20 billion parameters total or or above. um and that and if you but if you are operating a smaller model and you're operating on very large input contexts then it can show up as a as a bottleneck.

32:05 So that's maybe a good rough guide. I would say like I actually have not had to think about the tokenizer latency very much. It's generally been prefills and substantially decodes that are your the like latency bottleneck. And so decodes if you're producing 1,000 output tokens a second, each one is like one millisecond on average. Um and the um but up to like 10 milliseconds, 20 milliseconds um or even more if you have a um lower decode throughput like a large that you're running at lower interactivity.

32:40 And so that very the tokenizers are definitely millisecond scale or less and only happens once um per request. Yeah. Also notable that nobody talks about tokenizer caching. People talk about KV caching but not tokenizer caching even though you could do it. Um I'm going to keep going just to make sure we get through more stuff but these have been great questions and so if it's still relevant please ask.

33:01 Um another important question to ask wait why are we doing multipprocessing? The key reason is they want we want separate threads of control for each subcomponent for our server IO for our tokenization and detokenization for each model runner. Um we want them to be able to sort of operate as independently of each other as possible. Um they are many of them the tokenizer and the detokenizer are doing relatively computensive work relative to most of what happens maybe in a in a in a web server.

33:29 Um the model runner is um like interfacing with the GPU or accelerator. It interacts with that in a way that's kind of similar to the way you would interact with a nick um in like a web server, which is to say you offload work onto it. Um and uh so that one you you're giving it its own event loop. Maybe in theory that could be part of like a broader event loop.

33:52 Um uh but it's uh certainly like cleaner to have these separate threads of control and optimize their uh you know optimize their latencies and there's not sufficient uh interprocess communication to require them to like say live in the same thread. Um but why not threads uh for each of these instead of processes um and one of the fundamental reasons is that Python is single threaded per process.

34:18 This has recently started to change. there's still a lot of work for compatibility and performance to be done there. Um so that uh because Python has a global interpreter lock for all the threads in the same process, you would like frequently block on that. Um so it's so that is another reason to sort of split these up. Um but that just uh begs the question um uh or I guess begets the question why Python um and the answer is Python has excellent model or GPU support.

34:45 Uh so it's like easy to run code on accelerators with Python. So specifically for that model um uh model for pass component um and so wait but why wait why does Python have excellent support here? Partly it's that it's a great language for researchers because it's like ease of use and these uh models are still coming out of a sort of researchy data sciency environment.

35:06 But there's a deeper reason which is that Python has really has kept like really deep C and C++ interop as one of its core design principles and that allows for like much easier communication uh with these accelerators and uh with like accelerated code in general um than uh like a maybe equally ergonomic language like JavaScript. Um but why not like rewrite it in Rust?

35:30 Um the model worker processes again they're just trying to like tell the GPU what to do. So you could like the code on the GPU is written in a compiled language and and like carefully optimized like to within an inch of its life. Um but the model worker processes are just organizing that and there's a few tricks like CUDA graphs um that like really just pull all the work out of the worker processes and into sort of fast stuff running on the GPU.

35:57 Um but there is theuler on the other hand. So theuler just needs to stay faster than that work. Um, but as GPUs get faster and as the amount of work on the GPU goes down, the pressure on theuler component goes up. Um, so this is actually like a pretty good target for a rewrite it and rust type moment. Um, because this part is host side. It has to manipulate a small number of GPU resources, but it doesn't necessarily have to know stuff about models and pietorrch.

36:22 Um, and so it's like a thinner rewrite. It's like specific per engine. um and uh uh and host side only. So like in the sort of arc of future inference engines, I'd expect it to be a big target um making that faster now that the the sort of iterative development speed is slowing down a bit. Okay, so now let's look inside of the like what what is going on inside of these components.

36:50 Um um mostly mostly focused on the um the model forward pass piece. So at the very we're going sort of like inside out. So at the very bottom there are kernels that run operations for as part of the model architecture in general. Um these are mostly deferred by the engine which is to say like the engine doesn't implement all of the kernels for every operation it needs to support.

37:13 Um it has backends that implement the uh these kernels and it consumes them from external libraries of kernels. This is one reason why there isn't much differentiation between the engines on sort of raw GPU performance. Um, it's mostly about how good are you at getting out of the way of the GPU. Um, there are a few special kernels per engine like tree attention for treebased speculators and SG lang.

37:37 Um, and so there's like little bits where an engine will have a specific kernel, but like if it's useful, if it's popular, it ends up in some kernel libraries so it can be used in lots of places and then they consume it from an external library. Um then uh the kernels are selected by this sort of model forward pass code that does control flow of and sort of like selection of which kernels to operate.

38:02 Um and this is this is implemented in each engine in general and it's implemented per model like if you want to add support for a new model architecture this needs to be implemented in each engine. um there's generally a reference implementation that comes from a library like transformers um but that is insufficiently optimized um the like the setup there.

38:22 So there's um I guess in suang and vlm there is now a path where you can just use that transformers implementation and they'll maybe do a little bit tiny bit of surgery to see what they can automatically kernel fuse or automatically optimize or rewrite. Um but in general for most things people want to run you're going to go through this hand handwritten as handwritten as software is these days uh path for um like selecting which kernels to run.

38:48 Um then uh wrapping those model forward passes is our batch construction to select the inputs to that model forward pass. Um this is kind of one of the primary sources of special sauce per engine. you need to set up this um this sort of like queuing setup. Um so you might want to be able to for example not operate separate prefill and decode cues but operate one sort of queue where you can mix prefill and decodes together.

39:14 Um that might have some performance implications. So you might also want to be able to like dynamically switch whether you're running pure prefills, pure decodes or mixes. Um and then you want to do this I think in the ideal operation of an inference server there is no queuing. Um, so as somebody who operates these things, I like I treat queuing as a problem.

39:33 Um, but the engine should be able to kind of do it. Um, or may maybe it's too extreme to say that queuing is a problem, but um, you don't want this to be a like yeah a growing queue. Um, you want it to be flushed as quickly as possible. Um, and then lastly wrapping that you have uh detokenization and tokenization. Um, this is like mostly boring I'll say.

39:56 um it's like solved in the existing tokenization libraries outside of the engine. I think the exception to this being multimodal inputs and outputs um we now have pretty good open weights multimodal input models. Um and if you take a look at some traces you'll see that there's like still many operate opportunities for optimization on this path. Um, tokenization of an image is just like a harder problem than tokenization of a um of a string of Unicode bytes like conditioned on already having a a tokenizer.

40:28 There's more more work to be done. You might need like ffmpeg for instance to like pull out raw frames to go in to something that turns it into a tensor. Um, and so yeah, so that is maybe the most interesting part there. Um the one way to sort of conceptualize the structure around the model forward codes is that the model forward code is um like going to call into a couple of different types of libraries or directly sort of write its own code to run on the GPU.

41:00 Um so a lot of it goes through like PyTorch code which has its own kernel libraries and and can run things on the GPU. Um then there are like specific kernel libraries that you would draw from that aren't in the engine. And then finally the engine has its own uh set of like kernels that can directly um you know run code on the GPU. Um so picking that apart a little bit the um uh the kernels that exist in their backends.

41:29 The core of language modeling architectures these days look something like this. Um there's a sequence of layers uh that you take an input, you pass it, do some cross token computation on the sequence, then you do some per token computation to enrich the sequence. Um and then you merge that back in with your representation and you iterate this layer after layer.

41:50 Um and so this splits nicely into the two kind of classes of like really critical fast kernels that are available especially via the backends. Um so the cross token computation is your attention or linear attention. Um there's a bunch of different backends for attention. The flash attention kernels from tree and jaw and others. Uh flash and fur from Nvidia.

42:15 Uh cutless also from Nvidia. Um the Triton uh implementations of these kernels. That's OpenAI's kernel authoring library. Um, the TRTLM, the engine has also split out its kernels and they're available um, as as backends for attention. I think you'll see many of the same names that show up as attention backends also show up as backends for your matrix multiplications inside of your um, your uh, per token computation which is your MLP layer.

42:48 It's where the biggest uh, gems are. Uh it's also in some architectures it's this is a dense MLP a classic neural network. In other architectures it's these sparse mixture of experts which basically just says don't send every token to the same neural network. Per token decide which smaller neural network to send it to. Um so the um the dense MLPs there's a there's a nice uh library for this from uh deepseek deep gem that gets used a lot in sglang as kind of like the default.

43:17 Um in VLM there's some quantization focused backends for this the Marlin and bits and bytes backends on the uh for the group gem you can split this out into first like there we have this like routing or communication problem of where do tokens go and how do I bring stuff together at the end. Um so that can be done with a separate set of kernels like the uh for this all to all communication and routing like the another one from deepseek DP.

43:45 Um and that could be sort of factored out from the like matrix multiplication step. Um that had that can have performance penalties and so you can also do that as like one big kernel um a mega kernel uh which is what is done in the deepseek v4 architecture. Um okay so these uh with VLM and SGA lang these are things that you can sort of set and control in SG lang it's like kind of pushed up into your face in the form of configuration at start with smart defaults VLM is a lot more focused on the like smart defaults side

44:23 and then you sort of like patch change things more with um environment variables. Um, at least last I checked, you know, these things change every like few days. Um, but yes, >> questions are like who decides how much? >> Yeah. So the question was basically about what's happening with concurrency with really large numbers of requests. Um so that goes back so that's uh much higher level than the kernels and backends.

45:08 Um that's happening sort of at our like batch construction and scheduler level. So basically the scheduler is going to do what the best that it can to decide like um which requests to run emp generally emphasizing requests that already have KV cache allocated to them um because then you don't need to recomputee um the if you need to evict something from the cache to make space for another request um yeah that will generally get like put out to CPU memory and then disk um and I would say Yeah, to my point earlier about

45:45 queuing, I think of that as essentially like a failure state of the engine. Um, that we like I've deployed my engine improperly. If I get to the case where I have KV cache thrashing or like things being flushed to disk in the same way that there's like swap memory in your operating system, like you can swap memory to disk, but at that point your MacBook becomes unusable, right?

46:04 Um, so it's it's a useful feature to have, but almost like um so that you don't go hard down when you have a serving problem. Um yeah, how is mobile? >> Yeah. So the question was about like serverless deployments and cold starts. I'm not going to talk about that this much because I promised Swix I wouldn't do a vendor talk. Um but please talk to me afterwards.

46:37 There's a blog post called truly serverless GPUs that talks about how we solve that problem. Memory snapshotting, GPU memory snapshotting, custom file system, bunch of pieces to make the cold starts faster. Um, yes. Um, cool. All right. So, let's talk about some of the key techniques that you need uh to implement to in inside an inference engine to make sure like uh your performance is good.

47:02 Um so attention is fundamentally quadratic at least like the uh loadbearing part of the attention in in every architecture that's popular. Um and the um so the solution to that is to store things um after you've computed them and you do like a little space-time trade-off for linear computational complexity in exchange for linear storage. um that runs into problems with how much can you store in GPU memory because it's actually often slower to store this somewhere else uh and then load it back in as opposed to

47:35 recomputing it. GPUs are fast. Um so you want to be efficient with your use of GPU memory. Um and then you often have requests that share prefixes. So for example like all these commandments start with thou shalt not. Um, and so you can reuse the computation of everywhere up to here. Um, as soon as a single token differs, you no longer can share the prefix.

47:58 So it's kind of this like tree or try shaped structure. Um, but you want to be able to make use of that, especially if you have like a giant system prompt for like an agent system that's tens of thousands of tokens. Um, it's like quite nice to be able to share that across users. Um and so the solution to this is like KV caching with pages or or radices.

48:18 Uh so you store the information in this um in like a a page cache data structure like you'd have in an operating system and then you just need to reconstruct these tensors sort of on the fly to pass them to attention kernels. So that that job is now it used to be the job of the engine to solve that problem. That's now incorporated into the attention kernels like flash attention 4.

48:40 something that our team worked on the um page size one um and sort of like RAD XC form of KV caching you can see on the right um and uh yeah so this is this has moved a little bit I put this in here because it used to be like a defining feature of the engine but now it's kind of actually part of the kernels and really it's the management of KV cache capacity and um and the like layout of the KB cache that's the problem of the engine um another problem I already talked about how this the scheduler can be this um host

49:14 side bottleneck. What you want when you're running stuff on the GPU like both at the sort of level and then inside of models is that you want that just like at every moment some kernel is running on the GPU and things happen on the CPU to decide what that uh should run on the GPU. But they shouldn't ever block uh progress. But it's very easy to accidentally write some kind of host device sync.

49:36 There does have to be information passed between them. the GPU can't just like send the response directly to the user. Um, it wouldn't be that good at it if it did run an HTTP server. Um, and so you you do need this CPU work. You just want to make sure it doesn't block. Um, one of the key techniques for this is CUDA capture. So you take all of these kernels that get run in the forward pass of the model and you track all of them and you turn that into one big data structure a DAG of operations of like this kernel runs

50:03 and and like you know mutates the data at this pointer. Once that is done this kernel can now run to mutate what was present at that same pointer and mix it with data from another pointer. So this big DAG of C of of uh operations is your CUDA graph. And this um allows you to go from launching like each one of these GPU kernels having a little bit of CPU work to decide what to run, how to run it, uh what you know which pointer to put um you turn it into just one CPU side work to decide which graph to launch and then all

50:38 of it gets launched from that single graph which is what's depicted in this um uh peretto trace. Uh so this allows you to relax quite a bit about what is going on in the model uh forward pass component at least. Um another problem decode here is sequential um and every time you run it you have to load like either all of the parameters of the model or all the active parameters in ane architecture um and you need to do that over and over again um by default for each token.

51:08 So you're like on behalf of your users, you might be loading like a terabyte of data or 100 gigabytes of data per token and you want to produce a token every couple milliseconds. Um so now you're looking at that's pabyte per second scale uh memory bandwidth. That's uh pretty hard to achieve. Um so one solution is key technique for solving this is speculative decoding where you take um a separate model.

51:36 The speculator guess what the next couple of tokens will be and then pass it through the target model. Um and then the target model sort of grades all of those in parallel. Um and if you apply the right sampling techniques, rejection sampling, um then you can uh guarantee that uh the output from the target model is the same as you would have gotten if you'd run it sequentially up to numeric um the same output uh but you've now run it in parallel.

52:03 You sort of turn your decodes into tiny prefills which is nice. Um and so the um I'll skip over a little any details about the algorithm. There's a great we wrote a blog post about this uh just this past week or two. The key reason why this is really important is that as you improve the quality of the speculator, as you improve the number of tokens it can guess in a row, um you actually get basically a linear speed up in your decode throughput.

52:29 So like many other optimizations that you can do on your inference engine are like the it's like a game of inches of like a percent 5% 10% like each time and you have to do the like very like difficult performance engineering for each of those you know percentage points. Um but then on the other hand you could sit down and train a better speculator model um and to first approximation throw data and compute at it um and then get a 2x speed up or a 4x speed up or an 8x speed up um on top of the the the baseline.

53:03 Um so this is has emerged as one of the like key techniques to allow you to operate say like a thousand tokens per second on um on Nvidia hardware without have for actually fairly large models uh without having to go to um like a a novel accelerator like an LPU or Cabris. Um I'm actually going to there's a lot of parallelism techniques but in the interest of time I'm going to kind of skip over this one.

53:31 um you want to be able to split work onto accelerators. Um a lot of this comes from the underlying kernels. You need good kernels that can operate with this parallelism. But within the engine in the sort of model forward passcode which is written on a sort of per engine basis um you need to do all of the communications around that. Um and so the communications is where a lot of the problems come in in these uh parallelism uh techniques.

53:59 Um all right. So, uh, observability with our last couple minutes here. Um, inference services like all services will have some bugs in them. Um, you have everything from application level bugs from models misbehaving. Um, uh, which is really not the engine's problem. It's the app developer's problem, but they'll show up on your dashboard sometimes. Uh, model quality bugs.

54:21 These are the problem of the engine. Are you actually doing your inference correctly? When you applied this performance optimization, did you actually guarantee it didn't change model behavior? Um, one of the most common ones here is actually on model release. Tokenizers are often slightly bugged. They both tokenize and apply chat templates and there's often some like uh some issues there.

54:37 Um, they're beaten down in a couple weeks after model release, but it's something you always have to be like on guard for. Um, and then performance bugs or engine performance issues. The kind of the worst ones are regressions over time. as you operate a replica for an extended period, you might see performance degradation from issues that only arise once the thing has been running for an extended period.

54:58 Um then you also might observe, especially in a heterogeneous cloud deployment environment, differences across replicas in their performance um that then need to be investigated. So the goal with these as with all kinds of like production debugging is observability. Um if you can log enough information uh to debug just from those logs, then you don't have to spin up a development server and take a bunch of extra time.

55:22 Um and so this runs from application level stuff um to I think focusing on the engine problems, model quality. You'll you want to be able you want to be good at running evals for model quality as part of your deployment. Um so you want to hit your actual deployment with like evaluation benchmarking type scripts. Um there's scripts in SUAN, there scripts in VLM, there's external benchmarking tools.

55:45 Um take those uh calculate some numbers, log traces somewhere, log as much data as you can about those so you can debug later. Um and then in prod log traces and try to get feedback from users on correctness that you can thread through all the way back to those requests. Um for tokenizer bugs, hot tip, log the token IDs. Set your engine up so that the token IDs are logged as well.

56:07 Um and then finally for performance, performance bugs can show up anywhere. So the more metrics that you log uh the better like more metrics than you think you need um because you never know which crossorrelation will re will make the bug super easy to discover instead of super hard. Um so rather than going through those I'm going to talk a little bit about this.

56:28 This is an endpoint dashboard on modal and what like how you might read it. Um I was operating that um the uh art project uh that I showed at the beginning and I just actually spun up like three different replicas of it uh to see what would happen like under increased load. Um so there was a traffic spike which you can see down here that turns into increasing uh first time to first token as things start to cue uh on the prefill side and and don't enter prefill um until uh the things have finished.

56:58 uh and then also shows up as well as inter token latency as uh decodes get first get a little bit slower and then a lot slower with this queuing as there's sort of contention on the underlying GPU. Um and so the solution in this case with going back to the question about serverless deployments is this is detected as like increased load. You spin up new replicas.

57:20 Those replicas come online. They start taking on some of this request load um and then the um like the amount cued um uh decreases back down to baseline and things are um you your congestion is solved. So I think um that's the like primary way. I think in many cases, you know, we scale up additional resources to solve congestion and queuing. Um, so yeah.

57:44 Um, when you need to go deeper, I'll skip over this. Um, but there's the NSI systems or torch profiler is one of your key tools here. It allows you to see the whole engine at once. Um, allows you to see all the processes that are running, all the work that's going on on the GPU. Um, and pick out issues like, oh, this GPU is running uh like everything's slower than the others.

58:07 turned out to be a NUMA awareness problem. Um, okay. So, uh, closing out here, um, if you want to learn more, mini SGA and Nano VLM are a really great place to start. These are these simple implementations of the engine architecture, um, by the two teams. They're great for humans. Um, they're incredible for agents. They can fit all the context in, you know, in their context window.

58:31 Uh, Alexa Gordic did a great walkthrough blog of VLM. uh and uh share these slides later. Um also work through that in a notebook so you can play with it and then actually uh you can point your coding agent at the coding base and ask questions. Cognition al also operates this thing called deep wiki which basically crawls these a bunch of repos including sing and vlm and writes nice architecture diagrams and makes this like wiki style um page where you can sort of read and ask and then ask their um agent questions which

59:04 is a nice thing to pair with your own agent setup. Um so that's everything. I'm out of time. So I'll just quickly say if you're interested in deploying and owning your own inference, we we have a a new product on modal to help you like get started with deployment really easily. Get optim get an initial baseline of reasonably optimized uh inference engine deployment that you can work off of with modal endpoints.

59:27 Um so check that out and if you're interested in working on deploying inference for hundreds of of teams, uh check out modal.j jobs. Thank you. >> [applause] [music]