← All transcripts

Kubernetes GPU Autoscaling: Why Scale to Zero Costs You 10 Minutes Per Request Transcript, AI Summary & Key Points

DevOps & AI Toolkit · 3 days ago · Science & Technology · 36:52 · EN

Watch on YouTube

AI Summary

Kubernetes GPU inference autoscaling trades idle GPU cost against cold-start latency. Scaling a model to zero removes the pod and, after roughly 10 minutes, the GPU node, but a later request takes over 10 minutes to produce its first token instead of a few seconds. The cold start includes about 1 minute for KEDA to react, 1.25 minutes to boot a GPU node and driver, 4 minutes to pull a 9 GB container image, and about 5 minutes to load and initialize the model. A persistent volume containing 16 GB of model data removes the model download but improves the cold start by under 1 minute because network storage, Python and CUDA startup, KV-cache construction, and CUDA graph compilation still consume most of the time. A placeholder response can make clients retry quickly, but counting in-flight concurrency then prevents KEDA from detecting demand; request rate correctly counts refused arrivals and wakes the model. Scaling from one replica to two also needs advance planning: queue depth detects saturation only after a replica is needed, while a second GPU replica takes about 12 minutes to become ready. Inference autoscaling therefore requires forecasting demand or accepting slow responses, and scaling signals should reflect engine saturation rather than only arriving request counts.

Key Points

  • Idle GPU cost — deleting an inference pod does not reduce the bill while its GPU node remains allocated; the cluster autoscaler later removes the empty node, reducing GPU cost to approximately zero.
  • Scale to zero — an HTTP interceptor catches requests when no model pod exists, holds or answers the request, and supplies the signal that allows KEDA to scale from zero.
  • Scale-to-zero delay — the cluster autoscaler takes about 10 minutes to reclaim an idle GPU node, creating roughly 10 idle GPU minutes every time the model goes quiet.
  • Cold start — the same question that previously took a few seconds takes over 10 minutes to produce its first token after the GPU node and model have been removed.
  • Cold-start phases — KEDA reacts in about 1 second, GPU node and driver startup takes about 1.25 minutes, the 9 GB container image takes about 4 minutes to pull, and model loading takes about 5 minutes.
  • Persistent-volume cache — caching 16 GB of model data removes the model download but buys back under 1 minute, or about 8% of the 10-minute wall.
  • Persistent-volume explanation — reading the model from network storage takes about 1.5 minutes, while Python and CUDA startup takes about 1.25 minutes and KV-cache construction plus CUDA graph compilation takes another 1.25 minutes.
  • Image optimization — a registry in the same region, a slimmer runtime image, or pre-pooling the image on nodes can reduce the container-image portion of the cold start.

Tools & resources

3 items

HNo. 0412
AIAINotes.us AI product

Hugging Face

Open source · huggingface

Hugging Face is an AI and machine-learning collaboration platform that hosts and distributes models, datasets, and applications. Its Hub supports public model, dataset, and application repositories across text, image, video, audio, and 3D workloads, while its open-source stack includes Transformers, Diffusers, Safetensors, Tokenizers, TRL, Transformers.js, smolagents, and PEFT. The platform provides a unified API for accessing models from AI providers, GPU-based compute and Inference Endpoints, and paid team and enterprise features such as private datasets, access controls, single sign-on, audit logs, and dedicated support.

Mentioned in
8 videos
Kind
AI
KNo. 4569
AIAINotes.us Tool

KEDA

keda.sh

KEDA (Kubernetes Event-driven Autoscaling) is a lightweight Kubernetes component for scaling containerized workloads according to event-driven signals. It works alongside the Kubernetes Horizontal Pod Autoscaler and uses ScaledObjects and scalers to map applications to event sources, including queues, databases, metrics systems, cloud services, and HTTP traffic. KEDA supports deployments, jobs, and custom resources with a /scale sub-resource, including scaling workloads down to zero and back up when events require processing. Its HTTP add-on provides an interceptor that can receive or hold requests while an application has no active pods and help trigger scale-up. The project provides built-in scalers, supports custom and community-maintained scalers, and is vendor-agnostic across cloud providers and products.

Mentioned in
1 video
Kind
Other
KNo. 0464
AIAINotes.us Tool

Kubernetes

Open source · kubernetes/kubernetes

Kubernetes (K8s) is an open-source container orchestration system hosted by the Cloud Native Computing Foundation. It manages containerized applications across multiple hosts, providing mechanisms for deploying, maintaining, scaling, and scheduling workloads. Applications can be configured with YAML and deployed across hybrid-cloud environments; the platform is also used for production operations involving storage, databases, networking, and security hardening.

Mentioned in
9 videos
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Kubernetes GPU Autoscaling: Why Scale to Zero Costs You 10 Minutes Per Request — DevOps & AI Toolkit (36:52). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by DevOps & AI Toolkit. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:07 the same question to the same model on the same cluster. 4 seconds one time and over 10 minutes the next with nothing nothing broken in between and nobody having touched a line of configuration. That is the price of turning a GPU off when nobody's using it. and turning it off is the only way to stop paying for it because the cloud charges you for the machine whether or not anything is running on it.

00:32 The same 10 minutes governs the other direction too. A second replica has to be asked for long long long before the traffic that needs it arrives because it will not turn up in time to serve it. So this is that trade measured rather than argued. what scale to zero actually saves, what those 10 minutes are made of, which parts of them you can attack, and when you have to ask for more.

00:59 Now, notice what this video is not going to do, teach you any of this from the beginning. It assumes you already know what a GPU costs, what a pod is, and why an autoscaler exists. That is not a question with an answer. It is a subject. And the box you type questions into has never been much good at subjects. That is where the sponsor of this video comes in.

01:23 Master.dev, a learning platform built on a simple idea. Whether you have been coding for two months or 20 years, you have always been a developer. Tools change. You keep building. The people teaching there are building as well. At companies like Entropic, OpenAI, Netflix, Google, and Stripe. You get hundreds of courses and live workshops. Quiz mode assesses your skills and shows you which areas to focus on.

01:49 A learning agent answers your questions inside the player. And the AI learning pots were just expanded to cover the current tools and where they fit into a real workflow. Their builder sale runs through September 23rd. New members get 100 bucks off a yearly membership. The link is in the description. Big thanks to Master Dev for sponsoring this video.

02:10 And now let's go back to that GPU with a model sitting on it. Nobody asking it anything and the meter still running. We have a language model running on a GPU in Kubernetes. It works. It answers questions. That part is finished. And it is not what this video is about. This video is about the time when nobody's asking it anything and the meter is still running.

02:37 Now, one disclaimer because it will otherwise negate anyone doing this seriously. What I'm running here is one small model on a single GPU which is a video budget rather than a production deployment. Every number in this video gets worse at real scale. The shape of the problem does not. So before we change a single thing, let's look at what we are actually paying for.

03:02 Starting with the machines, three ordinary machines running uh the boring parts of the cluster and one one one only one with an MVDL4 bolted to it. That last one is where all the money goes. The other three cost roughly what a cup of coffee costs and this is what is running on the expensive one. The model server has been running on the GPU node for a while.

03:27 So it is warm caliente. Now, let's ask it something and time the answer. Look at that. A few seconds start to finish. Remember that because by the end of this video, the exact same question will take a very, very different amount of time. So, what was the GPU doing for the rest of that minute? Almost all of the cards memory is spoken for and the processor is doing absolutely nothing.

03:54 Those two numbers together are the entire problem. The memory is reserved so the model can answer the instant somebody asks. And reserving it costs exactly the same whether anybody asks or not. Right? I'm not going to tell you what my little demo cost because it is beside the point. This however is the cheapest card you would seriously serve from running the smallest model worth serving.

04:20 The GPU people in real production actually run large models very often uh that cost roughly 17 times as much per hour and it almost never arrives on its own. Now put eight of them in one machine and you're into tens tens of thousands a month for the GPUs alone before anything anything else comes to the bill. That is the number that should be bothering you.

04:46 And every hour of it that goes unused, he's gone. And that rises the obvious question. If nobody's using it, why not just turn it off? The obvious move is a horizontal port autocaler. And as of Kubernetes 137, it can finally finally finally scale to zero. It still will not help us here just to be clear because coming back up needs a signal that survives having no pods which rules out CPU and memory and ours is not a cue sitting there waiting to be read.

05:24 Ours is an HTTP request arriving for a model that does not exist and Kubernetes services do not hold on to those. They don't. So something has to catch that request, keep the caller waiting and start the model on their behalf. That something is an interceptor and everything everything that follows is built around it. Here is how you describe one of those that over there is an interceptor route and there are three things in it.

05:58 Which requests belong to this model? where to send them once something is actually running and how many requests one replica should be expected to juggle at a time. Now remember that last one, it looks like the least interesting field on the screen and it is going to cause us a great deal of trouble a few minutes from now. Now that is the routing half.

06:21 There is also a half that does the scaling and that is a scaled object. the route counts requests. This decides what to do about the count. Min replica count set to zero is the whole point. And the count it acts on comes from the interceptor rather than from the pods which is what makes zero reachable at all. Cool down period is how long the model has to sit idle before ka puts it away.

06:51 And I have set it deliberately intentionally low so that we are not all staring over there at the terminal right uh besides me fast forwarding. Now in production you would want it far higher for reasons that will be extremely extremely obvious in about few minutes probably. Now let's put both of those into the cluster. Applying those two does nothing.

07:17 nothing so far because traffic still goes straight at the model. It has to go to the interceptor instead. Otherwise, there is nobody home when the model is asleep. That means moving the ingress and moving it into a different name space because an ingress can only point at a service that lives besides it and the interceptor lives in the indicace. Now, let's take a look at the replacement.

07:46 Same host, different namespace and a different service name. Take the old one out and put this one in. That's what it does. Now, that is the whole configuration. We do what do we do? We do nothing, which is the the only part of this video that requires no effort. We just wait for the cool down to expire or maybe fast forward. Now, let's see what KDA makes of that.

08:09 Active is false. that is scared of telling us in its own way that nobody has asked this model anything for a while and it sees no reason to keep it around and the model itself and the model server is gone. No pod, no process, nothing nothing holding that GPU. We asked for a system that costs nothing when nobody's using it and we appear to have got one in about 20 lines of YAML something like that.

08:41 So solved. Let's check the bill. Now there is one thing we have not looked at since we started deleting things and it is the thing the money is actually attached to. The pot is gone. That's fine. The GPU node is still there, right? It's still running, still charging full price to host absolutely nothing. Now let's be clear. Google does not bill you for pods.

09:08 It bills you for machines. Deleting the workload, change the cluster and it did not change the invoice by a single scent. Now, nobody's coming to fix that on your behalf. Allocated means build, used or not. In every cloud, there is. So, capacity you forget to hand back is essentially close to the most profitable thing anyone can rent you because they're not doing anything at all.

09:38 and they're not doing anything for it. Your failure to reclaim it is somebody's else's margin and handing it back is entirely your job. Is that it? Does it just sit there indefinitely charging us for a machine with nothing on it? Let's give it a few minutes and look again. Look at that. Look at that over there. The GPU node is gone. Three ordinary machines left.

10:01 No accelerator anywhere in the cluster. And now the cost is generally zero for GPU at least. A model nobody was using has stopped costing anything at all. So something was watching after all that something is the cluster outcaler and it is the second of the two here working in series in tandem with the first. Getting rid of the pod was necessary because a node with a workload on it is a node nobody can take away.

10:32 But it was never what stops the matter. The cluster autoscaler had to look at the empty machine, decide it was not needed and hand it back. It takes its time over that, right? And reasonably so because deleting a node that turns out to be needed is far worse mistake than keeping one a few minutes too long, right? This is also one place the two clouds are not interchangeable.

11:00 On Google, the cluster autoscaler is part of the control plane and you get it by asking do not pull for it. On EKS, it is simply not there. You deploy the cluster autocaler into the cluster yourself or you run carpenter instead and until you do a node group stays at whatever size you set it to, no matter how empty it gets, the two of them run on completely separate clocks and the gap between them is part you just paid for.

11:28 the pod disappeared and then we sat there for several minutes at full price for a GPU with nothing on it and that delay is not yours to set. The cluster autoscaler waits a fixed stretch before it will accept that the node is generally spare and both times I run this it came to about 10 minutes give or take. So every time your model goes quiet, you buy something like 10 idle GPU minutes for the privilege of it going away.

12:01 And no amount of tuning on your side shortens them. So that is scaled to zero working. The model is gone. The machine is gone. The bill is gone. Solved. Well, only if nobody ever asks it anything again. One request. That is all it takes to find out what those savings actually cost. I'm sending a single question, the same kind we asked at the start when the answer came back in a few seconds.

12:30 And I'm going to leave it running in the background while we watch what happens behind it. And there we go. Nothing comes back. The interceptor has caught that request and is holding the connection open. And the person on the other end is looking at the spinner. So while they wait, let's find out what the cluster is doing on their behalf. So here we go.

12:56 It's good news in a sense. There is a pod can notice the request and ask for the model back. The bad news is that it's pending which is Kubernetes for I would very much like to run this and there is nowhere to put it. That's how pending is interpreted in Kubernetes. So why not the events attached to that pod spell it out four lines and there are four different failures.

13:23 It starts with no GPU anywhere in the cluster which is exactly what we asked for and is now our problem. That failure is what wakes the node out of scaler and gets a machine ordered like give me one. Then a fourth node comes into existence and the port still will not go on it because the node is not finished being born. And then the best one, the node is ready and it still reports no GPU.

13:51 The machine physically has the card in it but Kubernetes cannot see the card yet because the driver is still installing. That over there is three separate kinds of waiting and we have not started the model just yet. Eventually, it does land and eventually the request comes back and here's the whole sequence. Now, ignore the last line. That is just a readiness probe knocking on the door.

14:15 The engine has not finished building yet. Everything above it is the real story. C reacting in about a second, then a node. Then a very large image coming down the wire. Then the container finally starting. And now the number that matters. How long did that person with the spinner actually wait? Read that as minutes rather than milliseconds, right? Just not to be confused.

14:39 Over 10 minutes before the first token came out. 10 minutes. The identical question at the start of this video took a few seconds. Now, nothing failed here. Nothing was misconfigured. Nothing crashed. No quarter was hit. Every component did precisely what it was designed to do. in order competently. 10 minutes is what it costs us to start a language model that was not already running.

15:06 That is the wall. And 10 minutes to start something is not on its own alarming. If this were a release going out on Tuesday afternoon, nobody would mention it. What makes it alarming is where it sits. This is not a deployment. It is a request. Somebody asked a question. The machinery that answers questions had to be built from nothing before it could hear them with that person holding the line for every second of it.

15:32 And when it finally finally finishes, they still do not have an answer. They have a model that is finally ready to start writing one. And ours, the one I'm using in this video is the friendly version of it. A small model, a single car, a fast cloud network with nothing else competing for it. Everything in that wall that scales with the size of the model only gets longer from here.

15:55 More weights to fetch, more weights to move uh onto more cards, more cards that have to be found before any of it can start. Nobody's frontier model comes up quicker than the tiny one I just uh spin up. Now, let's break it apart because the four pieces are not equally your problem. Kada reacting in one second. a rounding error. Just ignore it. Getting a machine booted with a working GPU driver is a minute and a quarter which is uh quick more or less.

16:28 Loading the model is the biggest slice at 5 minutes and that is roughly what anybody would uh expect for a tiny one. Imagine a bigger one. Now the surprise is the container image at 4 minutes. 9 GB of CUDA libraries and Python dragged across the network before anybody reads a single model weight and it accounts for well over a third of the that wall that we are climbing on its own.

16:53 Model landing gets optimized constantly. Inference images mostly don't. So solved? Obviously not. But now we know where the minutes go which means we can start attacking them. So pick a target. Of those four phases, the model is the fattest one. And the instinct about what to do with it is almost universal. I had it as well. The model got downloaded from the internet.

17:24 It had already been downloaded from the internet an hour earlier. Why are we doing that twice? Give the thing a disc. Just give it a disc and then you speed it up. A persistent volume claim. one environment variable telling hugging face to cache into it and the label saying cached so we can tell the two variants apart that's all it takes that is the entire change 20 lines no new components nothing clever this is the fix everybody reaches for and it is generally the right instinct so let's put it in the disc starts out

17:54 empty so somebody has to pay uh for the download exactly once to fill it might just as well be me right now right one request does that. Now note that one is quicker and uh note why the node was already running this time and the image was already sitting on it. So what we just measured is the model on its own fetching it putting it on the card starting the engine.

18:20 That is the biggest single slice of the wall and it is the slice the disc is supposed to delete. Now we let it go quiet right and we wait and we wait for the node to be taken away again so that the comparison is fair same starting position or it proves nothing right so checking that the GPU node has gone again and there we go it's gone and it took more or less the same 10 minutes it took the first time that is not the cluster being slow it is the cluster autoscale as default a node has to sit unwanted for a fixed

18:54 stretch before anything will remove it and that stretch is set by the autoscaler, not by you, not by me, not by how much the machine is costing you while it runs down. Now, let's go back to where we were before. No replicas, no GPU, and an empty cluster with a full disc sitting in it. Same request as last time. And there it is. First token, milliseconds, barely moved, nothing.

19:18 Same thing. We gave it to disk. It does the same thing. So, we added the persistent disc. We preloaded 16 GB of model onto it. We removed an entire download from the critical path and we bought back under a minute of 10 minutes wall something like 8%. If you had shipped that to your team as a cold start improvement, you would have been quietly laughed at.

19:40 Maybe not quietly, you would be laughed at. A result that small deserves suspicion. And the first suspect is the cash itself. Maybe it silently did nothing. So before drawing any conclusions, let's let's prove it worked. It worked perfectly. There is no download in that log at all. The engine read the whole model straight off the volume and reading it took most of a minute and a half all by itself.

20:07 And then then it spent another minute of building the KV cache and compiling CUDA graphs before it would answer anything at all. Which explains the disappointing result. We did not remove a slow step. We swapped it for another one that takes about just as long. A persistent volume is not the fast local storage the word disk might make you imagine. It is storage touched over the network and pulling 16 GB across it is its own slow work.

20:39 The cache was never going to be a shortcut. It was a different route of roughly the same length. So break open the phase we actually attacked and it turns out to be three things rather than one. Python and CUDA starting up a minute and a quarter reading 15 gigabytes of the volume a minute and a half building the KVK cache and compiling graphs another minute and a quarter.

21:02 The weights were never the problem and that's the finding and I I did not expect it to be honest. Reading the model is about a third of that and no storage decision touches other two/3s. We went after the biggest block without checking what was inside it. So most of what is left is is fixed. You're not going to make Python import faster. You're not going to talk CUDA out of compiling graphs.

21:26 And the only way to skip getting a machine is to keep one running, which is the thing we are trying to avoid. At least I'm trying to avoid paying at least. And that over there leaves the container image as the biggest slice you could actually attack. We pulled ours from Docker Hub across the public internet, which is both the worst case and the default.

21:44 Uh so a registry in the same region a slimmer runtime or prepooling into the node would all buy something back. Now not all of it is network though. Some of those four minutes is unpacking 9 GB onto a machine and that cost the same uh wherever the bites come from. And the image is the one part of this that does not grow with your model. serve something 10 times this size and the weight swamp everything else at which point the biggest slice is also the one you can do least about right so is it solved absolutely not we

22:20 chip the minute off and learn that the wall is mostly loadbearing if you cannot make it fast the only thing left is to stop pretending it it is fast everything so far has assumed the caller is willing to hold the connection open for 10 minutes, hours. Mine was open for that much time more or less because uh mine was a script with no feelings. A browser is not a load balancer is not.

22:50 Your API gateway will give up long before that and so will the person clicking the button. So if we cannot make the wall short, it can at least stop lying about that. Interceptor is allowed to answer on the model's behalf. It can do that. So let's see what that takes, how it looks like. So fun addition call start block says that when nothing is ready do not hold the line do not hold it answer straight away instead and we chose the status code the headers and the body the caller gets so we can hand back something a

23:23 client actually understands rather than than a timeout. So let's let's try that one. The model is is asleep. So let's ask it something now. Look at that. a fraction of a second instead of 10 minutes retry after header the client can actually obey an error body shaped like the API's own errors so whatever is calling on can branch on it instead of choking on a dead socket so the wall has not moved we have converted one 10 minutes silence into a series of polite refusals but that is a genuine improvement because a caller

24:00 can do something with refusal and can do nothing at with with the spinner, right? So, is this now solved? Let's keep asking and watch the cluster instead of the response. 10 refusals. No, no, no, no, no. All delivered promptly and politely. And now the part that matters, which is what the cluster did about them. 10 requests, 10 refusals. And now look at what did not happen.

24:24 Active is still false. There is still no bloody pod. There is no pod. The model is not starting. It is never going to start. Every single caller will be told to come back in 60 seconds and in 60 seconds they will be told exactly the same thing. Nothing in this cluster is reporting an error. Every component is healthy. The service is simply dead. Politely dying at scale.

24:46 The cause is field. I told you to remember concurrency counts requests that are in flight. And being in flight is precisely what a placeholder stops a request from doing. It is answered and gone in a third of a second. So the count never leaves zero. So K is never told never told that anybody wants this model. The feature we added to make the cold start survivable is the same feature that destroys the evidence that a cold start is needed.

25:16 The fix is to count arrivals instead of counting time. Request rate counts arrivals over a window no matter how fast each one is answered. So a refused request still counts as somebody asking. Let's apply that. And the wakeup starts working again. And then we ask again exactly as before. Look at that. 10 refusals again. 10 refusals. It's not working from outside the cluster.

25:42 Nothing has changed at all. The broken version and the fixed one are indistinguishable to every caller. The difference is all on the inside. Active. Look at it. Active flips to true. A pod appears and the wall starts running down in the background while callers get told to retry having just for one of those fail quietly. The obvious worry is whether this one has a threshold as well.

26:06 Does a trickle of traffic sit below the target forever and never wake anything up? It does not. And the reason is a distinction uh maybe worth exploring for a second, right? Waking up and sizing up are two different decisions. The target rate governs how many replicas to run once something is running. Whether to run anything at all is a separate question and the answer to it is simply whether anybody asked.

26:30 Now one thing about that number because it will matter in a second. It is only deciding whether the model exists at all. It is a perfectly good wakeup signal and a generally terrible way to decide how many replicas you need for reasons we are about to see right now. Everything up to here has been about the gap between zero and one. Getting rid of the model when nobody wants it and getting it back when somebody does.

27:00 That is one half of autoscaling and it is a perfectly sensible thing to want. The other half is the direction the word usually brings to mind. Traffic climbs past what a single replica can absorb and you want the second one. Second replica. You need both. And most people reach for this one first. This is a different problem and it turns out to be actually the same problem and it ends worse.

27:26 Now first some housekeeping that disc we added is read write once which means one node at a time. Each of our GPU machines holds exactly one card and the first replica is already using the card on the node the disk is attached to. So a second replica needs a different node and on a different node it cannot have the disk which is effectively worse than the cache merely stopping being useful.

27:50 The cash stops you scaling out at all. The second replica would just sit there forever and ever starting until somebody took the volume away. So storage that fixes your cold start and storage that survives a second replica are not the same storage and you have to choose and we are choosing the second one. So we go back to downloading two changes here and the other matters more than it looks.

28:17 Uh the scaled object goes first because it is the thing that pins a replica in a place. swap the deployment while the old object still permits zero and the new pod gets scaled away in the middle of its own download which is it's it's really infuriating right it's infuriating 10 minutes or more to to lose so here's what changes mean replica count goes to one and max replica count to two scale to zero done it's over from here on there is always one replica running and at most one more and the trigger changes this as well.

28:53 And this is where inference stops behaving like a web app. Requests per second is a fine signal when every request costs roughly the same. Here, one request might be 20 tokens and the next 2,000. So, counting them tells you almost nothing about how hard the GPU is working. The engine already knows the answer. So, we ask it scrapes PLM's own metrics endpoint directly.

29:17 uh and that second field is is is maybe useful for a second look because it names a text format rather than a dependency. No primitive server is involved in this decision. Ker reads the number straight of the engine. So the thing deciding when to scale is not sitting downstream of a metric pipeline. You have to keep alive as well. So scale the object first then the deployment in that order and the deployment we are going back to is the original one with no volume attached.

29:47 So the model downloads on every start again the claim itself stays behind in the cluster unused and still built. That's a separate story. I'll tear it down later. So once it is serving again here's what KD is now watching. The one that matters is VLM n requests waiting. Those are requests that the engine has accepted and has no room to advance. Not requests arriving, not requests in progress, requests stuck.

30:15 And that is about as close as you can get to a machine telling you it is full, right? So both are zero right now because nothing is happening. So uh let's change that. And the obvious thing is to fire a pile of requests at it and see what happens. Uh I did that first and it taught me something about my own testing rather than uh about Kubernetes. The burst was over in under a minute and the replica takes minutes to exist.

30:42 So the traffic had gone long before help arrived. What I measured was my own impatience probably. So the schedule has to outlast a cold start. Same requests spread out arriving steadily for longer than it takes to build a replica. and order the rate matters more than the total. It has to sit above what one replica can serve or no Q builds and nothing ever scales.

31:09 It also has to sit below what two can serve or the queue never drains and the second replica changes. Nothing you can see between those two numbers is where the handover is visible and neither is as fixed as it sounds. A saturated engine is slower than a merrily busy one because a full batch has its sequences competing for the same cache. Now measure that ceiling gently, gently gently and you will pick a rate the replica quietly absorbs with nothing ever queuing for Ker to see.

31:38 So with that settled, we point it at the model and leave it going. And while that runs, we look uh let me look at uh look in on the replicas. So the first snapshot is where we started the one replica taking all of it. By the second KDO has already seen the queue and asked for another and Kubernetes has already created it and that happened uh in about a minute and then nothing happens for 11 minutes.

32:02 A node has to be found. 9 GB of image pulled into it. A model loaded and none none of it goes faster because there is a queue waiting on the other side. And the third snapshot is the far side of that weight. Two replicas, both ready 12 minutes after the first request arrived. Every request in that middle stretch was served by a single replica that had already been told it needs help.

32:28 The help was ordered on time. It just could not get there. And the requests themselves recorded what all of that felt like from the outside. For example, if we read TTFD P50MS down the column, it starts at two and a half seconds, which is about what one of these answers costs when nothing is ceued ahead of it. Then it climbs and it climbs every single minute for 11 minutes until the median median caller is waiting 74 seconds to see the first token come back.

33:02 It's 30 times worse than when we started. And every every one of those minutes is a minute in which help had already been asked for. And over there you can see that 12th minute is where it turns right. That is the second replica coming ready. From there the line falls 69 seconds 65 48 and it is still falling and the traffic stops. Never get back to 2 and 1/2 because a backlog that took 11 minutes to build does not clear in 8, right?

33:33 But it's falling. And now consider what that means for traffic that is not this patient. We watched a steady stream get help. Eventually a spike shorter than the wall gets none at all. The replica is requested. A node is bought. An image is pulled. And by the time any of it is ready, the spike is a line on a graph somebody looks at afterwards. You paid for the machine and served nobody with it.

34:01 There is a nuance about how those requests found two replicas because it it will bother anyone who has thought about this properly. The interceptor forwards to the service. So they were spread by ordinary round robin. Nothing and nothing in this setup knows which replica has a shorter queue or a warmer cache. And you can watch that cost you. Now while the first replica was uh still 100 requests deep more or less, the second one's skew set at zero.

34:29 it was taking its half of the new arrivals and nothing else because roundroin has no mechanism for handing back a backlog. So routing that actually knows is a generally interesting problem and it is a that's a different video. So the verdict right and it is um narrower uh one than the word autoscaling makes it sound scaling out works. What it cannot do is work reactively and our wrong configuration is the clearest example of why we scaled on the queue which is decent signal that tells us that replica is full and full

35:08 is already too late and the fix takes minutes to arrive. Now notice what full did not mean though. Nothing failed. Nothing was dropped. Nothing timed out. Nobody got an error page. The engine batches. So the overflow queued up inside it and got work through one piece at a time. Right. And every answer arrived late late but arrived and uh it worked and that is a far softer landing than most systems give you at capacity.

35:39 And it is interesting knowing before anybody panics probably. So this is a choice rather than a catastrophe, right? You can run hot and accept that some stretch of your traffic is slow while a replica is built and that costs nothing and irritates some people. Or you can scale before you're full on sequences in flight against the engine's limit or on how much of the KV cache has gone.

36:03 Both of which start moving well before anything cues. And that gets you a replica that arrives on time and it charges you every time the signal turns out to be a false alarm which is a problem as well. So that's the real shape of it. Autoscaling and inference workload is not a reaction to traffic. It is a forecast and the lead time is however long your model takes to come up. So the choice is between paying for wrong guesses and or living with slow minutes. Thank you for watching. See you in the next one. Cheers.