AutoClusters is a Crusoe Cloud Kubernetes add-on for detecting and remediating hardware failures in GPU cluster workloads. It continuously monitors nodes for configured hardware errors; when a failure is detected, it can gracefully terminate workloads, restart or replace the unhealthy node, and reschedule pods on a healthy node. In Crusoe's Managed Slurm workflow, the remediation flow drains and replaces a failed node, requeues Slurm jobs, and allows checkpointed training to resume without manual intervention. Remediation is configured per issue type, including GPU and interconnect errors, with actions of REPLACE_NODE or OFF; all remediation actions default to OFF until explicitly enabled. Workloads can opt out with a pod label, and the documentation recommends preStop hooks so applications can save state before termination. Support is limited to specified Crusoe GPU instance types and compatible Crusoe Managed Kubernetes versions, and does not cover multi-tenant VM types where node replacement cannot be safely determined.
Crusoe is an AI infrastructure and cloud-compute company providing compute, storage, networking, and GPU capacity for AI workloads. Its managed orchestration stack runs Slurm on Kubernetes: Slurm supplies gang scheduling, topology awareness, and familiar sbatch workflows, while Kubernetes provides dynamic resource management, node health handling, and observability. Crusoe's AutoClusters automates GPU-failure remediation by notifying, draining, replacing, requeuing, and resuming workloads, including checkpointed distributed PyTorch training; the platform also supports sharing GPUs between training and inference as demand changes. Crusoe describes its infrastructure as using an energy-first approach and offering 24/7 support.
Crusoe Cloud is Crusoe's cloud infrastructure platform for AI workloads, providing compute, storage, networking, GPU node pools, and a console for managing and observing GPU infrastructure. Crusoe describes its offering as AI infrastructure and cloud compute with an energy-first approach. Its Managed Slurm runs on Kubernetes: Slurm provides gang scheduling, topology awareness, and familiar sbatch workflows, while Kubernetes supplies dynamic resource management, node health handling, and observability. The platform's AutoClusters can notify operators, drain and replace a failed GPU node, requeue the workload, and resume training from a checkpoint without manual intervention; the demonstrations also cover sharing GPUs between training and inference as demand changes.
Crusoe Managed Kubernetes (CMK) is a managed Kubernetes platform from Crusoe that provides the infrastructure layer for Crusoe's managed Slurm service. In the described architecture, Kubernetes supplements Slurm with dynamic resource management, node health handling, and observability while Slurm retains gang scheduling, topology awareness, and familiar sbatch workflows. The platform supports GPU-oriented workloads and underlies automation that can notify about a GPU failure, drain and replace the affected node, requeue the job, and resume training from a checkpoint.
Crusoe Managed Slurm is a managed high-performance computing cluster orchestration service on Crusoe Cloud. It provisions complete Slurm clusters on Crusoe's GPU-optimized infrastructure through a CLI command, UI form, or Kubernetes add-on, combining Slurm job scheduling with Crusoe Managed Kubernetes. A cluster includes the Slurm controller and database, SSH-accessible login nodes, shared storage mounted at `/home`, topology discovery, and GPU worker node sets. The Crusoe Slurm Operator manages the Slurm components as Kubernetes resources, while Slinky runs Slurm daemons in Kubernetes pods. The service supports topology-aware scheduling, multi-user access, shared storage, and management through Slurm workflows or Kubernetes CRDs. Managed Slurm also includes AutoClusters for hardware remediation. When a critical GPU or HCA failure is detected, the affected node is taken down, running jobs are cancelled and requeued to healthy nodes, and the failed node is replaced automatically.
Kubernetes (K8s) is an open-source container orchestration system hosted by the Cloud Native Computing Foundation. It manages containerized applications across multiple hosts, providing mechanisms for deploying, maintaining, scaling, and scheduling workloads. Applications can be configured with YAML and deployed across hybrid-cloud environments; the platform is also used for production operations involving storage, databases, networking, and security hardening.
PyTorch is a Python machine-learning library for tensor computation and deep neural networks, with CPU and GPU execution. Its tensor library provides NumPy-like operations, while its tape-based autograd system uses reverse-mode automatic differentiation for differentiable tensor operations and dynamically defined networks. The project also includes neural-network, compilation, multiprocessing, and data-loading components, and supports extensions through Python packages such as NumPy, SciPy, and Cython.
sbatch is Slurm Workload Manager's command for launching batch jobs on a Slurm cluster. In the cited workflow, it launches a PyTorch training script; Slurm provides the surrounding workload-management features, including scheduling and support for distributed training workloads.
This is a Crusoe engineering blog guide to deploying fault-tolerant distributed PyTorch training on Crusoe Managed Kubernetes. It combines Slurm, deployed through Crusoe's Slurm operator, with AutoClusters for hardware-failure detection and node replacement, Command Center for cluster topology, GPU metrics, notifications, and remediation visibility, and checkpoint-based recovery. The example uses a Slurm batch script with torchrun and a dynamic rendezvous backend to coordinate one process per GPU across multiple nodes; when a GPU or node failure is detected, the workflow notifies operators, drains and replaces the node, requeues the Slurm job, and resumes training from its last checkpoint. The guide reports recovery in under 15 minutes and covers cluster setup, NCCL and InfiniBand configuration, distributed PyTorch code, checkpointing, and failure testing. It is published as a dated Crusoe blog resource; the page does not state a separate maintenance policy.
Slinky is an open-source Kubernetes operator for running Slurm on Kubernetes. The project provides the foundation for Crusoe's managed Slurm service, combining Slurm-based scheduling with Kubernetes infrastructure for GPU-oriented workloads.
Slurm is a workload orchestration system used to schedule synthetic-data generation, model-training, and evaluation jobs on a separate GPU cluster.
Searchable transcript of GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe — AI Engineer (16:54). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 [music] All right. Hello everyone. Uh, my name is Connor. I'm a developer advocate at Crusoe and I'm joined by my colleagues Nikil and Young. And we're going to be talking today about the infrastructure we've built at Crusoe and how it handles critical hardware errors. And so our customers when they're running extremely largecale training workloads across thousands of GPUs, GPU failures are inevitable.
00:34 And so at scale, manual remediation is completely unsustainable. And so what we'll cover today is how by using both Slurm and Kubernetes, we can provide a robust platform with automated remediation that maintains ease of use for both machine learning engineers running jobs and the platform teams managing the underlying clusters. So to provide a little bit more context on what Cruso cloud is for those of you who haven't heard of us um we provide infrastructure as a service consisting of compute storage networking and
01:07 the latest gen GPUs from Nvidia and AMD but the focus of this talk will be on the tooling that we've built on top of that to orchestrate these large fleets of GPUs including our managed Kubernetes service called CMK and our manage slurm service which is built on top of CMK and built off of Slinky which is SCADMD's official open source project for running slurm on top of Kubernetes and this allows us to get the high performance job scheduling of slurm with the infrastructure resilience features and observability of
01:43 Kubernetes including one of the systems we've developed called autoclusters which is our system for automatically detecting and replacing failed GPU nodes when a GPU failure has been detected. And so to explain how and why we landed on this architecture, it's helpful to understand where traditional slurm both excels and falls short, which Young will get into now.
02:13 >> Thanks, Connor. Um, traditionally, I mean, looking at it from the tra traditional perspective, >> sorry, can can you hear me now? Okay, cool. I I'll just do it like this. But basically, traditionally, slurm was built slurm was built by the researchers for the researchers, universities, labs over 20 years ago. Uh built for high performance computing uh which as we've observed with our customers and others that are using slurm today translates very well to modern AI specifically around training workloads.
02:45 this idea that when you're running a multi-node training job, you need this idea of tight collective communication across the ranks that that you're using. Uh including, you know, specific network built for training workloads today. Um concepts like gang scheduling, topology awareness, uh prologue and epilog for cluster validations. These are all the features that are readily available with slurm today that helps researchers run their jobs.
03:19 Uh essentially running their jobs with tools that are already familiar to them. Scripts, especi cluster today. Now, where we feel like this is a bit of a shortcoming or short falls a little bit short around the modern training is the things that you see on the screen right now. Um AI training workloads today are very dynamic and it doesn't just include training.
03:46 Uh there's a lot of other aspects that your AI labs and other companies are running as part of your your AI workloads including training post-training eval inferences some of which may or may not even land on your slurm cluster today. But because of of what we've observed as as a GPU capacity constraints and other things kind of limits you to how you need to be able to be more dynamic around the resources that you have.
04:12 Uh slurm traditionally is pretty static and it's partitioning and other other uh functionalities. The other thing I want to mention is this idea of health checks and maintaining your nodes. Uh you know you can make sure your nodes are healthy or not healthy. GPUs do fail, how do you manage that? Uh, SlurpM obviously has features around it to be able to, you know, reue the jobs, uh, detect bad nodes, uh, using things like prologue and other open source tools that are now available to be able to do this.
04:44 It just adds an operational burden, uh, that we've seen from customers today. you know, this idea that you have to detect if nodes are bad, you have to drain the nodes when they're bad, you have to kind of work around that. That could either be manual or you need some sort of an automation or scripting to be able to do all of this. And the last thing is observability.
05:06 Um, you know, jobs can fail. Jobs could also come with degraded performances for any number of reasons. I mean, GPU failure is just one of them. there's other things like network flaps and other things within your your switches that may be causing degraded or failed performance. Uh slurm knows when a job fails. It doesn't necessarily be able to tell you everything about why it failed.
05:28 And that's where the observability comes into key uh key play. So all of this kind of encompasses the reason why we we built this product that Nikil is going to go over and sort of describe some of the specifics around manage slur. Um all right uh thanks young. So yeah, I'm going to be talking about our Cruso manage slurm product and uh like like we've mentioned before um there's a lot of benefits uh that slurm can bring to the uh training teams and uh researchers that are building the next best models but there are
06:12 some issues with you know setting up slurm managing it keeping it reliable uh and on the other hand you have kubernetes which is the OS of the cloud It's widespread use and familiarity. It's a very mature ecosystem with lots of tools, extensions, and plugins for whatever you need to add, whether it be observability, networking, security. It helps uh maintain services with there, you know, built-in safe healing, uh load balancing and autoscaling that keeps everything running smoothly.
06:43 And so we often see that uh customers will uh come with us starting with say inference running on Kubernetes and then they want to expand to training uh or vice versa. And so maybe a Kubernetes familiar team struggles with deploying and managing a slurm cluster. Um or uh the incoming slurm native training team is unfamiliar with the existing Kubernetes frameworks uh and adds you know additional friction that just slows down your team and with you know how fast everyone's trying to get to uh the next best model and next
07:20 best product. um teams get set up with the tools they're familiar with and you end up with two separate infrastructure stacks which has its own host of extra operational burdens uh and extra costs. So we decided to build out our Cruso managed slurm uh on top of Kubernetes. So we have our Cruso slurm operator uh which coordinates with our Cruso managed Kubernetes platform and the broader Cruso cloud.
07:47 So including storage, networking etc. And uh cso manages uh slurm users, slurm partitions, configurations and storage. So you don't have to do it peacemeal. You just you know one one command you create your slurm cluster and there it is existing in your kubernetes cluster. So the benefit is for the slurm users um it's just another slurm cluster. They just need uh an IP address they can SSH with their uh user and now to them it's a slurm cluster with all the bells and whistles that they need.
08:18 They don't even need to know that Kubernetes is running underneath. For your platform team and your uh infrastructure team, uh the benefit is slurm is now just another service that plugs into your existing infrastructure. You don't need to set up an entirely new infrastructure stack, new observability tools, new you know on call teams, new runbooks, nothing like that.
08:38 It's all built into the same thing. It runs alongside all your other services and um resources in slurps. any jobs make use of the uh built-in topology aware scheduling um and the platform team just sees it as another service um and now since all your GPUs and hardware are in one cluster they can move around seamlessly uh it enables uh more interesting uh tooling and dynamic availability like when you need to burst your uh inference service when you're reaching uh high uh users high number of users you can take away
09:15 some some of your or even reallocate some of your hardware from your training cluster to your inference service or vice versa. You know, when your user uh load goes down, you have all these GPUs sitting idle. That's the perfect time to up your training job or do a a larger hero run across more GPUs. Um this is something that's only possible when you have everything in a single uh infrastructure stack like uh cru managed slurm.
09:39 Um an additional benefit is uh now crucial managetorm can make use of our autoclusters product which is our auto remediation for additional hardware errors. So Connor is uh going to walk through that and show a demo of uh of how that works. So this is an overview of what exactly happens when autoclusters detects a critical hardware error. in this case an XID79 error where one of the GPUs is completely unusable.
10:13 And so the first thing that happens is that the user is notified although there's no action required from them. Then the slurm operator sets that that specific node to down and cancels any running jobs. There's then a uh the process re receives a sig term signal before being terminated and you have up to two minutes to handle that sig term to either save your checkpoints or flush your logs before that node is eventually replaced.
10:41 The cancel job is automatically reued and then autoclusters start. So first it verifies that no pods on the node have a label that disables automatic replacement. Assuming that node can be replaced, the Kubernetes node is then cordoned and drained. And then the unhealthy node is removed from the node pool and replaced with a new healthy node from spare capacity.
11:09 And then a record of the alert and the remediation is logged for the user's reference. Uh but there's still no action required from them. And then once that healthy node comes back online, then the slurm job starts. Your application code runs to load your model and load your checkpoint and resume training. And so here's what that looks like um in a demo.
11:35 And so here on the left, I'm just using sbatch to launch a PyTorch training script. On the right is the Cruso cloud console where I have a GPU node pool of two A100 nodes. And I'm going to launch this as normal. Um, and then what I'm going to do is run a script in to tell our backend internal monitoring system that there's been an XID79 error on one of the GPUs.
12:01 And this will trigger autoclusters. And so we can see immediately the GPU utilization drops. And so we know that that that GPU has gone offline and autocluster's node replacement starts immediately. And the full process from detecting a critical hardware error to getting a healthy node back into the node pool takes roughly five minutes with autoclusters.
12:29 The rest of the time is mainly application level code. And so in this example, it's a little under 15 minutes um from end to end. And here I'm just showing some of the alerts that uh the user receives when there's a critical hardware error like this. So as you'll see um there's no action required from the user. The slurm job will restart. It'll load the checkpoint and continue training where it left off.
12:56 And we can see the GPU utilization returns back to normal. And the total downtime is less than 15 minutes. And this is um much better than having an engineer have to log on in the middle of the night and spend hours debugging something. All of this is handled automatically on our platform. Now, historically, creating this type of environment has been challenging to merge both Slurm and Kubernetes.
13:23 And so we're really excited to announce what we call one-click slurm where everything that we've talked about in this presentation, this entire environment can be provisioned with with just a single command. And so this one command includes provisioning the underlying Kubernetes cluster, the slurm controller, login nodes and storage. And then adding a GPU node pool is only one command beyond that.
13:50 And so this means that engineers can simply SSH in and start running jobs immediately and autoclusters is enabled by default. So wrapping up what we covered um starting with the limitations of traditional slurm it gives you the high performance job scheduling but not the automated the automated healing that modern AI infrastructure requires. And so by building slurm on top of Kubernetes, we can enable these advanced advanced features without having either team change uh their workflows.
14:21 Then we covered the full node remediation flow from an error being detected to the slurm job resuming and the checkpoint and the training resuming. And so the key takeaways here first is that traditionally slurm and the infrastructure that it runs on are two separate systems. But together with Kubernetes, both systems stay in sync and so a failed GPU node can get cord in automatically and that signal will propagate directly to the slurm operator.
14:58 Also, like I mentioned earlier, neither the platform teams nor the machine learning teams have to change their workloads. Researchers can still use the same commands to launch slurm. And although Kubernetes is running underneath, researchers never have to SSH into a pod or worry about containers, but the Kubernetes team still get access to the same level of telemetry and observability that they're used to.
15:27 And then lastly, like Nicole mentioned, the GPU nodes are Kubernetes nodes first. And that means that slurm is just a workload running on top. And when a slurm job finishes, those GPU nodes don't have to sit idle waiting for another job. You can schedule inference pods on those immediately through Kubernetes because from the platform's perspective, they're still just nodes in a pool.
15:55 And so the core idea here is that failures are inevitable. Um and so in today's AI landscape, the architecture of infrastructure should be designed in a way so that when there is a critical error, all the right actions are handled autonomously and engineers can focus on building their application and not worrying about their infrastructure. So hopefully this was an informative talk.
16:22 Um, if you have any questions, feel free to um, connect with us at booth ug12 or on social at crusadev. And uh, thank you for listening and feel free to ask any questions. Thanks.