AI product
Crusoe Managed Slurm is a managed high-performance computing cluster orchestration service on Crusoe Cloud. It provisions complete Slurm clusters on Crusoe's GPU-optimized infrastructure through a CLI command, UI form, or Kubernetes add-on, combining Slurm job scheduling with Crusoe Managed Kubernetes.
A cluster includes the Slurm controller and database, SSH-accessible login nodes, shared storage mounted at `/home`, topology discovery, and GPU worker node sets. The Crusoe Slurm Operator manages the Slurm components as Kubernetes resources, while Slinky runs Slurm daemons in Kubernetes pods. The service supports topology-aware scheduling, multi-user access, shared storage, and management through Slurm workflows or Kubernetes CRDs.
Managed Slurm also includes AutoClusters for hardware remediation. When a critical GPU or HCA failure is detected, the affected node is taken down, running jobs are cancelled and requeued to healthy nodes, and the failed node is replaced automatically.
1 use taken from transcripts — each links to the moment in the video.
Crusoe's managed Slurm service provides Slurm scheduling for large training workloads while running on top of Kubernetes. It combines familiar Slurm workflows with Kubernetes resilience, observability, and shared GPU capacity.
1 in the library.