AI product
Crusoe Cloud is Crusoe's cloud infrastructure platform for AI workloads, providing compute, storage, networking, GPU node pools, and a console for managing and observing GPU infrastructure. Crusoe describes its offering as AI infrastructure and cloud compute with an energy-first approach.
Its Managed Slurm runs on Kubernetes: Slurm provides gang scheduling, topology awareness, and familiar sbatch workflows, while Kubernetes supplies dynamic resource management, node health handling, and observability. The platform's AutoClusters can notify operators, drain and replace a failed GPU node, requeue the workload, and resume training from a checkpoint without manual intervention; the demonstrations also cover sharing GPUs between training and inference as demand changes.
1 use taken from transcripts — each links to the moment in the video.
Provides the cloud infrastructure, including compute, storage, networking, GPU node pools, and the console used to observe the GPU-failure demonstration.
1 in the library.