AI product
Crusoe is an AI infrastructure and cloud-compute company providing compute, storage, networking, and GPU capacity for AI workloads. Its managed orchestration stack runs Slurm on Kubernetes: Slurm supplies gang scheduling, topology awareness, and familiar sbatch workflows, while Kubernetes provides dynamic resource management, node health handling, and observability. Crusoe's AutoClusters automates GPU-failure remediation by notifying, draining, replacing, requeuing, and resuming workloads, including checkpointed distributed PyTorch training; the platform also supports sharing GPUs between training and inference as demand changes. Crusoe describes its infrastructure as using an energy-first approach and offering 24/7 support.
1 use taken from transcripts — each links to the moment in the video.
Crusoe is the company providing infrastructure as a service, including compute, storage, networking, and GPUs. The talk focuses on its tooling for orchestrating large GPU fleets.
1 in the library.