AI product
This is a Crusoe engineering blog guide to deploying fault-tolerant distributed PyTorch training on Crusoe Managed Kubernetes. It combines Slurm, deployed through Crusoe's Slurm operator, with AutoClusters for hardware-failure detection and node replacement, Command Center for cluster topology, GPU metrics, notifications, and remediation visibility, and checkpoint-based recovery. The example uses a Slurm batch script with torchrun and a dynamic rendezvous backend to coordinate one process per GPU across multiple nodes; when a GPU or node failure is detected, the workflow notifies operators, drains and replaces the node, requeues the Slurm job, and resumes training from its last checkpoint. The guide reports recovery in under 15 minutes and covers cluster setup, NCCL and InfiniBand configuration, distributed PyTorch code, checkpointing, and failure testing. It is published as a dated Crusoe blog resource; the page does not state a separate maintenance policy.
1 use taken from transcripts — each links to the moment in the video.
This is a Crusoe blog post about automatically recovering distributed PyTorch training after GPU failures by replacing the failed node and resuming from a checkpoint.
1 in the library.