AI product
AutoClusters is a Crusoe Cloud Kubernetes add-on for detecting and remediating hardware failures in GPU cluster workloads. It continuously monitors nodes for configured hardware errors; when a failure is detected, it can gracefully terminate workloads, restart or replace the unhealthy node, and reschedule pods on a healthy node. In Crusoe's Managed Slurm workflow, the remediation flow drains and replaces a failed node, requeues Slurm jobs, and allows checkpointed training to resume without manual intervention.
Remediation is configured per issue type, including GPU and interconnect errors, with actions of REPLACE_NODE or OFF; all remediation actions default to OFF until explicitly enabled. Workloads can opt out with a pod label, and the documentation recommends preStop hooks so applications can save state before termination. Support is limited to specified Crusoe GPU instance types and compatible Crusoe Managed Kubernetes versions, and does not cover multi-tenant VM types where node replacement cannot be safely determined.
1 use taken from transcripts — each links to the moment in the video.
AutoClusters automatically detects critical GPU hardware errors, drains and replaces unhealthy nodes, requeues Slurm jobs, and enables training to resume from a checkpoint without user intervention.
1 in the library.