SLURM Deployment
Running DeepSpeed on SLURM-managed HPC clusters — the submission model, the resource flags that matter, and how to launch multi-node jobs correctly.
CoreWeave Setup
Working on CoreWeave's SLURM-managed cluster: the shared-cluster model, environment setup that survives job boundaries, and the pre-fetching discipline air-gapped compute nodes require.
RunPod Setup
Single-tenant GPU pods: immediate access, no scheduler, and a billing model that rewards different habits than a shared cluster.
Hardware Requirements
Sizing hardware from parameter count — and understanding which specification actually constrains you.