Nemo Mbridge Multi Node Slurm

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive…

About this skill

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive allocation.

Maintained by NVIDIA. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01First Answer Checklist
  • 02Two Approaches: srun-native vs uv run torch.distributed
  • 03Cluster Environment
  • 04Log Directory
  • 05srun-native Approach (Preferred)
  • 06SBATCH Headers

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

nvidia/skills / nemo-mbridge-multi-node-slurm

Source reviewed October 2, 2026 · Apache-2.0 / CC-BY-4.0; see source notices