Mcore Run On Slurm

How to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDADEVICEMAXCONNECTIONS rules across hardware and parallelism modes…

About this skill

How to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDADEVICEMAXCONNECTIONS rules across hardware and parallelism modes, container conventions, monitoring, and per-rank failure diagnosis.

Maintained by NVIDIA. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Answer-First Constants
  • 02Prerequisites
  • 03Minimal sbatch script
  • 04Multi-node rules
  • 05CUDA_DEVICE_MAX_CONNECTIONS
  • 06Containers

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

nvidia/skills / mcore-run-on-slurm

Source reviewed October 2, 2026 · Apache-2.0 / CC-BY-4.0; see source notices