NVIDIA skills collection

Browse 392 official skills from NVIDIA. Open a skill to read its instructions and access the source.

NVIDIA skills

392 skills in this collection

NVIDIA logo

Nemo Mbridge Perf Activation Recompute

Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute. Use for activation memory OOMs or regressions involving recomputegranularity, recomputenumlayers…

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Cpu Offloading

Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf CUDA Graphs

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Expert Parallel Overlap

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Hierarchical Context Parallel

Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Megatron Fsdp

Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Memory Tuning

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Moe Comm Overlap

MoE expert-parallel communication overlap in Megatron Bridge. Covers dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Moe Dispatcher Selection

Choose the right MoE token dispatcher (alltoall, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage. Summarizes patterns from DSV3, Qwen3, Qwen3-Next, and VLM bring-up work.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Moe Hardware Configs

Representative, point-in-time MoE training playbooks by hardware and model family. Use them as candidate seeds, then revalidate the exact runtime, semantics, topology, and steady-state throughput.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Moe Long Context

Long-context MoE training guidance for Megatron Bridge. Covers CP sizing, selective recompute, dispatcher choices, and practical patterns from DSV3, Qwen3, and Qwen3-Next long-context experiments.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Moe Optimization Workflow

Evidence-gated workflow for MoE performance optimization in Megatron Bridge. Covers measurement contracts, the Three Walls framework, parallel folding, profiling, matched A/B tuning, and final validation.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Moe Vlm Training

Practical guidance for training MoE VLMs in Megatron Bridge. Compares FSDP and 3D-parallel approaches, using rounded lessons from Qwen3-VL, Qwen3-Next, and other multimodal experiments.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Parallelism Strategies

Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Sequence Packing

Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Perf Tp Dp Comm Overlap

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Recipe Recommender

Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Resiliency

Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.

NVIDIAOfficial
NVIDIA logo

Nemo Relay Debug Runtime Integration

Guidance for tasks where NeMo Relay is installed or imported but application-side runtime behavior is missing or incorrect, including load failures, inactive scopes, missing events, and plugin or adaptive wiring problems.

NVIDIAOfficial
NVIDIA logo

Nemo Relay Get Started

Guidance for tasks where first-time NeMo Relay users want to try Relay, choose the least-complex supported quick start, or verify initial value through the CLI, a maintained integration, or direct Python, Node.js, or Rust instrumentation…

NVIDIAOfficial

201 to 220 of 392 skills