NVIDIA skills collection

Browse 392 official skills from NVIDIA. Open a skill to read its instructions and access the source.

NVIDIA skills

392 skills in this collection

NVIDIA logo

Kermt Monitor

Check progress for a detached KERMT run (pretrain, finetune, or any kermtrundetached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).

NVIDIAOfficial
NVIDIA logo

Kermt Pretrain Scratch

Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrainddp.py inside the kermt container (detached for long…

NVIDIAOfficial
NVIDIA logo

Kermt Setup

Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test inside the container. Every other kermt…

NVIDIAOfficial
NVIDIA logo

Launch Nemo Rl

Playbook for launching, monitoring, stopping, and debugging NeMo-RL recipes on a Kubernetes cluster via the nrl-k8s CLI. Covers ephemeral vs long-lived RayCluster modes, iterating on runs, and debugging hung or failed training jobs.

NVIDIAOfficial
NVIDIA logo

Mcore Create Issue

Investigate a failing GitHub Actions run or job and create a GitHub issue for the failure.

NVIDIAOfficial
NVIDIA logo

Mcore Linting And Formatting

Linting and formatting for Megatron-LM. Covers running autoformat.sh, tools (ruff, black, isort, pylint, mypy), and code style rules.

NVIDIAOfficial
NVIDIA logo

Mcore Run On Slurm

How to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDADEVICEMAXCONNECTIONS rules across hardware and parallelism modes…

NVIDIAOfficial
NVIDIA logo

Mcore Split Pr

Split a PR into multiple PRs to reduce the number of required CODEOWNERS reviewer groups.

NVIDIAOfficial
NVIDIA logo

Mcore Testing

Test system for Megatron-LM. Covers test layout, recipe YAML structure, adding and running unit and functional tests, golden values, marker filters, and CI parity.

NVIDIAOfficial
NVIDIA logo

Medtech Model Evidence Export

Exports sanitized metadata, parameters, reproducibility details, quality metrics, and optional review artifacts from Medical AI inference runs or evidence packs to MLflow. Use after inference, including NV-Generate runs; not for live…

NVIDIAOfficial
NVIDIA logo

Molmim Nim

Generate and optimize small molecules with NVIDIA MolMIM, molecular embeddings, latent decoding, and seed SMILES sampling.

NVIDIAOfficial
NVIDIA logo

Msa Structure Prediction Pipeline

NOTE: your protein sequence and the retrieved MSA alignment are transmitted to external NVIDIA-hosted APIs (health.api.nvidia.com) on every call. Use local NIM containers for confidential or proprietary sequences. Run a complete protein…

NVIDIAOfficial
NVIDIA logo

Nemo Automodel Distributed Training

Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.

NVIDIAOfficial
NVIDIA logo

Nemo Automodel Launcher Config

Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.

NVIDIAOfficial
NVIDIA logo

Nemo Automodel Model Onboarding

Guide for onboarding new model architectures into NeMo AutoModel, including architecture discovery, implementation patterns, registration, and validation.

NVIDIAOfficial
NVIDIA logo

Nemo Automodel Recipe Development

Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.

NVIDIAOfficial
NVIDIA logo

Nemo Fabric Build Adapter

Build, migrate, review, and maintain third-party NVIDIA NeMo Fabric adapters against the public adapter contract.

NVIDIAOfficial
NVIDIA logo

Nemo Fabric Integrate

Guidance for tasks where integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own application, job, or deployment config into an…

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Mlm Bridge Training

Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. Covers correlation testing, available recipes, and multi-GPU examples.

NVIDIAOfficial
NVIDIA logo

Nemo Mbridge Multi Node Slurm

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive…

NVIDIAOfficial

181 to 200 of 392 skills