NVIDIA skills collection

Browse 392 official skills from NVIDIA. Open a skill to read its instructions and access the source.

NVIDIA skills

392 skills in this collection

NVIDIA logo

Tao Train Dino

DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Transformer-based detector with denoising training, multi-scale features, and optional distillation support.

NVIDIAOfficial
NVIDIA logo

Tao Train Dinov3

DINOv3 continual self-supervised pre-training. Domain-adapts public DINOv3 ViT backbones on unlabeled images via teacher-student self-distillation (DINO + iBOT + KoLeo, optional Gram anchoring) and converts the EMA teacher into a…

NVIDIAOfficial
NVIDIA logo

Tao Train Fast Foundation Stereo

Real-time stereo depth estimation using FastFoundationStereo (FFS), the distilled bp2 commercial variant of FoundationStereo. Predicts disparity maps from stereo image pairs with ~10× lower latency than full FoundationStereo.

NVIDIAOfficial
NVIDIA logo

Tao Train Foundation Stereo

Stereo depth estimation using FoundationStereo. Predicts disparity maps from stereo image pairs for 3D reconstruction.

NVIDIAOfficial
NVIDIA logo

Tao Train Grounding Dino

Grounding DINO for open-set object detection. Combines DINO-style detection with a BERT text encoder for language-guided detection — detects objects described by text prompts without a fixed class vocabulary.

NVIDIAOfficial
NVIDIA logo

Tao Train Image Classification

PyTorch-based TAO image classification. Supports a wide range of backbones (FAN, EfficientNet, ResNet, etc.) with distillation and quantization for deployment.

NVIDIAOfficial
NVIDIA logo

Tao Train Mask Auto Encoder

Masked Auto-Encoder (MAE) for self-supervised pretraining and fine-tuning. Masks random patches and reconstructs them to learn visual representations; supports pretrain and finetune stages.

NVIDIAOfficial
NVIDIA logo

Tao Train Mask Auto Label

MAL (Mask Auto-Label) for weakly-supervised segmentation. Produces segmentation masks from minimal annotations (point or box annotations) using a ViT-MAE backbone.

NVIDIAOfficial
NVIDIA logo

Tao Train Mask Grounding Dino

Mask Grounding DINO for grounded instance segmentation. Extends Grounding DINO with a mask-prediction head for open-set segmentation guided by text prompts.

NVIDIAOfficial
NVIDIA logo

Tao Train Mask2former

Mask2Former for universal image segmentation (panoptic, instance, and semantic). Transformer-based with masked attention for high-quality segmentation results.

NVIDIAOfficial
NVIDIA logo

Tao Train Metric Learning Recognition

Metric-learning recognition (ml-recog) for fine-grained visual recognition. Learns embeddings for retrieval-based matching (e.g., retail product recognition) using triplet / contrastive losses.

NVIDIAOfficial
NVIDIA logo

Tao Train Nvdinov2

NVDINOv2 for self-supervised visual representation learning. Trains vision transformers via self-distillation (teacher-student) without labels and produces general-purpose visual features.

NVIDIAOfficial
NVIDIA logo

Tao Train Nvpanoptix3d

NVPanoptix3D for panoptic 3D scene reconstruction from posed RGB images. Produces 3D panoptic segmentation (semantic, instance, and panoptic masks) with occupancy completion. Built on a VGGT backbone with a Mask2Former-style head and 3D…

NVIDIAOfficial
NVIDIA logo

Tao Train Ocdnet

OCDNet for scene text detection. Detects arbitrary-oriented text regions in natural images using a differentiable binarization approach.

NVIDIAOfficial
NVIDIA logo

Tao Train Ocrnet

OCRNet for scene text recognition. Recognizes text content from cropped text-region images and supports CTC and attention-based decoders.

NVIDIAOfficial
NVIDIA logo

Tao Train Oneformer

OneFormer for universal image segmentation. Unifies panoptic, instance, and semantic segmentation with a single architecture using task-conditioned queries.

NVIDIAOfficial
NVIDIA logo

Tao Train Optical Inspection

Optical Inspection for defect detection using Siamese networks. Compares image pairs to detect manufacturing defects, anomalies, or quality issues.

NVIDIAOfficial
NVIDIA logo

Tao Train Pointpillars

PointPillars for 3D object detection from LiDAR point clouds. Encodes point clouds into a pseudo-image via a pillar-based representation, then applies 2D detection — used in autonomous driving and robotics.

NVIDIAOfficial
NVIDIA logo

Tao Train Pose Classification

Pose classification using ST-GCN (Spatial Temporal Graph Convolutional Network). Classifies skeleton sequences into action categories from pose-keypoint data.

NVIDIAOfficial
NVIDIA logo

Tao Train Reid

Person re-identification (ReID). Learns discriminative embeddings to match the same person across different camera views, based on metric learning.

NVIDIAOfficial

341 to 360 of 392 skills