NVIDIA skills collection

Browse 392 official skills from NVIDIA. Open a skill to read its instructions and access the source.

NVIDIA skills

392 skills in this collection

NVIDIA logo

Tao Train Rtdetr

RT-DETR (Real-Time DEtection TRansformer) for 2D object detection. Designed for real-time inference with competitive accuracy and supports distillation and quantization for deployment optimization.

NVIDIAOfficial
NVIDIA logo

Tao Train Segformer

SegFormer for semantic segmentation. Lightweight transformer-based architecture with hierarchical feature extraction, efficient for real-time segmentation tasks.

NVIDIAOfficial
NVIDIA logo

Tao Train Single Step

Standard single-step train/eval/export workflow for any TAO model.

NVIDIAOfficial
NVIDIA logo

Tao Train Sparse4d

Sparse4D for multi-camera temporal 3D object detection and tracking. Uses sparse queries with deformable attention across camera views and time for end-to-end 3D perception, with an instance bank for temporal tracking.

NVIDIAOfficial
NVIDIA logo

Tao Train Visual Changenet

Visual ChangeNet for binary image classification and segmentation in AOI defect detection.

NVIDIAOfficial
NVIDIA logo

Tao Validate Dataset Format

Run tao-daft validate to check NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors. Do not use for non-DAFT formats.

NVIDIAOfficial
NVIDIA logo

Tao Validate Recipe Transfer

Port a published computer vision paper's official code and training recipe onto a customer's own dataset, or diagnose why such a transfer produced bad numbers. Use this whenever someone wants to reproduce a CV paper, run a paper's repo on…

NVIDIAOfficial
NVIDIA logo

Tilegym Adding Cutile Kernel

Add a new cuTile GPU kernel operator to TileGym. Covers dispatch registration in ops.py, cuTile backend implementation, init.py exports, test creation, and benchmark in tests/benchmark.

NVIDIAOfficial
NVIDIA logo

Tilegym Converting Cutile To Julia

Converts cuTile Python GPU kernels (@ct.kernel) to cuTile.jl Julia equivalents. Handles kernel syntax translation, 0-indexed to 1-indexed conversion, broadcasting differences, memory layout (row-major to column-major), type system…

NVIDIAOfficial
NVIDIA logo

Tilegym Converting Cutile To Triton

Converts cuTile GPU kernels (@ct.kernel) to Triton (@triton.jit). Handles standard in-repo conversion, debugging (cudaErrorIllegalAddress, shape mismatch, numerical mismatch), and mapping cuTile idioms (ct.load/ct.store, ct.Constant…

NVIDIAOfficial
NVIDIA logo

Tilegym Cutile Autotuning

Guidance for adding, modifying, optimizing, or debugging CuTile autotuning code. Trigger signals: exhaustivesearch / replacehints / hintsfn / cuda.tile.tune in code, autotune in filenames, or correctness/performance issues in autotuned…

NVIDIAOfficial
NVIDIA logo

Tilegym Cutile Python

Expert cuTile programming assistant. Write high-performance GPU kernels using cuTile's tile-based programming model with proper validation and optimization. Supports deep agent orchestration for complex multi-kernel tasks.

NVIDIAOfficial
NVIDIA logo

Tilegym Improve Cutile Kernel Perf

Iteratively optimize cuTile kernel performance through systematic profiling, bottleneck analysis, IR comparison, and targeted tuning. Covers tile sizes, occupancy, autotune configs, TMA, latency hints, persistent scheduling, numctas…

NVIDIAOfficial
NVIDIA logo

Tilegym Monkey Patch Kernels To Transformers

Integrate TileGym kernels into Hugging Face transformers models by replacing the library's submodule(s) and certain class(es)' implementations, and patching certain class(es)' init/forward/load weight methods prior to instantiating…

NVIDIAOfficial
NVIDIA logo

Vss Ask Video

Use this skill to ask the VSS agent's videounderstanding tool a fresh visual question about a recorded clip. Not for prior tool output, search hits, or metadata-answerable questions.

NVIDIAOfficial
NVIDIA logo

Vss Deploy Dense Captioning

Guidance for deploying standalone RT-VLM dense captioning or calling its REST API (uploads, captions, streams, chat-completions, Kafka). Not for VSS profile deploy or video-search ingestion.

NVIDIAOfficial
NVIDIA logo

Vss Deploy Detection Tracking 2d

Guidance to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice. Trigger when the user says things like 'deploy rtvi-cv', 'start warehouse 2d', 'add a stream', 'check rtvi-cv health', or…

NVIDIAOfficial
NVIDIA logo

Vss Deploy Detection Tracking 3d

Deploy and operate the RTVI-CV-3D microservice as MV3DT (MODE=mv3dt): per-camera DeepStream perception plus BEV Fusion over calibrated cameras. Supports the bundled sample dataset, custom video files, and RTSP streams, and chains to…

NVIDIAOfficial
NVIDIA logo

Vss Deploy Profile

Use to select, configure, deploy, verify, debug, or tear down a VSS profile (base, search, lvs, warehouse, edge). Not for standalone microservices — use the vss-deploy- skill.

NVIDIAOfficial
NVIDIA logo

Vss Deploy Video Embedding

Guidance for tasks where deploying, operating, or integrating the VSS 3.2 GA RT-Embed Video Embedding microservice. Covers Docker Compose bring-up, GPU and storage prerequisites, the /v1 REST API (file uploads, text and video embeddings…

NVIDIAOfficial

361 to 380 of 392 skills