Tao Finetune Clip

CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

About this skill

CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

Maintained by NVIDIA. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Train Action Policy
  • 02Instructions
  • 03Training Requirements
  • 04Supported Models
  • 05Per-Action Dataset Requirements
  • 06Typical Spec Overrides

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

nvidia/skills / tao-finetune-clip

Source reviewed October 2, 2026 · Apache-2.0 / CC-BY-4.0; see source notices