Tao Train Grounding Dino

Grounding DINO for open-set object detection. Combines DINO-style detection with a BERT text encoder for language-guided detection — detects objects described by text prompts without a fixed class vocabulary.

About this skill

Grounding DINO for open-set object detection. Combines DINO-style detection with a BERT text encoder for language-guided detection — detects objects described by text prompts without a fixed class vocabulary.

Maintained by NVIDIA. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Dataclass Schemas
  • 02Train Action Policy
  • 03Training Requirements
  • 04Per-Action Dataset Requirements
  • 05Typical Spec Overrides
  • 06Eval Dataset

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

nvidia/skills / tao-train-grounding-dino

Source reviewed October 2, 2026 · Apache-2.0 / CC-BY-4.0; see source notices