Tao Generate Referring Expressions

Four-step image referring-expression pipeline: turns images plus KITTI bounding-box labels into region descriptions, scene captions, grounded referring expressions, and (optionally) verified expressions via VLM distillation.

About this skill

Four-step image referring-expression pipeline: turns images plus KITTI bounding-box labels into region descriptions, scene captions, grounded referring expressions, and (optionally) verified expressions via VLM distillation.

Maintained by NVIDIA. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Purpose
  • 02Pipeline Architecture
  • 03Instructions
  • 04Initial setup
  • 05Running the pipeline
  • 06Recommended pilot workflow

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

nvidia/skills / tao-generate-referring-expressions

Source reviewed October 2, 2026 · Apache-2.0 / CC-BY-4.0; see source notices