Jetson Speculative Decoding

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

About this skill

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

Maintained by NVIDIA. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Purpose
  • 02When to use
  • 03When NOT to use
  • 04Prerequisites
  • 05Instructions
  • 06Jetson-specific tuning rules

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

nvidia/skills / jetson-speculative-decoding

Source reviewed October 2, 2026 · Apache-2.0 / CC-BY-4.0; see source notices