Jetson Inference Mem Tune

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

About this skill

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

Maintained by NVIDIA. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Purpose
  • 02When to use
  • 03Prerequisites
  • 04Available Scripts
  • 05Instructions
  • 06Expected workflow

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

nvidia/skills / jetson-inference-mem-tune

Source reviewed October 2, 2026 · Apache-2.0 / CC-BY-4.0; see source notices