Tilegym Converting Cutile To Triton

Converts cuTile GPU kernels (@ct.kernel) to Triton (@triton.jit). Handles standard in-repo conversion, debugging (cudaErrorIllegalAddress, shape mismatch, numerical mismatch), and mapping cuTile idioms (ct.load/ct.store, ct.Constant…

About this skill

Converts cuTile GPU kernels (@ct.kernel) to Triton (@triton.jit). Handles standard in-repo conversion, debugging (cudaErrorIllegalAddress, shape mismatch, numerical mismatch), and mapping cuTile idioms (ct.load/ct.store, ct.Constant, ct.launch) to Triton equivalents. Covers dual-kernel layout flags (e.g. transpose=True/False + autotune grid via META) per translations/advanced-patterns.md.

Maintained by NVIDIA. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Instructions
  • 02Workflow Selection
  • 03Pre-flight Analysis (Run BEFORE converting)
  • 04Conversion Checklist
  • 05Gotchas (Most Common Translation Errors) {#gotchas-most-common-translation-errors}
  • 06Performance Gotchas (10-50x Regression Risk) {#performance-gotchas-10-50x-regression-risk}

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

nvidia/skills / tilegym-converting-cutile-to-triton

Source reviewed October 2, 2026 · Apache-2.0 / CC-BY-4.0; see source notices