nvidia/tensorrt-llm · Archived
perf-torch-cuda-graphs
Apply CUDA Graphs to PyTorch workloads — API selection (torch.compile, PyTorch make_graphed_callables, TE make_graphed_callables, MCore CudaGraphManager,…
Installation
npx skills add https://github.com/nvidia/tensorrt-llm
Stronger alternatives
This repository is archived — consider an actively maintained alternative.
Analyze ncu (NVIDIA Nsight Compute) profiling output: SOL% bottleneck classification, roofline …
25 installsONLY for OpenAI Triton (@triton.jit) kernel development. NEVER use for CUDA C++ kernels, TileIR…
19 installsPerformance optimization coordination playbook. Contains specialist routing table, TileIR two-s…
17 installs>- Identify and eliminate host-device synchronizations in PyTorch code. Detects sync points (.i…
12 installsSimilar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
DEPRECATED redirect — this skill was renamed to agent-graphs. Do not use this skill; invoke age…
2.3K installsValidate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs a…
1.8K installsCreate and manage agent graphs — directed graphs of configs connected by edges with handoff log…
1.7K installsFlamegraph generation and interpretation skill. Use when converting perf, Valgrind Callgrind, o…
456 installsLoad this skill whenever the project contains charts, graphs, data visualizations, infographics…
112 installsAlso in this package
Other skills from nvidia/tensorrt-llm.
npx skills add https://github.com/nvidia/tensorrt-llm
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Repository health
main
History
- First seen on skills.sh
- First recorded snapshot · 13 installs