Source

nvidia/model-optimizer

18 skills · 157 combined installs

Skills from this source

#
Skill
Source
8W Activity
Installs
1
accessing-mlflow Query and browse evaluation results stored in MLflow. Use when the user wants to look up runs by invocation ID, compa…
nvidia/model-optimizer
12
2
ptq Use when the user asks to "quantize a model", "run PTQ", "post-training quantization", "NVFP4 quantization", "FP8 qua…
nvidia/model-optimizer
12
3
compare-results Establish baseline-vs-candidate evaluation plans, delegate missing evaluations, compare validated results, and decide…
nvidia/model-optimizer
11
4
day0-release Deterministic end-to-end driver for day-0 quantized-checkpoint releases — chains PTQ → evaluation → comparison with e…
nvidia/model-optimizer
11
5
launching-evals Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher. Covers running evaluations, checking st…
nvidia/model-optimizer
11
6
quant-recipe-search Use when the user asks to find, search for, or optimize the best quantization recipe for a model, including direct re…
nvidia/model-optimizer
11
7
debug Run commands inside a remote Docker container via the file-based command relay (tools/debugger). Use when the user sa…
nvidia/model-optimizer
10
8
evaluation Evaluates accuracy of quantized or unquantized LLMs using NeMo Evaluator Launcher (NEL). Triggers on "evaluate model"…
nvidia/model-optimizer
10
9
monitor Monitor submitted jobs (PTQ, evaluation, deployment) on SLURM clusters. Use when the user asks "check job status", "i…
nvidia/model-optimizer
10
10
deployment Serve a quantized or unquantized LLM checkpoint as an OpenAI-compatible API endpoint using vLLM, SGLang, or TRT-LLM. …
nvidia/model-optimizer
8
11
eagle3-new-model Add a new model to the EAGLE3 offline pipeline. Generates an hf_offline_eagle3.yaml launcher config for a new model c…
nvidia/model-optimizer
7
12
eagle3-review-logs Review EAGLE3 pipeline experiment logs from the launcher's experiments/ directory. Summarizes pass/fail status for al…
nvidia/model-optimizer
7
13
eagle3-triage Triage a failed EAGLE3 pipeline run. Identifies which step failed (data synthesis, hidden state dump, training, or be…
nvidia/model-optimizer
7
14
eagle3-validate Validate that an EAGLE3 pipeline run completed successfully end-to-end. Checks all 4 steps produced expected artifact…
nvidia/model-optimizer
7
15
release-cherry-pick Cherry-pick merged PRs labeled for a release branch into that branch, then open a PR and apply the cherry-pick-done l…
nvidia/model-optimizer
7
16
benchmark-model-kernels Inspect Hugging Face decoder layers on meta tensors and plan or run per-rank BF16, FP8, and NVFP4 GEMM or fused-MoE m…
nvidia/model-optimizer
6
17
common Shared ModelOpt support files. Use only when another ModelOpt skill directs you here.
nvidia/model-optimizer
5
18
qad Run explicitly requested ModelOpt Quantization-Aware Distillation (QAD) on Slurm through Megatron Bridge to recover a…
nvidia/model-optimizer
5