nvidia-nemo/speech · Archived

nemo-speech-asr-finetune

Guide NeMo Speech users through ASR fine-tuning with container setup and Lhotse training.

First seen May 29, 2026

Installation

$ npx skills add nvidia-nemo/speech --skill nemo-speech-asr-finetune

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from nvidia-nemo/speech.

npx skills add nvidia-nemo/speech

Browse all from nvidia-nemo/speech

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 18.4K
License LICENSE
Default branch main
Open issues 134
Status Archived

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 4,474 B
  • docs SUMMARY.md 121 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 14 installs

SKILL.md

NeMo Speech ASR Fine-Tuning

Use this skill when a user wants to fine-tune a NeMo Speech ASR model, choose a checkpoint, adapt a tokenizer, configure Lhotse dataloading, train, average checkpoints, or evaluate a fine-tuned ASR .nemo checkpoint. Also use it for post-run refinement planning after fine-tuning.

Default posture:

  • Use the NeMo container unless the user explicitly asks for local execution.
  • Prefer Lhotse for train and validation dataloaders.
  • Use trainer.maxsteps, not trainer.maxepochs.
  • Use val_wer as the checkpoint monitor for validation.
  • By default, evaluate WER without capitalization and punctuation effects. Change that only when the user explicitly

asks for raw/cased/punctuated scoring.

  • Report final quality from standalone evaluation, not only in-training validation logs.

Staged Workflow

Load only the reference file needed for the current stage:

  1. Setup and checkpoint selection: read references/setup-checkpoints.md.
  2. Data prep, transcript-style preflight, Lhotse, bucketing, validation dataloader, and blends: read

references/data-lhotse.md.

  1. Architecture detection, tokenizer changes, and AED/Canary multitask metrics: read

references/architecture-tokenizer-metrics.md.

  1. Training, checkpoint averaging, and evaluation: read references/training-evaluation.md and, when reporting WER,

references/evaluation-style-contract.md.

  1. Post-run refinement, error analysis, curriculum, and general-vs-domain evaluation: read

references/refinement-iteration.md.

If the user explicitly asks for parallel/sub-agent work, split the work by these same stages. Keep each agent scoped to one stage and have the main agent integrate the final command/config.

Core Commands

Generic fine-tuning uses examples/asr/speechtotext_finetune.py. For architecture-specific recipes, route to:

  • CTC: examples/asr/asrctc/speechtotextctc_bpe.py
  • RNNT: examples/asr/asrtransducer/speechtotextrnnt_bpe.py
  • Hybrid RNNT/CTC or TDT/CTC: examples/asr/asrhybridtransducerctc/speechtotexthybridrnntctc_bpe.py
  • AED/Canary: examples/asr/speechmultitask/speechtotextaed.py

Always check the current repo docs before giving version-sensitive claims:

  • README.md
  • docs/source/asr/fine_tuning.rst
  • docs/source/asr/datasets.rst
  • docs/source/dataloaders.rst
  • docs/source/asr/featured_models.rst
  • docs/source/asr/asr_checkpoints.rst
  • nemo/collections/common/data/lhotse/dataloader.py

Non-Negotiable Pitfalls

  • When changing Lhotse batch modes, explicitly null conflicting options. For OOMptimizer profiles, set

batchsize=null, batchduration=null, and quadraticduration=null when adding bucketbatch_size.

  • Set model.validationds.uselhotse=true, but prefer static validation batch_size with bucketing disabled.
  • Do not use fused loss/WER or tune fusedbatchsize for RNNT/TDT fine-tuning guidance from this skill.
  • Run the first OOMptimizer pass with default CLI settings; lower --memory-fraction only after a real training OOM.
  • Run preflight checks before long jobs: disk space, free GPUs, manifest validity, and duration/text sanity.
  • Before any fine-tuning, audit transcript style within and across all fine-tuning/validation/test sources. Do not

train on mixed casing, punctuation, inverse-text-normalization, or symbol conventions; choose and fix one target style first, and compare it with the original checkpoint's prediction style when applicable.

  • For small domain adaptation, start with a lower LR than large-data fine-tuning; do not blindly use 1e-4.
  • Do not train a tokenizer on validation or test transcripts.
  • Do not ignore silent Lhotse filtering from minduration, maxduration, mintps, and maxtps.
  • Do not use amp=true for inference/evaluation; use amp=false compute_dtype=bfloat16.
  • Unless the user asks otherwise, report the default WER with capitalization and punctuation removed, and record any raw

WER separately when it helps diagnose transcript-style mismatch.

  • For AED/Canary, configure multitaskmetricscfg so ASR and translation/task-specific samples are evaluated with

the right constrained metrics.

  • If checkpoint averaging is used, evaluate the averaged checkpoint and keep it only if it beats the best individual

checkpoint.