nvidia-nemo/speech · Archived

migrate-to-resumable-dataloader

This skill should be used when the user asks to "migrate to the resumable dataloader", "switch to indexed Lhotse", "adopt the indexed + resumable pipeline", "make my training resumable", "set up StatefulDataLoader for NeMo/Lhotse", "use index packs", "use AIStore GetBatch", or "convert this YAML to the resumable path". Walks a NeMo training YAML and optional launcher, data blend, and runtime context through the indexed + resumable Lhotse migration; lints interacting fields; auto-patches safe YA…

First seen Aug 13, 2026

Installation

$ npx skills add nvidia-nemo/speech --skill migrate-to-resumable-dataloader

Summary

  • This skill should be used when the user asks to "migrate to the resumable dataloader", "switch to indexed Lhotse", "adopt the indexed + resumable pipeline", "make my training resumable", "set up StatefulDataLoader for NeMo/Lhotse", "use index packs", "use AIStore GetBatch", or "convert this YAML to the resumable path".
  • Walks a NeMo training YAML and optional launcher, data blend, and runtime context through the indexed + resumable Lhotse migration; lints interacting fields; auto-patches safe YAML changes; emits a migration report, pre-flight checklist, and index-build command.
  • Static analysis only; never launches training.

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from nvidia-nemo/speech.

npx skills add nvidia-nemo/speech

Browse all from nvidia-nemo/speech

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 18.4K
License LICENSE
Default branch main
Open issues 134
Status Archived

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 9,539 B
  • docs SUMMARY.md 669 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

Migrate a NeMo training YAML to indexed + resumable Lhotse

Use this skill to port a NeMo training config from streaming/replay-style Lhotse loading to indexed access plus torchdata.StatefulDataLoader checkpoint/restore. The migration is fragile because YAML flags, launcher seed policy, index paths, storage backend, and resume topology all interact.

Core concepts

  • Indexed sources need .idx sidecars for random access into JSONL, tar, and

supported Shar-style data. Build these once per blend/source set.

  • For datasets with very many shards, an optional dataset-level .idxpack

combines existing sidecars into one memory-mapped catalog. Prefer one pack per independently configured outer input_cfg; do not create one global pack for an entire mixture.

  • usestatefuldataloader: true lets Lightning checkpoint the dataloader

iterator state, but only if seeds, worker counts, and distributed topology are stable across chunks.

  • Training configs must use forcemapdataset: false so indexed sources

partition across data-parallel ranks and workers without map-style sampler overhead. Treat forcemapdataset: true for training as not launch-ready unless the user explicitly approves a temporary exception; every source in the training iteration graph must be indexed and partition-compatible before launch.

  • Remote audio on AIStore/S3 generally needs USEAISGET_BATCH=true so audio

fetches are deferred to sample time instead of constructing eager tar readers for every shard.

Inputs

input required source purpose
Training YAML yes argument or --config= Inspect data.trainds, data.validationds, trainer, exp_manager, and any model fields that affect resume.
Launcher script no argument or auto-detect from project conventions Check per-chunk seed policy, resume topology invariance, Python path setup, AIStore env vars, and optional index staging.
Data-blend YAML no resolved from data.trainds.inputcfg when possible Check indexability: compressed paths, non-seekable paths, unsupported extrafields, slicelength, and mixed indexed/non-indexed chains.
Runtime context no argument, config file, or user-provided notes Detect storage backend, AIStore endpoint availability, container constraints, and index mirror destination.

Outputs

Every output lands in migrate-resumable/<config-stem>/ in the current repo:

output purpose
migration-report.md Findings, rationale, patched fields, and unresolved blockers.
<config-stem>-resumable.yaml Patched training config when safe automatic edits are possible.
<blend-stem>-resumable.yaml Patched blend, only when a blend was inspected and safe changes are possible.
pre-flight-checklist.md User-run steps before submitting training.
build-indexes-cmd.sh One-shot sidecar-build command and, when packs are selected, one conversion command per outer dataset. Use a project wrapper when available, otherwise the generic NeMo scripts.

Workflow

1. Discover and parse inputs

  1. Resolve the training YAML path and read it with OmegaConf or a

comment-preserving YAML parser.

  1. Resolve any referenced blend YAMLs from data.*.input_cfg. Prefer project

conventions when obvious, but fall back to paths relative to the config.

  1. If a launcher path is supplied, read it. Otherwise inspect likely project

launchers (train.py, pretrain.py, shell wrappers, or raw torchrun / python commands) and pick the closest match.

  1. If runtime context is supplied, read it for container image, environment

variables, filesystem mounts, worker counts, and AIStore endpoint settings.

  1. Detect remote storage from source paths (s3://, ais://, http(s)://) and

local filesystem storage from ordinary absolute or relative paths.

2. Run lint pipeline

Run every relevant check in:

  • references/option-reference.md
  • references/conflict-matrix.md
  • references/failure-modes.md
  • references/aistore-vs-non-aistore.md when remote storage is present

Each finding should include severity, field/path, current value, recommended value, and a short rationale.

Severities:

  • fatal: automatic patching is not possible; user must preprocess data or

change the source layout.

  • error: automatic patching is safe and should be applied.
  • warning: context-dependent; emit a report item and optional YAML comment.
  • note: informational; no patch.

3. Emit patched YAML and blend

Apply safe error-severity patches. Preserve comments when possible with ruamel.yaml; otherwise serialize with OmegaConf/YAML and rely on the report for rationale. For blend edits, never silently drop data: leave an explicit report entry and comment for every excluded or rewritten source.

4. Generate migration-report.md

Use templates/migration-report.md. Include:

  1. Summary of storage workflow, counts by severity, and readiness.
  2. Inputs inspected.
  3. Findings table.
  4. Walkthrough for train data, validation data, trainer/exp manager, launcher,

and storage backend.

  1. Data-blend audit.
  2. Verification and pre-flight steps.

5. Generate pre-flight-checklist.md

Use templates/pre-flight-checklist.md when present. Required steps:

  • Build .idx sidecars for every training/validation/test blend involved.
  • When startup would open many loose sidecars, build and validate one .idxpack

per supported outer dataset after the sidecars exist. Record the owning input_cfg entry and output filename explicitly.

  • Verify indexes_root points at the same stable mirror used by the runtime, or

that explicit node-local index staging populates it before training starts.

  • If AIStore is in play: verify aistore SDK availability, AIS_ENDPOINT, and

whether USEAISGETBATCH or USEAISINDIVIDUALGETS is required.

  • Verify one invariant seed across resumable chunks.
  • Verify numworkers, worldsize, and relevant distributed topology do not

change across resume boundaries.

  • Recommend a small smoke ladder: single-node single chunk, single-node resume,

then full topology.

6. Generate build-indexes-cmd.sh

Prefer a project-provided wrapper when one is clearly present. Otherwise emit a generic command using:

python <NeMo>/scripts/dataloading/build_indexes.py \
    --indexes-root <shared-index-mirror> \
    --workers <N> \
    <blend>.yaml [<validation-blend>.yaml ...]

When an outer dataset is supported and has enough shards to benefit from one memory map, append a command for that dataset (repeat for every independently configured outer input_cfg):

python <NeMo>/scripts/dataloading/convert_indexes_to_idxpack.py \
    --indexes-root <shared-index-mirror> \
    --output <index-pack-root>/<dataset-name>.idxpack \
    <dataset-input-cfg>.yaml

Patch the owning outer entry with indexpack: <dataset-name>.idxpack and set indexpack_root at the dataloader level. Never infer a pack by filename: an explicit declaration is part of the runtime validation contract. If the converter rejects a type, keep that adapter on loose sidecars and report it.

If running through a managed runtime or container wrapper, include comments for required container image, mounts, environment variables, worker count, and any CPU/GPU container-hook workaround the project requires.

7. Print final summary to chat

Keep the final chat response under 10 lines: output directory, finding counts, report path, and the next command the user should run.

Knowledge base

  • references/option-reference.md: field-by-field reference for YAML and

launcher settings.

  • references/failure-modes.md: known failure signatures, triggers, and fixes.
  • references/conflict-matrix.md: incompatible option pairs.
  • references/best-practices.md: priority-ordered checklist.
  • references/aistore-vs-non-aistore.md: storage workflow selection.
  • templates/migration-report.md: report template.
  • templates/pre-flight-checklist.md: checklist template, when present.
  • scripts/analyze.py: optional static-analysis helper, when present.

Constraints

  • Prefer static analysis. Do not launch training, build indexes, prefetch data, or

modify external runtime state unless the user explicitly asks.

  • Cross-check recommendations against the actual NeMo/Lhotse code in the user's

checkout when paths are available. Relevant areas are common Lhotse dataloader config, indexed adapters, lhotse.indexing, AIStore batch loading, and NeMo dataloader construction.

  • Treat project wrappers as optional conveniences, not as part of the generic

migration contract.

  • When evidence is missing, say so. Do not encode project-specific run history

or local experiment names as general guidance.