nvidia/cosmos-framework · Archived

cosmos3-post-training

Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired launch shell recommended, raw `torchrun` as an alternative), running T2V/I2V/V2V inference with the trained DCP checkpoint, and optionally exporting it to Hugging Face safetensors. Use when the user asks how to post-train Cosmos3, fine-tune on a custom video dataset, export a trained checkpoint, or inv…

First seen Jul 13, 2026

Installation

$ npx skills add nvidia/cosmos-framework --skill cosmos3-post-training

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from nvidia/cosmos-framework.

npx skills add nvidia/cosmos-framework

Browse all from nvidia/cosmos-framework

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 500
License LICENSE
Default branch main
Open issues 12
Status Archived

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 11,032 B
  • docs SUMMARY.md 882 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 4 installs

SKILL.md

Cosmos3 Post-Training (SFT)

When to use this skill

  • User wants to fine-tune Cosmos3-Nano (or Cosmos3-Super via LoRA) on the example Bridge video dataset or a custom video dataset (SFT)
  • User asks which fields in a recipe TOML to override ([model.parallelism].dataparallelsharddegree, [dataloadertrain].maxsamplesperbatch, [optimizer].lr, [trainer].maxiter, [checkpoint].load_path, ...) or which experiment SKU to pick
  • User wants to convert a base Hugging Face checkpoint to DCP, or convert a trained DCP back to safetensors
  • For installation, --group=cu130-train / cu128-train, or LDLIBRARYPATH issues, hand off to cosmos3-setup
  • For inference parameters, parallelism presets, or online serving, hand off to cosmos3-inference
  • For raw-video captioning or assembling a SFT JSONL, see docs/dataset_jsonl.md (the captioning flow has moved out of docs/training.md)

Path convention

All paths below are relative to the cosmos3 package root (../../../ from this skill file). All uv run / python / torchrun / bash commands should also be run from there.

Where to find answers

The canonical reference is docs/training.md. Use this table to route questions:

User question Go to
Full step-by-step SFT workflow docs/training.md
Which install group? (cu130-train vs cu128-train) docs/setup.md § CUDA Variants
Which recipes exist? (Vision SFT / Reasoner) docs/training.md § Step 1 - Prepare data and config
How do I download the example dataset / Wan VAE? docs/training.md § Step 1 - Prepare data and config
How do I convert a base HF checkpoint to DCP? docs/training.md § Step 2 — Prepare checkpoint
How do I launch training (paired shell, recommended)? docs/training.md § Step 3 → Generator Post-Training → Option A
How do I launch training with raw torchrun? docs/training.md § Step 3 → Generator Post-Training → Option B
How do I override DATASETPATH / BASECHECKPOINTPATH / WANVAE_PATH? docs/training.md § Step 3 → Generator Post-Training → Overriding the defaults
Which TOML keys are commonly tuned? docs/training.md § Config
LoRA knobs (loraenabled / lorarank / lora_alpha) docs/training.md § Config — [model] block (VFM only)
How do I validate the config without actually training? cosmos_framework/scripts/train.py --dryrun flag (not in docs)
How do I export the trained DCP back to safetensors? docs/training.md § Export checkpoint to Hugging Face safetensors
How do I run inference with the trained checkpoint? cosmos3-inference skill (point at $RUNDIR/checkpoints/iter<N>)
Where do training artifacts land? docs/training.md § Outputs
How do I caption raw videos / build a SFT JSONL? docs/dataset_jsonl.md

Workflow at a glance

  1. Setup — install the training extras: uv sync --all-extras --group=cu130-train (or cu128-train on older drivers), then source .venv/bin/activate && export LDLIBRARYPATH=.
  2. Step 1 - Prepare data and config — for the recipe you're running, download the HF dataset to examples/data/<dataset>/ and the Wan2.2 VAE to examples/checkpoints/wan22vae/Wan2.2VAE.pth via uvx hf@latest download …. The Reasoner recipe streams its dataset from HF Hub at startup — no Step 1 download needed.
  3. Step 2 — Prepare checkpoint — set BASECHECKPOINTNAME (Cosmos3-Nano or Cosmos3-Super, matching the recipe) and run python -m cosmosframework.scripts.convertmodeltodcp -o examples/checkpoints/$BASECHECKPOINTNAME --checkpoint-path $BASECHECKPOINTNAME. Skip for the Reasoner recipe (the Qwen3-VL backbone is fetched from HF Hub at startup).
  4. Step 3 — Run training (Option A, recommended) — from the repo root, bash examples/launchsft<recipe>.sh (e.g. launchsftvisionnano.sh). The launcher resolves DATASETPATH, BASECHECKPOINTPATH, WANVAEPATH from the default examples/ locations populated by Steps 1+2; export any of them in the shell first to override.
  5. Step 3 — Run training (Option B, raw torchrun) — export the env vars yourself, then IMAGINAIREOUTPUTROOT=outputs/train PYTHONPATH=. torchrun --nprocpernode=8 -m cosmosframework.scripts.train --sft-toml=examples/toml/sftconfig/<recipe>.toml. Unlike Option A, raw torchrun does NOT auto-resolve the env-var trio from examples/ — they must come from the shell, or you must hand-edit the TOML to inline literal paths.
  6. Outputs — $RUNDIR = $IMAGINAIREOUTPUTROOT/<job.project>/<job.group>/<job.name>. DCP checkpoints land under $RUNDIR/checkpoints/iter<N>/; the latest iter name is in $RUNDIR/checkpoints/latestcheckpoint.txt. $RUNDIR/config.yaml next to the checkpoints is what inference consumes.
  7. Inference — point cosmosframework.scripts.inference at $RUNDIR/checkpoints/iter<N> together with --config-file $RUNDIR/config.yaml (see cosmos3-inference skill for presets / input formats).
  8. Export (optional) — python -m cosmosframework.scripts.exportmodel --checkpoint-path $RUNDIR/checkpoints/$(cat $RUNDIR/checkpoints/latestcheckpoint.txt) --config-file $RUNDIR/config.yaml -o $RUNDIR/model writes a portable HF safetensors checkpoint to $RUNDIR/model.

Things not obvious from the docs

  • Training extras are a separate group: SFT requires the cu130-train / cu128-train install group, not the inference-only cu130 / cu128. Re-running uv sync with the wrong group silently leaves training deps uninstalled.
  • Recipe = paired examples/launchsft<r>.sh + examples/toml/sftconfig/<r>.toml: the .sh declares TOMLFILE directly (full repo-relative path) plus : "${DATASETPATH:=…}" / : "${BASECHECKPOINTPATH:=…}" defaults that line up with where Steps 1+2 land, then sources examples/sftlaunchercommon.sh. exported values in the user's shell win over the defaults. The helper forwards into cosmosframework.scripts.train --sft-toml=$TOMLFILE, with any TAILOVERRIDES bash-array entries appended after -- as Hydra-style key.path=value overrides (applied last on top of the pydantic-validated TOML schema in cosmosframework/configs/tomlconfig/sftconfig.py). MASTER_PORT defaults to 50012 in the helper; set it in the launcher (or export) only if you need to co-launch multiple jobs on one node.
  • Option A (paired launch shell) vs Option B (raw torchrun): Option A resolves the env-var trio (DATASETPATH / BASECHECKPOINTPATH / WANVAEPATH) from examples/ defaults so unset env runs out of the box. Option B requires you to either export those env vars yourself or hand-edit the TOML to inline the paths; the TOMLs use ${ENV:DATASETPATH} interpolation that's resolved at TOML load time.
  • IMAGINAIREOUTPUTROOT controls the entire output tree: setting IMAGINAIREOUTPUTROOT=outputs/train makes everything land under outputs/train/<job.project>/<job.group>/<job.name>/ (logs, config.yaml, checkpoints/iter_<N>, callback outputs). Unset, training falls back to /tmp/imaginaire4-output/....
  • W&B is disabled by default: every recipe TOML sets [job].wandbmode = "disabled". To log to W&B, flip it to "online" in the TOML and export WANDBAPI_KEY before launching.
  • Inference uses the DCP checkpoint directly: the standard flow points cosmosframework.scripts.inference at $RUNDIR/checkpoints/iter<N> together with --config-file $RUNDIR/config.yaml. The Hugging Face safetensors export ($RUN_DIR/model) is optional — only needed if you want a portable single-file checkpoint.
  • Parallelism degree must match topology: in the TOML [model.parallelism] block, dataparallelsharddegree × dataparallelreplicatedegree × contextparallelsharddegree must equal WORLDSIZE. -1 autoselects dataparallelshard_degree from torchrun world size. Mismatch → FSDP init failure.
  • --dryrun: cosmos_framework.scripts.train accepts --dryrun to validate the config end-to-end without launching training. Use it whenever iterating on TOML keys or Hydra overrides.

Related skills

Skill When to use
../cosmos3-setup/SKILL.md Initial install, CUDA variant selection, container/LDLIBRARYPATH setup
../cosmos3-inference/SKILL.md Inference parameters, parallelism presets, input JSON format, online serving
../cosmos3-codebase-nav/SKILL.md Locating configs, scripts, and defaults inside the package
../cosmos3-env-troubleshoot/SKILL.md Debugging environment / runtime errors during training