SKILL.md
Cosmos3 Post-Training (SFT)
When to use this skill
- User wants to fine-tune Cosmos3-Nano (or Cosmos3-Super via LoRA) on the example Bridge video dataset or a custom video dataset (SFT)
- User asks which fields in a recipe TOML to override (
[model.parallelism].dataparallelsharddegree,[dataloadertrain].maxsamplesperbatch,[optimizer].lr,[trainer].maxiter,[checkpoint].load_path, ...) or which experiment SKU to pick - User wants to convert a base Hugging Face checkpoint to DCP, or convert a trained DCP back to safetensors
- For installation,
--group=cu130-train/cu128-train, or LDLIBRARYPATH issues, hand off to cosmos3-setup - For inference parameters, parallelism presets, or online serving, hand off to cosmos3-inference
- For raw-video captioning or assembling a SFT JSONL, see
docs/dataset_jsonl.md(the captioning flow has moved out ofdocs/training.md)
Path convention
All paths below are relative to the cosmos3 package root (../../../ from this skill file). All uv run / python / torchrun / bash commands should also be run from there.
Where to find answers
The canonical reference is docs/training.md. Use this table to route questions:
| User question | Go to |
|---|---|
| Full step-by-step SFT workflow | docs/training.md |
Which install group? (cu130-train vs cu128-train) |
docs/setup.md § CUDA Variants |
| Which recipes exist? (Vision SFT / Reasoner) | docs/training.md § Step 1 - Prepare data and config |
| How do I download the example dataset / Wan VAE? | docs/training.md § Step 1 - Prepare data and config |
| How do I convert a base HF checkpoint to DCP? | docs/training.md § Step 2 — Prepare checkpoint |
| How do I launch training (paired shell, recommended)? | docs/training.md § Step 3 → Generator Post-Training → Option A |
How do I launch training with raw torchrun? |
docs/training.md § Step 3 → Generator Post-Training → Option B |
How do I override DATASETPATH / BASECHECKPOINTPATH / WANVAE_PATH? |
docs/training.md § Step 3 → Generator Post-Training → Overriding the defaults |
| Which TOML keys are commonly tuned? | docs/training.md § Config |
LoRA knobs (loraenabled / lorarank / lora_alpha) |
docs/training.md § Config — [model] block (VFM only) |
| How do I validate the config without actually training? | cosmos_framework/scripts/train.py --dryrun flag (not in docs) |
| How do I export the trained DCP back to safetensors? | docs/training.md § Export checkpoint to Hugging Face safetensors |
| How do I run inference with the trained checkpoint? | cosmos3-inference skill (point at $RUNDIR/checkpoints/iter<N>) |
| Where do training artifacts land? | docs/training.md § Outputs |
| How do I caption raw videos / build a SFT JSONL? | docs/dataset_jsonl.md |
Workflow at a glance
- Setup — install the training extras:
uv sync --all-extras --group=cu130-train(orcu128-trainon older drivers), thensource .venv/bin/activate && export LDLIBRARYPATH=. - Step 1 - Prepare data and config — for the recipe you're running, download the HF dataset to
examples/data/<dataset>/and the Wan2.2 VAE toexamples/checkpoints/wan22vae/Wan2.2VAE.pthviauvx hf@latest download …. The Reasoner recipe streams its dataset from HF Hub at startup — no Step 1 download needed. - Step 2 — Prepare checkpoint — set
BASECHECKPOINTNAME(Cosmos3-NanoorCosmos3-Super, matching the recipe) and runpython -m cosmosframework.scripts.convertmodeltodcp -o examples/checkpoints/$BASECHECKPOINTNAME --checkpoint-path $BASECHECKPOINTNAME. Skip for the Reasoner recipe (the Qwen3-VL backbone is fetched from HF Hub at startup). - Step 3 — Run training (Option A, recommended) — from the repo root,
bash examples/launchsft<recipe>.sh(e.g.launchsftvisionnano.sh). The launcher resolvesDATASETPATH,BASECHECKPOINTPATH,WANVAEPATHfrom the defaultexamples/locations populated by Steps 1+2; export any of them in the shell first to override. - Step 3 — Run training (Option B, raw
torchrun) — export the env vars yourself, thenIMAGINAIREOUTPUTROOT=outputs/train PYTHONPATH=. torchrun --nprocpernode=8 -m cosmosframework.scripts.train --sft-toml=examples/toml/sftconfig/<recipe>.toml. Unlike Option A, rawtorchrundoes NOT auto-resolve the env-var trio fromexamples/— they must come from the shell, or you must hand-edit the TOML to inline literal paths. - Outputs —
$RUNDIR = $IMAGINAIREOUTPUTROOT/<job.project>/<job.group>/<job.name>. DCP checkpoints land under$RUNDIR/checkpoints/iter<N>/; the latest iter name is in$RUNDIR/checkpoints/latestcheckpoint.txt.$RUNDIR/config.yamlnext to the checkpoints is what inference consumes. - Inference — point
cosmosframework.scripts.inferenceat$RUNDIR/checkpoints/iter<N>together with--config-file $RUNDIR/config.yaml(seecosmos3-inferenceskill for presets / input formats). - Export (optional) —
python -m cosmosframework.scripts.exportmodel --checkpoint-path $RUNDIR/checkpoints/$(cat $RUNDIR/checkpoints/latestcheckpoint.txt) --config-file $RUNDIR/config.yaml -o $RUNDIR/modelwrites a portable HF safetensors checkpoint to$RUNDIR/model.
Things not obvious from the docs
- Training extras are a separate group: SFT requires the
cu130-train/cu128-traininstall group, not the inference-onlycu130/cu128. Re-runninguv syncwith the wrong group silently leaves training deps uninstalled. - Recipe = paired
examples/launchsft<r>.sh+examples/toml/sftconfig/<r>.toml: the.shdeclaresTOMLFILEdirectly (full repo-relative path) plus: "${DATASETPATH:=…}"/: "${BASECHECKPOINTPATH:=…}"defaults that line up with where Steps 1+2 land, then sourcesexamples/sftlaunchercommon.sh.exported values in the user's shell win over the defaults. The helper forwards intocosmosframework.scripts.train --sft-toml=$TOMLFILE, with anyTAILOVERRIDESbash-array entries appended after--as Hydra-stylekey.path=valueoverrides (applied last on top of the pydantic-validated TOML schema incosmosframework/configs/tomlconfig/sftconfig.py).MASTER_PORTdefaults to50012in the helper; set it in the launcher (orexport) only if you need to co-launch multiple jobs on one node. - Option A (paired launch shell) vs Option B (raw
torchrun): Option A resolves the env-var trio (DATASETPATH/BASECHECKPOINTPATH/WANVAEPATH) fromexamples/defaults so unset env runs out of the box. Option B requires you to either export those env vars yourself or hand-edit the TOML to inline the paths; the TOMLs use${ENV:DATASETPATH}interpolation that's resolved at TOML load time. IMAGINAIREOUTPUTROOTcontrols the entire output tree: settingIMAGINAIREOUTPUTROOT=outputs/trainmakes everything land underoutputs/train/<job.project>/<job.group>/<job.name>/(logs,config.yaml,checkpoints/iter_<N>, callback outputs). Unset, training falls back to/tmp/imaginaire4-output/....- W&B is disabled by default: every recipe TOML sets
[job].wandbmode = "disabled". To log to W&B, flip it to"online"in the TOML and exportWANDBAPI_KEYbefore launching. - Inference uses the DCP checkpoint directly: the standard flow points
cosmosframework.scripts.inferenceat$RUNDIR/checkpoints/iter<N>together with--config-file $RUNDIR/config.yaml. The Hugging Face safetensors export ($RUN_DIR/model) is optional — only needed if you want a portable single-file checkpoint. - Parallelism degree must match topology: in the TOML
[model.parallelism]block,dataparallelsharddegree × dataparallelreplicatedegree × contextparallelsharddegreemust equalWORLDSIZE.-1autoselectsdataparallelshard_degreefrom torchrun world size. Mismatch → FSDP init failure. --dryrun:cosmos_framework.scripts.trainaccepts--dryrunto validate the config end-to-end without launching training. Use it whenever iterating on TOML keys or Hydra overrides.
Related skills
| Skill | When to use |
|---|---|
../cosmos3-setup/SKILL.md |
Initial install, CUDA variant selection, container/LDLIBRARYPATH setup |
../cosmos3-inference/SKILL.md |
Inference parameters, parallelism presets, input JSON format, online serving |
../cosmos3-codebase-nav/SKILL.md |
Locating configs, scripts, and defaults inside the package |
../cosmos3-env-troubleshoot/SKILL.md |
Debugging environment / runtime errors during training |