nvidia/skills · Official

tao-train-depth-anything-v2

Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures. Predicts per-pixel depth from single RGB images. Use when training, evaluating, exporting, or running inference for a TAO monocular depth model. Trigger phrases include "train monocular depth", "DepthAnything v2", "metric depth from single image", "monocular depth estimation".

All-time #7495 First seen Jun 8, 2026
8-week activity · all time api

Installation

$ npx skills add nvidia/skills --skill tao-train-depth-anything-v2

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from nvidia/skills · top by installs.

npx skills add nvidia/skills

Browse all from nvidia/skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 3.2K
License LICENSE-APACHE
Default branch main
Open issues 5
Status Active

Skill metadata

Parsed from SKILL.md frontmatter.

Version0.1.0
LicenseApache-2.0
CompatibilityRequires docker + nvidia-container-toolkit.
Allowed toolsRead Bash
More metadata
version
0.1.0
author
NVIDIA Corporation

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 14,915 B
  • docs SUMMARY.md 409 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1,543 installs

SKILL.md

Depth Net Mono

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures. Predicts per-pixel depth from single RGB images.

Pretrained checkpoint loading varies by model variant and use case — see the Pretrained checkpoint loading — use case matrix in references/parameters.md.

The mono and stereo skills both invoke the unified TAO depthnet CLI inside the container; the mono/stereo family is selected via model.modeltype (see references/parameters.md).

For TAO Deploy TensorRT actions (gentrtengine, TensorRT evaluate, and TensorRT inference), read references/tao-deploy-depth-anything-v2.md first. The deploy spec template lives in this skill's references/spectemplatedeploy.yaml.

PyT actions packaged by this model skill: train, evaluate, inference, export, and quantize. The PyT depthnet entrypoint does not accept a PyT-side gentrtengine action in the current TAO image. The gentrt_engine action metadata must run with the TAO Deploy container, and the deploy workflow remains the deploy-specific entrypoint.

Train Action Policy

This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skillinfo.yaml and resolve the run override from either an explicit automlpolicy value or the user's workflow request. Use automlpolicy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automlpolicy: off for this run only. When automlpolicy: on, automlenabled: true, and both schemas/train.schema.json and references/spectemplatetrain.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skilldir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automlpolicy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.

Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.

Workflow

Prerequisites — data accessibility

Your dataset (RGB images + GT depth files) must be reachable from inside the container:

  • SDK runner: place files at the S3 paths the runner resolves (the S3TRAIN / S3EVAL placeholders shown in Typical Spec Overrides). The runner handles S3 → container-path mounting transparently.
  • Direct docker run (e.g. local testing): mount the host dataset root read-only at the same in-container path:
docker run ... -v <host_data_root>:<host_data_root>:ro <container> ...

The same accessibility requirement applies to the <output_dir> written by all actions.

Step 1 — Annotation file

Per-line annotation file referenced by datasources[*].datafile:

Columns Format Use
1 <image> Mono inference (no GT)
2 <image> <gt_depth> Mono with GT

Do not pass stereo annotation rows such as <leftimage> <rightimage> <gt_depth> directly to mono train/evaluate/inference. If only a stereo depth dataset is available, derive a mono annotation file by keeping the left image and GT depth columns, then mount or stage the image/depth archive at the same container paths referenced by that derived annotation file.

If you already have one, point to it. Otherwise generate via depth_net convert:

depth_net convert -e <convert_spec.yaml>

convert_spec.yaml template:

results_dir: <directory where generated annotation files are written>
data_root: <directory whose immediate children are scene/sample folders that contain your image+depth files; convert walks data_root recursively but expects per-scene subdirectories at one level below>
image_dir_pattern: [<substring matching left/RGB image paths>]
depth_dir_pattern: [<substring matching GT depth paths>]
image_extension: ''     # optional .endswith filter, e.g. '.jpg'
depth_extension: ''     # optional, swapped during depth derivation, e.g. '.png'
split_ratio: 0.0        # 0.0/1.0 = test-only; 0.8 = 80/20 train+val

convert walks dataroot recursively, selects paths whose path-string contains all substrings in imagedirpattern (AND-filter), then derives the depth path by replacing imagedirpattern[0] with depthdirpattern[0] and imageextension with depthextension. Inspect your dataset's directory layout and identify the substring distinguishing RGB images from depth files (e.g. rgb vs syncdepth).

dataroot must point at the parent that contains the per-scene subdirectories (e.g. for NYU eval, use /data/nyuv2/eval/test, not /data/nyuv2/eval/test/bathroom — the latter limits the walk to a single scene). Always include the leading dot in imageextension / depth_extension (e.g. '.jpg' not 'jpg'); the substring swap is form-sensitive and a mismatch silently corrupts derived paths.

Step 2 — Pair modeltype and datasetname based on your data

Default — generic class for each task:

Data category model_type dataset_name
Disparity-encoded data (pixels) RelativeDepthAnything RelativeMonoDataset
Metric depth (meters) MetricDepthAnything MetricMonoDataset
Mono inference (no GT, any image) matches train choice RelativeMonoDataset or MetricMonoDataset

Dataset-specific class — switch when the data needs preprocessing the generic class does not perform:

Special case model_type dataset_name What the class adds
NYU syncdepth*.png (raw uint16 millimetres) — relative RelativeDepthAnything NYUDV2Relative mm→m unit conversion + Eigen evaluation crop
NYU syncdepth*.png (raw uint16 millimetres) — metric MetricDepthAnything NYUDV2 same

Using a generic class on data that requires unit conversion (e.g. raw NYU uint16 PNGs) results in an empty valid mask and silent train_loss = NaN. Match the class to your data's encoding.

For relative mono data (RelativeMonoDataset or NYUDV2Relative), leave dataset.mindepth and dataset.maxdepth unset or set both to null. Non-null metric depth ranges are passed into the relative dataset constructor and fail with BaseRelativeMonoDataset.init() got an unexpected keyword argument 'min_depth'.

Step 3 — Write spec yaml from Typical Spec Overrides

Copy the action block from Typical Spec Overrides (references/spec-overrides.md). Replace:

  • model.model_type from Step 2
  • dataset.<...>.datasources[*].datasetname from Step 2
  • datasources[*].datafile with the path from Step 1 (S3 path under SDK runner, host path for direct docker)
  • For metric finetune: additionally apply the Metric Variant Finetuning Recipe in references/finetuning-recipes.md.

For mono training set train.precision: fp32 (recommended) or bf16 (Ampere SM80+, alternative).

Step 4 — Run

Create writable home/cache directories inside the mounted output path before using --user. Some TAO containers do not have an /etc/passwd entry for the host UID, and PyTorch / matplotlib need writable cache paths when running as that UID.

mkdir -p <output_dir>/home \
         <output_dir>/.cache/matplotlib \
         <output_dir>/.cache/torchinductor \
         <output_dir>/.cache/xdg
docker run --gpus 'device=0' --shm-size 16G --ipc=host \
  --user "$(id -u):$(id -g)" \
  -e USER="$(id -un)" \
  -e LOGNAME="$(id -un)" \
  -e HOME=<output_dir>/home \
  -e MPLCONFIGDIR=<output_dir>/.cache/matplotlib \
  -e TORCHINDUCTOR_CACHE_DIR=<output_dir>/.cache/torchinductor \
  -e XDG_CACHE_HOME=<output_dir>/.cache/xdg \
  -v <data_root>:<data_root>:ro \
  -v <output_dir>:<output_dir> \
  <container> \
  depth_net <action> -e <spec.yaml>

Without --user "$(id -u):$(id -g)" the container writes outputs as nobody:nogroup, blocking host-side cleanup and retry.

Step 5 — Verify

  • Container exit code 0
  • status.json kpi block populated
  • For train: inspect per-step trainloss directly — the entrypoint reports Execution status: PASS even when trainloss = NaN (see the Metric Variant Finetuning Recipe → Sanity-run PASS criteria in references/finetuning-recipes.md)
  • For evaluate / inference: artifacts under results_dir

For TAO Deploy TensorRT actions (gentrtengine, TensorRT evaluate, and TensorRT inference), read references/tao-deploy-depth-anything-v2.md first. Deploy spec templates live in this skill's references/ folder with the spectemplatedeploy_*.yaml prefix.

Training Requirements

  • Valid datasetname values for mono datasources (case-insensitive): ThreeDVLM, FSD, NvCLIP, IssacStereo, Crestereo, Middlebury, NYUDV2, NYUDV2Relative, RelativeMonoDataset, MetricMonoDataset. NYUDV2 carries metric depth GT (meters) — pair with MetricDepthAnything; NYUDV2Relative is the same data with relative-depth conventions — pair with RelativeDepthAnything.
  • Monitoring metric: val/d1, val/loss
  • For AutoML sanity runs on the packaged relative-depth smoke data, use val/d1 as the primary monitor. val/loss can be emitted as NaN even when the trainer exits successfully and writes a usable checkpoint, so it is not a reliable AutoML objective unless the run's status metrics show a finite value.

Per-Action Dataset Requirements

Action Spec Key Source Files List?
evaluate dataset.testdataset.datasources eval_dataset datafile: annotations.txt + datasetname Yes
inference dataset.inferdataset.datasources inference_dataset datafile: annotations.txt + datasetname Yes
quantize dataset.traindataset.datasources train_datasets datafile: annotations.txt + datasetname Yes
quantize dataset.valdataset.datasources eval_dataset datafile: annotations.txt + datasetname Yes
quantize dataset.quantcalibrationdataset.images_dir train_datasets images.tar.gz No
train dataset.traindataset.datasources train_datasets datafile: annotations.txt + datasetname Yes
train dataset.valdataset.datasources eval_dataset datafile: annotations.txt + datasetname Yes

Typical Spec Overrides

Data source overrides are mandatory for every action — construct data source paths from the Per-Action Dataset Requirements table above and include them in specoverrides. Each datasources entry is a dict with two mandatory fields: datafile and datasetname. See references/spec-overrides.md for the full per-action override blocks (train, evaluate, export, inference, quantize), the S3TRAIN / S3EVAL placeholders, the relative-variant precision recommendation, and the quantize known-issue note.

Eval Dataset

Optional. Val dataset configured via dataset.valdataset.datasources (each entry needs datafile and datasetname).

Important Parameters

See references/parameters.md for the full parameter glossary (model, train, dataset, export, and inference keys with options, defaults, and sources) and the Pretrained checkpoint loading — use case matrix.

Finetuning Recipes

See references/finetuning-recipes.md for:

  • Relative Variant Finetuning Recipe — finetune from a TAO-trained RelativeDepthAnything checkpoint (lr 5e-6, LambdaLR, sanity-vs-convergent guidance, deploy LSQ alignment note).
  • Metric Variant Finetuning Recipe — checkpoint compatibility, required overrides, the dataset normalization block (normalizedepth/mindepth/max_depth) required in train AND export specs, trainer-enforced defaults, precision, the 1-epoch sanity-run override, and the Sanity-run PASS criteria with the NaN-mitigation order.

Multi-GPU / Multi-Node

Launch method: Lightning-managed (single python process, Lightning spawns workers).

Spec Key Description Default
train.num_gpus Number of GPUs 1
train.gpu_ids GPU device indices [0]
train.num_nodes Number of nodes 1
train.distributed_strategy ddp or fsdp ddp
  • ddp with activation checkpointing: findunusedparameters=False
  • ddp without: findunusedparameters=True
  • fsdp forces precision to FP16

Multi-node env vars (set by orchestrator): WORLDSIZE, NODERANK, MASTERADDR, MASTERPORT, NUMGPUPER_NODE.

Export / TRT Defaults

  • TRT data types: FP32, BF16 (Ampere SM80+). FP16 is not supported for the ViT-L mono backbone.
  • Fresh-install TRT precision: fp32. BF16 is supported on Ampere SM80+ hardware, but keep smoke tests on FP32 unless the user explicitly requests BF16.

Hardware

Minimum 1 GPU(s), recommended 2 GPU(s). 24GB+ VRAM per GPU. ViT-Large encoder is memory intensive. Use fp32 (recommended) or bf16 (Ampere SM80+, alternative) for training. Activation checkpointing is available for larger inputs.

Error Patterns

See references/troubleshooting.md for the full error-pattern catalog (depth range mismatch, relative dataset rejecting mindepth, missing pretrained weights, encoder key location, datasetname not in struct, depthnetmono not found, metric variant hyperparameter sourcing, and export ONNX overwrite).

Spec Param / Parent Model Inference

See references/spec-param-inference.md for the model-specific inference mappings (the TAO Core depthnetmono.config.json action table), checkpoint-file naming under <resultsdir>/train/, the dnmodellatest.pth policy, the parent-gentrtengine rationale, and the parentmodel / parentjobid resolution rules.

Deployment

  • [tao-deploy-depth-anything-v2](references/tao-deploy-depth-anything-v2.md)