npx skills add nvidia/skills --skill tao-generate-image-grounding
promptingcompany/nv-skills
tao-generate-image-grounding
Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Use when the user wants to ground captions to bboxes, generate phrase-grounded annotations, auto-label images for grounding, or run the image_grounding pipeline. Triggers include 'image grounding', 'phrase grounding', 'ground captions', 'auto-label image grounding', 'image_grounding'."
Installation
npx skills add promptingcompany/nv-skills --skill tao-generate-image-grounding
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs a…
1.6K installsGrounding DINO for open-set object detection. Combines DINO-style detection with a BERT text en…
1.5K installsMask Grounding DINO for grounded instance segmentation. Extends Grounding DINO with a mask-pred…
1.5K installsUse when starting any non-trivial change, investigating a bug, or working in unfamiliar code - …
5.3K installsCompose Mapbox MCP tools to produce grounded, cited location-aware responses from live data ins…
953 installs理论溯源与案例重释。审查 PPT、课程、文案或方法论中的经验判断,?
181 installsAlso in this package
Other skills from promptingcompany/nv-skills · top by installs.
npx skills add promptingcompany/nv-skills
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Also listed on
Alternate registries and mirrors of this skill.
Repository health
main
Skill metadata
Parsed from SKILL.md frontmatter.
Read Bash WriteMore metadata
- author
- NVIDIA Corporation
- version
- 0.1.0
Package contents
Files included with this skill beyond the listing page.
-
skill md
SKILL.md7,829 B -
docs
SUMMARY.md473 B
History
- First seen on skills.sh
- First recorded snapshot · 39 installs
SKILL.md
Image Grounding Pipeline
Turn (image, caption) pairs into per-image grounded annotations: cleaned captions, referring expressions with character spans, and pixel-space bounding boxes for each expression. A single VLM (Gemini or any OpenAI-compatible endpoint) handles both steps.
Purpose
Generate phrase-grounded training data for referring-expression and grounding models. The VLM acts as a "teacher" annotator: Step 0 extracts referring expressions from the caption while looking at the image; Step 1 returns one bbox set per expression for each image.
Pipeline Architecture
Step 0: Expression extraction → VLM cleans caption, extracts referring expressions + char spans
Step 1: Phrase grounding → VLM returns pixel bboxes + scores per expression
Steps are individually selectable via workflow.steps. Each step writes a per-sample checkpoint to step<N>*/.ckpt/<sampleid>.json and skips already-processed records on re-run. Set workflow.forcereprocess: true to ignore checkpoints and reprocess from scratch.
Instructions
Initial setup
When a user wants to run this pipeline, walk through these steps:
- Input JSONL: Ask for the JSONL path. Each line must be one object like
{"imagepath": "...", "caption": "..."}.imagepathcan be absolute or relative. - Image root: If any
imagepathvalues are relative, setdata.imagerootto the directory they should resolve from. - API access: Ask the user which VLM endpoint they want to use. Present these five options and act on the choice:
1. Gemini — set vlm.backend: "gemini"; require GOOGLEAPIKEY (env var or vlm.gemini.apikey). 2. NIM (e.g. https://inference-api.nvidia.com/v1) — set vlm.backend: "openai"; collect baseurl, modelname, and apikey. 3. TAO inference microservice (self-hosted, OpenAI-compatible). Confirm whether the server is already running: - Running — collect baseurl, modelname, and (optionally) apikey; set vlm.backend: "openai". - Not running — guide the user through the skills/applications/tao-run-inference-service skill, which stands up a local TAO inference microservice with an OpenAI-compatible API. Before promising a specific model, check skills/applications/tao-run-inference-service/references/service.yaml for validnetworkarchconfigbasenames. Once the server is up, collect baseurl, modelname, and (optionally) apikey; set vlm.backend: "openai". 4. vLLM (self-hosted, OpenAI-compatible). Confirm whether the server is already running: - Running — collect baseurl, modelname, and (optionally) apikey; set vlm.backend: "openai". - Not running — follow [references/vllmserver.md](references/vllmserver.md) to install and launch a vLLM server, then collect baseurl, modelname, and (optionally) apikey; set vlm.backend: "openai". 5. Custom (any other OpenAI-compatible endpoint) — set vlm.backend: "openai"; collect baseurl, modelname, and (optionally) api_key.
If the user has no endpoint and does not want to set one up, stop and help resolve API access first.
- Workflow steps: Choose one of:
- Full pipeline: ["0", "1"] - Expression extraction only: ["0"] - Grounding only: ["1"], which requires existing step-0 output at resultsdir/step0expressionextraction/annotations.jsonl
- Resume vs fresh run: By default, the workflow reuses checkpoints and skips completed records. To reprocess everything, set
imagegrounding.workflow.forcereprocess=true.
Running the pipeline
The pipeline runs inside the TAO Toolkit container via the auto_label CLI:
auto_label generate -e /path/to/spec.yaml \
results_dir=/results \
image_grounding.data.input_jsonl=/data/captions.jsonl \
image_grounding.data.image_root=/data/images \
image_grounding.vlm.gemini.api_key=$GOOGLE_API_KEY
Generate a default spec: autolabel defaultspecs resultsdir=/results modulename=autolabel, then set autolabeltype: "image_grounding". All fields support Hydra dot-notation overrides on the command line.
See [references/configuration.md](references/configuration.md) for the full YAML structure, all parameters, model/endpoint setup, and error patterns.
Recommended pilot workflow
- Run on 5-10 images with both steps
- Inspect
step0expressionextraction/annotations.jsonl— arecleanedcaptionandexpressions[]accurate? Are the right noun phrases captured? - Inspect
step1grounding/annotations.jsonl— do the bboxes inexpressions[].instances[]look right? Are confidence scores reasonable? - If quality is insufficient, switch the VLM to a stronger model (e.g.
gemini-2.5-pro) or raisemediaresolution/maxoutputtokens, then re-run withforcereprocess=true. - Scale to the full dataset once satisfied.
Configuration
Key configuration fields (full reference in [references/configuration.md](references/configuration.md)):
| Field | Default | Description |
|---|---|---|
workflow.steps |
["0","1"] |
Which pipeline steps to execute ("0" = expressions, "1" = grounding) |
workflow.max_workers |
4 |
Parallel threads per step (watch API rate limits) |
workflow.force_reprocess |
false |
Ignore per-sample checkpoints and reprocess from scratch |
vlm.backend |
"gemini" |
"gemini" or "openai" (OpenAI-compatible endpoint) |
data.input_jsonl |
required | Path to input JSONL with image_path + caption per line |
data.image_root |
"" |
Optional prefix for resolving relative image_path entries |
Inputs
A single JSONL file at data.input_jsonl. One JSON object per line:
| Field | Required | Description |
|---|---|---|
image_path |
yes | Absolute path, or relative path resolved against data.image_root |
caption |
yes | Free-text caption for the image |
image_id |
no | Stable identifier; auto-derived from the filename if missing |
width, height |
no | Image dimensions in pixels; default to 1920×1080 for bbox clamping if missing |
Outputs
All outputs go to results_dir/:
step0expressionextraction/annotations.jsonl— per-record output enriched withcleanedcaptionandexpressions[](each withtext,expressionid,charspan,noun_chunk, emptyinstances[]).step1grounding/annotations.jsonl— same records withexpressions[].instances[]filled in (each instance hasbbox: [x1,y1,x2,y2]in pixel space,scorein[0.0, 1.0], andbbox_id).results_dir/annotations.jsonl— copy of the last step's output for convenience.step<N>*/.ckpt/<sample_id>.json— per-sample checkpoints used for resume.
Prerequisites
- Container:
nvcr.io/nvidia/tao/tao-toolkit:6.26.3-pyt - API access: At least one VLM endpoint (Gemini API key or OpenAI-compatible endpoint capable of image input)