Use when the user wants to search, query, extract, transcribe, describe, quote, filter, or aggregate across documents — PDFs, scanned forms / images (`.jpg` `.png` `.tiff`), Office (`.docx` `.pptx`), text (`.html` `.txt`), audio (`.mp3` `.wav` `.m4a`), or video (`.mp4` `.mov`).
All-time #5861Trending #8375First seen May 30, 2026
Use when the user wants to search, query, extract, transcribe, describe, quote, filter, or aggregate across documents — PDFs, scanned forms / images (`.jpg` `.png` `.tiff`), Office (`.docx` `.pptx`), text (`.html` `.txt`), audio (`.mp3` `.wav` `.m4a`), or video (`.mp4` `.mov`).
Prefer this over native Read / Grep for multi-file or non-PDF corpora.
Not for: editing files, web browsing, single-file plain-text lookups, fine-tuning.
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Claude CodeNot declared
CursorNot declared
CodexNot declared
GitHub CopilotNot declared
WindsurfNot declared
Gemini CLINot declared
ClineNot declared
OpenCodeNot declared
Repository health
Stars3.2K
LicenseLICENSE-APACHE
Default branchmain
Open issues5
Status
Active
Skill metadata
Parsed from SKILL.md frontmatter.
LicenseApache-2.0
Allowed toolsBash Write Read
Package contents
Files included with this skill beyond the listing page.
skill mdSKILL.md3,417 B
docsSUMMARY.md456 B
History
First seen on skills.sh
First recorded snapshot · 2,149 installs
SKILL.md
nemo-retriever
The retriever CLI indexes a folder of PDFs into LanceDB (retriever ingest) and serves vector search over it (retriever query). For any task about searching/answering questions across a folder of PDFs, use this CLI — do not write a custom RAG.
Beyond PDFs and beyond semantic search.retriever ingest also handles images, Office, HTML, TXT, audio, and video — see references/setup.md for the per-format recipe and references/install.md for the install extras ([multimedia], libreoffice, ffmpeg). For non-semantic operations — page filter, verbatim quote with citation, corpus-level aggregate, chart/image caption hits — see references/query.md. Don't fall back to native Read/Grep/Python on non-PDF inputs.
Install (if retriever is missing)
If command -v retriever returns nothing, follow references/install.md to install the NeMo Retriever Library before proceeding. It prints RETRIEVERVENV=<path>; substitute that path for <RETRIEVERVENV> in every example in this skill (setup, query, troubleshooting, and the CLI references).
Workflow — read the reference for the current phase, then execute
Query turn (every subsequent turn — user asks a question)
references/query.md
One retriever query call
Anything errored or returned empty
references/troubleshooting.md
Apply the named recovery; do not improvise
For the full retriever ingest / retriever query CLI specs, see references/cli/ingest.md and references/cli/query.md. You do not need these for routine turns — <RETRIEVER_VENV>/bin/retriever <subcommand> --help is faster.
Before ingesting a mixed folder, inventory extensions (find <dir> -name '.' | sed 's/.*\.//' | sort -u) — --input-type=auto silently drops anything outside the supported set. See references/troubleshooting.md "Unsupported file types".
Hard limits (apply to every turn)
Setup turn: build the index in one shell command (see references/setup.md). STOP after the index lands.
Query turn: at most 2 Bash calls — 1 retriever query, +1 optional targeted text-extract per references/query.md. Reply and then STOP.
No narration between tool calls. Tokens you emit between calls become input + cached input for every later turn — quadratic cost. Go straight from reading the summary to writing the JSON file.
Long query turns (5+ tool calls, 1M+ cache-read tokens) cost ~5× a disciplined turn and almost always still produce the wrong answer. Answering partially beats timing out.