tavily-ai/skills · Official

tavily-extract

Extract clean markdown or text content from specific URLs via the Tavily CLI. Use this skill when the user has one or more URLs and wants their content, says "extract", "grab the content from", "pull the text from", "get the page at", "read this webpage", or needs clean text from web pages. Handles JavaScript-rendered pages, returns LLM-optimized markdown, and supports query-focused chunking for targeted extraction. Can process up to 20 URLs in a single call.

All-time #1191 Trending #1230 Hot #6469 First seen Mar 16, 2026
8-week activity · all time api

Installation

$ npx skills add tavily-ai/skills --skill tavily-extract

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from tavily-ai/skills · top by installs.

npx skills add tavily-ai/skills

Browse all from tavily-ai/skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 477
License LICENSE
Default branch main
Open issues 4
Status Active

Skill metadata

Parsed from SKILL.md frontmatter.

Allowed toolsBash(tvly *)

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 3,475 B
  • docs SUMMARY.md 485 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 14,046 installs

SKILL.md

tavily extract

Extract clean markdown or text content from one or more URLs.

Before running

Run extract directly when tvly is available. Extract supports capped keyless access, so do not look for an API key or authenticate before the first request.

If tvly is missing, follow the [tavily-cli setup](../tavily-cli/SKILL.md#setup) before retrying. If the keyless cap is reached in an interactive session, run tvly login to open browser OAuth, then retry the original extraction once. In an unattended environment, report the cap and authentication options instead of starting an interactive flow. Do not start a second login immediately after guided setup has completed.

When to use

  • You have a specific URL and want its content
  • You need text from JavaScript-rendered pages
  • Step 2 in the [workflow](../tavily-cli/SKILL.md): search → extract → map → crawl → research

Quick start

# Single URL
tvly extract "https://example.com/article" --json

# Multiple URLs
tvly extract "https://example.com/page1" "https://example.com/page2" --json

# Query-focused extraction (returns relevant chunks only)
tvly extract "https://example.com/docs" --query "authentication API" --chunks-per-source 3 --json

# JS-heavy pages
tvly extract "https://app.example.com" --extract-depth advanced --json

# Save to file
tvly extract "https://example.com/article" -o article.json

Options

Option Description
--query Rerank chunks by relevance to this query
--chunks-per-source Chunks per URL (1-5, requires --query)
--extract-depth basic (default) or advanced (for JS pages)
--format markdown (default) or text
--include-images Include image URLs
--timeout Max wait time (1-60 seconds)
-o, --output Save the JSON response to a file
--json Structured JSON output

Extract depth

Depth When to use
basic Simple pages, fast — try this first
advanced JS-rendered SPAs, dynamic content, tables

Tips

  • Max 20 URLs per request — batch larger lists into multiple calls.
  • Use --query + --chunks-per-source to get only relevant content instead of full pages.
  • Try basic first, fall back to advanced if content is missing.
  • Set --timeout for slow pages (up to 60s).
  • Inspect failed_results even after exit code 0. A successful request can

still return no extracted pages. Retry the affected URL with advanced when appropriate, otherwise report the per-URL failure instead of treating the request as complete.

  • If search results already contain the content you need (via --include-raw-content), skip the extract step.

See also

  • [tavily-search](../tavily-search/SKILL.md) — find pages when you don't have a URL
  • [tavily-crawl](../tavily-crawl/SKILL.md) — extract content from many pages on a site