smithery/ali

trafilatura

Extract clean article text from websites using trafilatura. Use when you need to read web pages, articles, blog posts, or documentation. Better than WebFetch for article content. Triggers: 'read this URL', 'fetch article', 'extract text from', 'what does this page say'.

Installation

$ npx skills add smithery/ali --skill trafilatura

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery/ali.

npx skills add smithery/ali

Browse all from smithery/ali

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,469 B
  • docs SUMMARY.md 289 B

History

  1. First recorded snapshot · 0 installs

SKILL.md

Trafilatura - Web Content Extraction

Extract clean, readable text from web pages. Use this instead of WebFetch when you need the actual article content (not a summary).

Installation

uv tool install trafilatura

Basic Usage

Extract article as markdown:

trafilatura -u "https://example.com/article" --markdown

Extract as plain text:

trafilatura -u "https://example.com/article"

With metadata (title, author, date):

trafilatura -u "https://example.com/article" --markdown --with-metadata

When to Use

Scenario Tool
Read an article/blog post trafilatura
Get verbatim page content trafilatura
Quick summary of a page WebFetch
API docs / technical pages trafilatura
News articles trafilatura
Pages behind auth Neither (need browser)

Common Patterns

Read and analyze an article:

trafilatura -u "https://blog.example.com/post" --markdown

Then discuss the content with the user.

Research a topic across multiple URLs:

# Run for each URL, compare findings
trafilatura -u "https://site1.com/article" --markdown
trafilatura -u "https://site2.com/article" --markdown

Extract academic paper info:

trafilatura -u "https://arxiv.org/abs/2512.12345" --markdown --with-metadata

Options Reference

Flag Purpose
-u URL URL to fetch
--markdown Output as markdown (recommended)
--with-metadata Include title, author, date
--no-comments Exclude comment sections
--no-tables Exclude tables
-o FILE Write to file instead of stdout

Troubleshooting

Empty output: Page may be JS-rendered (trafilatura can't handle SPAs) Timeout: Large pages may need --timeout 60 Encoding issues: Add --encoding utf-8

vs WebFetch

  • trafilatura: Gets full article text, better for reading/analysis
  • WebFetch: Gets AI summary, better for quick lookups
  • yomu skill: Similar to trafilatura but with different extraction engine

Use trafilatura when you need the actual content, not a summary.