rookie-ricardo/erduo-skills

web-to-markdown

Convert a web URL into cleaned Markdown with deterministic routing.

First seen Apr 7, 2026

Installation

$ npx skills add rookie-ricardo/erduo-skills --skill web-to-markdown

Summary

  • Convert a web URL into cleaned Markdown with deterministic routing.
  • Use when Codex needs to read article-like content from links and should apply source-aware fetch strategies: default to r.jina.ai for general pages (including X/Twitter), use defuddle.md for YouTube links, and use browser-impersonated extraction for WeChat/Zhihu/Feishu pages with Mozilla Readability cleanup.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from rookie-ricardo/erduo-skills.

npx skills add rookie-ricardo/erduo-skills

Browse all from rookie-ricardo/erduo-skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 935
License LICENSE
Default branch main
Open issues 0
Status Active

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 3,576 B
  • docs SUMMARY.md 3,529 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 246 installs

SKILL.md

Web To Markdown

Convert URLs into usable Markdown by applying domain-aware fetching routes, then return the cleaned content directly.

Quick Workflow

  1. Normalize and validate the input URL.
  2. Select route:
  • r.jina.ai: general web + X/Twitter.
  • defuddle.md: YouTube transcript/content extraction.
  • special-browser-fetch: WeChat/Zhihu/Feishu.
  1. Return markdown text (or JSON metadata if needed).

For generic URLs (non-YouTube, non-WeChat/Zhihu/Feishu), use this fallback chain:

  • try r.jina.ai first,
  • if it fails, fallback to direct HTTP fetch + Readability,
  • if direct fetch still fails or returns shell-like content, fallback to browser extraction.

Commands

Run from this skill directory (skills/web-to-markdown):

npm install
node scripts/url_to_markdown.mjs <url>

Return metadata with markdown:

node scripts/url_to_markdown.mjs <url> --json

Force special-site browser extraction:

node scripts/fetch_special_sites.mjs <url> --json

Routing Policy

  • Default route: https://r.jina.ai/<url>.
  • YouTube (youtube.com, youtu.be): https://defuddle.md/<url>.
  • X/Twitter (x.com, twitter.com): https://r.jina.ai/<url>.
  • WeChat/Zhihu/Feishu: run scripts/fetchspecialsites.mjs.
  • If input is already proxy-formatted (https://defuddle.md/https://... or https://r.jina.ai/https://...), normalize back to the original URL and re-apply routing.

Special-Site Extraction Behavior

Use a two-stage strategy for WeChat/Zhihu/Feishu:

  1. Try cuimp HTTP/TLS impersonation first, then clean HTML with Mozilla Readability.
  2. If stage 1 fails or returns blocked/shell content, fallback to puppeteer-extra browser impersonation.
  • HTTP stage impersonates modern Chrome TLS/HTTP profile via cuimp.
  • Browser stage impersonates a modern Chrome user agent and standard sec-ch-ua headers.
  • Remove known login modals and backdrop overlays (best effort).
  • Scroll the page to trigger lazy-loaded article blocks.
  • Parse cleaned document with Mozilla Readability.
  • Convert extracted HTML body to Markdown via Turndown.
  • Resolve browser executable from CHROME_PATH first, then system Chrome/Chromium/Edge paths.

If special-site extraction fails due to anti-bot checks, account-only pages, or network limits, report failure clearly and ask for fallback input (for example raw page text).

Output Contract

For normal usage, output markdown only.

When --json is used, return:

  • source: backend source (r.jina.ai, defuddle, cuimp, browser-readability).
  • strategy: selected route (r-jina, defuddle, special-http-fetch, special-browser-fetch-fallback).
  • requestedUrl: original input.
  • resolvedUrl: normalized/final URL.
  • markdown: extracted markdown body.

Resources

  • [references/routing-and-notes.md](references/routing-and-notes.md): domain routing rules and operational caveats.
  • scripts/urltomarkdown.mjs: primary entrypoint.
  • scripts/fetchspecialsites_http.mjs: WeChat/Zhihu/Feishu HTTP impersonation fetcher (cuimp JS).
  • scripts/fetchspecialsites.mjs: two-stage extractor (HTTP-first, browser-fallback).