harehare/mq · Archived

web-scraping

Fetches web pages and extracts structured data using mq's toolchain (mq-crawl for fetching/crawling/JS rendering, mq for HTML-to-Markdown selector-based extraction, http() for in-query requests).

First seen Jul 12, 2026

Installation

$ npx skills add harehare/mq --skill web-scraping

Summary

  • Fetches web pages and extracts structured data using mq's toolchain (mq-crawl for fetching/crawling/JS rendering, mq for HTML-to-Markdown selector-based extraction, http() for in-query requests).
  • Use when the user wants to scrape a URL, pull structured data out of a webpage, crawl a site, or turn HTML into Markdown/JSON — an alternative to a one-off curl-plus-parser script.

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from harehare/mq.

npx skills add harehare/mq

Browse all from harehare/mq

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 1.0K
License LICENSE
Default branch main
Open issues 12
Status Archived

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 5,382 B
  • docs SUMMARY.md 398 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 2 installs

SKILL.md

Web Scraping with mq

Three tools — pick the smallest one that fits:

  • mq-crawl — fetch/crawl a URL to Markdown, with JS rendering, multi-page crawling, robots.txt handling.
  • mq — filter/transform Markdown or HTML (-I html) with jq-like selectors. Full selector/function reference: processing-markdown skill.
  • mq's http() builtin (--allow-net) — make the request inside a query, for cases mq-crawl can't (POST, custom headers/auth).

mq-crawl sends logs/stats to stderr — redirect with 2>/dev/null when scripting.

Fetch & Extract

mq-crawl --depth 0 https://example.com 2>/dev/null                       # page as Markdown (--depth 0 = don't follow links)
mq-crawl --depth 0 -q '.h | to_text()' https://example.com 2>/dev/null   # inline query, e.g. an outline

-q output is always Markdown/plain text, regardless of -f (which only formats the stats block on stderr, not content). For JSON/table/grep, skip -q and pipe the page into mq instead:

mq-crawl --depth 0 https://example.com 2>/dev/null | mq -F json 'select(.link)'                            # structured rows
mq-crawl --depth 0 https://example.com 2>/dev/null | mq -F json '.[][]'                                     # table cells
mq-crawl --depth 0 https://example.com 2>/dev/null | mq -F grep --context 2 'select(contains("Pricing"))'   # locate text

Filter with select(...) like jq: select(.h.level <= 2), select(.code.lang == "js"), select(contains("keyword")).

JS-Rendered Pages & Multi-Page Crawls

mq-crawl --depth 0 --headless --headless-network-idle https://spa.example.com 2>/dev/null

mq-crawl https://docs.example.com --depth 2 --allowed-domains docs.example.com \
  --concurrency 4 -o ./crawled 2>/dev/null
  • --headless needs local Chrome (or --webdriver-url for a remote WebDriver).
  • --allowed-domains is an exact-match domain list, not a wildcard — list every subdomain you want crawled.
  • -o DIR saves one Markdown file per page instead of stdout; aggregate them with mq -I null --allow-read 'collection("./crawled") | map(self, fn(page): get(page, "title");)' /dev/null (collection reads from disk, so it requires --allow-read).

In-Query HTTP (--allow-net)

For POST/custom headers, or to keep fetch+extract in a single query:

mq --allow-net -I null 'http("get", "https://example.com")' /dev/null
mq --allow-net -I null 'http("post", "https://api.example.com/submit", "field=value")' /dev/null

http(method, url, body?, headers?) — headers is a dict, e.g. set(dict(), "Authorization", "Bearer …"). HTTPS-only; loopback/private/link-local addresses are always blocked. Errors unless --allow-net is passed. (--allow-write is the equivalent gate for write_file().)

To extract from the response, pipe through fromhtml() — its result is a single Array, so selectors like .link/.h map element-wise and leave nulls for non-matches instead of dropping them. Use compactmap() with an explicit-arg accessor (geturl(), totext(), attr()) instead:

mq --allow-net -I null -F json 'http("get", "https://example.com") | from_html(self) | compact_map(self, fn(n): get_url(n);)' /dev/null

Prefer mq-crawl for plain page-reading (JS rendering, crawling, robots.txt come free); reach for http() only when mq-crawl can't make the request you need.

Already Have HTML?

curl -s https://example.com | mq -I html '.h | to_text()'
curl -s https://example.com | mq -I html 'select(.link) | .link.url'

-I html converts to Markdown first — use Markdown selectors (.link, .h, .text), not HTML tags. Container elements (div/span/section) and their class/id/data-* attributes are discarded in that conversion.

To query the raw HTML directly instead — e.g. .price divs, a.download[href], attribute-driven card UIs — use the css()/csstext()/cssattr() builtins on the pre-conversion HTML string (needs the CLI built with the css-selector feature, on by default):

curl -s https://example.com | mq -I text 'css_text(self, ".price")'                 # text of every .price match
curl -s https://example.com | mq -I text 'css_attr(self, "a.download", "href")'     # href of every a.download match
curl -s https://example.com | mq -I text 'css(self, ".card") | map(self, fn(card): css_text(card, ".card-title") | get(0);)'  # filter by selector, then extract from each match

cssattr returns null for elements missing that attribute instead of dropping them, so array positions line up across parallel css* calls.

Limitations

  • No HTTP status/headers/timing from mq/mq-crawl — use curl -sI / curl -w for that.
  • For JSON/XML APIs, skip mq-crawl: curl -s <url> | mq -I json '...'.
  • Cap output size with --limit N / --skip N, or pipe through head.

Run mq-crawl --help / mq --help for full flag lists.