zytedata/claude-skills · Archived

scrape-codegen-generate

Generate web-poet page object code from per-page extraction analyses

First seen Jun 15, 2026

Installation

$ npx skills add zytedata/claude-skills --skill scrape-codegen-generate

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from zytedata/claude-skills · top by installs.

npx skills add zytedata/claude-skills

Browse all from zytedata/claude-skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 25
License LICENSE.md
Default branch main
Open issues 0
Status Archived

Skill metadata

Parsed from SKILL.md frontmatter.

Allowed toolsSkill, Bash, Read, Write

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 6,821 B
  • docs SUMMARY.md 99 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 4 installs

SKILL.md

You are generating web-poet page object code. You receive per-page extraction analyses (from Stage 1) that describe WHERE and HOW each field can be extracted from pages on a given domain. Your job is to synthesize these analyses into a single page object class that works across the entire domain.

Input

The raw argument string is $ARGUMENTS. Split it into 3 whitespace-separated positional arguments:

  1. work_path: directory containing Stage 1 analysis files, e.g. .scrape/.work/spec
  2. outputpath: where to save the generated page object, e.g. .scrape/spec/pageobject.py
  3. spec_path: path to spec.json file, e.g. .scrape/spec/spec.json

Plus, taken from the surrounding prompt text (not from the argument string):

  • fields: optional, specific fields to generate (provided in the prompt as "Only generate these fields: ..."). When set, only generate @field methods for those fields. When not set, generate all fields found in the analyses.

Process

1. Read inputs

Read web-poet.md and docs-access.md from ${CLAUDESKILLDIR}/../scrape/references/.

Read the schema from {specpath} — use the properties object inside schema. Also read datatype from {spec_path}.

If datatype ends with -list, set islist_type = true.

Read all Stage 1 analysis files from {work_path}/codegen-analyze/:

{work_path}/codegen-analyze/list-1.json
{work_path}/codegen-analyze/list-2.json
...

(or detail-*.json for non-list specs)

2. Build consensus across pages

For each field in the schema, review all per-page analyses together:

  • Identify common patterns: Do all pages use the same CSS selector / JSON-LD path? If so, that's a strong signal.
  • Resolve disagreements: If analyses recommend different approaches for different pages, determine:

- Is one approach more general? (e.g., a selector that works on all pages vs one that only works on some) - Should you combine approaches? (e.g., try JSON-LD first, fall back to CSS) - Are there structural differences between pages that require conditional logic?

  • Validate target values: Check that the recommended extraction method would produce the target values across all pages.
  • Decide on data source: For each field, decide whether to use HTML selectors, JSON-LD, microdata, URL, or a combination.

3. Generate page object code

No content filtering — ever. Even if the user's prompt asks to filter, exclude, or limit results by value, do NOT implement that logic in the page object. Page objects extract; spiders filter. Mention in your summary that filtering belongs at the spider level.

Generate a complete, self-contained Python module following the web-poet reference. The code must:

  • Work across the domain: not just for the analyzed pages. Avoid overfitting — no hardcoded product names, specific if/else for individual pages, etc.
  • Be simple: prefer the simplest approach that works. Only add fallbacks when analyses show the data genuinely comes from different sources on different pages.
  • Use BrowserPage as base class if a browser response is needed.
  • Handle missing data: return None when a field is not present (never empty string or []).
  • Match schema types: return values matching the JSON schema types.
  • Use recommended libraries: extruct for JSON-LD/microdata, price_parser for prices, jmespath for JSON queries.
  • Include all imports: the module must be self-contained and runnable.

For non-list types (islisttype = false), standard structure:

from web_poet import WebPage, field
# ... other imports as needed

class PageObject(WebPage[dict]):
    # shared helpers as @cached_property if multiple fields need them

    @field
    def field_name(self) -> type | None:
        # extraction logic
        ...

For list types (islisttype = true), use web-poet's SelectorExtractor (see the "Processors for nested fields" section of the web-poet fields reference). Define a private SelectorExtractor subclass for per-item extraction, then iterate over the container selector in the items field. Each analysis file has container_selector and per-field relative selectors.

import logging
from web_poet import WebPage, Returns, SelectorExtractor, field, handle_urls
from {project}.items import Product, ProductListItems

logger = logging.getLogger(__name__)


class _ProductExtractor(SelectorExtractor[Product]):
    @field
    def title(self) -> str | None:
        return self.css("h3 a::attr(title)").get()

    @field
    def price(self) -> str | None:
        return self.css("p.price_color::text").get()


@handle_urls("example.com")
class ProductListPage(WebPage, Returns[ProductListItems]):

    @field
    async def items(self) -> list[Product] | None:
        result = []
        for el in self.css("article.product_pod"):
            try:
                result.append(await _ProductExtractor(el).to_item())
            except Exception:
                logger.error("Failed to extract item from %s", self.url, exc_info=True)
        return result or None

Key rules for list-type code:

  • Define a <ItemClass>Extractor class (e.g. ProductExtractor) that extends

SelectorExtractor[ItemClass] and has one @field method per item field

  • The items field must be async def because to_item() is a coroutine
  • Wrap each await extractor.to_item() in try/except Exception with a

logger.error(..., exc_info=True) call — this isolates extraction failures to individual items and logs a full traceback without aborting the whole list

  • Use return result or None (never return [])
  • The container selector must be consistent across all analyzed pages; if analyses

show different selectors, choose the most general one

  • Field selectors in the extractor are relative to each container element

(self.css(...) on the extractor operates on el, not the full page)

  • Import both ItemClass and WrapperClass from the project's items module

4. Save and report

Save the generated code to {output_path}.

Return a summary of what was generated:

Generated page object with N fields:
  name: CSS h1.product-title::text
  price: JSON-LD offers.price, fallback to CSS span.price
  description: CSS div.product-description (text join)
  rating: JSON-LD aggregateRating.ratingValue
  image_url: CSS img.product-image::attr(src) + urljoin

Include notes on any fields where consensus was difficult or where extraction may be fragile.