smithery/jhd3197

scaffold-extraction

Scaffold a complete extraction pipeline for a new domain — Pydantic model, field definitions, example script, and tests. Use when building a new extraction use case like medical records, invoices, or reviews.

Installation

$ npx skills add smithery/jhd3197 --skill scaffold-extraction

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery/jhd3197.

npx skills add smithery/jhd3197

Browse all from smithery/jhd3197

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Skill metadata

Parsed from SKILL.md frontmatter.

Version1.0
More metadata
author
prompture
version
1.0

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,445 B
  • docs SUMMARY.md 237 B

History

  1. First recorded snapshot · 0 installs

SKILL.md

Scaffold an Extraction Pipeline

Builds the full pipeline for a new extraction domain: model, fields, example, and tests.

Before Starting

Ask the user:

  • Domain / use case — what kind of data to extract
  • Fields — list with types, or infer from a sample text
  • Provider/model — default ollama/gpt-oss:20b
  • Method — one-shot or stepwise (see selection guide below)

Method Selection Guide

Scenario Method
Simple model, < 8 fields extractwithmodel (1 LLM call)
Complex model, 8+ fields stepwiseextractwith_model (N calls, more accurate)
No Pydantic model, raw schema extractandjsonify
Structured input (CSV, JSON) extractfromdata (TOON, 45-60% token savings)
DataFrame input extractfrompandas
Non-JSON output render_output

Deliverables

1. Pydantic Model

from pydantic import BaseModel, Field

class InvoiceData(BaseModel):
    vendor_name: str = Field(description="Company that issued the invoice")
    total_amount: float = Field(description="Total amount due")
    currency: str = Field(default="USD", description="ISO 4217 currency code")

Rules:

  • Field(description=...) on every field — these become LLM instructions
  • Type-appropriate defaults: "", 0, 0.0, []
  • Optional[T] = None only when a field genuinely might not exist

2. Field Definitions (optional)

Register in prompture/field_definitions.py only if fields are general-purpose. Use the add-field skill.

3. Example Script

Create examples/{domain}extractionexample.py. Use the add-example skill.

4. Tests

Unit test for the Pydantic model + integration test for end-to-end extraction. Use the add-test skill.

Sample Call

from prompture import extract_with_model

result = extract_with_model(
    model_cls=InvoiceData,
    text=invoice_text,
    model_name="ollama/gpt-oss:20b",
    instruction_template="Extract all invoice details from the following document:",
)

invoice = result["model"]    # Pydantic instance
usage = result["usage"]      # Token counts and cost