tiangong-ai/skills

document-granular-decompose

Upload local documents to TianGong AI Unstructure `/mineru_with_images` API for fine-grained parsing and return only plain fulltext content.

First seen Mar 9, 2026

Installation

$ npx skills add tiangong-ai/skills --skill document-granular-decompose

Summary

  • Upload local documents to TianGong AI Unstructure `/mineru_with_images` API for fine-grained parsing and return only plain fulltext content.
  • Use when a task needs document fulltext extraction with `return_txt=true`, strict file-type allowlist validation, API base URL/auth token from environment variables, and optional provider/model overrides.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from tiangong-ai/skills · top by installs.

npx skills add tiangong-ai/skills

Browse all from tiangong-ai/skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 9
License LICENSE
Default branch main
Open issues 3
Status Active

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 3,894 B
  • docs SUMMARY.md 380 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 103 installs

SKILL.md

Document Granular Decompose

Core Goal

  • Parse a local document through POST /mineruwithimages.
  • Always force return_txt=true.
  • Read environment variables for endpoint, request identity, and model routing:

- UNSTRUCTUREDAPIBASEURL (example: https://your-unstructured-host:7770) - UNSTRUCTUREDAUTHTOKEN - UNSTRUCTUREDPROVIDER (optional) - UNSTRUCTURED_MODEL (optional)

  • Return only plain fulltext (prefer API txt; fallback to joined result[].text).

Triggering Conditions

  • Need robust document fulltext extraction for PDF/Office/image files.
  • Need image-aware MinerU parsing but only textual output for downstream chunking/search/summarization.
  • Need to standardize provider/model/token input via environment variables instead of ad-hoc command parameters.

Workflow

  1. Prepare environment variables.
export UNSTRUCTURED_AUTH_TOKEN="your-fastapi-bearer-token"
export UNSTRUCTURED_API_BASE_URL="https://your-unstructured-host:7770"
# Optional routing overrides. Omit them to let the server choose its defaults.
export UNSTRUCTURED_PROVIDER="vllm"
export UNSTRUCTURED_MODEL="Qwen/Qwen3.5-122B-A10B-FP8"
  1. Run extraction and print fulltext to stdout.
python3 scripts/mineru_fulltext_extract.py \
  --file "/absolute/path/to/document.pdf"
  1. Save fulltext to a local file when needed.
python3 scripts/mineru_fulltext_extract.py \
  --file "/absolute/path/to/document.pdf" \
  --output "/absolute/path/to/fulltext.txt"

Request Contract

  • Endpoint resolution:

- --api-url if provided - else UNSTRUCTUREDAPIBASEURL + /mineruwith_images - else fail fast with missing environment variable error

  • Method: POST multipart form.
  • Query params:

- Force return_txt=true (always set by script).

  • Form fields sent:

- file (required) - provider (optional, from UNSTRUCTUREDPROVIDER when set) - model (optional, from UNSTRUCTUREDMODEL when set)

  • Header sent:

- Authorization: Bearer $UNSTRUCTUREDAUTHTOKEN

Supported File Types (Strict)

  • Supported file types:

- .bmp, .doc, .docm, .docx, .dot, .dotx, .gif, .jp2, .jpeg, .jpg, .odp, .odt, .pdf, .png, .pot, .potx, .pps, .ppsx, .ppt, .pptm, .pptx, .tiff, .webp, .xls, .xlsm, .xlsx, .xlt, .xltx

  • Office formats:

- .doc, .docm, .docx, .dot, .dotx, .odp, .odt, .pot, .potx, .pps, .ppsx, .ppt, .pptm, .pptx, .xls, .xlsm, .xlsx, .xlt, .xltx

  • Any other extension is rejected before sending API requests.

Output Rules

  • Success output must be plain text fulltext only.
  • Normalize the confirmed upstream Markdown underscore escape (\ to )

in both supported response paths; preserve other backslashes and escapes.

  • Fulltext source priority:

1. response.txt 2. join non-empty response.result[].text by blank lines

  • Do not output chunk metadata/json unless the user explicitly requests debugging.

Error Handling

  • Missing required env vars (UNSTRUCTUREDAPIBASEURL, UNSTRUCTUREDAUTH_TOKEN): fail fast with actionable message.
  • Missing UNSTRUCTUREDPROVIDER or UNSTRUCTUREDMODEL: omit the form field and let the service choose its default.
  • HTTP 401/403: report token/auth issue.
  • HTTP 4xx/5xx: print status and API error body if available.
  • Missing text in response: fail with explicit schema mismatch error.

References

  • references/env.md
  • references/request-response.md

Assets

  • assets/config.example.env

Scripts

  • scripts/minerufulltextextract.py