smithery.ai

agent-survey-corpus

Download a small corpus of open-access arXiv survey/review PDFs about agentic systems and extract text for style learning.

First seen May 3, 2026

Installation

$ npx skills add https://smithery.ai

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery.ai · top by installs.

npx skills add https://smithery.ai

Browse all from smithery.ai

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Skill metadata

Parsed from SKILL.md frontmatter.

Declared agents codex

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,529 B
  • docs SUMMARY.md 593 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

Agent Survey Corpus (arXiv PDFs → text extracts)

Triggers & routing

  • Trigger: agent survey corpus, ref corpus, download surveys, 学习综述写法, 下载 survey.
  • Use when: you want to study how real agent surveys structure sections (6–8 H2), size subsections, and write evidence-backed comparisons.

Goal: create a small, local reference library so you can learn from real agent surveys when refining:

  • C2 outline structure (paper-like sectioning)
  • C4 tables/claims organization
  • C5 writing style and density

This is intentionally not part of the pipeline; it is an optional, repo-level toolkit.

Inputs

  • ref/agent-surveys/arxiv_ids.txt

Outputs

  • ref/agent-surveys/pdfs/
  • ref/agent-surveys/text/
  • ref/agent-surveys/STYLE_REPORT.md (tracked; auto-generated summary)

Workflow

  1. Edit ref/agent-surveys/arxiv_ids.txt (one arXiv id per line).
  2. Run the downloader to fetch PDFs and extract the first N pages to text.
  3. Skim the extracted text under ref/agent-surveys/text/:

- look at section counts (H2), subsection granularity (H3), and how they transition between chapters. - identify repeated rhetorical patterns you want the pipeline writer to imitate.

Script

Quick Start

  • uv run python .codex/skills/agent-survey-corpus/scripts/run.py --help
  • uv run python .codex/skills/agent-survey-corpus/scripts/run.py --workspace . --max-pages 20

All Options

  • --workspace <dir> (use . to write into repo root)
  • --inputs <semicolon-separated> (default: ref/agent-surveys/arxiv_ids.txt)
  • --max-pages <N> (default: 20)
  • --sleep <seconds> (default: 1.0)
  • --overwrite (re-download + re-extract)

Examples

  • Download/extract into repo root ref/:

- uv run python .codex/skills/agent-survey-corpus/scripts/run.py --workspace . --max-pages 20

  • Download/extract into a specific folder (treated as workspace root):

- uv run python .codex/skills/agent-survey-corpus/scripts/run.py --workspace /tmp/surveys --max-pages 30

Troubleshooting

  • Download fails / timeout: rerun with a larger --sleep, or try fewer ids.
  • Text extract is empty: the PDF may be scanned; try another survey or increase --max-pages.
  • Files showing up in git status: PDFs/text are ignored via .gitignore (ref//pdfs/, ref//text/).