nomadamas/autorag-research · Archived

create-retrieval-plugin

Guide developers through creating a custom retrieval pipeline plugin for AutoRAG-Research. Walks through scaffolding, implementing BaseRetrievalPipeline methods, writing YAML configs, testing, and installing. Use when building a new search/retrieval strategy (e.g., Elasticsearch, ColBERT, custom vector search).

First seen Jun 20, 2026

Installation

$ npx skills add nomadamas/autorag-research --skill create-retrieval-plugin

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from nomadamas/autorag-research.

npx skills add nomadamas/autorag-research

Browse all from nomadamas/autorag-research

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 141
License LICENSE
Default branch main
Open issues 29
Status Archived

Skill metadata

Parsed from SKILL.md frontmatter.

Allowed toolsBash, Read, Write, Edit

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,931 B
  • docs SUMMARY.md 343 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

Create Retrieval Plugin

Workflow

1. Scaffold

autorag-research plugin create my_search --type=retrieval

Read the generated pipeline.py, pyproject.toml, YAML config, and test file to understand the structure.

2. Implement

For the shared pipeline implementation and testing rules, read:

  • aiinstructions/pipelineimplementer.md
  • aiinstructions/pipelinetest_writer.md
  • aiinstructions/pipelinearchitecture_mapper.md

Implement the two abstract methods in the pipeline class:

  • retrievebyid(queryid, top_k) — retrieve using query ID (query exists in DB with stored embedding)
  • retrievebytext(querytext, top_k) — retrieve using raw text (may need on-the-fly embedding)

Both must return list[dict[str, Any]] with doc_id (chunk ID) and score keys.

DO NOT add your own asyncio.gather, asyncio.Semaphore, or any concurrency control.
The base pipeline's run() already handles parallel execution of all queries via
runwithconcurrencylimit() (semaphore + gather), controlled by the maxconcurrency
config parameter. Your method is called once per single query — just implement the
retrieval logic for that one query.

Custom parameters: Add fields to your config class and pass them via getpipelinekwargs() → accept them in the pipeline constructor. See bm25.py for a real example.

3. Write tests and install

cd my_search_plugin
pip install -e .   # or: uv pip install -e .
cd .. && autorag-research plugin sync

Verify: ls configs/pipelines/retrieval/my_search.yaml

Key Files

Purpose Path
Base config class autorag_research/config.py → BaseRetrievalPipelineConfig
Base pipeline class autorag_research/pipelines/retrieval/base.py → BaseRetrievalPipeline
Service layer autoragresearch/orm/service/retrievalpipeline.py → RetrievalPipelineService
Plugin entry point discovery autoragresearch/pluginregistry.py

Examples

Study these existing implementations for patterns:

  • autorag_research/pipelines/retrieval/bm25.py — BM25 retrieval (simple)
  • autoragresearch/pipelines/retrieval/vectorsearch.py — Vector similarity search
  • autorag_research/pipelines/retrieval/hybrid.py — Hybrid (BM25 + vector)
  • autorag_research/pipelines/retrieval/hyde.py — HyDE (Hypothetical Document Embeddings)
  • YAML configs: configs/pipelines/retrieval/bm25.yaml, configs/pipelines/retrieval/vector_search.yaml