SKILL.md
GAIK Toolkit
Current PyPI version: !python ${CLAUDESKILLDIR}/scripts/fetchpypireadme.py --version
Python toolkit for knowledge extraction, capture, and generation. Use when working with:
- Structured data extraction from documents, PDFs, images, or audio
- Schema generation from natural language requirements
- Document parsing (PDF, DOCX, images)
- Audio/video transcription with Whisper + local Whisper backends (Finnish fine-tuned model)
- Transcript enhancement — two-pass LLM error correction
- Parallel transcription with FFmpeg chunking
- Text-to-speech generation
- Document classification
- Text-to-SQL: natural-language querying of PostgreSQL databases and CSV/Excel/Parquet files (DuckDB)
- RAG pipelines: embedder, vector store (Chroma / PostgreSQL), retriever, answer generator
- End-to-end pipelines: AudioToStructuredData, DocumentsToStructuredData, RAGWorkflow
Related Skills
This skill is the overview / reference. Two sibling skills handle workflows:
| Task | Skill |
|---|---|
Create a new installable component package (source + pyproject.toml + extras) |
build-software-component |
| Add an example, then optionally publish (docs → demo app → PyPI tag) | gaik-add-examples |
| Understand the toolkit: components, config, repo layout, docs update map | this skill |
Quick Links
- Documentation: https://gaik-project.github.io/gaik-toolkit/
- Live Demo: https://gaik-demo.2.rahtiapp.fi/ (registration required)
- GitHub: https://github.com/GAIK-project/gaik-toolkit
- Source Code:
implementation_layer/src/gaik/ - PyPI: https://pypi.org/project/gaik/
Repository Structure
| Path | Description |
|---|---|
implementation_layer/src/gaik/ |
Python package source (building blocks + software modules) |
implementationlayer/toolkitdemo_app/ |
Next.js + FastAPI interactive demo app (bun + uv) |
guidance_layer/website/ |
Documentation website (Fumadocs/Next.js, deployed to GitHub Pages) |
guidance_layer/website/content/docs/ |
Documentation source (.mdx files) |
implementation_layer/no-code-assets/ |
Prompt templates and agent skills for no-code usage |
strategy_layer/ |
Value evaluation framework, AI maturity assessment |
business_layer/ |
GenAI product canvas templates |
Toolkit Demo App
Interactive web app at implementationlayer/toolkitdemo_app/. Next.js 16 + FastAPI (bun + uv).
- Live: https://gaik-demo.2.rahtiapp.fi/ (registration required)
- Dev:
bun run dev:all(runs both frontend and API) - See: [Demo App Reference](references/demo-app.md) for full architecture, routes, and conventions
Documentation Website
Fumadocs/Next.js site at guidance_layer/website/. Content in .mdx files under content/docs/.
- Live: https://gaik-project.github.io/gaik-toolkit/
- Dev:
pnpm dev(fromguidance_layer/website/-- uses pnpm, not bun) - See: [Docs Website Reference](references/docs-website.md) for content structure and editing guide
Installation
Install via pip with optional extras: pip install "gaik[extract]", pip install "gaik[all-cpu]", etc. See [Installation Reference](references/installation.md) for all available extras and setup.
Environment Variables
Azure OpenAI (recommended):
AZURE_API_KEY=your-key
AZURE_ENDPOINT=https://your-resource.openai.azure.com/
AZURE_DEPLOYMENT=gpt-5.4
AZURE_API_VERSION=2025-03-01-preview
OpenAI:
OPENAI_API_KEY=your-key
OPENAI_MODEL=gpt-5.4
Configuration Pattern
Two parallel surfaces ship since gaik>=0.3.21. Pick the simpler one for OpenAI/Azure-only use cases; pick the multi-provider one when the same code needs to switch between OpenAI, Azure, Anthropic, or Google.
Legacy surface (OpenAI/Azure only — bit-for-bit unchanged):
from gaik.software_components.config import get_openai_config, create_openai_client
config = get_openai_config(use_azure=True) # Azure OpenAI
config = get_openai_config(use_azure=False) # Standard OpenAI
client = create_openai_client(config) # OpenAI/AzureOpenAI client
Multi-provider surface (Anthropic, Google, OpenAI, Azure):
from gaik.software_components.llm import get_llm_config, create_llm_client
config = get_llm_config("google") # or "anthropic", "openai", "azure"
client = create_llm_client(config) # ProviderClient with chat/chat_parsed/chat_stream/embed
gaik[llm-anthropic] and gaik[llm-google] extras pull in the provider SDKs on demand. Audio components (transcriber, TTS) and vision parsing only support OpenAI/Azure — they raise NotImplementedError for native Anthropic/Google. For multi-provider vision, use MultimodalParser. For Gemini-via-OpenAI-compat-endpoint, set OPENAIBASEURL=https://generativelanguage.googleapis.com/v1beta/openai/ on a standard OpenAI config — every component then routes through the legacy path.
Building Blocks
Core classes in gaik.software_components.*. For detailed API and constructor parameters, see [Building Blocks Reference](references/building-blocks.md).
| Component | Import | Key Method |
|---|---|---|
| SchemaGenerator | from gaik.software_components.extractor import SchemaGenerator |
generateschema(userrequirements) |
| DataExtractor | from gaik.software_components.extractor import DataExtractor |
extract(extraction_model, requirements, ...) |
| VisionExtractor | from gaik.softwarecomponents.visionextractor import VisionExtractor |
extract(filepaths, userrequirements, extractionmodel=None, requirements=None, schemadir=None) → VisionExtractionResult (single-pass PDF/image → structured data; OpenAI / Claude / Google) |
| VisionParser | from gaik.software_components.parsers import VisionParser |
convert_pdf(path) → list[str] per page |
| PyMuPDFParser | from gaik.software_components.parsers import PyMuPDFParser |
parsepdf(path) → str · parsedocument(path) → dict |
| DocxParser | from gaik.software_components.parsers import DocxParser |
parsedocx(path) → str · parsedocument(path) → dict |
| DoclingParser | from gaik.software_components.parsers import DoclingParser |
parse_document(path) → dict |
| VisionPlusParser | from gaik.software_components.parsers import VisionPlusParser |
parse_document(path) → dict (markdown + per-element metadata) |
| DoclingApiClientParser | from gaik.software_components.parsers import DoclingApiClientParser |
parse_document(path) → dict (remote Docling result) |
| MultimodalParser | from gaik.software_components.parsers import MultimodalParser |
parse(pdf_path) → ParseResult (OpenAI / Claude / Gemini) |
| Transcriber | from gaik.software_components.transcriber import Transcriber |
transcribe(path) → TranscriptionResult |
| TranscriptEnhancer | from gaik.softwarecomponents.enhancetranscript import TranscriptEnhancer |
enhancetext(text) / enhancefile(path) |
| ParallelTranscriber | from gaik.softwarecomponents.paralleltranscriber import ParallelTranscriber |
transcribe(path) → TranscriptionResult |
| TextToSpeech | from gaik.softwarecomponents.textto_speech import TextToSpeech |
synthesize(text) → SpeechSynthesisResult |
| DocumentClassifier | from gaik.softwarecomponents.docclassifier import DocumentClassifier |
classify(fileordir, classes) |
| FormUnderstander | from gaik.softwarecomponents.formunderstander import FormUnderstander |
cleanlabels(fields, languagehint="fi") → dict[str, str] (cryptic ASP.NET / generated form ids → readable labels) |
| PostgresAgent | from gaik.softwarecomponents.postgresagent import PostgresAgent |
ask(question) → AnswerResult (text-to-SQL agent: introspects schema, generates validated read-only SQL, runs it, answers; also getschema() / generatesql() / query() / run_sql(); install gaik[postgres-agent]) |
| TabularAgent | from gaik.softwarecomponents.tabularagent import TabularAgent |
ask(question) → AnswerResult (text-to-SQL agent for files: loads CSV/Excel/Parquet/JSON into DuckDB, profiles columns, generates validated read-only SQL, answers; cleans up messy report sheets — title rows, subtotals, Nordic comma-decimals; one table per Excel sheet so cross-sheet joins work; same getschema() / runsql() tool surface as PostgresAgent; install gaik[tabular-agent]) |
| LLMJudge | from gaik.software_components.validators import LLMJudge |
validate(sourcepages, extracted, rubric) → ValidationResult (rubric scoring; Likert 1-5 via rubric.scoringmode="likert15") / detecthallucinations(source, extracted) → schema-agnostic post-validator / judgetext_pair(a, b) → text-vs-text equivalence (multi-provider) |
| LLMJudgePanel | from gaik.software_components.validators import LLMJudgePanel |
validate(source_pages, extracted, rubric) → JudgePanelResult (3+ judges, majority vote, agreement metric) |
| compare_pairwise | from gaik.softwarecomponents.validators import comparepairwise |
comparepairwise(judge, pages, a, b, swapand_average=True) → PairwiseResult (A/B with position-bias mitigation) |
| calibrateagainsthuman_labels | from gaik.softwarecomponents.validators import calibrateagainsthumanlabels |
calibrateagainsthuman_labels(judge, dataset) → CalibrationReport (Pearson r vs. human raters) |
| FinnishTextProcessor | from gaik.softwarecomponents.RAG.finnishtext_processor import FinnishTextProcessor |
lemmatize(text) / totsvectortext(text) / expand_query(text) (Finnish lemmatization + compound splitting; backends: voikko / spacy / uralic / simple) |
| ExtractionEvaluator | from gaik.software_components.evaluators import ExtractionEvaluator |
evaluatedataset(dataset, extractedoutputs) → ExtractionEvaluationResult (field-level P/R/F1 + hallucination rate; optional semantic mode via LLMJudge) |
| RAGEvaluator | from gaik.software_components.evaluators import RAGEvaluator |
evaluatedataset(items) → RAGEvaluationResult (RAGAS-style faithfulness / answerrelevance / contextprecision / contextrecall via LLMJudge) |
| BatchEvaluationRunner | from gaik.software_components.evaluators import BatchEvaluationRunner |
run(dataset) → RunnerResult (applies a pipeline callable over a dataset; on_error="skip" tolerates failures) |
Parser notes
Every parser ships a class and a module-level convenience function, and the two do not agree on return type — the class method gives you the text, the function gives you a metadata dict. Reaching for the shorter name is the easy mistake:
parser = PyMuPDFParser()
text = parser.parse_pdf("doc.pdf") # -> str
result = parse_pdf("doc.pdf") # -> dict, text lives under result["text_content"]
The same split applies to DocxParser.parsedocx / parsedocx, and every parse_document variant returns a dict on both the class and the function.
Transcriber notes
- Models:
"whisper","whisper-1","gpt-4o-transcribe","whisper_local" enhanced_transcript=Trueruns output through TranscriptEnhancer (two-pass LLM correction)whisperlocalrequireslocalapibase+localapi_key;language="fi"selects Finnish fine-tuned model- ParallelTranscriber uses FFmpeg chunking; requires
ffmpeg+ffprobeon$PATH
SRT/VTT Utilities
from gaik.software_components.transcriber import segments_to_srt, segments_to_vtt, parse_srt, chunk_segments
Video Search Helpers
from gaik.software_components.RAG.pg_vector_store import PgVectorStore, ingest_video_segments, format_search_results
RAG Building Blocks
Core RAG classes in gaik.software_components.RAG.*. For full API, see [RAG Reference](references/rag.md).
| Component | Import | Key Method |
|---|---|---|
| Embedder | from gaik.software_components.RAG.embedder import Embedder |
embed(docs), embed_query(text) |
| VectorStore | from gaik.softwarecomponents.RAG.vectorstore import VectorStore |
add(docs, embeddings), search(vec, top_k) |
| PgVectorStore | from gaik.softwarecomponents.RAG.pgvector_store import PgVectorStore |
searchhybrid(vec, text, topk) |
| Retriever | from gaik.software_components.RAG.retriever import Retriever |
search(query, topk, hybridsearch, re_rank) |
| Ranker | from gaik.software_components.RAG.ranker import Ranker |
fuse(*lists, weights) → weighted RRF; also rerank(query, results), orderby(results, field, direction) for asc/desc, todocuments(results); reorders lists you already have, no IO; install gaik[ranker] (cross-encoder needs gaik[ranker-rerank]) |
| AnswerGenerator | from gaik.softwarecomponents.RAG.answergenerator import AnswerGenerator |
generate(query, documents, stream) |
| VisionRagParser | from gaik.softwarecomponents.RAG.ragparser_vision import VisionRagParser |
convertdoctochunkswith_vision(path) |
| DoclingRagParser | from gaik.softwarecomponents.RAG.ragparser_docling import DoclingRagParser |
convertpdftochunkswith_metadata(path) |
End-to-End Pipelines
Composed pipelines in gaik.software_modules.*. For full API, see [Software Components Reference](references/software-components.md).
| Pipeline | Flow | Import |
|---|---|---|
| AudioToStructuredData | Audio → Transcript → Schema → JSON | from gaik.softwaremodules.audiotostructureddata import AudioToStructuredData |
| DocumentsToStructuredData | PDF/DOCX → Parse → Schema → JSON | from gaik.softwaremodules.documentstostructureddata import DocumentsToStructuredData |
| RAGWorkflow | PDF → Parse → Embed → Store → Retrieve → Answer | from gaik.softwaremodules.RAGworkflow import RAGWorkflow |
| MultiSourceReportGenerator | Mixed files (PDF/DOCX/Excel/audio/images) → Normalize → Sectioned Markdown report | from gaik.softwaremodules.multisourcereportgenerator import MultiSourceReportGenerator |
The first three pipelines follow: pipeline = Pipeline(useazure=True) → result = pipeline.run(filepath, user_requirements, ...). MultiSourceReportGenerator instead takes a set of source files plus a report structure (section titles + per-section instructions) and returns the assembled Markdown report with a per-section breakdown.
Architecture Overview
| Level | Concept | Examples |
|---|---|---|
| Service | Logical capability | speechtotext, documentparsing, informationextraction, rag |
| Building block | Atomic toolkit class/function | Transcriber, ParallelTranscriber, TranscriptEnhancer, TextToSpeech, SchemaGenerator, DataExtractor, VisionParser, Embedder, VectorStore, PgVectorStore, Retriever, AnswerGenerator |
| Software component | Composed, workflow-ready unit | AudioToStructuredData, DocumentsToStructuredData, RAGWorkflow, MultiSourceReportGenerator |
Observability
Token usage, execution time, and provider-specific pricing for all LLM calls. A shared UsageRecord type ensures all components report data in the same format regardless of provider (OpenAI / Azure / Anthropic / Google).
from gaik.observability import (
UsageRecord, build_usage_record, # uniform usage shape
compute_cost_usd, lookup_price, # cost from prompt/completion tokens
measure_duration, # context-manager timing helper
openai_usage_to_dict, # OpenAI-shape → dict normalizer
OPENAI_PRICING_PER_M, ANTHROPIC_PRICING_PER_M, GEMINI_PRICING_PER_M,
)
Use when building a dashboard, logging pipeline, or compliance reporter that needs a unified cost/duration report across providers.
Use Cases
Documented in guidance_layer/website/content/docs/use-cases/: incident reporting, dental transcription & captioning, semantic dental video search, construction diary, dental learning assistant, purchase order processing, report writing, sales proposal generation, customer onboarding.
When to Update Documentation
When adding or modifying a component, update both documentation locations:
| What changed | Update |
|---|---|
| New/modified building block or pipeline | guidancelayer/docs/softwarecomponents/ or guidancelayer/docs/softwaremodules/ |
| New/modified building block or pipeline | guidance_layer/website/content/docs/toolkit/software-components.mdx or software-modules.mdx |
| New use case or example | guidance_layer/website/content/docs/use-cases/ (new .mdx file) |
| New examples added | implementation_layer/examples/ + README updated |
guidance_layer/docs/: Technical Markdown docs (API-level details, constructor params)guidance_layer/website/content/docs/: User-facing MDX for the Fumadocs website- Run
pnpm devfromguidance_layer/website/to preview website changes - For the gated, step-by-step publish flow (docs → demo app → PyPI tag), use the
gaik-add-examplesskill Step 6 — the canonical follow-up workflow
Gotchas
Non-obvious things that cause real mistakes in this repo. Check here before assuming.
- Docs website uses
pnpm, notbun. Everything else intoolkitdemoapp/usesbun. Runningbun devinsideguidance_layer/website/silently installs a second lockfile and breaks Fumadocs build. - Fumadocs needs
meta.jsonupdates. When adding a new.mdxpage undercontent/docs/, also add it to the parent directory'smeta.json, or it will not appear in the navigation. ParallelTranscriberrequiresffmpeg+ffprobeon$PATH. On Windows that means installing ffmpeg and adding itsbin/to PATH — there is no Python wheel fallback.whisperlocalmodel needslocalapibase+localapikey.language="fi"switches to the Finnish fine-tuned model. Leavinglocalapi_baseunset fails with an unhelpful OpenAI-style error.- Never edit
versionstrings by hand. The package version is derived from the git tag bysetuptools-scm. Manual edits desync the wheel and break the PyPI publish workflow's version validation. - **CORSORIGINS must be valid JSON, not
"*".** In the OpenShift API deployment,CORSORIGINS='[""]' works; plaincrashloops (pydantic-settings parses the env var as a list[str]). - OpenAI structured outputs can't use
additionalProperties. Prefer an explicit list-of-entries model (seeFormUnderstander.LabelEntry) over a free-form dict.
Detailed References
- [Building Blocks API](references/building-blocks.md) - Constructor params, return types, all options
- [RAG Building Blocks](references/rag.md) - RAG components: Embedder, stores, Retriever, AnswerGenerator
- [Software Components](references/software-components.md) - Pipeline patterns, schema persistence, batch processing
- [Evaluators](references/evaluators.md) - ExtractionEvaluator, RAGEvaluator, BatchEvaluationRunner (LLMJudge v2 -based)
- [Examples](references/examples.md) - Complete working examples (invoice extraction, RAG, parallel transcription, etc.)
- [Demo App](references/demo-app.md) - Demo app architecture, routes, env vars, deployment
- [Docs Website](references/docs-website.md) - Documentation site structure and editing guide
- [Installation](references/installation.md) - All pip install extras and system dependencies
- [Maintenance](references/maintenance.md) - Skill maintenance and PyPI fetch script