smithery/vamseeachanta

data-pipeline-processor

Process data files through transformation pipelines with validation, cleaning, and export. Use for CSV/Excel/JSON data processing, encoding handling, batch operations, and data transformation workflows.

Installation

$ npx skills add smithery/vamseeachanta --skill data-pipeline-processor

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery/vamseeachanta · top by installs.

npx skills add smithery/vamseeachanta

Browse all from smithery/vamseeachanta

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Skill metadata

Parsed from SKILL.md frontmatter.

Version1.1.0

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 4,199 B
  • docs SUMMARY.md 233 B

History

  1. First recorded snapshot · 0 installs

SKILL.md

Data Pipeline Processor

Quick Start

import pandas as pd
from pathlib import Path

# Simple pipeline: Load -> Transform -> Export
df = pd.read_csv("data/raw/source.csv")

# Transform
df = df[df['value'] > 0]  # Filter
df['date'] = pd.to_datetime(df['date'])  # Convert types
df = df.sort_values('date')  # Sort

# Export
Path("data/processed").mkdir(parents=True, exist_ok=True)
df.to_csv("data/processed/cleaned.csv", index=False)

print(f"Processed {len(df)} rows")

When to Use

  • Processing CSV/Excel/JSON files with validation
  • Data cleaning and transformation workflows
  • Batch file processing with aggregation
  • Handling encoding issues (UTF-8, Latin-1 fallback)
  • ETL (Extract, Transform, Load) operations
  • Data quality checks and reporting

Core Pattern

Input (CSV/Excel/JSON) -> Validate -> Transform -> Analyze -> Export

Implementation

Data Reader with Encoding Detection

import pandas as pd
from pathlib import Path
from typing import Any, Dict, List, Optional, Union
import logging
import chardet

logger = logging.getLogger(__name__)



*See sub-skills for full details.*
### Data Validator

from dataclasses import dataclass, field from typing import Callable, List, Dict, Any

@dataclass class ValidationResult: """Result of data validation.""" isvalid: bool errors: List[str] = field(defaultfactory=list)

See sub-skills for full details.

Data Transformer

class DataTransformer:
    """Apply transformations to data."""

    def __init__(self, df: pd.DataFrame):
        self.df = df.copy()

    def rename_columns(self, mapping: Dict[str, str]) -> 'DataTransformer':
        """Rename columns."""
        self.df = self.df.rename(columns=mapping)

*See sub-skills for full details.*
### Data Exporter

class DataExporter: """Export data to various formats."""

@staticmethod def tocsv(df: pd.DataFrame, path: str, kwargs) -> str: """Export to CSV.""" Path(path).parent.mkdir(parents=True, existok=True) df.to_csv(path, index=False, kwargs) return path

See sub-skills for full details.

Pipeline Orchestrator

from dataclasses import dataclass
from typing import List, Dict, Any, Optional


@dataclass
class PipelineConfig:
    """Configuration for data pipeline."""
    input_path: str
    output_path: str

*See sub-skills for full details.*

## YAML Configuration Format

### Basic Pipeline Config

config/pipelines/data_clean.yaml

input: path: data/raw/source.csv options: delimiter: "," skiprows: 1

validation:

See sub-skills for full details.

Aggregation Pipeline

# config/pipelines/monthly_summary.yaml

input:
  path: data/processed/daily_data.csv

validation:
  required_columns:
    - date
    - category

*See sub-skills for full details.*

## Related Skills

- [yaml-workflow-executor](../yaml-workflow-executor/SKILL.md) - Workflow orchestration
- [engineering-report-generator](../engineering-report-generator/SKILL.md) - Report generation
- [parallel-file-processor](../parallel-file-processor/SKILL.md) - Parallel file operations

---

## Version History

- **1.1.0** (2026-01-02): Upgraded to SKILL_TEMPLATE_v2 format with Quick Start, Error Handling, Metrics, Execution Checklist, additional examples
- **1.0.0** (2024-10-15): Initial release with DataReader, DataValidator, DataTransformer, pipeline orchestration

## Sub-Skills

- [Example 1: Simple CSV Processing (+3)](example-1-simple-csv-processing/SKILL.md)
- [Do (+6)](do/SKILL.md)

## Sub-Skills

- [Error Handling](error-handling/SKILL.md)
- [Execution Checklist](execution-checklist/SKILL.md)
- [Metrics](metrics/SKILL.md)