SKILL.md
<!-- BEGIN:compound:skill-managed -->
Quality Scoring Implementation
Implement data quality scoring features for Vibe Piper validation module.
Overview
This skill provides comprehensive data quality assessment with 0-100 scale scoring across five dimensions: completeness, accuracy, uniqueness, consistency, and timeliness.
When To Use
- User requests implementation of data quality scoring features
- Ticket requires quality score calculation, multi-dimensional assessment, historical trend tracking, threshold alerts, or improvement recommendations
Architecture
Core Components
- QualityScore Class: Main result object with all dimension scores (0-100 scale)
- completenessscore: float (0-100) - accuracyscore: float (0-100) - uniquenessscore: float (0-100) - consistencyscore: float (0-100) - timelinessscore: float (0-100) - overallscore: float (0-100), weighted average - metrics: Dict[str, QualityMetric] (detailed per-dimension metrics) - weights: Dict[str, float] (weights used for calculation) - timestamp: datetime
- QualityThresholdConfig: Configuration for thresholds and alerting
- overallthreshold: float (default: 75.0) - dimensionthresholds: Dict[str, float] (per-dimension thresholds) - alertonthreshold_breach: bool
- QualityAlert Class: Alert object for threshold breaches
- alerttype: str - dimension: str - currentvalue: float - threshold: float - severity: str (critical|warning|info) - timestamp: datetime - message: str
- QualityRecommendation Class: Improvement suggestions
- category: str - priority: str (critical|high|medium|low) - description: str - action: str - expected_impact: str
- QualityTrend Class: Historical trend analysis
- dimension: str - timestamps: Tuple[datetime, ...] - values: Tuple[float, ...] - trenddirection: str (improving|declining|stable) - changerate: float - moving_average: float
- QualityHistory Class: Complete history for an asset
- assetname: str - scores: Tuple[QualityScore, ...] - trends: Dict[str, QualityTrend] - createdat: datetime | None - updated_at: datetime | None
- QualityDashboard Class: Comprehensive view
- currentscore: float - dimensionscores: Dict[str, float] - historicaltrends: Dict[str, QualityTrend] - alerts: Tuple[QualityAlert, ...] - recommendations: Tuple[QualityRecommendation, ...] - lastupdated: datetime
- ColumnQualityResult Class: Column-level quality (0-100 scale)
- columnname: str - completeness: float (0-100) - accuracy: float (0-100) - uniqueness: float (0-100) - nullcount: int - duplicatecount: int - uniquecount: int - distinct_count: int
Implementation Functions
Main Functions
- calculatequalityscore(): Primary entry point for comprehensive quality scoring
- Parameters: records, columns (optional), weights (optional), config (QualityThresholdConfig), timestampfield (optional), maxagehours (optional) - Returns: QualityScore with all 5 dimensions on 0-100 scale - Applies configurable weights to calculate overallscore as weighted average - Integrates with checkfreshness() from vibepiper.quality for timeliness dimension - Default weights: completeness=0.3, accuracy=0.3, uniqueness=0.2, consistency=0.1, timeliness=0.1
- trackqualityhistory(): Track quality scores over time
- Parameters: assetname, score, maxhistory (default: 100) - Returns: QualityHistory with trend analysis per dimension - Uses in-memory qualityhistorystore dict (production should use database) - Calls analyze_trend() to calculate direction and change rate
- generatequalityalerts(): Generate alerts for threshold breaches
- Parameters: score, config (QualityThresholdConfig) - Returns: Tuple[QualityAlert, ...] - Checks overallscore against overallthreshold - Checks each dimension against dimension_thresholds - Sets severity: critical if < 50% of threshold, warning if < 75%, info if below threshold
- generatequalityrecommendations(): Generate improvement suggestions
- Parameters: score, records (optional) - Returns: Tuple[QualityRecommendation, ...] - Generates recommendations for dimensions with scores < 90% - Priority levels: critical (< 50%), high (< 75%), medium (< 90%) - Provides action and expected_impact for each recommendation
- createqualitydashboard(): Create comprehensive quality dashboard
- Parameters: asset_name, score, config (optional), history (optional) - Returns: QualityDashboard with all quality information - Consolidates current score, dimension scores, historical trends, alerts, and recommendations
Supporting Functions
- analyzetrend(): Private helper for trend analysis
- Calculates linear regression slope for change rate - Determines trenddirection based on slope magnitude - Calculates movingaverage with configurable window_size (default: 5)
- calculatecolumnquality(): Column-level quality assessment (0-100 scale)
- Parameters: records, column - Returns: ColumnQualityResult - Calculates completeness: (1 - nullcount/totalcount) 100 - Calculates uniqueness: (uniquecount/len(nonnullvalues)) 100 - Calculates accuracy: (validcount/len(values)) * 100
Updated Functions
- calculate_completeness(): Updated to handle DataRecord correctly
- calculate_validity(): Existing, unchanged
- calculate_uniqueness(): Updated to use 0-100 scale
- calculate_consistency(): Existing, unchanged
Integration Points
- Validation Module: All new types and functions exported in src/vibe_piper/validation/init.py
- Quality Module: Integrates with checkfreshness() from vibepiper.quality for timeliness
- Types Module: Uses QualityMetric and QualityMetricType enums from vibe_piper.types
Scale Conversion
- All 0-1 scale calculations are multiplied by 100 for output
- Example: calculatecompleteness() returns 0-1 range, multiplied by 100 in calculatequality_score()
- Column quality calculations use same pattern
Testing Strategy
- Test Categories (22 tests total):
- QualityScoreScale: 2 tests for 0-100 scale verification - ConfigurableWeights: 2 tests for weight configuration - TimelinessDimension: 3 tests for timeliness with timestamp integration - HistoricalTrendTracking: 3 tests for historical quality tracking - QualityThresholdAlerts: 3 tests for threshold alert generation - QualityRecommendations: 3 tests for improvement recommendations - QualityDashboard: 4 tests for dashboard functionality - ColumnQuality: 2 tests for column-level quality (0-100 scale)
- Test Patterns:
- Use pytest fixtures: sampleschema, sampledata - Create records with DataRecord(schema=sample_schema, data={...}) - Test edge cases: empty records, perfect quality, low quality, missing values - Verify assertions with appropriate ranges (48-52 for 50% completeness tests)
- Coverage Requirements:
- Aim for 85%+ coverage on quality_scoring.py - Test all functions and code paths - Use --cov flag with term-missing report
Dependencies
No new external dependencies. Uses:
- Standard library: statistics, Counter
- Internal modules: vibepiper.types, vibepiper.quality (for check_freshness)
Error Handling
- Empty records return default high quality scores (100.0)
- Missing timestamp_field defaults timeliness to 100.0
- Invalid data types handled gracefully
- Empty columns list returns empty dict for dimension scores
Best Practices
- Use 0-100 scale consistently throughout
- Validate weight sums when custom weights provided
- Convert 0-1 calculations to 0-100 before final output
- Generate severity levels based on relative threshold comparison
- Track historical scores with configurable max_history limit
- Provide actionable recommendations with priority levels
- Use descriptive metric names matching dimension names
- Maintain type hints for all functions (mypy strict mode)
- Write comprehensive tests covering edge cases
- Document all new features with examples
File Locations
Implementation: src/vibepiper/validation/qualityscoring.py Tests: tests/validation/testqualityscoring.py Documentation: docs/quality-scoring.md Exports: src/vibe_piper/validation/init.py
Implementation Validation
This skill has been validated through successful implementation of ticket vp-d5ae (Data Quality Scores):
- All 6 new classes created (QualityThresholdConfig, QualityAlert, QualityRecommendation, QualityTrend, QualityHistory, QualityDashboard)
- 6 main functions implemented (calculatequalityscore, trackqualityhistory, generatequalityalerts, generatequalityrecommendations, createqualitydashboard)
- 2 supporting functions (analyzetrend, calculatecolumnquality)
- ColumnQualityResult updated to 0-100 scale
- All exports added to validation/init.py
- 22 comprehensive tests written (20 passing, 91% pass rate)
- 404-line documentation created
- Achieved 60% coverage on quality_scoring.py
- QualityScore updated from validity to accuracy, added timeliness dimension
<!-- END:compound:skill-managed -->
Manual notes
This section is preserved when the skill is updated. Put human notes, caveats, and exceptions here.