databricks/databricks-agent-skills

databricks-spark-structured-streaming

Comprehensive guide to Spark Structured Streaming for production workloads.

First seen Jun 26, 2026

Installation

$ npx skills add databricks/databricks-agent-skills --skill databricks-spark-structured-streaming

Summary

  • Comprehensive guide to Spark Structured Streaming for production workloads.
  • Use when building streaming pipelines, working with Kafka ingestion, implementing Real-Time Mode (RTM), configuring triggers (processingTime, availableNow), handling stateful operations with watermarks, optimizing checkpoints, performing stream-stream or stream-static joins, writing to multiple sinks, or tuning streaming cost and performance.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from databricks/databricks-agent-skills · top by installs.

npx skills add databricks/databricks-agent-skills

Browse all from databricks/databricks-agent-skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 304
License LICENSE
Default branch main
Open issues 31
Status Active

Skill metadata

Parsed from SKILL.md frontmatter.

Version0.1.0
CompatibilityRequires databricks CLI (>= v1.0.0)
More metadata
version
0.1.0

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 3,826 B
  • docs SUMMARY.md 465 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 172 installs

SKILL.md

Spark Structured Streaming

Production-ready streaming pipelines with Spark Structured Streaming. This skill provides navigation to detailed patterns and best practices.

Quick Start

from pyspark.sql.functions import col, from_json

# Basic Kafka to Delta streaming
df = (spark
    .readStream
    .format("kafka")
    .option("kafka.bootstrap.servers", "broker:9092")
    .option("subscribe", "topic")
    .load()
    .select(from_json(col("value").cast("string"), schema).alias("data"))
    .select("data.*")
)

df.writeStream \
    .format("delta") \
    .outputMode("append") \
    .option("checkpointLocation", "/Volumes/catalog/checkpoints/stream") \
    .trigger(processingTime="30 seconds") \
    .start("/delta/target_table")

Core Patterns

Pattern Description Reference
Kafka Streaming Kafka to Delta, Kafka to Kafka, Real-Time Mode See [references/kafka-streaming.md](references/kafka-streaming.md)
Real-Time Mode (RTM) Sub-second E2E latency — cluster setup, slot math, supported ops (incl. stream-stream inner join on DBR 18+), transformWithState, observability, error classes, delivery semantics See [references/real-time-mode.md](references/real-time-mode.md)
Lakebase Sink Write streaming records into Lakebase Postgres with transactional upserts. Native format("postgresql") sink (DBR 18.3+) and manual foreach sink as a fallback See [references/lakebase-sink-python.md](references/lakebase-sink-python.md)
Stream Joins Stream-stream joins, stream-static joins See [references/stream-stream-joins.md](references/stream-stream-joins.md), [references/stream-static-joins.md](references/stream-static-joins.md)
Multi-Sink Writes Write to multiple tables, parallel merges See [references/multi-sink-writes.md](references/multi-sink-writes.md)
Merge Operations MERGE performance, parallel merges, optimizations See [references/merge-operations.md](references/merge-operations.md)

Configuration

Topic Description Reference
Checkpoints Checkpoint management and best practices See [references/checkpoint-best-practices.md](references/checkpoint-best-practices.md)
Stateful Operations Watermarks, state stores, RocksDB configuration See [references/stateful-operations.md](references/stateful-operations.md)
Trigger & Cost Trigger selection, cost optimization, RTM See [references/trigger-and-cost-optimization.md](references/trigger-and-cost-optimization.md)

Best Practices

Topic Description Reference
Production Checklist Comprehensive best practices See [references/streaming-best-practices.md](references/streaming-best-practices.md)

Production Checklist

  • Checkpoint location is persistent (UC volumes, not DBFS)
  • Unique checkpoint per stream
  • Fixed-size cluster (no autoscaling for streaming)
  • Monitoring configured (input rate, lag, batch duration)
  • Exactly-once verified (txnVersion/txnAppId)
  • Watermark configured for stateful operations
  • Left joins for stream-static (not inner)