duyet/clickhouse-monitoring · Archived

cluster-operations

Cluster management: distributed tables, ON CLUSTER DDL, node lifecycle, resharding, load balancing, and Keeper migration.

First seen Apr 19, 2026

Installation

$ npx skills add duyet/clickhouse-monitoring --skill cluster-operations

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from duyet/clickhouse-monitoring.

npx skills add duyet/clickhouse-monitoring

Browse all from duyet/clickhouse-monitoring

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 241
License LICENSE
Default branch main
Open issues 5
Status Archived

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,557 B
  • docs SUMMARY.md 147 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 3 installs

SKILL.md

Cluster Operations

Distributed Tables

  • CREATE TABLE dist ENGINE = Distributed(cluster, db, localtable, shardingkey)
  • Sharding key: rand() for even distribution, cityHash64(user_id) for user affinity
  • Reads: query all shards in parallel; Writes: route to correct shard or write locally

ON CLUSTER DDL

  • ALTER TABLE t ON CLUSTER '{cluster}' ADD COLUMN col Type — propagate schema to all replicas
  • CREATE TABLE t ON CLUSTER '{cluster}' AS templatedb.templatetable — clone across shards
  • distributedddloutput_mode: throw (fail on error), null (ignore), none, active
  • Check status: SELECT * FROM system.distributedddlqueue

Load Balancing and Read Routing

  • loadbalancing: random, inorder, firstorrandom, nearest_hostname
  • maxreplicadelayfordistributed_queries — skip lagging replicas
  • fallbacktostalereplicasfordistributedqueries=1 — use stale when all delayed

Adding Nodes

  1. Install ClickHouse on new node
  2. Configure Keeper/ZooKeeper connection
  3. Update cluster config (remote_servers) on all nodes
  4. Create local tables on new node
  5. ReplicatedMergeTree: data syncs automatically; non-replicated: copy or re-insert

Removing Nodes

  1. Stop writes, wait for replication queue to drain
  2. SYSTEM DROP REPLICA for replicated tables
  3. Remove from cluster config, restart remaining nodes

Resharding

  • No native online resharding — create new distributed table with new sharding scheme
  • INSERT INTO newdist SELECT * FROM olddist or clickhouse-copier for large migrations

Cluster Recovery

  • SYSTEM RESTART REPLICA ON CLUSTER '{cluster}' — restart replication across all nodes
  • SYSTEM SYNC REPLICA ON CLUSTER '{cluster}' — force sync from ZooKeeper
  • SYSTEM FETCH PARTS ON CLUSTER '{cluster}' — pull missing parts from other replicas

Monitoring Clusters

  • system.clusters — topology; system.distributedddlqueue — DDL status; system.replicas — replication
  • Cross-shard queries: use Distributed table or remote() function

Keeper Migration (ZooKeeper to ClickHouse Keeper)

  1. Deploy ClickHouse Keeper alongside ZooKeeper
  2. Snapshot data: clickhouse-keeper-converter or zk-dump.sh
  3. Configure Keeper with converted snapshot
  4. Update zookeeper config, restart one node at a time
  5. Verify replication recovers, then remove ZooKeeper