duyet/clickhouse-monitoring · Archived

replication-guide

ReplicatedMergeTree operations, failover procedures, lag diagnosis, quorum writes, and Keeper management.

First seen Apr 19, 2026

Installation

$ npx skills add duyet/clickhouse-monitoring --skill replication-guide

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from duyet/clickhouse-monitoring.

npx skills add duyet/clickhouse-monitoring

Browse all from duyet/clickhouse-monitoring

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 241
License LICENSE
Default branch main
Open issues 5
Status Archived

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,246 B
  • docs SUMMARY.md 130 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 3 installs

SKILL.md

Replication Guide

ReplicatedMergeTree Basics

  • Engine: ReplicatedMergeTree('/clickhouse/tables/{shard}/{database}/{table}', '{replica}')
  • Requires ZooKeeper or ClickHouse Keeper
  • All replicas are equal — any replica can accept writes
  • Replication is asynchronous by default

Monitoring Replication

  • system.replicas — per-table status: absolutedelay, queuesize, isleader, isreadonly
  • system.replication_queue — pending operations: fetches, merges, mutations
  • Key health indicators:

- absolutedelay = 0 — fully caught up - isreadonly = 0 — accepting writes - queuesize < 10 — healthy queue - activereplicas = total_replicas — all replicas online

Failover Procedures

  1. Check replica status: SELECT * FROM system.replicas WHERE is_readonly = 1
  2. Verify Keeper connectivity: SELECT * FROM system.zookeeper WHERE path = '/'
  3. If replica is readonly due to Keeper disconnect, it auto-recovers when connection restores
  4. For permanent failures: SYSTEM DROP REPLICA 'replica_name' FROM TABLE db.table
  5. Corrupted metadata: SYSTEM RESTORE REPLICA db.table

Quorum Writes

  • SET insert_quorum = 2 — wait for N replicas to confirm
  • SET insertquorumparallel = 1 — parallel quorum inserts (v21.8+)
  • SET insertquorumtimeout = 60000 — timeout in ms
  • Use for critical data that must survive node failures

Keeper Management

  • ClickHouse Keeper is the recommended replacement for ZooKeeper
  • Monitor: system.zookeeper table for browsing ZK tree
  • Key paths: /clickhouse/tables/ for table metadata
  • Check Keeper health: SELECT * FROM system.asynchronous_metrics WHERE metric LIKE '%Keeper%'

Common Issues

  • Split brain: Multiple leaders — usually Keeper issue, restart Keeper
  • Readonly replica: Lost Keeper session — check network, Keeper logs
  • Queue buildup: Slow fetches — check network bandwidth, disk I/O
  • Diverged replicas: SYSTEM SYNC REPLICA db.table to force sync
  • Load the troubleshooting skill for OOM-related replica failures and disk-full scenarios.