kubesense-ai/kubesense-mcp-skills

kubesense-infra

Inventory and topology of the monitored Kubernetes estate via KubeSense MCP — clusters, nodes, pods, workloads, detected issues, infra failures (OOM/CrashLoop/probe/scheduling), and recent deploys/scaling changes.

First seen Aug 11, 2026

Installation

$ npx skills add kubesense-ai/kubesense-mcp-skills --skill kubesense-infra

Summary

  • Inventory and topology of the monitored Kubernetes estate via KubeSense MCP — clusters, nodes, pods, workloads, detected issues, infra failures (OOM/CrashLoop/probe/scheduling), and recent deploys/scaling changes.
  • Use for "what is running", "what is broken at the k8s layer", and "what changed".

Also in this package

Other skills from kubesense-ai/kubesense-mcp-skills.

npx skills add kubesense-ai/kubesense-mcp-skills

Browse all from kubesense-ai/kubesense-mcp-skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

License MIT
Default branch main
Open issues 0
Status Active

Skill metadata

Parsed from SKILL.md frontmatter.

Version2.0.0
More metadata
version
2.0.0
author
kubesense
repository
https://github.com/kubesense-ai/kubesense-mcp-skills
tags
kubesense,kubernetes,inventory,topology,pods,nodes,workloads,issues,oomkill,crashloop,changes,deploys

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 9,332 B
  • docs SUMMARY.md 320 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 30 installs

SKILL.md

KubeSense Infrastructure & Inventory

Telemetry tells you a service is unhealthy. These tools tell you what exists, what is failing at the Kubernetes layer, and what changed — the context that turns a symptom into a cause.

Requires the KubeSense MCP server. See [kubesense-mcp](../kubesense-mcp/SKILL.md) for connection and auth.

Routing — Read This First

The user is asking… Call Not
"which clusters / what cluster names can I use?" list-clusters guessing a cluster name — every other tool's clusters filter is free-text with no validation, so a typo silently returns nothing
"what's running / how many pods / pod inventory" list-pods search-logs — logs only show pods that logged
"why is this pod restarting / crashlooping / OOMKilled" get-infra-issues list-pods (gives the restart count, not the reason)
"what changed / did someone deploy / why did it break at 14:00" get-recent-changes anything else — check this first, it is the cheapest root cause
"what's broken across the estate" list-issues (platform's own detection) re-deriving issues from raw traces
"node pressure / eviction / scheduling failures" list-nodes, then get-node-detail analyze-metrics (use it to confirm, not to discover)
"service health: RPS, p95, error rate per app" list-workloads analyze-traces (right answer for arbitrary slicing, overkill for a health sweep)
"container exit code / last state / restart reason" get-pod-detail list-pods

Discovery order. list-clusters → list-workloads / list-pods → get-*-detail. Names flow forward: never type a cluster, namespace, workload, or pod name you have not read out of a previous call.

Tools

Tool Purpose Required args
list-clusters Clusters this deployment monitors —
list-nodes Nodes with CPU/mem/disk pressure and readiness —
get-node-detail One node: conditions, taints, addresses, capacity name
list-pods Pods with phase, restarts, usage —
get-pod-detail One pod: container statuses, exit codes, last state, QoS name, namespace, cluster
list-workloads Workloads with golden signals (RPS, p95, error rate) —
get-workload-detail One workload: replicas, 4xx/5xx counts, RPS workload, namespace
list-issues Problems KubeSense already detected —
get-infra-issues K8s-layer failures: OOM, crashes, probes, scheduling, image pull —
get-recent-changes Image updates (deploys) and replica scaling —

Every tool defaults to the last 1 hour and all accessible clusters. Pass fromtime/totime (RFC3339) to widen; pass clusters to narrow.

list-clusters

Call this first when you need a cluster name. It is the only discovery for the clusters filter that every other tool accepts.

{ "search": "prod" }

Returns TSV: name, source, envtype, updatedat.

Nodes

{ "clusters": ["prod-us"], "sort_by": "memory_usage_precent", "sort_direction": "DESC" }

[!WARNING]
sortby for nodes is spelled precent, not _percent — the accepted values
are memoryusageprecent, diskusageprecent, cpuusageprecent. The output
columns use the correct spelling (memoryusagepercent). Passing _percent to
sort_by does not sort.

Output: name, cluster, ready, runningpods, cpuusagepercent, memoryusagepercent, diskusagepercent, kubeletversion, instancetype, uid, creationtimestamp.

get-node-detail takes name (required) and optional clusters to disambiguate a name present in several clusters. It returns conditions, addresses, labels, taints, and capacity/usage as a record with ## conditions / ## addresses sub-tables.

Pods

{ "namespace": "production", "status": "Pending", "sort_by": "memory_usage", "sort_direction": "DESC" }
  • status filters on pod phase: Running, Pending, Failed, Succeeded, Unknown.
  • sortby accepts memoryusage or cpu_usage only.

Output: name, namespace, cluster, status, ownerkind, owner, node, restarts, cpuusage, memoryusage, creationtime.

get-pod-detail needs all three of name, namespace, cluster. It is the tool that answers why a pod restarted — container statuses carry the restart reason, exit code, and last terminated state, which list-pods does not.

Workloads

The application-level view — one row per Deployment/StatefulSet/DaemonSet with its golden signals already joined from traces.

{ "namespace": "production", "kind": "Deployment", "sort_by": "error_rate", "sort_direction": "DESC" }

sortby accepts rps, p95, errors, errorrate, restarts.

Output: name, namespace, cluster, kind, ready, desired, restarts, rps, p95, errorrate, issuecount, service_protocols.

Use this for a health sweep — it is one call for what would otherwise be an analyze-traces per workload. Drop to analyze-traces when you need a slice the summary doesn't have (by endpoint, by status code, by attribute).

list-issues

KubeSense's own issue detection — errors, latency, and connectivity problems between workloads. Prefer it over re-deriving problems from raw telemetry.

Output: issueid, kind, cluster, primaryworkload, primarynamespace, associateworkload, issuereason, returncode, subtype, sumissuecount, maxlastseen.

associateworkload is the other side of a connectivity issue — the dependency that primaryworkload failed to reach.

get-infra-issues

Kubernetes-layer failures, from cluster events. This is the OOM / CrashLoop / probe-failure tool.

{ "namespace": "production", "workload": "checkout", "reason": "OOMKilling" }

Known reason values: OOMKilling, BackOff, Unhealthy, FailedScheduling, FailedMount, ImagePullBackOff, NodeNotReady. Omit reason to see all failure modes — narrow only once you know which one you're chasing.

Output: reason, type, workload, object, namespace, exitcode (crashes only), message, count, lastseen. Defaults to limit 50, max 200.

Layer routing: infra failures → this tool. Application errors and latency → analyze-traces / search-traces. What changed → get-recent-changes.

get-recent-changes

Image updates (deploys) and replica scaling, from cluster events. Usually the first thing to check in any investigation — "what changed right before it broke" is the cheapest hypothesis you will test.

{ "namespace": "production", "workload": "checkout", "kind": "image_update" }

kind is image_update or scaling; omit for both. Output: kind, workload, namespace, object, detail, when — where detail reads like kubesense: dev-v1.2.1279 -> dev-v1.2.1280 or replicas 1 -> 0. Defaults to limit 50, max 200.

Scope with namespace + workload. Both are optional but strongly recommended — an unscoped call across a busy cluster returns mostly noise.

Investigation Pattern

Symptom → context → cause, cheapest signal first:

1. get-recent-changes    namespace+workload → did we just deploy or scale?
2. get-infra-issues      same scope        → is the platform killing it (OOM, probes, sched)?
3. list-pods / get-pod-detail               → restart counts, exit codes, last state
4. analyze-traces / search-logs              → application-level errors and latency
5. analyze-metrics                           → confirm resource pressure quantitatively

Steps 1–2 are two cheap calls that resolve a large share of incidents outright. Reach for telemetry once they come back clean.

Rules

  1. list-clusters before any clusters filter. The filter is unvalidated free

text — a wrong name returns an empty result that looks exactly like "no data".

  1. Names flow forward from list → detail calls. Never invent a pod, node, workload, or

namespace name.

  1. get-pod-detail requires name + namespace + cluster; get-workload-detail

requires workload + namespace. Discover them with the matching list-* call.

  1. Node sortby uses the misspelling precent; output columns use _percent.
  2. Prefer list-issues / get-infra-issues over re-deriving problems from raw

telemetry — the platform already did the detection.

  1. Scope get-recent-changes and get-infra-issues with namespace + workload;

unscoped calls are noisy.

  1. Default window is 1 hour. If a result set looks empty, widen the window before

concluding nothing happened.