npx skills add smithery/aj-geddes --skill infrastructure-monitoring
aj-geddes/useful-ai-prompts
infrastructure-monitoring
Set up comprehensive infrastructure monitoring with Prometheus, Grafana, and alerting systems for metrics, health checks, and performance tracking.
Installation
npx skills add aj-geddes/useful-ai-prompts --skill infrastructure-monitoring
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Import existing Azure resources into Terraform using Azure CLI discovery and Azure Verified Mod…
6.9K installs>- Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastruct…
1.9K installsCore infrastructure providing backend connection configuration, storage client, and React app e…
18.5K installsShip Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — `k8s-monito…
2.9K installsHelp users build and scale internal platforms and technical infrastructure. Use when someone is…
1.6K installsReference guide for Agentica multi-agent infrastructure APIs
470 installsAlso in this package
Other skills from aj-geddes/useful-ai-prompts · top by installs.
npx skills add aj-geddes/useful-ai-prompts
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Also listed on
Alternate registries and mirrors of this skill.
Repository health
main
Package contents
Files included with this skill beyond the listing page.
-
skill md
SKILL.md2,270 B -
docs
SUMMARY.md907 B
History
- First seen on skills.sh
- First recorded snapshot · 446 installs
SKILL.md
Infrastructure Monitoring
Table of Contents
- [Overview](#overview)
- [When to Use](#when-to-use)
- [Quick Start](#quick-start)
- [Reference Guides](#reference-guides)
- [Best Practices](#best-practices)
Overview
Implement comprehensive infrastructure monitoring to track system health, performance metrics, and resource utilization with alerting and visualization across your entire stack.
When to Use
- Real-time performance monitoring
- Capacity planning and trends
- Incident detection and alerting
- Service health tracking
- Resource utilization analysis
- Performance troubleshooting
- Compliance and audit trails
- Historical data analysis
Quick Start
Minimal working example:
# prometheus.yml
global:
scrape_interval: 15s
evaluation_interval: 15s
external_labels:
monitor: "infrastructure-monitor"
environment: "production"
# Alertmanager configuration
alerting:
alertmanagers:
- static_configs:
- targets:
- localhost:9093
# Rule files
rule_files:
- "alerts.yml"
- "rules.yml"
scrape_configs:
# Prometheus itself
- job_name: "prometheus"
static_configs:
- targets: ["localhost:9090"]
// ... (see reference guides for full implementation)
Reference Guides
Detailed implementations in the references/ directory:
| Guide | Contents |
|---|---|
| [Prometheus Configuration](references/prometheus-configuration.md) | Prometheus Configuration |
| [Alert Rules](references/alert-rules.md) | Alert Rules |
| [Alertmanager Configuration](references/alertmanager-configuration.md) | Alertmanager Configuration |
| [Grafana Dashboard](references/grafana-dashboard.md) | Grafana Dashboard |
| [Monitoring Deployment](references/monitoring-deployment.md) | Monitoring Deployment |
Best Practices
✅ DO
- Follow established patterns and conventions
- Write clean, maintainable code
- Add appropriate documentation
- Test thoroughly before deploying
❌ DON'T
- Skip testing or validation
- Ignore error handling
- Hard-code configuration values