CO

collecting-infrastructure-metrics

Automates the collection and aggregation of infrastructure performance metrics across all system layers.

Install

mkdir -p .claude/skills/collecting-infrastructure-metrics && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8719" && unzip -o skill.zip -d .claude/skills/collecting-infrastructure-metrics && rm skill.zip

Installs to .claude/skills/collecting-infrastructure-metrics

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Collect comprehensive infrastructure performance metrics across compute,
72 charsno explicit “when” trigger
Intermediate

Key capabilities

  • Identify infrastructure layers for monitoring
  • Configure agents for metric collection from various infrastructure components
  • Aggregate collected metrics centrally
  • Generate dashboards for health monitoring and performance analysis
  • Configure alerts based on critical metric thresholds

How it works

The skill automates setting up infrastructure metrics collection by identifying layers, configuring agents like Prometheus or Datadog, aggregating metrics, and creating dashboards for visualization and alerting.

Inputs & outputs

You give it
User request for infrastructure metrics collection, specific infrastructure components (e.g., web server, database)
You get back
Metrics collection configuration files, agent installation scripts, dashboard definitions, metric export configurations, and alert rules

When to use collecting-infrastructure-metrics

  • Setting up system performance dashboards
  • Monitoring container health
  • Tracking database performance metrics
  • Centralizing infrastructure logs

About this skill

Infrastructure Metrics Collector

Collect and centralize infrastructure metrics across compute, storage, network, containers, load balancers, and databases using Prometheus, Datadog, or CloudWatch.

Overview

This skill automates the process of setting up infrastructure metrics collection. It identifies key performance indicators (KPIs) across various infrastructure layers, configures agents to collect these metrics, and assists in setting up central aggregation and visualization.

How It Works

  1. Identify Infrastructure Layers: Determines the infrastructure layers to monitor (compute, storage, network, containers, load balancers, databases).
  2. Configure Metrics Collection: Sets up agents (Prometheus, Datadog, CloudWatch) to collect metrics from the identified layers.
  3. Aggregate Metrics: Configures central aggregation of the collected metrics for analysis and visualization.
  4. Create Dashboards: Generates infrastructure dashboards for health monitoring, performance analysis, and capacity tracking.

When to Use This Skill

This skill activates when you need to:

  • Monitor the performance of your infrastructure.
  • Identify bottlenecks in your system.
  • Set up dashboards for real-time monitoring.

Examples

Example 1: Setting up basic monitoring

User request: "Collect infrastructure metrics for my web server."

The skill will:

  1. Identify compute, storage, and network layers relevant to the web server.
  2. Configure Prometheus to collect CPU, memory, disk I/O, and network bandwidth metrics.

Example 2: Troubleshooting database performance

User request: "I'm seeing slow database queries. Can you help me monitor the database performance?"

The skill will:

  1. Identify the database layer and relevant metrics such as connection pool usage, replication lag, and cache hit rates.
  2. Configure Datadog to collect these metrics and create a dashboard to visualize performance trends.

Best Practices

  • Agent Selection: Choose the appropriate agent (Prometheus, Datadog, CloudWatch) based on your existing infrastructure and monitoring tools.
  • Metric Granularity: Balance the granularity of metrics collection with the storage and processing overhead. Collect only the essential metrics for your use case.
  • Alerting: Configure alerts based on thresholds for key metrics to proactively identify and address performance issues.

Integration

This skill can be integrated with other plugins for deployment, configuration management, and alerting to provide a comprehensive infrastructure management solution. For example, it can be used with a deployment plugin to automatically configure metrics collection after deploying new infrastructure.

Prerequisites

  • Access to infrastructure monitoring systems (Prometheus, Datadog, CloudWatch)
  • System permissions for metrics agent installation
  • Network access to monitored infrastructure components
  • Storage for metrics data in ${CLAUDE_SKILL_DIR}/metrics/

Instructions

  1. Identify infrastructure layers to monitor (compute, storage, network, databases)
  2. Select appropriate metrics collection agent based on environment
  3. Configure agent with target endpoints and metric types
  4. Set up central aggregation for collected metrics
  5. Create dashboards for visualization
  6. Configure alerts for critical metrics thresholds

Output

  • Metrics collection configuration files
  • Agent installation and setup scripts
  • Dashboard definitions for infrastructure monitoring
  • Metric export configurations
  • Alert rules for critical thresholds

Error Handling

If metrics collection fails:

  • Verify agent installation and permissions
  • Check network connectivity to targets
  • Validate authentication credentials
  • Review firewall and security group rules
  • Confirm metric endpoint availability

Resources

  • Prometheus documentation for metric collection
  • Datadog agent configuration guides
  • AWS CloudWatch metrics reference
  • Infrastructure monitoring best practices

When not to use it

  • When monitoring is not required
  • When troubleshooting issues unrelated to infrastructure performance

Prerequisites

Access to infrastructure monitoring systems (Prometheus, Datadog, CloudWatch)System permissions for metrics agent installationNetwork access to monitored infrastructure componentsStorage for metrics data in ${CLAUDE_SKILL_DIR}/metrics/

Limitations

  • Requires appropriate agent selection based on existing infrastructure
  • Balancing metric granularity with storage and processing overhead is necessary
  • Alerting relies on configured thresholds for key metrics

How it compares

This skill automates the setup of a complete infrastructure monitoring system, contrasting with manual configuration of individual metrics and dashboards.

Compared to similar skills

collecting-infrastructure-metrics side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
collecting-infrastructure-metrics (this skill)025dReviewIntermediate
observability02moNo flagsBeginner
performance-monitor04moNo flagsIntermediate
railway-metrics17moReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

Search skills

Search the agent skills registry