This skill collects and aggregates AWS metrics, logs, and events to support dashboard creation and alarm triggers.

Install

mkdir -p .claude/skills/cloudwatch && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/6616" && unzip -o skill.zip -d .claude/skills/cloudwatch && rm skill.zip

Installs to .claude/skills/cloudwatch

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

AWS CloudWatch monitoring for logs, metrics, alarms, and dashboards. Use when setting up monitoring, creating alarms, querying logs with Insights, configuring metric filters, building dashboards, or troubleshooting application issues.
234 chars✓ has a “when” trigger
Intermediate

Key capabilities

  • →Create and manage metric alarms
  • →Query application logs using Insights
  • →Build operational dashboards
  • →Configure metric filters for log patterns
  • →Publish custom metrics to CloudWatch

How it works

It utilizes AWS CLI commands or boto3 to interact with CloudWatch APIs for collecting metrics, logs, and events to enable observability.

Inputs & outputs

You give it
AWS resource monitoring requirements or log query parameters
You get back
CloudWatch alarm definitions, log query results, or dashboard JSON

When to use cloudwatch

  • →Create CPU utilization alarms
  • →Query application logs with Insights
  • →Build operational dashboards
  • →Troubleshoot AWS resource performance

About this skill

AWS CloudWatch

Amazon CloudWatch provides monitoring and observability for AWS resources and applications. It collects metrics, logs, and events, enabling you to monitor, troubleshoot, and optimize your AWS environment.

Table of Contents

Core Concepts

Metrics

Time-ordered data points published to CloudWatch. Key components:

  • Namespace: Container for metrics (e.g., AWS/Lambda)
  • Metric name: Name of the measurement (e.g., Invocations)
  • Dimensions: Name-value pairs for filtering (e.g., FunctionName=MyFunc)
  • Statistics: Aggregations (Sum, Average, Min, Max, SampleCount, pN)

Logs

Log data from AWS services and applications:

  • Log groups: Collections of log streams
  • Log streams: Sequences of log events from same source
  • Log events: Individual log entries with timestamp and message

Alarms

Automated actions based on metric thresholds:

  • States: OK, ALARM, INSUFFICIENT_DATA
  • Actions: SNS notifications, Lambda functions, Auto Scaling, EC2 actions, Systems Manager OpsItems
  • Types: metric (incl. metric math, anomaly detection, Metrics Insights), composite, log (scheduled Logs Insights query evaluated M-out-of-N on query results)
  • Mute rules: scheduled windows that mute actions while alarms keep evaluating

Common Patterns

Create a Metric Alarm

AWS CLI:

# CPU utilization alarm for EC2
aws cloudwatch put-metric-alarm \
  --alarm-name "HighCPU-i-1234567890abcdef0" \
  --metric-name CPUUtilization \
  --namespace AWS/EC2 \
  --statistic Average \
  --period 300 \
  --threshold 80 \
  --comparison-operator GreaterThanThreshold \
  --evaluation-periods 2 \
  --dimensions Name=InstanceId,Value=i-1234567890abcdef0 \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts \
  --ok-actions arn:aws:sns:us-east-1:123456789012:alerts

boto3:

import boto3

cloudwatch = boto3.client('cloudwatch')

cloudwatch.put_metric_alarm(
    AlarmName='HighCPU-i-1234567890abcdef0',
    MetricName='CPUUtilization',
    Namespace='AWS/EC2',
    Statistic='Average',
    Period=300,
    Threshold=80.0,
    ComparisonOperator='GreaterThanThreshold',
    EvaluationPeriods=2,
    Dimensions=[
        {'Name': 'InstanceId', 'Value': 'i-1234567890abcdef0'}
    ],
    AlarmActions=['arn:aws:sns:us-east-1:123456789012:alerts'],
    OKActions=['arn:aws:sns:us-east-1:123456789012:alerts']
)

Lambda Error Rate Alarm

aws cloudwatch put-metric-alarm \
  --alarm-name "LambdaErrorRate-MyFunction" \
  --metrics '[
    {
      "Id": "errors",
      "MetricStat": {
        "Metric": {
          "Namespace": "AWS/Lambda",
          "MetricName": "Errors",
          "Dimensions": [{"Name": "FunctionName", "Value": "MyFunction"}]
        },
        "Period": 60,
        "Stat": "Sum"
      },
      "ReturnData": false
    },
    {
      "Id": "invocations",
      "MetricStat": {
        "Metric": {
          "Namespace": "AWS/Lambda",
          "MetricName": "Invocations",
          "Dimensions": [{"Name": "FunctionName", "Value": "MyFunction"}]
        },
        "Period": 60,
        "Stat": "Sum"
      },
      "ReturnData": false
    },
    {
      "Id": "errorRate",
      "Expression": "errors/invocations*100",
      "Label": "Error Rate",
      "ReturnData": true
    }
  ]' \
  --threshold 5 \
  --comparison-operator GreaterThanThreshold \
  --evaluation-periods 3 \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts

Query Logs with Insights

# Find errors in Lambda logs
aws logs start-query \
  --log-group-name /aws/lambda/MyFunction \
  --start-time $(date -d '1 hour ago' +%s) \
  --end-time $(date +%s) \
  --query-string '
    fields @timestamp, @message
    | filter @message like /ERROR/
    | sort @timestamp desc
    | limit 50
  '

# Get query results
aws logs get-query-results --query-id <query-id>

boto3:

import boto3
import time

logs = boto3.client('logs')

# Start query
response = logs.start_query(
    logGroupName='/aws/lambda/MyFunction',
    startTime=int(time.time()) - 3600,
    endTime=int(time.time()),
    queryString='''
        fields @timestamp, @message
        | filter @message like /ERROR/
        | sort @timestamp desc
        | limit 50
    '''
)

query_id = response['queryId']

# Wait for results
while True:
    result = logs.get_query_results(queryId=query_id)
    if result['status'] == 'Complete':
        break
    time.sleep(1)

for row in result['results']:
    print(row)

Create Metric Filter

Extract metrics from log patterns:

# Create metric filter for error count
aws logs put-metric-filter \
  --log-group-name /aws/lambda/MyFunction \
  --filter-name ErrorCount \
  --filter-pattern "ERROR" \
  --metric-transformations \
    metricName=ErrorCount,metricNamespace=MyApp,metricValue=1,defaultValue=0

Create a Log Alarm

Alarm directly on a Logs Insights query (no metric filter needed). CloudWatch creates and manages the underlying scheduled query. The ScheduledQueryRoleARN role must trust logs.amazonaws.com and allow logs:StartQuery and logs:GetQueryResults on the log group ARNs, plus logs:StopQuery and logs:DescribeLogGroups on "Resource": "*" (these two support no resource-level permissions, so a log-group-scoped statement never matches them).

# ALARM when >100 errors in 3 of the last 5 query runs
aws cloudwatch put-log-alarm \
  --alarm-name "HighErrorCount" \
  --comparison-operator GreaterThanThreshold \
  --threshold 100 \
  --query-results-to-evaluate 5 \
  --query-results-to-alarm 3 \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:alerts \
  --scheduled-query-configuration '{
    "QueryString": "fields @timestamp, @message | filter @message like /ERROR/",
    "LogGroupIdentifiers": ["/aws/lambda/MyFunction"],
    "ScheduledQueryRoleARN": "arn:aws:iam::123456789012:role/ScheduledQueryRole",
    "AggregationExpression": "count(*)",
    "ScheduleConfiguration": {
      "ScheduleExpression": "rate(10 minutes)",
      "StartTimeOffset": 600
    }
  }'

# Log alarms are omitted from describe-alarms unless requested
aws cloudwatch describe-alarms --alarm-types LogAlarm

Publish Custom Metrics

import boto3

cloudwatch = boto3.client('cloudwatch')

cloudwatch.put_metric_data(
    Namespace='MyApp',
    MetricData=[
        {
            'MetricName': 'OrdersProcessed',
            'Value': 1,
            'Unit': 'Count',
            'Dimensions': [
                {'Name': 'Environment', 'Value': 'Production'},
                {'Name': 'OrderType', 'Value': 'Standard'}
            ]
        }
    ]
)

Create Dashboard

cat > dashboard.json << 'EOF'
{
  "widgets": [
    {
      "type": "metric",
      "x": 0, "y": 0, "width": 12, "height": 6,
      "properties": {
        "title": "Lambda Invocations",
        "metrics": [
          ["AWS/Lambda", "Invocations", "FunctionName", "MyFunction"]
        ],
        "period": 60,
        "stat": "Sum",
        "region": "us-east-1"
      }
    },
    {
      "type": "log",
      "x": 12, "y": 0, "width": 12, "height": 6,
      "properties": {
        "title": "Recent Errors",
        "query": "SOURCE '/aws/lambda/MyFunction' | filter @message like /ERROR/ | limit 20",
        "region": "us-east-1"
      }
    }
  ]
}
EOF

aws cloudwatch put-dashboard \
  --dashboard-name MyAppDashboard \
  --dashboard-body file://dashboard.json

CLI Reference

Metrics Commands

CommandDescription
aws cloudwatch put-metric-dataPublish custom metrics
aws cloudwatch get-metric-dataRetrieve metric values
aws cloudwatch get-metric-statisticsGet aggregated statistics
aws cloudwatch list-metricsList available metrics

Alarms Commands

CommandDescription
aws cloudwatch put-metric-alarmCreate or update alarm
aws cloudwatch put-log-alarmCreate or update log query alarm
aws cloudwatch describe-alarmsList alarms (--alarm-types LogAlarm for log alarms)
aws cloudwatch describe-alarm-contributorsShow breaching contributors of a multi-contributor alarm
aws cloudwatch put-alarm-mute-ruleCreate or update scheduled mute window
aws cloudwatch list-alarm-mute-rulesList mute rules (--statuses SCHEDULED ACTIVE EXPIRED)
aws cloudwatch delete-alarm-mute-ruleDelete mute rule (unmutes immediately)
aws cloudwatch set-alarm-stateManually set alarm state
aws cloudwatch delete-alarmsDelete alarms

Logs Commands

CommandDescription
aws logs create-log-groupCreate log group
aws logs put-log-eventsWrite log events
aws logs filter-log-eventsSearch log events
aws logs start-queryStart Insights query
aws logs put-metric-filterCreate metric filter
aws logs put-retention-policySet log retention

Best Practices

Metrics

  • Use dimensions wisely — too many creates metric explosion
  • Aggregate before publishing — batch custom metrics
  • Use high-resolution metrics (1-second) only when needed
  • Set meaningful units for custom metrics

Alarms

  • Use composite alarms for complex conditions
  • Set appropriate evaluation periods to avoid flapping
  • Include OK actions to track recovery
  • Use anomaly detection for dynamic thresholds
  • Use log alarms instead of metric filter + metric alarm for query-based conditions; set --treat-missing-data notBreaching for sparse errors, breaching to detect logs that stop arriving
  • Stagger log alarm schedules — concurrent scheduled query executions per account are

Content truncated.

When not to use it

  • →When monitoring non-AWS resources
  • →When high-resolution metrics are not required

Limitations

  • →Metric explosion from excessive dimensions
  • →Potential for alarm flapping with short evaluation periods

How it compares

This approach automates the creation of complex alarm logic and log queries compared to manual console configuration.

Compared to similar skills

cloudwatch side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
cloudwatch (this skill)18moReviewIntermediate
mlops-engineer35moNo flagsAdvanced
genkit-infra-expert12moReviewAdvanced
huawei-event-driven-architecture-review04moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

mlops-engineer

sickn33

Build comprehensive ML pipelines, experiment tracking, and model registries with MLflow, Kubeflow, and modern MLOps tools. Implements automated training, deployment, and monitoring across cloud platforms. Use PROACTIVELY for ML infrastructure, experiment management, or pipeline automation.

333

genkit-infra-expert

jeremylongshore

Execute use when deploying Genkit applications to production with Terraform. Trigger with phrases like "deploy genkit terraform", "provision genkit infrastructure", "firebase functions terraform", "cloud run deployment", or "genkit production infrastructure". Provisions Firebase Functions, Cloud Run services, GKE clusters, monitoring dashboards, and CI/CD for AI workflows.

15

huawei-event-driven-architecture-review

Raishin

Review Huawei Cloud event-driven architecture designs — DMS Kafka dead-letter configuration, ROMA Connect integration flow capacity, FunctionGraph event trigger idempotency, SMN delivery retry policy, consumer group lag monitoring, cross-region event replication, and retry storm prevention.

00

ecs-runtime-debug-playbook

talolard

Debug AWS ECS or Fargate deployments where CI or workflow status does not match live behavior, especially when an old task definition keeps serving traffic, a new task exits during startup, health checks are false-green, Alembic or schema state may be inconsistent with physical tables, or AWS profil

00

tencentcloud-lighthouse-skill

LJT-520

Manage Tencent Cloud Lighthouse (轻量应用服务器) — auto-setup mcporter + MCP, query instances, monitoring & alerting, self-diagnostics, firewall, snapshots, remote command execution (TAT). Use when user asks about Lighthouse or 轻量应用服务器. NOT for CVM or other cloud server types.

00

check-prod

learntocloud

Check Azure production health: app status, errors, latency, database, dependencies. Use when user says "check prod", "how''s prod", "hows prod doing", "is prod up", "prod status", "health check", "any errors?", "how''s the app doing?", or "check Azure".

00

Search skills

Search the agent skills registry