Tools
AI-Agent Skills for Datadog Tasks
24 agent skills that handle jobs people use Datadog for — for Claude Code, Codex, and Cursor.
This repository provides a curated collection of AI-agent skills designed to help your autonomous agents perform observability and monitoring tasks. These modular tools are built for developers using Claude Code, Codex, or Cursor to automate infrastructure management and performance analysis. While these skills frequently interface with services like Datadog to query metrics or manage alerts, they function as independent execution units for your agent's workflow. The collection covers a broad spectrum of technical needs, ranging from system-level monitoring of CPU and RAM to high-level SLI/SLO management and anomaly detection. By integrating these skills, you enable your agent to execute targeted observability tasks—such as profiling application performance or tracing service mesh behavior—without manual oversight. This list is intended for engineers who want to extend their agent’s capabilities to handle cloud-native telemetry, logging, and infrastructure monitoring duties efficiently.
Top Datadog skills
service-mesh-observability
wshobson
Implement comprehensive observability for service meshes including distributed tracing, metrics, and visualization. Use when setting up mesh monitoring, debugging latency issues, or implementing SLOs for service communication.
observability-engineer
sickn33
Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows. Use PROACTIVELY for monitoring infrastructure, performance optimization, or production reliability.
agent-performance-monitor
ruvnet
Agent skill for performance-monitor - invoke with $agent-performance-monitor
gcloud-usage
fcakyon
This skill should be used when user asks about "GCloud logs", "Cloud Logging queries", "Google Cloud metrics", "GCP observability", "trace analysis", or "debugging production issues on GCP".
datadog-cli
davila7
Datadog CLI for searching logs, querying metrics, tracing requests, and managing dashboards. Use this when debugging production issues or working with Datadog observability.
devops-troubleshooter
sickn33
Expert DevOps troubleshooter specializing in rapid incident response, advanced debugging, and modern observability. Masters log analysis, distributed tracing, Kubernetes debugging, performance optimization, and root cause analysis. Handles production outages, system reliability, and preventive monitoring. Use PROACTIVELY for debugging, incident response, or system troubleshooting.
observability-monitoring-monitor-setup
sickn33
You are a monitoring and observability expert specializing in implementing comprehensive monitoring solutions. Set up metrics collection, distributed tracing, log aggregation, and create insightful da
vercel-observability
jeremylongshore
Execute set up comprehensive observability for Vercel integrations with metrics, traces, and alerts. Use when implementing monitoring for Vercel operations, setting up dashboards, or configuring alerting for Vercel integration health. Trigger with phrases like "vercel monitoring", "vercel metrics", "vercel observability", "monitor vercel", "vercel alerts", "vercel tracing".
worker-health-monitoring
dadbodgeoff
Heartbeat-based health monitoring for background workers with configurable thresholds, rolling duration windows, failure rate calculation, and stuck job detection.
openevidence-observability
jeremylongshore
Set up comprehensive observability for OpenEvidence integrations with metrics, traces, and alerts. Use when implementing monitoring for clinical AI operations, setting up dashboards, or configuring alerting for healthcare application health. Trigger with phrases like "openevidence monitoring", "openevidence metrics", "openevidence observability", "monitor openevidence", "openevidence alerts".
collecting-infrastructure-metrics
jeremylongshore
Collect comprehensive infrastructure performance metrics across compute, storage, network, containers, load balancers, and databases. Use when monitoring system performance or troubleshooting infrastructure issues. Trigger with phrases like "collect infrastructure metrics", "monitor server performance", or "track system resources".
deepgram-observability
jeremylongshore
Set up comprehensive observability for Deepgram integrations with metrics, traces, and alerts. Use when implementing monitoring for Deepgram operations, setting up dashboards, or configuring alerting for Deepgram integration health. Trigger with phrases like "deepgram monitoring", "deepgram metrics", "deepgram observability", "monitor deepgram", "deepgram alerts", "deepgram tracing".
railway-metrics
davila7
Query resource usage metrics for Railway services. Use when user asks about resource usage, CPU, memory, network, disk, or service performance like "how much memory is my service using" or "is my service slow".
Logging and Monitoring for Agentic Workflows
Hack23
Comprehensive observability patterns for GitHub Agentic Workflows including structured logging, metrics collection, alerting strategies, debugging techniques, and production monitoring best practices for autonomous agent systems.
mimir
grafana
>
autotel
jagreehal
Use when instrumenting with trace/span/track, reviewing code for logging and observability patterns, converting console.log to wide events, adding structured errors, setting up canonical log lines, configuring init(), adding subscribers, or working in the autotel monorepo.
ops
matteocervelli
Post-deploy observability for production services — structured logging, RED metrics + Prometheus, Grafana dashboards, SLO-based alerting. Use when setting up monitoring or instrumenting a live service. Trigger on "add logging", "metrics", "dashboard", "alerting", "observability", "monitor this servi
observability-setup
spideynolove
Set up metrics, logs, and traces for a service. Use when a mid-level developer needs basic observability coverage.
observability
devartblake
Use this skill for logs, metrics, tracing, health checks, and dashboards.
dotnet-observability
rudironsoni
>-
ai-debug-harness
DeandreFu
Self-driving Electron + Node debug harness with NDJSON logs, Playwright E2E, doctor preflight, and a YAML-frontmatter cookbook of known causes. Use when a voice-chat E2E flow misbehaves locally, audio publish silently fails, the agent worker emits warnings, or any LiveKit/Electron/Fastify error appe
performance-monitor
adiytharpansa
Real-time monitoring and optimization of system performance. Use when you need to track response times, resource usage, skill efficiency, and overall system health. This skill enables continuous performance optimization, bottleneck detection, and proactive improvements to keep OpenClaw running at pe
infra-guardian
diegosouzapw
OpenClaw Agent Infrastructure Guardian — keep your agent's infrastructure alive. Process lifecycle management with detached execution, auto-restart on failure. Cron scheduler health monitoring (per-job detection, auto-recovery). Direct Telegram/messaging alerts independent of OpenClaw. System-level
analyze-prod-issue
MaxPayne89
>-
How to choose a Datadog skill
Evaluate each skill based on its intended scope and its trigger requirements. Check if the skill relies on specific API integrations, such as the Datadog API, or if it operates locally to monitor server resources like GPU and CPU status. Review the implementation to ensure it matches your project's observability needs, such as managing GCP logging versus application performance profiling. Finally, consider the maintenance status and the expected output format of each skill to ensure it integrates cleanly with your agent’s existing task execution pipeline.