Tools

AI-Agent Skills for Datadog Tasks

24 agent skills that handle jobs people use Datadog for — for Claude Code, Codex, and Cursor.

This repository provides a curated collection of AI-agent skills designed to help your autonomous agents perform observability and monitoring tasks. These modular tools are built for developers using Claude Code, Codex, or Cursor to automate infrastructure management and performance analysis. While these skills frequently interface with services like Datadog to query metrics or manage alerts, they function as independent execution units for your agent's workflow. The collection covers a broad spectrum of technical needs, ranging from system-level monitoring of CPU and RAM to high-level SLI/SLO management and anomaly detection. By integrating these skills, you enable your agent to execute targeted observability tasks—such as profiling application performance or tracing service mesh behavior—without manual oversight. This list is intended for engineers who want to extend their agent’s capabilities to handle cloud-native telemetry, logging, and infrastructure monitoring duties efficiently.

Top Datadog skills

service-mesh-observability

wshobson

Implement comprehensive observability for service meshes including distributed tracing, metrics, and visualization. Use when setting up mesh monitoring, debugging latency issues, or implementing SLOs for service communication.

574

observability-engineer

sickn33

Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows. Use PROACTIVELY for monitoring infrastructure, performance optimization, or production reliability.

1242

agent-performance-monitor

ruvnet

Agent skill for performance-monitor - invoke with $agent-performance-monitor

335

gcloud-usage

fcakyon

This skill should be used when user asks about "GCloud logs", "Cloud Logging queries", "Google Cloud metrics", "GCP observability", "trace analysis", or "debugging production issues on GCP".

14

datadog-cli

davila7

Datadog CLI for searching logs, querying metrics, tracing requests, and managing dashboards. Use this when debugging production issues or working with Datadog observability.

12

devops-troubleshooter

sickn33

Expert DevOps troubleshooter specializing in rapid incident response, advanced debugging, and modern observability. Masters log analysis, distributed tracing, Kubernetes debugging, performance optimization, and root cause analysis. Handles production outages, system reliability, and preventive monitoring. Use PROACTIVELY for debugging, incident response, or system troubleshooting.

12

observability-monitoring-monitor-setup

sickn33

You are a monitoring and observability expert specializing in implementing comprehensive monitoring solutions. Set up metrics collection, distributed tracing, log aggregation, and create insightful da

12

vercel-observability

jeremylongshore

Execute set up comprehensive observability for Vercel integrations with metrics, traces, and alerts. Use when implementing monitoring for Vercel operations, setting up dashboards, or configuring alerting for Vercel integration health. Trigger with phrases like "vercel monitoring", "vercel metrics", "vercel observability", "monitor vercel", "vercel alerts", "vercel tracing".

21

worker-health-monitoring

dadbodgeoff

Heartbeat-based health monitoring for background workers with configurable thresholds, rolling duration windows, failure rate calculation, and stuck job detection.

12

openevidence-observability

jeremylongshore

Set up comprehensive observability for OpenEvidence integrations with metrics, traces, and alerts. Use when implementing monitoring for clinical AI operations, setting up dashboards, or configuring alerting for healthcare application health. Trigger with phrases like "openevidence monitoring", "openevidence metrics", "openevidence observability", "monitor openevidence", "openevidence alerts".

11

collecting-infrastructure-metrics

jeremylongshore

Collect comprehensive infrastructure performance metrics across compute, storage, network, containers, load balancers, and databases. Use when monitoring system performance or troubleshooting infrastructure issues. Trigger with phrases like "collect infrastructure metrics", "monitor server performance", or "track system resources".

01

deepgram-observability

jeremylongshore

Set up comprehensive observability for Deepgram integrations with metrics, traces, and alerts. Use when implementing monitoring for Deepgram operations, setting up dashboards, or configuring alerting for Deepgram integration health. Trigger with phrases like "deepgram monitoring", "deepgram metrics", "deepgram observability", "monitor deepgram", "deepgram alerts", "deepgram tracing".

10

railway-metrics

davila7

Query resource usage metrics for Railway services. Use when user asks about resource usage, CPU, memory, network, disk, or service performance like "how much memory is my service using" or "is my service slow".

10

Logging and Monitoring for Agentic Workflows

Hack23

Comprehensive observability patterns for GitHub Agentic Workflows including structured logging, metrics collection, alerting strategies, debugging techniques, and production monitoring best practices for autonomous agent systems.

00

mimir

grafana

>

00

autotel

jagreehal

Use when instrumenting with trace/span/track, reviewing code for logging and observability patterns, converting console.log to wide events, adding structured errors, setting up canonical log lines, configuring init(), adding subscribers, or working in the autotel monorepo.

00

ops

matteocervelli

Post-deploy observability for production services — structured logging, RED metrics + Prometheus, Grafana dashboards, SLO-based alerting. Use when setting up monitoring or instrumenting a live service. Trigger on "add logging", "metrics", "dashboard", "alerting", "observability", "monitor this servi

00

observability-setup

spideynolove

Set up metrics, logs, and traces for a service. Use when a mid-level developer needs basic observability coverage.

00

observability

devartblake

Use this skill for logs, metrics, tracing, health checks, and dashboards.

00

dotnet-observability

rudironsoni

>-

00

ai-debug-harness

DeandreFu

Self-driving Electron + Node debug harness with NDJSON logs, Playwright E2E, doctor preflight, and a YAML-frontmatter cookbook of known causes. Use when a voice-chat E2E flow misbehaves locally, audio publish silently fails, the agent worker emits warnings, or any LiveKit/Electron/Fastify error appe

00

performance-monitor

adiytharpansa

Real-time monitoring and optimization of system performance. Use when you need to track response times, resource usage, skill efficiency, and overall system health. This skill enables continuous performance optimization, bottleneck detection, and proactive improvements to keep OpenClaw running at pe

00

infra-guardian

diegosouzapw

OpenClaw Agent Infrastructure Guardian — keep your agent's infrastructure alive. Process lifecycle management with detached execution, auto-restart on failure. Cron scheduler health monitoring (per-job detection, auto-recovery). Direct Telegram/messaging alerts independent of OpenClaw. System-level

00

analyze-prod-issue

MaxPayne89

>-

00

How to choose a Datadog skill

Evaluate each skill based on its intended scope and its trigger requirements. Check if the skill relies on specific API integrations, such as the Datadog API, or if it operates locally to monitor server resources like GPU and CPU status. Review the implementation to ensure it matches your project's observability needs, such as managing GCP logging versus application performance profiling. Finally, consider the maintenance status and the expected output format of each skill to ensure it integrates cleanly with your agent’s existing task execution pipeline.

Frequently asked

Are these skills official Datadog products?
No. These are independent AI-agent skills designed by the developer community. While some skills provide functionality to interface with the Datadog API for managing dashboards or alerts, they are not developed or endorsed by Datadog. Always verify the individual skill implementation before connecting it to your production systems.
How do I know which skill to activate for my specific monitoring task?
Match your requirement to the skill's trigger and primary function. If you need local system data, use the system-monitor skill. For cloud-specific observability like Google Cloud logs, select the gcloud-usage skill. Each skill is designed for narrow, specific tasks, so choose the one that maps directly to the infrastructure layer you are currently investigating.

Skills for other tools

Search skills

Search the agent skills registry