Tools

AI-Agent Skills for Grafana Tasks

15 agent skills that handle jobs people use Grafana for — for Claude Code, Codex, and Cursor.

This collection provides specific skills for AI agents to assist with tasks often managed through Grafana-based observability stacks. These skills are designed for DevOps engineers, site reliability engineers, and platform developers who need to automate complex monitoring and visualization workflows. Each skill adds targeted functionality to your AI agent—such as Claude Code, Codex, or Cursor—allowing it to handle infrastructure tasks like configuring database migration observability, setting up Databricks monitoring, or performing incident response. Instead of manually navigating interfaces, you can task your agent to deploy service mesh observability, manage dashboard configurations via API, or monitor local server resources like CPU and GPU. These tools act as specialized modules that bridge the gap between your intent and the technical execution of observability strategy. Whether you are debugging a production incident or standing up monitoring for a new data pipeline, these skills allow your agent to execute concrete, repeatable infrastructure operations with precision.

Top Grafana skills

service-mesh-observability

wshobson

Implement comprehensive observability for service meshes including distributed tracing, metrics, and visualization. Use when setting up mesh monitoring, debugging latency issues, or implementing SLOs for service communication.

574

observability-engineer

sickn33

Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows. Use PROACTIVELY for monitoring infrastructure, performance optimization, or production reliability.

1242

database-migrations-migration-observability

sickn33

Migration monitoring, CDC, and observability infrastructure

212

devops-troubleshooter

sickn33

Expert DevOps troubleshooter specializing in rapid incident response, advanced debugging, and modern observability. Masters log analysis, distributed tracing, Kubernetes debugging, performance optimization, and root cause analysis. Handles production outages, system reliability, and preventive monitoring. Use PROACTIVELY for debugging, incident response, or system troubleshooting.

12

observability-monitoring-monitor-setup

sickn33

You are a monitoring and observability expert specializing in implementing comprehensive monitoring solutions. Set up metrics collection, distributed tracing, log aggregation, and create insightful da

12

worker-health-monitoring

dadbodgeoff

Heartbeat-based health monitoring for background workers with configurable thresholds, rolling duration windows, failure rate calculation, and stuck job detection.

12

deepgram-observability

jeremylongshore

Set up comprehensive observability for Deepgram integrations with metrics, traces, and alerts. Use when implementing monitoring for Deepgram operations, setting up dashboards, or configuring alerting for Deepgram integration health. Trigger with phrases like "deepgram monitoring", "deepgram metrics", "deepgram observability", "monitor deepgram", "deepgram alerts", "deepgram tracing".

10

railway-metrics

davila7

Query resource usage metrics for Railway services. Use when user asks about resource usage, CPU, memory, network, disk, or service performance like "how much memory is my service using" or "is my service slow".

10

Logging and Monitoring for Agentic Workflows

Hack23

Comprehensive observability patterns for GitHub Agentic Workflows including structured logging, metrics collection, alerting strategies, debugging techniques, and production monitoring best practices for autonomous agent systems.

00

ops

matteocervelli

Post-deploy observability for production services — structured logging, RED metrics + Prometheus, Grafana dashboards, SLO-based alerting. Use when setting up monitoring or instrumenting a live service. Trigger on "add logging", "metrics", "dashboard", "alerting", "observability", "monitor this servi

00

observability-setup

spideynolove

Set up metrics, logs, and traces for a service. Use when a mid-level developer needs basic observability coverage.

00

observability

devartblake

Use this skill for logs, metrics, tracing, health checks, and dashboards.

00

dotnet-observability

rudironsoni

>-

00

grafana-lens

awsome-o

Grafana tools for data visualization, monitoring, alerting, security, SRE investigation, and data collection pipeline management via Alloy. Use grafana_query, grafana_query_logs, grafana_query_traces, grafana_create_dashboard, grafana_update_dashboard, grafana_create_alert, grafana_share_dashboard,

00

performance-monitor

adiytharpansa

Real-time monitoring and optimization of system performance. Use when you need to track response times, resource usage, skill efficiency, and overall system health. This skill enables continuous performance optimization, bottleneck detection, and proactive improvements to keep OpenClaw running at pe

00

How to choose a Grafana skill

Evaluate these skills based on your current infrastructure scope and maintenance needs. Review the contributor metadata to ensure the skill is updated for your agent's version. Focus on the specific output: some skills, like the observability engineer or devops troubleshooter, provide broader strategy and debugging support, while others are task-specific, such as managing local server monitoring or interacting with Grafana APIs. Check if the skill relies on direct API management versus higher-level strategy implementation to ensure it aligns with your existing automation permissions and technical requirements.

Frequently asked

Do these skills replace my existing Grafana installation?
No. These are agent skills that interact with your infrastructure to automate tasks associated with Grafana and observability. They act as a specialized interface for your AI agent to manage dashboards, alerts, and data sources via API, or to help you design observability strategies. They do not substitute the visualization engine itself.
How do I know which skill is right for my incident response workflow?
Choose skills based on the technical depth required. For rapid debugging and log analysis, the devops-troubleshooter skill is specialized for incident response. If you need to manage dashboard alerts or data source connections during an incident, the grafana or system-monitor skills provide the direct API control needed to adjust your monitoring setup.

Skills for other tools

Search skills

Search the agent skills registry