Tags

Best Observability Skills for AI Agents

147 Observability skills for AI coding assistants — ranked by popularity.

This collection provides modular SKILL.md files designed to turn your AI coding agent into an expert observability engineer. These skills allow agents like Claude Code, Codex, or Cursor to execute specific infrastructure tasks, ranging from setting up Prometheus metric collection and distributed tracing with Jaeger to managing complex service mesh environments. Developers can delegate the heavy lifting of production monitoring, logging, and SLI/SLO strategy to their agent, ensuring systems remain transparent and measurable. Whether you are troubleshooting request flows across microservices, configuring Langfuse for LLM evaluation, or using Arize Phoenix to trace AI model performance, these skills provide the direct context your agent needs. By integrating terminal awareness tools like terminal-context, agents gain the ability to map active processes and ports, bridging the gap between your code and your live runtime environment. These tools are built for engineers who need to automate the implementation of monitoring standards without sacrificing architectural control.

Top Observability skills

distributed-tracing

wshobson

Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks. Use when debugging microservices, analyzing request flows, or implementing observability for distributed systems.

577

service-mesh-observability

wshobson

Implement comprehensive observability for service meshes including distributed tracing, metrics, and visualization. Use when setting up mesh monitoring, debugging latency issues, or implementing SLOs for service communication.

574

terminal-context

aelaguiz

Complete Kitty terminal awareness + control for coding agents: list panes/tabs, read scrollback, map ports→processes, parse per-pane git/last-command metadata (from shell hooks), and send commands/focus panes. Use when user mentions "another terminal", "is the server running", "what failed", or you need to run/inspect commands across panes.

567

observability-engineer

sickn33

Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows. Use PROACTIVELY for monitoring infrastructure, performance optimization, or production reliability.

1242

prometheus-configuration

wshobson

Set up Prometheus for comprehensive metric collection, storage, and monitoring of infrastructure and applications. Use when implementing metrics collection, setting up monitoring infrastructure, or configuring alerting systems.

645

langfuse

davila7

Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production. Use when: langfuse, llm observability, llm tracing, prompt management, llm evaluation.

743

slo-implementation

wshobson

Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting. Use when establishing reliability targets, implementing SRE practices, or measuring service performance.

338

agent-performance-monitor

ruvnet

Agent skill for performance-monitor - invoke with $agent-performance-monitor

335

langsmith-observability

davila7

LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications.

430

opentelemetry-instrumentation-extension

docker

Extend OpenTelemetry instrumentation when new functionality is added to the MCP Gateway. Use when (1) new operations/functions are added, (2) reviewing code for missing instrumentation, (3) user requests otel/telemetry additions, or (4) working with state-changing operations. Analyzes git diff, suggests instrumentation points following project standards in docs/telemetry/README.md, implements with approval, writes tests, updates documentation, and verifies with debug logging and docker logs.

326

agent-v3-performance-engineer

ruvnet

Agent skill for v3-performance-engineer - invoke with $agent-v3-performance-engineer

323

error-diagnostics-smart-debug

sickn33

Use when working with error diagnostics smart debug

521

phoenix-observability

davila7

Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights.

323

debugging-toolkit-smart-debug

sickn33

Use when working with debugging toolkit smart debug

421

agent-performance-optimizer

ruvnet

Agent skill for performance-optimizer - invoke with $agent-performance-optimizer

318

mlops-observability

fmind

Guide to implement full stack observability including reproducibility, lineage, monitoring, alerting, and explainability.

216

jaeger-analysis

incidentfox

Jaeger distributed tracing analysis. Use when investigating request latency, tracing errors across services, finding slow spans, or understanding service dependencies.

611

agentation

benjitaylor

Add Agentation visual feedback toolbar to a Next.js project

69

log-analyzer

mikopbx

Анализ логов Docker контейнера для диагностики проблем и мониторинга здоровья системы. Использовать при отладке ошибок, отслеживании процессов воркеров, исследовании проблем API или мониторинге поведения системы после тестов.

213

database-migrations-migration-observability

sickn33

Migration monitoring, CDC, and observability infrastructure

212

typescript-sdk

comet-ml

TypeScript SDK patterns for Opik. Use when working in sdks/opik-typescript.

212

langsmith-fetch

ComposioHQ

Debug LangChain and LangGraph agents by fetching execution traces from LangSmith Studio. Use when debugging agent behavior, investigating errors, analyzing tool calls, checking memory operations, or examining agent performance. Automatically fetches recent traces and analyzes execution patterns. Requires langsmith-fetch CLI installed.

67

appinsights-instrumentation

github

Instrument a webapp to send useful telemetry data to Azure App Insights

66

logging

HoangNguyen0403

Standards for structured logging and observability in Golang.

210

How to choose a Observability skill

Evaluate these skills based on your specific stack and project requirements. Start by checking the intended use case for each skill, such as whether you need infrastructure-level monitoring with Prometheus or LLM-specific tracing via Langfuse or Arize Phoenix. Review the contributor notes to see if the skill provides actionable code generation or diagnostic analysis. Prioritize skills that match your current observability maturity, choosing targeted tools like slo-implementation for policy definition or terminal-context if you need the agent to interact directly with your live development environment.

More Observability skills

application-performance-performance-optimization
sickn33 · 2 installs
agent-session-monitor
alibaba · 2 installs
instantly-observability
jeremylongshore · 2 installs
session-investigator
evalstate · 2 installs
trulens-instrumentation
truera · 2 installs
k8s-cilium
rohitg00 · 1 installs
vastai-observability
jeremylongshore · 1 installs
effect-patterns-observability
PaulJPhilp · 1 installs
model-debugging
pollinations · 1 installs
customerio-observability
jeremylongshore · 1 installs
gcloud-usage
fcakyon · 1 installs
health-checks
dadbodgeoff · 3 installs
k8s-cost
rohitg00 · 3 installs
k8s-incident
rohitg00 · 1 installs
error-debugging-error-analysis
sickn33 · 1 installs
error-diagnostics-error-trace
sickn33 · 1 installs
k8s-diagnostics
rohitg00 · 1 installs
phoenix-tracing
Arize-ai · 1 installs
tldr-stats
parcadei · 1 installs
trace-claude-code
parcadei · 2 installs
apollo-observability
jeremylongshore · 1 installs
azure-monitor-ingestion-py
microsoft · 1 installs
azure-monitor-opentelemetry-exporter-py
microsoft · 1 installs
braintrust-analyze
parcadei · 1 installs
braintrust-tracing
parcadei · 1 installs
cloudwatch
itsmostafa · 1 installs
devops-troubleshooter
sickn33 · 1 installs
distributed-debugging-debug-trace
sickn33 · 1 installs
gamma-observability
jeremylongshore · 1 installs
k8s-troubleshoot
rohitg00 · 1 installs
langfuse-deploy-integration
jeremylongshore · 1 installs
observability
incidentfox · 1 installs
observability-monitoring-monitor-setup
sickn33 · 1 installs
posthog-observability
jeremylongshore · 1 installs
vercel-observability
jeremylongshore · 2 installs
clay-observability
jeremylongshore · 1 installs
clerk-observability
jeremylongshore · 1 installs
error-debugging-error-trace
sickn33 · 1 installs
fireflies-observability
jeremylongshore · 0 installs
ideogram-observability
jeremylongshore · 1 installs
langfuse-common-errors
jeremylongshore · 1 installs
langfuse-reference-architecture
jeremylongshore · 1 installs
log-focus-debug
solidSpoon · 1 installs
monitoring-database-transactions
jeremylongshore · 1 installs
observability-monitoring-slo-implement
sickn33 · 1 installs
openevidence-observability
jeremylongshore · 1 installs
replit-observability
jeremylongshore · 1 installs
coderabbit-observability
jeremylongshore · 1 installs
coralogix-analysis
incidentfox · 1 installs
evernote-observability
jeremylongshore · 1 installs
exa-observability
jeremylongshore · 1 installs
incident-responder
sickn33 · 1 installs
langfuse-core-workflow-a
jeremylongshore · 0 installs
langfuse-debug-bundle
jeremylongshore · 0 installs
langfuse-enterprise-rbac
jeremylongshore · 0 installs
langfuse-hello-world
jeremylongshore · 1 installs
langfuse-install-auth
jeremylongshore · 0 installs
langfuse-local-dev-loop
jeremylongshore · 1 installs
langfuse-prod-checklist
jeremylongshore · 1 installs
linear-observability
jeremylongshore · 1 installs
performance-engineer
sickn33 · 1 installs
perplexity-observability
jeremylongshore · 1 installs
railway-metrics
davila7 · 1 installs
sentry-observability
jeremylongshore · 0 installs
supabase-observability
jeremylongshore · 0 installs
tinybird-deploy
pollinations · 1 installs
groq-observability
jeremylongshore · 0 installs
langfuse-migration-deep-dive
jeremylongshore · 0 installs
langfuse-webhooks-events
jeremylongshore · 0 installs
lindy-observability
jeremylongshore · 0 installs
triage-logs
rdkcentral · 0 installs
Logging and Monitoring for Agentic Workflows
Hack23 · 0 installs
sail-voyage
sailresearchco · 0 installs
Jaeger Trace Explorer
agentskillexchange · 0 installs
maple-rust-style
Makisuo · 0 installs
Nginx Error Pattern Analyzer
agentskillexchange · 0 installs
mcp
Piebald-AI · 0 installs
lmt
SELISEdigitalplatforms · 0 installs
autotel
jagreehal · 0 installs
investigate-telemetry
kibertoad · 0 installs
obs-alerts
shepard-system · 0 installs
investigate
agatx · 0 installs
ops
matteocervelli · 0 installs
mission-control-event-digest
MN755 · 0 installs
observability-setup
spideynolove · 0 installs
update-otel-genai-conventions
dotnet · 0 installs
observability
devartblake · 0 installs
go-logging-observability
riordanpawley · 0 installs
dotnet-observability
rudironsoni · 0 installs
altinity-expert-clickhouse-metrics
ntk148v · 0 installs
yoyopod-logs
attmous · 0 installs
quaid-alerts
quaid-app · 0 installs
sre-practices
duylinhdang1998 · 0 installs
stop-monitor
akeemphilbert · 0 installs
homelab-investigator
macgregor · 0 installs
prompt-logging-hook-setup
YangSiJun528 · 0 installs

+27 more — browse all skills.

Frequently asked

Do I need to install anything on my server to use these skills?
These SKILL.md files act as instructions for your AI agent, not binary software. Your agent uses these files to understand how to interact with your existing tools like Prometheus, Jaeger, or Langfuse. You only need the relevant observability platforms already running in your environment for the agent to effectively configure or monitor them.
Can these skills help me debug LLM-specific issues?
Yes. If your project involves AI, you can select the Langfuse or Arize Phoenix skills. These are specifically designed to help your agent implement tracing, evaluate prompt performance, and manage datasets, allowing the agent to identify why a specific model output failed or performed poorly within your application flow.

Browse other tags

Search skills

Search the agent skills registry