Tags
Best Observability Skills for AI Agents
147 Observability skills for AI coding assistants — ranked by popularity.
This collection provides modular SKILL.md files designed to turn your AI coding agent into an expert observability engineer. These skills allow agents like Claude Code, Codex, or Cursor to execute specific infrastructure tasks, ranging from setting up Prometheus metric collection and distributed tracing with Jaeger to managing complex service mesh environments. Developers can delegate the heavy lifting of production monitoring, logging, and SLI/SLO strategy to their agent, ensuring systems remain transparent and measurable. Whether you are troubleshooting request flows across microservices, configuring Langfuse for LLM evaluation, or using Arize Phoenix to trace AI model performance, these skills provide the direct context your agent needs. By integrating terminal awareness tools like terminal-context, agents gain the ability to map active processes and ports, bridging the gap between your code and your live runtime environment. These tools are built for engineers who need to automate the implementation of monitoring standards without sacrificing architectural control.
Top Observability skills
distributed-tracing
wshobson
Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks. Use when debugging microservices, analyzing request flows, or implementing observability for distributed systems.
service-mesh-observability
wshobson
Implement comprehensive observability for service meshes including distributed tracing, metrics, and visualization. Use when setting up mesh monitoring, debugging latency issues, or implementing SLOs for service communication.
terminal-context
aelaguiz
Complete Kitty terminal awareness + control for coding agents: list panes/tabs, read scrollback, map ports→processes, parse per-pane git/last-command metadata (from shell hooks), and send commands/focus panes. Use when user mentions "another terminal", "is the server running", "what failed", or you need to run/inspect commands across panes.
observability-engineer
sickn33
Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows. Use PROACTIVELY for monitoring infrastructure, performance optimization, or production reliability.
prometheus-configuration
wshobson
Set up Prometheus for comprehensive metric collection, storage, and monitoring of infrastructure and applications. Use when implementing metrics collection, setting up monitoring infrastructure, or configuring alerting systems.
langfuse
davila7
Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production. Use when: langfuse, llm observability, llm tracing, prompt management, llm evaluation.
slo-implementation
wshobson
Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting. Use when establishing reliability targets, implementing SRE practices, or measuring service performance.
agent-performance-monitor
ruvnet
Agent skill for performance-monitor - invoke with $agent-performance-monitor
langsmith-observability
davila7
LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications.
opentelemetry-instrumentation-extension
docker
Extend OpenTelemetry instrumentation when new functionality is added to the MCP Gateway. Use when (1) new operations/functions are added, (2) reviewing code for missing instrumentation, (3) user requests otel/telemetry additions, or (4) working with state-changing operations. Analyzes git diff, suggests instrumentation points following project standards in docs/telemetry/README.md, implements with approval, writes tests, updates documentation, and verifies with debug logging and docker logs.
agent-v3-performance-engineer
ruvnet
Agent skill for v3-performance-engineer - invoke with $agent-v3-performance-engineer
error-diagnostics-smart-debug
sickn33
Use when working with error diagnostics smart debug
phoenix-observability
davila7
Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights.
debugging-toolkit-smart-debug
sickn33
Use when working with debugging toolkit smart debug
agent-performance-optimizer
ruvnet
Agent skill for performance-optimizer - invoke with $agent-performance-optimizer
mlops-observability
fmind
Guide to implement full stack observability including reproducibility, lineage, monitoring, alerting, and explainability.
jaeger-analysis
incidentfox
Jaeger distributed tracing analysis. Use when investigating request latency, tracing errors across services, finding slow spans, or understanding service dependencies.
agentation
benjitaylor
Add Agentation visual feedback toolbar to a Next.js project
log-analyzer
mikopbx
Анализ логов Docker контейнера для диагностики проблем и мониторинга здоровья системы. Использовать при отладке ошибок, отслеживании процессов воркеров, исследовании проблем API или мониторинге поведения системы после тестов.
database-migrations-migration-observability
sickn33
Migration monitoring, CDC, and observability infrastructure
typescript-sdk
comet-ml
TypeScript SDK patterns for Opik. Use when working in sdks/opik-typescript.
langsmith-fetch
ComposioHQ
Debug LangChain and LangGraph agents by fetching execution traces from LangSmith Studio. Use when debugging agent behavior, investigating errors, analyzing tool calls, checking memory operations, or examining agent performance. Automatically fetches recent traces and analyzes execution patterns. Requires langsmith-fetch CLI installed.
appinsights-instrumentation
github
Instrument a webapp to send useful telemetry data to Azure App Insights
logging
HoangNguyen0403
Standards for structured logging and observability in Golang.
How to choose a Observability skill
Evaluate these skills based on your specific stack and project requirements. Start by checking the intended use case for each skill, such as whether you need infrastructure-level monitoring with Prometheus or LLM-specific tracing via Langfuse or Arize Phoenix. Review the contributor notes to see if the skill provides actionable code generation or diagnostic analysis. Prioritize skills that match your current observability maturity, choosing targeted tools like slo-implementation for policy definition or terminal-context if you need the agent to interact directly with your live development environment.
More Observability skills
+27 more — browse all skills.