K8

k8s-troubleshoot

Automated diagnostics for Kubernetes clusters, covering pod logs, events, and node health.

Install

mkdir -p .claude/skills/k8s-troubleshoot && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/5528" && unzip -o skill.zip -d .claude/skills/k8s-troubleshoot && rm skill.zip

Installs to .claude/skills/k8s-troubleshoot

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Debug Kubernetes pods, nodes, and workloads. Use when pods are failing, containers crash, nodes are unhealthy, or users mention debugging, troubleshooting, or diagnosing Kubernetes issues.
188 chars✓ has a “when” trigger
Intermediate

Key capabilities

  • Inspect pod status and conditions
  • Retrieve and analyze pod logs
  • Check node health and resource pressure
  • Examine Kubernetes events
  • Debug DNS resolution issues

How it works

It aggregates data from pod states, events, and resource metrics to identify root causes for common failures like CrashLoopBackOff or OOMKilled.

Inputs & outputs

You give it
Pod or node name
You get back
Diagnostic report or log output

When to use k8s-troubleshoot

  • Debug CrashLoopBackOff states
  • Inspect logs for failing pods
  • Analyze kubernetes node health
  • Investigate pod resource usage

About this skill

Kubernetes Troubleshooting

Expert debugging and diagnostics for Kubernetes clusters using kubectl-mcp-server tools.

When to Apply

Use this skill when:

  • User mentions: "debug", "troubleshoot", "diagnose", "failing", "crash", "not starting", "broken"
  • Pod states: Pending, CrashLoopBackOff, ImagePullBackOff, OOMKilled, Error, Unknown
  • Node issues: NotReady, MemoryPressure, DiskPressure, NetworkUnavailable, PIDPressure
  • Keywords: "logs", "events", "describe", "why isn't working", "stuck", "not responding"

Priority Rules

PriorityRuleImpactTools
1Check pod status firstCRITICALget_pods, describe_pod
2View recent eventsCRITICALget_events
3Inspect logs (including previous)HIGHget_pod_logs
4Check resource metricsHIGHget_pod_metrics
5Verify endpointsMEDIUMget_endpoints
6Review network policiesMEDIUMget_network_policies
7Examine node statusLOWget_nodes, describe_node

Quick Reference

SymptomFirst ToolNext Steps
Pod Pendingdescribe_podCheck events, node capacity, resource requests
CrashLoopBackOffget_pod_logs(previous=True)Check exit code, resources, liveness probes
ImagePullBackOffdescribe_podVerify image name, registry auth, network
OOMKilledget_pod_metricsIncrease memory limits, check for memory leaks
ContainerCreatingdescribe_podCheck PVC binding, secrets, configmaps
Terminating (stuck)describe_podCheck finalizers, PDBs, preStop hooks

Diagnostic Workflows

Pod Not Starting

1. get_pods(namespace, label_selector) - Get pod status
2. describe_pod(name, namespace) - See events and conditions
3. get_events(namespace, field_selector="involvedObject.name=<pod>") - Check events
4. get_pod_logs(name, namespace, previous=True) - For crash loops

Common Pod States

StateLikely CauseTools to Use
PendingScheduling issuesdescribe_pod, get_nodes, get_events
ImagePullBackOffRegistry/authdescribe_pod, check image name
CrashLoopBackOffApp crashget_pod_logs(previous=True)
OOMKilledMemory limitget_pod_metrics, adjust limits
ContainerCreatingVolume/networkdescribe_pod, get_pvc

Node Issues

1. get_nodes() - List nodes and status
2. describe_node(name) - See conditions and capacity
3. Check: Ready, MemoryPressure, DiskPressure, PIDPressure
4. node_logs_tool(name, "kubelet") - Kubelet logs

Deep Debugging Workflows

CrashLoopBackOff Investigation

1. get_pod_logs(name, namespace, previous=True) - See why it crashed
2. describe_pod(name, namespace) - Check resource limits, probes
3. get_pod_metrics(name, namespace) - Memory/CPU at crash time
4. If OOM: compare requests/limits to actual usage
5. If app error: check logs for stack trace

Networking Issues

1. get_services(namespace) - Verify service exists
2. get_endpoints(namespace) - Check endpoint backends
3. If empty endpoints: pods don't match selector
4. get_network_policies(namespace) - Check traffic rules
5. For Cilium: cilium_endpoints_list_tool(), hubble_flows_query_tool()

Storage Problems

1. get_pvc(namespace) - Check PVC status
2. describe_pvc(name, namespace) - See binding issues
3. get_storage_classes() - Verify provisioner exists
4. If Pending: check storage class, access modes

DNS Resolution

1. kubectl_exec(pod, namespace, "nslookup kubernetes.default") - Test DNS
2. If fails: check coredns pods in kube-system
3. get_pods(namespace="kube-system", label_selector="k8s-app=kube-dns")
4. get_pod_logs(name="coredns-*", namespace="kube-system")

Multi-Cluster Debugging

All tools support context parameter for targeting different clusters:

get_pods(namespace="kube-system", context="production-cluster")
get_events(namespace="default", context="staging-cluster")
describe_pod(name="myapp-xyz", namespace="prod", context="prod-east")

Diagnostic Scripts

For comprehensive diagnostics, run the bundled scripts:

Decision Tree

See references/DECISION-TREE.md for visual troubleshooting flowcharts.

Common Errors Reference

See references/COMMON-ERRORS.md for error message explanations and fixes.

Related Tools

Core Diagnostics

  • get_pods, describe_pod, get_pod_logs, get_pod_metrics
  • get_events, get_nodes, describe_node
  • get_resource_usage, compare_namespaces

Advanced (Ecosystem)

  • Cilium: cilium_endpoints_list_tool, hubble_flows_query_tool
  • Istio: istio_proxy_status_tool, istio_analyze_tool

Related Skills

When not to use it

  • When cluster-level logs are not accessible
  • When the issue is outside the Kubernetes control plane

Prerequisites

kubectl

Limitations

  • Requires sufficient permissions to view logs
  • Limited to Kubernetes-native resources

How it compares

It provides a structured diagnostic workflow that correlates events and logs rather than just returning raw output.

Compared to similar skills

k8s-troubleshoot side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
k8s-troubleshoot (this skill)16moReviewIntermediate
k8s-incident16moReviewIntermediate
debug-cluster28moReviewIntermediate
k8s-cilium16moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by rohitg00

View all by rohitg00

Search skills

Search the agent skills registry