K8

k8s-incident

Diagnose and respond to Kubernetes outages using standardized runbooks and inspection tools.

Install

mkdir -p .claude/skills/k8s-incident && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/7137" && unzip -o skill.zip -d .claude/skills/k8s-incident && rm skill.zip

Installs to .claude/skills/k8s-incident

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Respond to Kubernetes incidents with runbooks and diagnostics. Use for outages, pod failures, node issues, network problems, and emergency response.
148 chars✓ has a “when” trigger
Intermediate

Key capabilities

  • Check control plane health
  • Assess node status
  • Gather event logs
  • Perform emergency pod deletion
  • Rollback deployments and Helm releases

How it works

The skill provides a structured triage framework using specific Kubernetes diagnostic tools to identify and resolve service outages or pod failures.

Inputs & outputs

You give it
Kubernetes incident context or symptom
You get back
Diagnostic report and resolution actions

When to use k8s-incident

  • Debugging pod crash loops
  • Assessing cluster-wide outages
  • Checking control plane component health
  • Gathering logs for failed deployments

About this skill

Kubernetes Incident Response

Runbooks and diagnostic workflows for common Kubernetes incidents.

When to Apply

Use this skill when:

  • User mentions: "incident", "outage", "emergency", "down", "not working"
  • Operations: emergency response, production issues, service degradation
  • Keywords: "urgent", "broken", "fix", "restore", "recover"

Priority Rules

PriorityRuleImpactTools
1Check control plane firstCRITICALget_pods(namespace="kube-system")
2Assess node healthCRITICALget_nodes
3Gather events before changesHIGHget_events
4Document timelineHIGHManual notes
5Rollback if safeMEDIUMrollback_deployment

Quick Reference

IncidentFirst ToolNext Steps
Pod failureget_pod_logs(previous=True)describe_pod, get_events
Node downdescribe_nodeCheck kubelet logs
Service unreachableget_endpointsget_network_policies
Control planeget_pods(namespace="kube-system")Check API server logs

Incident Triage

Quick Health Check

get_nodes()
get_pods(namespace="kube-system")
get_events(namespace)

Severity Assessment

IndicatorSeverityAction
Multiple nodes NotReadyCriticalEscalate immediately
kube-system pods failingCriticalControl plane issue
Single pod CrashLoopMediumDebug pod
High latencyMediumCheck resources

Runbook: Pod Failures

CrashLoopBackOff

get_pod_logs(name, namespace, previous=True)
describe_pod(name, namespace)
get_events(namespace, field_selector="involvedObject.name=<pod>")
get_pod_metrics(name, namespace)

Common Causes:

  • OOMKilled → Increase memory limits
  • Exit code 1 → Application error in logs
  • Exit code 137 → Killed by OOM or SIGKILL
  • Exit code 143 → Graceful SIGTERM

ImagePullBackOff

describe_pod(name, namespace)
get_secrets(namespace)

Pending Pod

describe_pod(name, namespace)
get_nodes()
get_events(namespace)

Runbook: Node Issues

Node NotReady

describe_node(name)
get_events(namespace="", field_selector="involvedObject.name=<node>")
node_logs_tool(name, "kubelet")

Node DiskPressure

describe_node(name)
get_pods(field_selector="spec.nodeName=<node>")

Runbook: Network Issues

Service Not Accessible

get_services(namespace)
get_endpoints(namespace)
get_pods(namespace, label_selector="<service-selector>")
get_network_policies(namespace)

DNS Resolution Failures

get_pods(namespace="kube-system", label_selector="k8s-app=kube-dns")
get_pod_logs("coredns-xxx", "kube-system")

With Cilium

cilium_status_tool()
cilium_endpoints_list_tool(namespace)
hubble_flows_query_tool(namespace)

With Istio

istio_analyze_tool(namespace)
istio_proxy_status_tool()

Runbook: Storage Issues

PVC Pending

describe_pvc(name, namespace)
get_storage_classes()
get_events(namespace)

Pod Stuck in ContainerCreating

describe_pod(name, namespace)
get_pvc(namespace)
get_events(namespace)

Runbook: Control Plane Issues

API Server Unavailable

get_pods(namespace="kube-system", label_selector="component=kube-apiserver")
get_events(namespace="kube-system")

etcd Issues

get_pods(namespace="kube-system", label_selector="component=etcd")
get_pod_logs("etcd-xxx", "kube-system")

Emergency Actions

Force Delete Pod

delete_pod(name, namespace, grace_period=0, force=True)

Rollback Deployment

rollback_deployment(name, namespace, revision=0)

Helm Rollback

rollback_helm_release(name, namespace, revision=1)

Diagnostic Collection Script

For comprehensive incident diagnostics, see scripts/collect-diagnostics.py.

Multi-Cluster Incident Response

Check all clusters:

for context in ["prod-1", "prod-2", "staging"]:
    get_nodes(context=context)
    get_pods(namespace="kube-system", context=context)
    get_events(namespace="kube-system", context=context)

Post-Incident

Document Timeline

  1. When did the incident start?
  2. What was the impact?
  3. What was the root cause?
  4. What fixed it?

Prevent Recurrence

  • Add monitoring/alerting
  • Improve resource limits
  • Add readiness probes
  • Document runbook

Related Skills

When not to use it

  • Non-Kubernetes infrastructure issues
  • Security-specific incident response

Limitations

  • Requires manual notes for timeline documentation
  • Does not provide automated root cause analysis

How it compares

It automates the selection of diagnostic tools based on the incident type rather than requiring manual selection of kubectl commands.

Compared to similar skills

k8s-incident side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
k8s-incident (this skill)16moReviewIntermediate
k8s-troubleshoot16moReviewIntermediate
debug-cluster28moReviewIntermediate
k8s-cilium16moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by rohitg00

View all by rohitg00

Search skills

Search the agent skills registry