WI

windsurf-incident-runbook

Provides a structured runbook for handling Windsurf service issues and production bugs caused by AI code.

Install

mkdir -p .claude/skills/windsurf-incident-runbook && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4813" && unzip -o skill.zip -d .claude/skills/windsurf-incident-runbook && rm skill.zip

Installs to .claude/skills/windsurf-incident-runbook

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Execute Windsurf incident response when AI features fail or cause production
76 chars✓ has a “when” trigger
Advanced

Key capabilities

  • Triage incidents based on severity levels
  • Revert deployments caused by AI-generated code
  • Identify Cascade-generated commits in git history
  • Execute team communication during service outages
  • Conduct post-incident reviews for AI workflows

How it works

It provides a structured triage process and specific playbooks for reverting AI-introduced bugs or managing service outages, including templates for team communication and postmortems.

Inputs & outputs

You give it
Incident report or service status
You get back
Mitigation steps and post-incident documentation

When to use windsurf-incident-runbook

  • Handle production incidents caused by AI-generated code
  • Triage Windsurf service outages
  • Execute post-incident reviews for AI workflows
  • Establish team communication for AI failure

About this skill

Windsurf Incident Runbook

Overview

Incident response procedures for Windsurf-related issues: Cascade service outages, AI-generated code causing bugs, and team workflow disruptions.

Prerequisites

  • Access to Windsurf dashboard and status page
  • Git access to affected repositories
  • Team communication channel (Slack, Teams)

Severity Levels

LevelDefinitionResponse TimeExamples
P1Production broken by AI code< 15 minCascade-generated code deployed with critical bug
P2Team workflow blocked< 1 hourWindsurf service outage, all Cascade down
P3Degraded AI features< 4 hoursSlow Cascade, Supercomplete intermittent
P4Minor inconvenienceNext business daySpecific model unavailable, feature regression

Quick Triage Decision Tree

Is Windsurf service itself down?
├─ YES: Check https://status.windsurf.com
│   ├─ Status page shows incident → WAIT for Windsurf to resolve
│   │   Action: Switch to manual coding, notify team
│   └─ Status page green → Local issue
│       Action: Restart Windsurf, check internet, re-authenticate
│
└─ NO: Did AI-generated code cause a production issue?
    ├─ YES → P1 INCIDENT
    │   1. Revert the deployment immediately
    │   2. Identify the Cascade-generated commit(s)
    │   3. Fix manually or with targeted Cascade prompt
    │   4. Post-incident: update review policy
    │
    └─ NO: Is Cascade giving bad suggestions?
        ├─ YES → Check .windsurfrules, start fresh Cascade session
        └─ NO → See windsurf-common-errors

P1 Playbook: AI Code Caused Production Bug

Step 1: Immediate Mitigation

set -euo pipefail
# Revert the deployment
git log --oneline -10  # Find the bad commit(s)

# If tagged with [cascade]:
git revert HEAD --no-edit  # Revert most recent commit
git push origin main       # Deploy revert

# If multiple Cascade commits:
git revert --no-commit HEAD~3..HEAD  # Revert last 3 commits
git commit -m "revert: undo cascade changes causing [issue]"
git push origin main

Step 2: Identify Root Cause

# Find all Cascade-generated commits
git log --all --oneline --grep="cascade" --since="1 week ago"
git log --all --oneline --grep="\[cascade\]" --since="1 week ago"

# Compare before/after
git diff [last-good-commit]..HEAD -- src/

# Common root causes:
# 1. Cascade modified shared utility used by many modules
# 2. Cascade changed error handling (swallowed exceptions)
# 3. Cascade "optimized" code that had intentional behavior
# 4. Cascade introduced dependency on newer API version

Step 3: Fix and Validate

set -euo pipefail
git checkout -b fix/cascade-revert
# Make targeted fix
npm test
npm run typecheck
# Deploy to staging first

P2 Playbook: Windsurf Service Outage

Step 1: Confirm and Communicate

# Check Windsurf status
curl -sf https://status.windsurf.com || echo "Status page unreachable"

Step 2: Team Notification

Team notification template:

Windsurf AI features are currently unavailable.
Status: https://status.windsurf.com

Impact: Cascade and Supercomplete are not working.
Workaround: Continue coding manually. Windsurf still works as a
standard VS Code editor — only AI features are affected.

ETA: Monitoring status page for updates.

Step 3: Workarounds During Outage

1. Windsurf still works as VS Code (file editing, terminal, git)
2. Extensions still work (ESLint, Prettier, debugger)
3. Only Cascade, Supercomplete, and Command mode are down
4. Continue coding manually until service restores
5. Do NOT switch to a different editor mid-task (context loss)

P3 Playbook: Degraded AI Features

Symptoms and fixes:

Slow Cascade → Start fresh session, reduce workspace size
No Supercomplete → Check status bar widget, verify enabled
Wrong model → Check credit balance, switch to available model
MCP disconnected → Restart MCP servers (Command Palette)
Indexing stuck → Reset indexing (Command Palette > "Codeium: Reset Indexing")

Post-Incident Actions

Evidence Collection

set -euo pipefail
# Collect relevant data
mkdir incident-$(date +%Y%m%d)
git log --since="1 day ago" --stat > incident-$(date +%Y%m%d)/commits.txt
cp .windsurfrules incident-$(date +%Y%m%d)/ 2>/dev/null || true
# See windsurf-debug-bundle for full diagnostic collection

Postmortem Template

## Incident: [Title]
**Date:** YYYY-MM-DD
**Duration:** X hours Y minutes
**Severity:** P[1-4]

### Summary
[1-2 sentence description]

### Timeline
- HH:MM — [Event]
- HH:MM — [Event]

### Root Cause
[Was this an AI-generated code issue? Windsurf service issue? Config issue?]

### What Went Wrong
- [ ] AI-generated code not reviewed thoroughly
- [ ] Missing tests for AI-modified code
- [ ] .windsurfrules didn't prevent the bad pattern
- [ ] Cascade modified shared code without constraint

### Action Items
- [ ] Update .windsurfrules to prevent this pattern
- [ ] Add test coverage for affected module
- [ ] Update team Cascade usage policy
- [ ] Add CI gate for AI-modified code

Error Handling

IssueImmediate ActionLong-Term Fix
AI code in prod broke featureGit revert + redeployEnforce test gates for Cascade commits
Windsurf service downCode manuallyNo action needed — external service
AI modified protected filesGit revert those filesAdd to .codeiumignore
Team lost work from CascadeRecover from git historyEnforce pre-Cascade git commit policy

Examples

Quick Health Check

curl -sf https://status.windsurf.com | head -5 || echo "WINDSURF STATUS UNREACHABLE"

Find Recent Cascade Commits

git log --all --oneline --since="7 days ago" | grep -i cascade

Resources

Next Steps

For data handling compliance, see windsurf-data-handling.

When not to use it

  • When the issue is a minor feature request
  • When the Windsurf service is fully operational and no AI-generated bug exists

Prerequisites

Access to Windsurf dashboard and status pageGit access to affected repositoriesTeam communication channel

Limitations

  • P1 response time required is less than 15 minutes
  • Requires manual coding if Windsurf service is down

How it compares

This runbook is specifically tailored for AI-integrated development environments, focusing on the unique risks of AI-generated code and service dependencies.

Compared to similar skills

windsurf-incident-runbook side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
windsurf-incident-runbook (this skill)127dReviewAdvanced
find-bugs57moNo flagsIntermediate
supabase-common-errors427dReviewIntermediate
tech-debt-analyzer59moReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

find-bugs

davila7

Find bugs, security vulnerabilities, and code quality issues in local branch changes. Use when asked to review changes, find bugs, security review, or audit code on the current branch.

529

supabase-common-errors

jeremylongshore

Execute diagnose and fix Supabase common errors and exceptions. Use when encountering Supabase errors, debugging failed requests, or troubleshooting integration issues. Trigger with phrases like "supabase error", "fix supabase", "supabase not working", "debug supabase".

430

tech-debt-analyzer

ailabs-393

This skill should be used when analyzing technical debt in a codebase, documenting code quality issues, creating technical debt registers, or assessing code maintainability. Use this for identifying code smells, architectural issues, dependency problems, missing documentation, security vulnerabilities, and creating comprehensive technical debt documentation.

522

static-analysis

gmh5225

Expertise in LLVM-based static analysis including dataflow analysis, pointer analysis, taint tracking, and program verification. Use this skill when implementing security scanners, bug finders, code quality tools, or performing program analysis research.

518

agent-code-analyzer

ruvnet

Agent skill for code-analyzer - invoke with $agent-code-analyzer

317

memory-safety-patterns

sickn33

Implement memory-safe programming with RAII, ownership, smart pointers, and resource management across Rust, C++, and C. Use when writing safe systems code, managing resources, or preventing memory bugs.

415

Search skills

Search the agent skills registry