ideogram-incident-runbook
Standardize response procedures for Ideogram outages and service degradation.
Install
mkdir -p .claude/skills/ideogram-incident-runbook && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8633" && unzip -o skill.zip -d .claude/skills/ideogram-incident-runbook && rm skill.zipInstalls to .claude/skills/ideogram-incident-runbook
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Execute Ideogram incident response with triage, mitigation, and postmortem.Key capabilities
- →Execute triage scripts for connectivity and auth
- →Enable fallback image modes during outages
- →Manage incident communication templates
- →Conduct post-incident reviews
How it works
It provides a structured decision tree and triage scripts to identify the cause of service failures, such as auth errors or rate limits, and offers commands to apply temporary mitigations like fallback modes.
Inputs & outputs
When to use ideogram-incident-runbook
- →Triage during API outages
- →Investigating error rates
- →Running post-incident reviews
- →Managing incident severity levels
About this skill
Ideogram Incident Response
Overview
Stabilize an Ideogram integration before chasing image quality or replaying paid requests. Preserve evidence, reduce blast radius, reconcile accepted work, protect sensitive media, and choose recovery actions by incident class.
Prerequisites
- Incident commander, environment, start time, impact statement, and communication channel.
- Read access to content-free application, queue, vendor-status, webhook, polling, storage, and billing evidence.
- Known controls for admission, concurrency, key revocation, traffic rollback, queue drain, and publication stop.
Current Contract
Incidents commonly cross separate boundaries: server-side Api-Key, prepaid shared team credit, default in-flight capacity, async generation_id, signed but non-guaranteed webhook delivery, item safety, expiring URLs, and application storage. Recovery must not collapse these into a generic retry.
Authentication
Never paste the API key into incident chat or commands. Verify secret provenance through metadata; if exposure is credible, stop new traffic, revoke under owner approval, rotate consumers, and audit historical leakage.
Instructions
- Declare severity, impact, affected environment and tenants, known time window, and incident commander.
- Freeze risky deployments and broad retries; preserve sanitized statuses, identifiers, queue state, and storage receipts.
- Classify credential, depleted credit, validation,
429pressure, vendor capacity, unsafe publication, webhook loss, expired URL, or storage failure. - Contain with the smallest control: close admission, reduce concurrency, stop publishing, switch to polling, isolate storage, or roll back traffic.
- Reconcile every accepted async identifier before resubmission and prevent duplicate asset publication.
- Recover with synthetic canaries, then restore bounded traffic while monitoring safety, durable completion, errors, latency, and spend.
- Record timeline, decisions, evidence, customer impact, cleanup, follow-ups, and final state.
Tool Discipline
Use Read, Glob, and Grep for evidence and known runbooks. Use Write and Edit for approved incident records or fixes. Do not revoke keys, add credit, delete assets, replay jobs, or deploy without incident authority.
Approval Boundaries
The incident commander approves containment and restoration; security owns credential response, billing owns credit, moderation owns unsafe publication, and data owners approve asset access or deletion. Separate reversible mitigation from destructive action.
Error Handling
- Do not regenerate solely because a vendor URL expired; confirm rights, budget, and absence of durable copy.
- Do not increase concurrency during
429pressure. - Missing webhook delivery should trigger polling reconciliation, not duplicate submission.
Output
Return incident class, impact, timeline, evidence IDs, containment, accepted-work reconciliation, spend and data exposure, recovery canary, owners, residual risk, and rollback or cleanup state. Exclude secrets and content.
Examples
- On webhook loss, keep submissions bounded, poll known generation IDs, and deduplicate late deliveries.
- On unsafe publication, stop the publisher while preserving the generation and safety decision evidence.
Validation
Confirm impact has stopped, all known generations and objects reconcile, synthetic canaries pass, and monitoring remains stable through the observation window. Test that rollback and emergency admission closure still work.
Resources
- Current first-party evidence map — use the dated endpoint, webhook, billing, team, and training links as the contract index for this workflow.
- Recheck the endpoint-specific page and current OpenAPI description before relying on an enum, limit, beta feature, or lifecycle claim.
- Record live observations as environment-specific evidence, not as universal vendor guarantees.
When not to use it
- →When the issue is a minor bug not affecting service availability
Prerequisites
Limitations
- →Requires pre-configured fallback infrastructure to be effective
- →Severity definitions are specific to Ideogram API integration
How it compares
It replaces ad-hoc troubleshooting with a standardized response process including defined severity levels and communication templates.
Compared to similar skills
ideogram-incident-runbook side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| ideogram-incident-runbook (this skill) | 0 | 2mo | Caution | Advanced |
| observability-engineer | 12 | 5mo | No flags | Advanced |
| mlops-engineer | 3 | 5mo | No flags | Advanced |
| senior-devops | 7 | 9mo | Review | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
observability-engineer
sickn33
Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows. Use PROACTIVELY for monitoring infrastructure, performance optimization, or production reliability.
mlops-engineer
sickn33
Build comprehensive ML pipelines, experiment tracking, and model registries with MLflow, Kubeflow, and modern MLOps tools. Implements automated training, deployment, and monitoring across cloud platforms. Use PROACTIVELY for ML infrastructure, experiment management, or pipeline automation.
senior-devops
davila7
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup, infrastructure as code, deployment automation, and monitoring. Use when setting up pipelines, deploying applications, managing infrastructure, implementing monitoring, or optimizing deployment processes.
server-management
davila7
Server management principles and decision-making. Process management, monitoring strategy, and scaling decisions. Teaches thinking, not commands.
debug-cluster
openshift
Provides systematic debugging approaches for HyperShift hosted-cluster issues. Auto-applies when debugging cluster problems, investigating stuck deletions, or troubleshooting control plane issues.
domain-cloud-native
actionbook
Use when building cloud-native apps. Keywords: kubernetes, k8s, docker, container, grpc, tonic, microservice, service mesh, observability, tracing, metrics, health check, cloud, deployment, 云原生, 微服务, 容器