vastai-prod-checklist
Provides a comprehensive pre-flight checklist for launching production GPU workloads on Vast.ai to ensure safety and stability.
Install
mkdir -p .claude/skills/vastai-prod-checklist && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8175" && unzip -o skill.zip -d .claude/skills/vastai-prod-checklist && rm skill.zipInstalls to .claude/skills/vastai-prod-checklist
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Execute Vast.ai production deployment checklist for GPU workloads.Key capabilities
- →Verify account balance and offer availability
- →Implement spot instance preemption handlers
- →Configure checkpointing to persistent storage
- →Monitor GPU utilization and instance health
- →Set budget and spending limits
How it works
The checklist provides a structured set of pre-flight verification steps and scripts to ensure that GPU workloads are configured for reliability, data safety, and cost management.
Inputs & outputs
When to use vastai-prod-checklist
- →Audit production readiness for GPU jobs
- →Prepare for large-scale model training
- →Validate instance reliability and disk configuration
- →Implement go-live procedures for AI workloads
About this skill
Vast.ai Production Go/No-Go
Overview
Production approval is an evidence bundle, not a list of optimistic assertions. Every gate needs an owner, artifact, and expiry; any failed critical gate yields NO-GO.
Prerequisites
- Release ID, image digest, dataset/model identity, and expected traffic or job profile
- Capacity, latency, reliability, recovery, security, and spend objectives
- Named incident, billing, data-recovery, and teardown owners
Instructions
Step 1: Verify account and permissions
Confirm account/team context, positive balance or approved autobilling, least-privilege keys, audit visibility, and no credential in release artifacts.
Step 2: Verify capacity policy
Demonstrate compliant offers or Serverless worker capacity across required GPU, VRAM, geography, reliability, and price constraints.
Step 3: Verify immutable execution
Pin image/template/model identity and prove startup, health, output contract, and workload-specific acceptance on a canary.
Step 4: Verify recovery
Restore from an external checkpoint or roll back a Serverless template. Prove that stop, outbid, offline, expiry, and zero-balance paths have owners.
Step 5: Verify operations
Show bounded retries, terminal-state handling, logs/metrics, signed webhook or polling coverage, cost alarms, and escalation contacts.
Step 6: Issue the decision
Record PASS, FAIL, owner, evidence URI, and expiry for each gate. Launch only on GO; retain the exact rollback and cleanup commands.
Authentication
Production keys must be named and scoped by function. Separate billing, team administration, deployment, workload storage, and monitoring authority.
Tool Discipline
Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.
Output
- Per-gate PASS/FAIL matrix with evidence and expiry
- GO or NO-GO decision with accepted residual risks
- Rollback, incident, billing, and teardown owner receipt
Return release ID, account context, immutable identities, gate results, decision, approvers, expiry, and rollback target.
Examples
A Serverless model release receives GO only after canary output parity, bounded scale testing, signed webhook acceptance, cost thresholds, and reverse rolling-update evidence are attached.
Error Handling
| Failure | Response |
|---|---|
| Critical evidence is missing | Issue NO-GO; an owner assertion is not a substitute. |
| Offer capacity is temporarily absent | Delay or use an approved alternate profile; do not weaken policy silently. |
| Rollback was not exercised | Run it on the canary before production approval. |
| Billing owner is unavailable | Issue NO-GO because zero balance can stop workloads and endanger data. |
Resources
When not to use it
- →Running production jobs without checkpointing
- →Ignoring billing alerts
Prerequisites
Limitations
- →Requires sufficient balance for job duration
- →Spot instances may be reclaimed
- →Data pipeline bottlenecks can cause low utilization
How it compares
This checklist provides a formal audit process for production stability compared to launching jobs without predefined safety and recovery measures.
Compared to similar skills
vastai-prod-checklist side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| vastai-prod-checklist (this skill) | 0 | 2mo | Review | Intermediate |
| ecs-runtime-debug-playbook | 0 | 6mo | No flags | Advanced |
| deploy-preflight | 0 | 2mo | Review | Intermediate |
| vibeops | 0 | 3mo | No flags | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
ecs-runtime-debug-playbook
talolard
Debug AWS ECS or Fargate deployments where CI or workflow status does not match live behavior, especially when an old task definition keeps serving traffic, a new task exits during startup, health checks are false-green, Alembic or schema state may be inconsistent with physical tables, or AWS profil
deploy-preflight
vincentxuu
Pre-deploy verification for the quidproquo Cloudflare Workers site — run lint / astro check / tests / check:references, sanity-check git state and wrangler config, and produce a go/no-go report before `pnpm deploy`. Does NOT deploy. Use when user says 準備 deploy / 上線前檢查 / preflight / 部署前看一下.
vibeops
rifatshampod
>
ops
denniszielke
>
gcp-cloud-run
aj-geddes
Deploy containerized applications on Google Cloud Run with automatic scaling, traffic management, and service mesh integration. Use for container-based serverless computing.
deployment-pipeline-design
wshobson
Design multi-stage CI/CD pipelines with approval gates, security checks, and deployment orchestration. Use when architecting deployment workflows, setting up continuous delivery, or implementing GitOps practices.