VA

vastai-prod-checklist

Provides a comprehensive pre-flight checklist for launching production GPU workloads on Vast.ai to ensure safety and stability.

Install

mkdir -p .claude/skills/vastai-prod-checklist && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8175" && unzip -o skill.zip -d .claude/skills/vastai-prod-checklist && rm skill.zip

Installs to .claude/skills/vastai-prod-checklist

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Execute Vast.ai production deployment checklist for GPU workloads.
66 charsno explicit “when” trigger
Intermediate

Key capabilities

  • →Verify account balance and offer availability
  • →Implement spot instance preemption handlers
  • →Configure checkpointing to persistent storage
  • →Monitor GPU utilization and instance health
  • →Set budget and spending limits

How it works

The checklist provides a structured set of pre-flight verification steps and scripts to ensure that GPU workloads are configured for reliability, data safety, and cost management.

Inputs & outputs

You give it
Production training configuration
You get back
Readiness verification report

When to use vastai-prod-checklist

  • →Audit production readiness for GPU jobs
  • →Prepare for large-scale model training
  • →Validate instance reliability and disk configuration
  • →Implement go-live procedures for AI workloads

About this skill

Vast.ai Production Go/No-Go

Overview

Production approval is an evidence bundle, not a list of optimistic assertions. Every gate needs an owner, artifact, and expiry; any failed critical gate yields NO-GO.

Prerequisites

  • Release ID, image digest, dataset/model identity, and expected traffic or job profile
  • Capacity, latency, reliability, recovery, security, and spend objectives
  • Named incident, billing, data-recovery, and teardown owners

Instructions

Step 1: Verify account and permissions

Confirm account/team context, positive balance or approved autobilling, least-privilege keys, audit visibility, and no credential in release artifacts.

Step 2: Verify capacity policy

Demonstrate compliant offers or Serverless worker capacity across required GPU, VRAM, geography, reliability, and price constraints.

Step 3: Verify immutable execution

Pin image/template/model identity and prove startup, health, output contract, and workload-specific acceptance on a canary.

Step 4: Verify recovery

Restore from an external checkpoint or roll back a Serverless template. Prove that stop, outbid, offline, expiry, and zero-balance paths have owners.

Step 5: Verify operations

Show bounded retries, terminal-state handling, logs/metrics, signed webhook or polling coverage, cost alarms, and escalation contacts.

Step 6: Issue the decision

Record PASS, FAIL, owner, evidence URI, and expiry for each gate. Launch only on GO; retain the exact rollback and cleanup commands.

Authentication

Production keys must be named and scoped by function. Separate billing, team administration, deployment, workload storage, and monitoring authority.

Tool Discipline

Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.

Output

  • Per-gate PASS/FAIL matrix with evidence and expiry
  • GO or NO-GO decision with accepted residual risks
  • Rollback, incident, billing, and teardown owner receipt

Return release ID, account context, immutable identities, gate results, decision, approvers, expiry, and rollback target.

Examples

A Serverless model release receives GO only after canary output parity, bounded scale testing, signed webhook acceptance, cost thresholds, and reverse rolling-update evidence are attached.

Error Handling

FailureResponse
Critical evidence is missingIssue NO-GO; an owner assertion is not a substitute.
Offer capacity is temporarily absentDelay or use an approved alternate profile; do not weaken policy silently.
Rollback was not exercisedRun it on the canary before production approval.
Billing owner is unavailableIssue NO-GO because zero balance can stop workloads and endanger data.

Resources

When not to use it

  • →Running production jobs without checkpointing
  • →Ignoring billing alerts

Prerequisites

Vast.ai account with creditsCheckpoint-based training pipeline

Limitations

  • →Requires sufficient balance for job duration
  • →Spot instances may be reclaimed
  • →Data pipeline bottlenecks can cause low utilization

How it compares

This checklist provides a formal audit process for production stability compared to launching jobs without predefined safety and recovery measures.

Compared to similar skills

vastai-prod-checklist side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
vastai-prod-checklist (this skill)02moReviewIntermediate
ecs-runtime-debug-playbook06moNo flagsAdvanced
deploy-preflight02moReviewIntermediate
vibeops03moNo flagsBeginner

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore →

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

Search skills

Search the agent skills registry