VA

vastai-core-workflow-b

Orchestrates multi-GPU instances and automates spot interruption recovery on Vast.ai.

Install

mkdir -p .claude/skills/vastai-core-workflow-b && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/3820" && unzip -o skill.zip -d .claude/skills/vastai-core-workflow-b && rm skill.zip

Installs to .claude/skills/vastai-core-workflow-b

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Execute Vast.ai secondary workflow: multi-instance orchestration, spot
70 charsno explicit “when” trigger
Advanced

Key capabilities

  • →Provision multiple GPU instances for distributed training
  • →Monitor instances and replace preempted spot instances
  • →Resume training from checkpoints after preemption
  • →Analyze spending and compute cost-per-GPU-hour
  • →Destroy GPU clusters to stop billing
  • →Search for GPU offers based on specified criteria

How it works

The skill provisions multiple GPU instances, monitors them for preemption, replaces lost instances, resumes training from checkpoints, and analyzes GPU spending to optimize costs.

Inputs & outputs

You give it
Number of nodes, GPU name, minimum VRAM, Docker image, checkpoint directory
You get back
Multi-node GPU cluster, automatic spot interruption recovery, cost analysis report, and cluster teardown

When to use vastai-core-workflow-b

  • →Provisioning clusters for distributed training
  • →Implementing checkpoint-based resume for spot instances
  • →Analyzing and reducing per-job GPU spend

About this skill

Vast.ai Serverless Endpoint Rollout

Overview

Replace ad hoc multi-instance orchestration with the provider Serverless control plane. Prove a template independently, establish worker and queue bounds from load evidence, then let the workergroup perform a graceful rolling update.

Prerequisites

  • Latency, error-rate, queue-time, concurrency, and cost objectives
  • Immutable model/template candidate and a separate canary endpoint
  • Initial, minimum, maximum, cold-worker, and inactivity policy

Instructions

Step 1: Prove the candidate template

Launch the new model or environment on a non-production endpoint and verify load, readiness, response schema, and representative outputs.

Step 2: Define scaling bounds

Set min_load, min_workers, max_workers, cold_workers, inactivity_timeout, target_queue_time, and max_queue_time from explicit SLO and budget assumptions.

Step 3: Exercise convergence

During initial rollout, drive representative load up to roughly twice expected capacity and back down three times so the engine can learn GPU cost/performance.

Step 4: Establish the pre-update baseline

Record endpoint latency, queue time, error rate, active/inactive workers, model identity, and spend before changing production.

Step 5: Trigger the rolling update

Save the new template, update the workergroup reference, and monitor inactive workers updating first while active workers drain in-flight requests.

Step 6: Accept or roll back

Verify every worker is on the candidate and compare SLOs. If it fails, point the workergroup back to the last verified template and observe the reverse rollout.

Authentication

Use a scoped key with the documented misc Serverless permissions and no billing-write authority. Keep model registry credentials in approved environment variables, separate from the Vast.ai key.

Tool Discipline

Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.

Output

  • Canary and production template identities
  • Scaling policy plus load-test and rollout timeline
  • SLO comparison, worker convergence, and rollback decision

Return endpoint/workergroup IDs, old and new template identities, scaling bounds, load profile, SLO delta, and final rollout state.

Examples

A vLLM endpoint validates a new model on a canary, applies bounded queue targets, then updates its workergroup; active requests drain while new requests move to updated workers, with the old template retained for rollback.

Error Handling

FailureResponse
Canary cannot load the modelDo not update production; fix image, model, or environment configuration.
Queue time breaches during rolloutPause acceptance, increase safe capacity within budget, or roll back the template.
Workers do not convergeInspect workergroup logs and configuration; do not claim zero-downtime completion.
New output contract regressesRoll back to the last verified template and preserve comparison evidence.

Resources

When not to use it

  • →When the user is not using Vast.ai for GPU instances
  • →When the user is not performing distributed training
  • →When the user does not need spot instance recovery

Prerequisites

Completed vastai-core-workflow-aUnderstanding of distributed training (PyTorch DDP, DeepSpeed)Checkpoint-based training pipeline

Limitations

  • →Insufficient offers for a cluster may occur if not enough matching GPUs are available.
  • →Node communication failure can happen if instances are not from the same datacenter.

How it compares

This skill automates the orchestration of multiple GPU instances and provides specific mechanisms for spot instance recovery and cost analysis, which is more specialized than manually managing individual instances.

Compared to similar skills

vastai-core-workflow-b side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
vastai-core-workflow-b (this skill)12moReviewAdvanced
machine-learning-ops-ml-pipeline45moNo flagsAdvanced
uv37moReviewBeginner
mflux-dev-env26moNo flagsBeginner

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore →

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

Search skills

Search the agent skills registry