vastai-core-workflow-b
Orchestrates multi-GPU instances and automates spot interruption recovery on Vast.ai.
Install
mkdir -p .claude/skills/vastai-core-workflow-b && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/3820" && unzip -o skill.zip -d .claude/skills/vastai-core-workflow-b && rm skill.zipInstalls to .claude/skills/vastai-core-workflow-b
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Execute Vast.ai secondary workflow: multi-instance orchestration, spotKey capabilities
- →Provision multiple GPU instances for distributed training
- →Monitor instances and replace preempted spot instances
- →Resume training from checkpoints after preemption
- →Analyze spending and compute cost-per-GPU-hour
- →Destroy GPU clusters to stop billing
- →Search for GPU offers based on specified criteria
How it works
The skill provisions multiple GPU instances, monitors them for preemption, replaces lost instances, resumes training from checkpoints, and analyzes GPU spending to optimize costs.
Inputs & outputs
When to use vastai-core-workflow-b
- →Provisioning clusters for distributed training
- →Implementing checkpoint-based resume for spot instances
- →Analyzing and reducing per-job GPU spend
About this skill
Vast.ai Serverless Endpoint Rollout
Overview
Replace ad hoc multi-instance orchestration with the provider Serverless control plane. Prove a template independently, establish worker and queue bounds from load evidence, then let the workergroup perform a graceful rolling update.
Prerequisites
- Latency, error-rate, queue-time, concurrency, and cost objectives
- Immutable model/template candidate and a separate canary endpoint
- Initial, minimum, maximum, cold-worker, and inactivity policy
Instructions
Step 1: Prove the candidate template
Launch the new model or environment on a non-production endpoint and verify load, readiness, response schema, and representative outputs.
Step 2: Define scaling bounds
Set min_load, min_workers, max_workers, cold_workers, inactivity_timeout, target_queue_time, and max_queue_time from explicit SLO and budget assumptions.
Step 3: Exercise convergence
During initial rollout, drive representative load up to roughly twice expected capacity and back down three times so the engine can learn GPU cost/performance.
Step 4: Establish the pre-update baseline
Record endpoint latency, queue time, error rate, active/inactive workers, model identity, and spend before changing production.
Step 5: Trigger the rolling update
Save the new template, update the workergroup reference, and monitor inactive workers updating first while active workers drain in-flight requests.
Step 6: Accept or roll back
Verify every worker is on the candidate and compare SLOs. If it fails, point the workergroup back to the last verified template and observe the reverse rollout.
Authentication
Use a scoped key with the documented misc Serverless permissions and no billing-write authority. Keep model registry credentials in approved environment variables, separate from the Vast.ai key.
Tool Discipline
Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.
Output
- Canary and production template identities
- Scaling policy plus load-test and rollout timeline
- SLO comparison, worker convergence, and rollback decision
Return endpoint/workergroup IDs, old and new template identities, scaling bounds, load profile, SLO delta, and final rollout state.
Examples
A vLLM endpoint validates a new model on a canary, applies bounded queue targets, then updates its workergroup; active requests drain while new requests move to updated workers, with the old template retained for rollback.
Error Handling
| Failure | Response |
|---|---|
| Canary cannot load the model | Do not update production; fix image, model, or environment configuration. |
| Queue time breaches during rollout | Pause acceptance, increase safe capacity within budget, or roll back the template. |
| Workers do not converge | Inspect workergroup logs and configuration; do not claim zero-downtime completion. |
| New output contract regresses | Roll back to the last verified template and preserve comparison evidence. |
Resources
When not to use it
- →When the user is not using Vast.ai for GPU instances
- →When the user is not performing distributed training
- →When the user does not need spot instance recovery
Prerequisites
Limitations
- →Insufficient offers for a cluster may occur if not enough matching GPUs are available.
- →Node communication failure can happen if instances are not from the same datacenter.
How it compares
This skill automates the orchestration of multiple GPU instances and provides specific mechanisms for spot instance recovery and cost analysis, which is more specialized than manually managing individual instances.
Compared to similar skills
vastai-core-workflow-b side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| vastai-core-workflow-b (this skill) | 1 | 2mo | Review | Advanced |
| machine-learning-ops-ml-pipeline | 4 | 5mo | No flags | Advanced |
| uv | 3 | 7mo | Review | Beginner |
| mflux-dev-env | 2 | 6mo | No flags | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
machine-learning-ops-ml-pipeline
sickn33
Design and implement a complete ML pipeline for: $ARGUMENTS
uv
mitsuhiko
Use `uv` instead of pip/python/venv. Run scripts with `uv run script.py`, add deps with `uv add`, use inline script metadata for standalone scripts.
mflux-dev-env
filipstrand
Set up and work in the mflux dev environment (arm64 expectation, uv, Makefile targets, lint/format/test).
hugging-face-jobs
patchy631
This skill should be used when users want to run any workload on Hugging Face Jobs infrastructure. Covers UV scripts, Docker-based jobs, hardware selection, cost estimation, authentication with tokens, secrets management, timeout configuration, and result persistence. Designed for general-purpose compute workloads including data processing, inference, experiments, batch jobs, and any Python-based tasks. Should be invoked for tasks involving cloud compute, GPU workloads, or when users mention running jobs on Hugging Face infrastructure without local setup.
mlops-initialization
fmind
Guide to initialize a new MLOps project with standard tools (uv, git, VS Code) and best practices.
mqtt-uns-feeder
clemensv
Use when adding MQTT/Unified-Namespace transport to a feeder in this repo — either alongside an existing Kafka feeder (transport split) or for a brand-new source where MQTT is the primary fit. Covers UNS topic-tree design, xRegistry MQTT messagegroup pattern, paho-mqtt v5 binary-mode CloudEvent emis