vastai-core-workflow-a
Manages the lifecycle of Vast.ai GPU instances, from marketplace search to instance teardown.
Install
mkdir -p .claude/skills/vastai-core-workflow-a && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8352" && unzip -o skill.zip -d .claude/skills/vastai-core-workflow-a && rm skill.zipInstalls to .claude/skills/vastai-core-workflow-a
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Execute Vast.ai primary workflow: GPU instance provisioning and jobKey capabilities
- →Search GPU marketplace with performance filters
- →Provision cloud GPU instances
- →Transfer training data to remote instances
- →Execute training or inference jobs
- →Collect artifacts and destroy instances
How it works
The skill automates the lifecycle of a GPU instance by searching for offers, provisioning the instance, running remote commands via SSH, and cleaning up resources to stop billing.
Inputs & outputs
When to use vastai-core-workflow-a
- →Searching for GPU instances via command line
- →Provisioning cloud GPUs for training jobs
- →Running inference tasks on remote GPU instances
- →Automating instance teardown after job completion
About this skill
Checkpointed Vast.ai Training Run
Overview
Treat a training run as a recoverable state machine, not an SSH session. Bind code and image identity, select an offer through policy, persist checkpoints outside the disposable root disk, export final evidence, and destroy.
Prerequisites
- Immutable image and code revision with deterministic training command
- GPU, VRAM, disk, reliability, geography, and price policy
- Checkpoint destination, resume test, runtime deadline, and cleanup owner
Instructions
Step 1: Freeze the run manifest
Record code revision, image digest, dataset version, command, seed, expected checkpoint cadence, budget, and output destination.
Step 2: Select and create
Search only verified rentable offers that meet the manifest, record price components, and create one labeled instance. Persist new_contract immediately.
Step 3: Reach readiness safely
Poll structured instance state with a deadline and terminal branches. Confirm image identity, disk headroom, GPU model, and CUDA visibility.
Step 4: Run with external checkpoints
Start the workload so checkpoints are uploaded or copied to durable storage at the declared cadence. A local checkpoint alone is not recovery evidence.
Step 5: Verify completion or resume
Validate artifact checksums and run metadata. For an interruption, provision a replacement from policy and prove resume from the last durable checkpoint.
Step 6: Close the cost boundary
Copy final logs and checksums, destroy the instance, confirm removal, and reconcile actual spend against the manifest.
Authentication
Use a scoped control-plane key for search and instance operations. Give the workload only the storage credential needed for its checkpoint prefix, with no Vast.ai billing or team authority.
Tool Discipline
Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.
Output
- Frozen run and offer-selection manifest
- State, checkpoint, resume, artifact, and spend evidence
- Confirmed instance destruction and discrepancy report
Return revision, image digest, offer and instance IDs, checkpoint URI/checksum, terminal result, actual spend, and cleanup confirmation.
Examples
A fine-tuning job checkpoints every ten minutes to a run-specific object prefix; after a simulated interruption, a replacement instance resumes from the last checksum and the original contract is destroyed.
Error Handling
| Failure | Response |
|---|---|
| No compliant offer exists | Pause the run and report the binding constraint; do not silently weaken reliability or price policy. |
| Checkpoint upload fails | Stop training before the recovery window is exceeded and repair storage access. |
| Host goes offline | Use the last external checkpoint on a different host and preserve the affected instance ID for support. |
| Artifact checksum fails | Do not mark the run complete; retain evidence and rerun from the last verified checkpoint. |
Resources
Prerequisites
Limitations
- →Instance may be preempted if using spot instances
- →Large Docker images can cause loading delays
How it compares
This workflow automates the entire lifecycle from search to teardown, preventing manual errors that lead to unnecessary billing.
Compared to similar skills
vastai-core-workflow-a side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| vastai-core-workflow-a (this skill) | 0 | 2mo | Review | Intermediate |
| modal | 5 | 9mo | Review | Intermediate |
| machine-learning-ops-ml-pipeline | 4 | 5mo | No flags | Advanced |
| hugging-face-jobs | 1 | 8mo | Caution | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
modal
davila7
Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.
machine-learning-ops-ml-pipeline
sickn33
Design and implement a complete ML pipeline for: $ARGUMENTS
hugging-face-jobs
patchy631
This skill should be used when users want to run any workload on Hugging Face Jobs infrastructure. Covers UV scripts, Docker-based jobs, hardware selection, cost estimation, authentication with tokens, secrets management, timeout configuration, and result persistence. Designed for general-purpose compute workloads including data processing, inference, experiments, batch jobs, and any Python-based tasks. Should be invoked for tasks involving cloud compute, GPU workloads, or when users mention running jobs on Hugging Face infrastructure without local setup.
senior-computer-vision
davila7
World-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems. Expertise in PyTorch, OpenCV, YOLO, SAM, diffusion models, and vision transformers. Includes 3D vision, video analysis, real-time processing, and production deployment. Use when building vision AI systems, implementing object detection, training custom vision models, or optimizing inference pipelines.
senior-prompt-engineer
davila7
World-class prompt engineering skill for LLM optimization, prompt patterns, structured outputs, and AI product development. Expertise in Claude, GPT-4, prompt design patterns, few-shot learning, chain-of-thought, and AI evaluation. Includes RAG optimization, agent design, and LLM system architecture. Use when building AI products, optimizing LLM performance, designing agentic systems, or implementing advanced prompting techniques.
senior-ml-engineer
davila7
World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. Expertise in PyTorch, TensorFlow, model deployment, feature stores, model monitoring, and ML infrastructure. Includes LLM integration, fine-tuning, RAG systems, and agentic AI. Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs into production systems.