vastai-deploy-integration
Streamlines deployment of GPU workloads to Vast.ai, covering Docker image optimization and automated provisioning.
Install
mkdir -p .claude/skills/vastai-deploy-integration && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/9307" && unzip -o skill.zip -d .claude/skills/vastai-deploy-integration && rm skill.zipInstalls to .claude/skills/vastai-deploy-integration
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Deploy ML training jobs and inference services on Vast.ai GPU cloud.Key capabilities
- →Build optimized Docker images for GPU workloads
- →Automate instance provisioning on Vast.ai
- →Manage data transfer via SCP and rsync
- →Perform post-deployment health checks
- →Filter GPU offers by price and reliability
How it works
The deployment script queries the Vast.ai API for available GPU offers matching specified constraints, provisions an instance, and verifies the environment using post-deploy health checks.
Inputs & outputs
When to use vastai-deploy-integration
- →Deploy GPU-accelerated training jobs
- →Optimize Docker containers for Vast.ai
- →Automate inference service deployment
- →Configure cloud GPU resources
About this skill
Immutable Vast.ai Instance Deployment
Overview
Make the deployment reproducible outside the console. Bind release bytes to an immutable image and versioned template hash, launch a canary under policy, validate service and GPU health, then promote or destroy.
Prerequisites
- Release commit, image digest, template hash, ports, environment names, and startup contract
- Offer policy, health checks, artifact/checkpoint destinations, and maximum rollout time
- Last-known-good image/template plus rollback and teardown owners
Instructions
Step 1: Freeze the deployment manifest
Record all non-secret template settings, image digest, on-start behavior, exposed ports, disk, label, and required GPU/host constraints.
Step 2: Validate outside production
Build and scan the image, run local contract tests, and launch a disposable canary from the exact template hash.
Step 3: Verify readiness
Bound state polling, resolve current endpoints, check process and GPU health, and execute a representative request or batch assertion.
Step 4: Promote deliberately
Create or update the production instance only after canary acceptance. Persist resource IDs and externalize important data before traffic or work begins.
Step 5: Observe the release
Track state, logs, request/job outcome, cost, disk, and checkpoint health through the acceptance window.
Step 6: Roll back or close
On regression, route work to the last-known-good template or recreate from it; copy evidence and destroy rejected or superseded instances.
Authentication
Give deployment automation only search, template, and instance permissions it needs. Keep registry, model, and storage secrets in approved runtime variables and out of template descriptions and logs.
Tool Discipline
Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.
Output
- Versioned deployment manifest and immutable identities
- Canary and production health/acceptance evidence
- Promotion, rollback, and superseded-resource cleanup receipt
Return release, image/template identities, offer and instance IDs, health result, acceptance window, decision, and cleanup status.
Examples
A model API is launched from a pinned template hash and image digest, passes GPU and request canaries, then replaces the previous instance; the rejected candidate is destroyed and the old template remains recorded for rollback.
Error Handling
| Failure | Response |
|---|---|
| Template resolves different bytes | Stop and pin a new immutable hash before provisioning. |
| Health passes but workload assertion fails | Reject the release and retain the last-known-good service. |
| Rollback data is only on local disk | Copy it externally before destructive recovery if the incident permits. |
| Superseded instance remains | Treat it as a cost leak and complete or escalate teardown. |
Resources
When not to use it
- →Deploying images larger than 10GB without multi-stage builds
- →Running without sufficient disk allocation
Prerequisites
Limitations
- →Docker pull timeouts for large images
- →SSH timeouts if instance loading is slow
- →CUDA version mismatches
How it compares
This automated workflow replaces manual instance searching and configuration with a scriptable, repeatable deployment process.
Compared to similar skills
vastai-deploy-integration side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| vastai-deploy-integration (this skill) | 0 | 2mo | Review | Intermediate |
| modal | 5 | 9mo | Review | Intermediate |
| machine-learning-ops-ml-pipeline | 4 | 5mo | No flags | Advanced |
| hugging-face-jobs | 1 | 8mo | Caution | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
modal
davila7
Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.
machine-learning-ops-ml-pipeline
sickn33
Design and implement a complete ML pipeline for: $ARGUMENTS
hugging-face-jobs
patchy631
This skill should be used when users want to run any workload on Hugging Face Jobs infrastructure. Covers UV scripts, Docker-based jobs, hardware selection, cost estimation, authentication with tokens, secrets management, timeout configuration, and result persistence. Designed for general-purpose compute workloads including data processing, inference, experiments, batch jobs, and any Python-based tasks. Should be invoked for tasks involving cloud compute, GPU workloads, or when users mention running jobs on Hugging Face infrastructure without local setup.
vastai-core-workflow-a
jeremylongshore
Execute Vast.ai primary workflow: Core Workflow A. Use when implementing primary use case, building main features, or core integration tasks. Trigger with phrases like "vastai main workflow", "primary task with vastai".
hosted-agents-v2-py
wegonbeok45
Build hosted agents using Azure AI Projects SDK with ImageBasedHostedAgentDefinition. Use when creating container-based agents in Azure AI Foundry.
customerio-deploy-pipeline
jeremylongshore
Deploy Customer.io integrations to production. Use when deploying to cloud platforms, setting up production infrastructure, or automating deployments. Trigger with phrases like "deploy customer.io", "customer.io production", "customer.io cloud run", "customer.io kubernetes".