VA

vastai-deploy-integration

Streamlines deployment of GPU workloads to Vast.ai, covering Docker image optimization and automated provisioning.

Install

mkdir -p .claude/skills/vastai-deploy-integration && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/9307" && unzip -o skill.zip -d .claude/skills/vastai-deploy-integration && rm skill.zip

Installs to .claude/skills/vastai-deploy-integration

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Deploy ML training jobs and inference services on Vast.ai GPU cloud.
68 charsno explicit “when” trigger
Intermediate

Key capabilities

  • →Build optimized Docker images for GPU workloads
  • →Automate instance provisioning on Vast.ai
  • →Manage data transfer via SCP and rsync
  • →Perform post-deployment health checks
  • →Filter GPU offers by price and reliability

How it works

The deployment script queries the Vast.ai API for available GPU offers matching specified constraints, provisions an instance, and verifies the environment using post-deploy health checks.

Inputs & outputs

You give it
GPU requirements and Docker image URI
You get back
Provisioned GPU instance connection details

When to use vastai-deploy-integration

  • →Deploy GPU-accelerated training jobs
  • →Optimize Docker containers for Vast.ai
  • →Automate inference service deployment
  • →Configure cloud GPU resources

About this skill

Immutable Vast.ai Instance Deployment

Overview

Make the deployment reproducible outside the console. Bind release bytes to an immutable image and versioned template hash, launch a canary under policy, validate service and GPU health, then promote or destroy.

Prerequisites

  • Release commit, image digest, template hash, ports, environment names, and startup contract
  • Offer policy, health checks, artifact/checkpoint destinations, and maximum rollout time
  • Last-known-good image/template plus rollback and teardown owners

Instructions

Step 1: Freeze the deployment manifest

Record all non-secret template settings, image digest, on-start behavior, exposed ports, disk, label, and required GPU/host constraints.

Step 2: Validate outside production

Build and scan the image, run local contract tests, and launch a disposable canary from the exact template hash.

Step 3: Verify readiness

Bound state polling, resolve current endpoints, check process and GPU health, and execute a representative request or batch assertion.

Step 4: Promote deliberately

Create or update the production instance only after canary acceptance. Persist resource IDs and externalize important data before traffic or work begins.

Step 5: Observe the release

Track state, logs, request/job outcome, cost, disk, and checkpoint health through the acceptance window.

Step 6: Roll back or close

On regression, route work to the last-known-good template or recreate from it; copy evidence and destroy rejected or superseded instances.

Authentication

Give deployment automation only search, template, and instance permissions it needs. Keep registry, model, and storage secrets in approved runtime variables and out of template descriptions and logs.

Tool Discipline

Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.

Output

  • Versioned deployment manifest and immutable identities
  • Canary and production health/acceptance evidence
  • Promotion, rollback, and superseded-resource cleanup receipt

Return release, image/template identities, offer and instance IDs, health result, acceptance window, decision, and cleanup status.

Examples

A model API is launched from a pinned template hash and image digest, passes GPU and request canaries, then replaces the previous instance; the rejected candidate is destroyed and the old template remains recorded for rollback.

Error Handling

FailureResponse
Template resolves different bytesStop and pin a new immutable hash before provisioning.
Health passes but workload assertion failsReject the release and retain the last-known-good service.
Rollback data is only on local diskCopy it externally before destructive recovery if the incident permits.
Superseded instance remainsTreat it as a cost leak and complete or escalate teardown.

Resources

When not to use it

  • →Deploying images larger than 10GB without multi-stage builds
  • →Running without sufficient disk allocation

Prerequisites

Vast.ai CLIDocker image in a registry

Limitations

  • →Docker pull timeouts for large images
  • →SSH timeouts if instance loading is slow
  • →CUDA version mismatches

How it compares

This automated workflow replaces manual instance searching and configuration with a scriptable, repeatable deployment process.

Compared to similar skills

vastai-deploy-integration side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
vastai-deploy-integration (this skill)02moReviewIntermediate
modal59moReviewIntermediate
machine-learning-ops-ml-pipeline45moNo flagsAdvanced
hugging-face-jobs18moCautionAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore →

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

modal

davila7

Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.

587

machine-learning-ops-ml-pipeline

sickn33

Design and implement a complete ML pipeline for: $ARGUMENTS

436

hugging-face-jobs

patchy631

This skill should be used when users want to run any workload on Hugging Face Jobs infrastructure. Covers UV scripts, Docker-based jobs, hardware selection, cost estimation, authentication with tokens, secrets management, timeout configuration, and result persistence. Designed for general-purpose compute workloads including data processing, inference, experiments, batch jobs, and any Python-based tasks. Should be invoked for tasks involving cloud compute, GPU workloads, or when users mention running jobs on Hugging Face infrastructure without local setup.

14

vastai-core-workflow-a

jeremylongshore

Execute Vast.ai primary workflow: Core Workflow A. Use when implementing primary use case, building main features, or core integration tasks. Trigger with phrases like "vastai main workflow", "primary task with vastai".

01

hosted-agents-v2-py

wegonbeok45

Build hosted agents using Azure AI Projects SDK with ImageBasedHostedAgentDefinition. Use when creating container-based agents in Azure AI Foundry.

00

customerio-deploy-pipeline

jeremylongshore

Deploy Customer.io integrations to production. Use when deploying to cloud platforms, setting up production infrastructure, or automating deployments. Trigger with phrases like "deploy customer.io", "customer.io production", "customer.io cloud run", "customer.io kubernetes".

12

Search skills

Search the agent skills registry