VA

vastai-migration-deep-dive

Facilitates the migration of GPU workloads to Vast.ai from other cloud providers using structured patterns.

Install

mkdir -p .claude/skills/vastai-migration-deep-dive && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4104" && unzip -o skill.zip -d .claude/skills/vastai-migration-deep-dive && rm skill.zip

Installs to .claude/skills/vastai-migration-deep-dive

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Migrate GPU workloads to or from Vast.ai, or between GPU providers.
67 charsno explicit “when” trigger
Advanced

Key capabilities

  • Compare GPU costs between providers
  • Adapt Dockerfiles for Vast.ai environment
  • Pass cloud storage credentials via environment variables
  • Validate migration with automated checks
  • Execute rollback procedures

How it works

It provides a framework to analyze cost savings and adapt existing Docker-based workloads to the Vast.ai environment. It replaces cloud-native IAM authentication with environment variable-based credential passing.

Inputs & outputs

You give it
Existing cloud GPU workload configuration
You get back
Adapted configuration and validation report for Vast.ai

When to use vastai-migration-deep-dive

  • Migrating workloads to Vast.ai
  • Switching GPU providers
  • Re-platforming ML infrastructure

About this skill

Vast.ai Migration Deep Dive

Current State

!vastai --version 2>/dev/null || echo 'vastai CLI not installed' !pip show vastai 2>/dev/null | grep Version || echo 'N/A'

Overview

Migrate GPU workloads to Vast.ai from hyperscaler providers (AWS, GCP, Azure) or other GPU clouds (Lambda, RunPod, CoreWeave). Also covers migrating between GPU types on Vast.ai and the reverse migration away from Vast.ai.

Prerequisites

  • Existing GPU workload with Docker image
  • Understanding of current GPU costs and utilization
  • Checkpoint-based training pipeline (for training migrations)

Instructions

Step 1: Cost Comparison Analysis

# Compare your current GPU costs against Vast.ai marketplace prices
PROVIDER_COSTS = {
    "aws_p4d.24xlarge":      {"gpu": "A100 40GB", "gpus": 8, "hourly": 32.77},
    "aws_p3.2xlarge":        {"gpu": "V100 16GB", "gpus": 1, "hourly": 3.06},
    "gcp_a2-highgpu-1g":     {"gpu": "A100 40GB", "gpus": 1, "hourly": 3.67},
    "azure_NC24ads_A100_v4": {"gpu": "A100 80GB", "gpus": 1, "hourly": 3.67},
    "lambda_1xA100":         {"gpu": "A100",      "gpus": 1, "hourly": 1.25},
}

VASTAI_TYPICAL = {
    "RTX_4090":  0.20,
    "A100":      1.50,
    "H100_SXM":  3.00,
}

def savings_analysis(current_provider, current_hourly, vastai_gpu, vastai_hourly):
    monthly_current = current_hourly * 730  # hours/month
    monthly_vastai = vastai_hourly * 730
    savings = monthly_current - monthly_vastai
    pct = (savings / monthly_current) * 100
    print(f"Current ({current_provider}): ${monthly_current:,.0f}/mo")
    print(f"Vast.ai ({vastai_gpu}): ${monthly_vastai:,.0f}/mo")
    print(f"Savings: ${savings:,.0f}/mo ({pct:.0f}%)")

savings_analysis("AWS p3.2xlarge", 3.06, "RTX_4090", 0.20)
# Output: Savings: $2,088/mo (93%)

Step 2: Docker Image Migration

# Most Docker images work unchanged on Vast.ai
# Key differences:
# - Vast.ai instances run as root
# - /workspace is the default working directory
# - SSH access (not IAM roles) for authentication

# Adapt your existing Dockerfile
cat << 'DOCKERFILE' > Dockerfile.vastai
FROM your-existing-image:latest

# Vast.ai instances use /workspace by default
WORKDIR /workspace

# Install any Vast.ai-specific tools
RUN pip install boto3  # for S3 checkpoint uploads

# Copy training code
COPY src/ /workspace/src/
COPY configs/ /workspace/configs/

CMD ["python", "src/train.py"]
DOCKERFILE

docker build -t ghcr.io/org/training:vastai -f Dockerfile.vastai .
docker push ghcr.io/org/training:vastai

Step 3: Adapt Cloud Storage Credentials

# On AWS/GCP: IAM roles provide automatic credentials
# On Vast.ai: Pass credentials explicitly via environment variables

# Create instance with env vars for cloud storage access
vastai create instance $OFFER_ID \
  --image ghcr.io/org/training:vastai \
  --disk 100 \
  --env "AWS_ACCESS_KEY_ID=AKIA... AWS_SECRET_ACCESS_KEY=... AWS_DEFAULT_REGION=us-east-1"

Step 4: Migration Validation

#!/bin/bash
set -euo pipefail
echo "Migration Validation Checklist"

# 1. Docker image runs on Vast.ai
vastai create instance $OFFER_ID --image ghcr.io/org/training:vastai --disk 50
# Wait for running...

# 2. GPU access works
ssh -p $PORT root@$HOST "nvidia-smi && python -c 'import torch; print(torch.cuda.is_available())'"

# 3. Cloud storage works
ssh -p $PORT root@$HOST "aws s3 ls s3://your-bucket/ | head -5"

# 4. Training runs and saves checkpoints
ssh -p $PORT root@$HOST "cd /workspace && python src/train.py --epochs 1 --checkpoint-dir /workspace/ckpt"

# 5. Checkpoints uploaded to cloud storage
ssh -p $PORT root@$HOST "aws s3 sync /workspace/ckpt/ s3://your-bucket/ckpt/"

# 6. Clean up
vastai destroy instance $INSTANCE_ID
echo "Migration validation complete"

Step 5: Rollback Plan

## Rollback Procedure
1. Stop all Vast.ai instances: `vastai show instances` → `vastai destroy instance ID`
2. Re-provision on original cloud provider
3. Resume training from cloud-stored checkpoint
4. Vast.ai Docker image remains available for future retry

Migration Comparison

FactorAWS/GCP/AzureVast.ai
PricingFixed, premiumVariable, 50-90% cheaper
GPU availabilityOn-demand guaranteedMarketplace (may sell out)
SLA99.9% uptimeNo SLA (spot instances)
IAM rolesNativeManual credential passing
NetworkingVPC, private subnetsPublic SSH only
StorageEBS/PD attachedLocal disk + cloud storage
SupportEnterprise supportCommunity/email

Output

  • Cost savings analysis comparing providers
  • Adapted Docker image for Vast.ai
  • Cloud credential migration pattern
  • Validation script for migration testing
  • Rollback procedure

Error Handling

ErrorCauseSolution
Docker image incompatibleRelies on IAM roles or cloud-specific APIsPass credentials via env vars
CUDA version mismatchDifferent CUDA on Vast.ai hostsFilter by cuda_max_good in search
Data transfer too slowLarge dataset over public internetStage data in cloud storage, download on instance
No matching offersSpecific GPU unavailableTry alternative GPU type or wait for availability

Resources

Next Steps

Review vastai-reference-architecture for best-practice project structure.

Examples

AWS to Vast.ai: Replace p3.2xlarge ($3.06/hr) with RTX 4090 ($0.20/hr) for a 93% cost reduction. Adapt the Dockerfile to pass AWS credentials via env vars for S3 checkpoint access.

Hybrid approach: Use Vast.ai for experimentation and hyperparameter search (cheap GPUs), then run final training on AWS for SLA guarantees.

When not to use it

  • Migrating workloads that require specific VPC or private subnet networking
  • Migrating workloads that rely on native cloud IAM roles

Prerequisites

Existing GPU workload with Docker imageUnderstanding of current GPU costs and utilizationCheckpoint-based training pipeline

Limitations

  • Vast.ai instances do not support native cloud provider IAM roles
  • No SLA for uptime compared to hyperscaler providers

How it compares

It provides a structured migration path that accounts for the lack of native cloud-provider IAM roles on Vast.ai, unlike a direct lift-and-shift approach.

Compared to similar skills

vastai-migration-deep-dive side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
vastai-migration-deep-dive (this skill)127dReviewAdvanced
bazel-build-optimization142moNo flagsAdvanced
linux-production-shell-scripts76moReviewIntermediate
machine-learning-ops-ml-pipeline44moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

Search skills

Search the agent skills registry