VA

vastai-migration-deep-dive

Facilitates the migration of GPU workloads to Vast.ai from other cloud providers using structured patterns.

Install

mkdir -p .claude/skills/vastai-migration-deep-dive && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4104" && unzip -o skill.zip -d .claude/skills/vastai-migration-deep-dive && rm skill.zip

Installs to .claude/skills/vastai-migration-deep-dive

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Migrate GPU workloads to or from Vast.ai, or between GPU providers.
67 charsno explicit “when” trigger
Advanced

Key capabilities

  • →Compare GPU costs between providers
  • →Adapt Dockerfiles for Vast.ai environment
  • →Pass cloud storage credentials via environment variables
  • →Validate migration with automated checks
  • →Execute rollback procedures

How it works

It provides a framework to analyze cost savings and adapt existing Docker-based workloads to the Vast.ai environment. It replaces cloud-native IAM authentication with environment variable-based credential passing.

Inputs & outputs

You give it
Existing cloud GPU workload configuration
You get back
Adapted configuration and validation report for Vast.ai

When to use vastai-migration-deep-dive

  • →Migrating workloads to Vast.ai
  • →Switching GPU providers
  • →Re-platforming ML infrastructure

About this skill

Evidence-Gated Migration to Vast.ai

Overview

Separate portability from cutover. First identify source-provider dependencies, then prove immutable image, storage, networking, secrets, GPU, and output behavior on a disposable Vast.ai canary before moving production work.

Prerequisites

  • Source inventory covering images, accelerators, storage, network, identity, schedules, and cost
  • Acceptance thresholds for correctness, throughput, latency, recovery, and total spend
  • Versioned data/checkpoint transfer, dual-run or drain plan, and rollback owner

Instructions

Step 1: Freeze source truth

Record source release, image digest, GPU profile, command, secrets interfaces, ports, persistent data, checkpoints, SLOs, and representative outputs.

Step 2: Map Vast.ai equivalents

Choose offer or Serverless profiles, template, disk/volume/cloud-copy route, scoped keys, SSH/network mode, and lifecycle semantics.

Step 3: Prove artifact parity

Run the same image and input sample on one disposable Vast.ai target; verify CUDA, dependencies, output schema, checksums, and external checkpoint recovery.

Step 4: Compare production characteristics

Measure startup, throughput, latency, reliability, bandwidth, storage, interruption recovery, and cost per accepted unit.

Step 5: Cut over reversibly

Quiesce or dual-run according to data semantics, move only verified state, switch a bounded slice, and monitor explicit acceptance gates.

Step 6: Accept or roll back

Promote only if every gate passes. Otherwise restore source routing, reconcile writes and checkpoints, and destroy rejected Vast.ai resources.

Authentication

Translate identity to scoped Vast.ai keys and separate storage/registry credentials. Do not export source-provider credentials into images or long-lived Vast.ai environment variables.

Tool Discipline

Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.

Output

  • Source-to-Vast dependency and control map
  • Canary parity, recovery, performance, and cost evidence
  • Cutover or rollback timeline with reconciled data and resource cleanup

Return source/target releases, immutable identities, data checkpoint, acceptance deltas, decision, rollback point, and destroyed resources.

Examples

A Runpod training job keeps its container contract, moves checkpoints to a versioned cloud prefix, proves resume on one Vast.ai canary, then shifts scheduled jobs while the source environment remains available for one rollback window.

Error Handling

FailureResponse
Source dependency has no target equivalentDesign and test an adapter before cutover.
Data or output checksums differStop migration and reconcile the semantic difference.
Target capacity violates policyDelay or approve a documented alternate; do not weaken constraints silently.
Rollback window closes earlyIssue NO-GO until source restoration remains provable.

Resources

When not to use it

  • →Migrating workloads that require specific VPC or private subnet networking
  • →Migrating workloads that rely on native cloud IAM roles

Prerequisites

Existing GPU workload with Docker imageUnderstanding of current GPU costs and utilizationCheckpoint-based training pipeline

Limitations

  • →Vast.ai instances do not support native cloud provider IAM roles
  • →No SLA for uptime compared to hyperscaler providers

How it compares

It provides a structured migration path that accounts for the lack of native cloud-provider IAM roles on Vast.ai, unlike a direct lift-and-shift approach.

Compared to similar skills

vastai-migration-deep-dive side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
vastai-migration-deep-dive (this skill)12moReviewAdvanced
bazel-build-optimization144moNo flagsAdvanced
linux-production-shell-scripts78moReviewIntermediate
machine-learning-ops-ml-pipeline45moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore →

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

Search skills

Search the agent skills registry