vastai-migration-deep-dive
Facilitates the migration of GPU workloads to Vast.ai from other cloud providers using structured patterns.
Install
mkdir -p .claude/skills/vastai-migration-deep-dive && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4104" && unzip -o skill.zip -d .claude/skills/vastai-migration-deep-dive && rm skill.zipInstalls to .claude/skills/vastai-migration-deep-dive
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Migrate GPU workloads to or from Vast.ai, or between GPU providers.Key capabilities
- →Compare GPU costs between providers
- →Adapt Dockerfiles for Vast.ai environment
- →Pass cloud storage credentials via environment variables
- →Validate migration with automated checks
- →Execute rollback procedures
How it works
It provides a framework to analyze cost savings and adapt existing Docker-based workloads to the Vast.ai environment. It replaces cloud-native IAM authentication with environment variable-based credential passing.
Inputs & outputs
When to use vastai-migration-deep-dive
- →Migrating workloads to Vast.ai
- →Switching GPU providers
- →Re-platforming ML infrastructure
About this skill
Evidence-Gated Migration to Vast.ai
Overview
Separate portability from cutover. First identify source-provider dependencies, then prove immutable image, storage, networking, secrets, GPU, and output behavior on a disposable Vast.ai canary before moving production work.
Prerequisites
- Source inventory covering images, accelerators, storage, network, identity, schedules, and cost
- Acceptance thresholds for correctness, throughput, latency, recovery, and total spend
- Versioned data/checkpoint transfer, dual-run or drain plan, and rollback owner
Instructions
Step 1: Freeze source truth
Record source release, image digest, GPU profile, command, secrets interfaces, ports, persistent data, checkpoints, SLOs, and representative outputs.
Step 2: Map Vast.ai equivalents
Choose offer or Serverless profiles, template, disk/volume/cloud-copy route, scoped keys, SSH/network mode, and lifecycle semantics.
Step 3: Prove artifact parity
Run the same image and input sample on one disposable Vast.ai target; verify CUDA, dependencies, output schema, checksums, and external checkpoint recovery.
Step 4: Compare production characteristics
Measure startup, throughput, latency, reliability, bandwidth, storage, interruption recovery, and cost per accepted unit.
Step 5: Cut over reversibly
Quiesce or dual-run according to data semantics, move only verified state, switch a bounded slice, and monitor explicit acceptance gates.
Step 6: Accept or roll back
Promote only if every gate passes. Otherwise restore source routing, reconcile writes and checkpoints, and destroy rejected Vast.ai resources.
Authentication
Translate identity to scoped Vast.ai keys and separate storage/registry credentials. Do not export source-provider credentials into images or long-lived Vast.ai environment variables.
Tool Discipline
Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.
Output
- Source-to-Vast dependency and control map
- Canary parity, recovery, performance, and cost evidence
- Cutover or rollback timeline with reconciled data and resource cleanup
Return source/target releases, immutable identities, data checkpoint, acceptance deltas, decision, rollback point, and destroyed resources.
Examples
A Runpod training job keeps its container contract, moves checkpoints to a versioned cloud prefix, proves resume on one Vast.ai canary, then shifts scheduled jobs while the source environment remains available for one rollback window.
Error Handling
| Failure | Response |
|---|---|
| Source dependency has no target equivalent | Design and test an adapter before cutover. |
| Data or output checksums differ | Stop migration and reconcile the semantic difference. |
| Target capacity violates policy | Delay or approve a documented alternate; do not weaken constraints silently. |
| Rollback window closes early | Issue NO-GO until source restoration remains provable. |
Resources
When not to use it
- →Migrating workloads that require specific VPC or private subnet networking
- →Migrating workloads that rely on native cloud IAM roles
Prerequisites
Limitations
- →Vast.ai instances do not support native cloud provider IAM roles
- →No SLA for uptime compared to hyperscaler providers
How it compares
It provides a structured migration path that accounts for the lack of native cloud-provider IAM roles on Vast.ai, unlike a direct lift-and-shift approach.
Compared to similar skills
vastai-migration-deep-dive side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| vastai-migration-deep-dive (this skill) | 1 | 2mo | Review | Advanced |
| bazel-build-optimization | 14 | 4mo | No flags | Advanced |
| linux-production-shell-scripts | 7 | 8mo | Review | Intermediate |
| machine-learning-ops-ml-pipeline | 4 | 5mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
bazel-build-optimization
wshobson
Optimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
linux-production-shell-scripts
davila7
This skill should be used when the user asks to "create bash scripts", "automate Linux tasks", "monitor system resources", "backup files", "manage users", or "write production shell scripts". It provides ready-to-use shell script templates for system administration.
machine-learning-ops-ml-pipeline
sickn33
Design and implement a complete ML pipeline for: $ARGUMENTS
windows-builder
hashicorp
Build Windows images with Packer using WinRM communicator and PowerShell provisioners. Use when creating Windows AMIs, Azure images, or VMware templates.
cicd-automation-workflow-automate
sickn33
You are a workflow automation expert specializing in creating efficient CI/CD pipelines, GitHub Actions workflows, and automated development processes. Design automation that reduces manual work, improves consistency, and accelerates delivery while maintaining quality and security.
create-worktree-skill
disler
Use when the user explicitly asks for a SKILL to create a worktree. If the user does not mention "skill" or explicitly request skill invocation, do NOT trigger this. Only use when user says things like "use a skill to create a worktree" or "invoke the worktree skill". Creates isolated git worktrees with parallel-running configuration.