VA

vastai-core-workflow-a

Manages the lifecycle of Vast.ai GPU instances, from marketplace search to instance teardown.

Install

mkdir -p .claude/skills/vastai-core-workflow-a && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8352" && unzip -o skill.zip -d .claude/skills/vastai-core-workflow-a && rm skill.zip

Installs to .claude/skills/vastai-core-workflow-a

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Execute Vast.ai primary workflow: GPU instance provisioning and job
67 charsno explicit “when” trigger
Intermediate

Key capabilities

  • →Search GPU marketplace with performance filters
  • →Provision cloud GPU instances
  • →Transfer training data to remote instances
  • →Execute training or inference jobs
  • →Collect artifacts and destroy instances

How it works

The skill automates the lifecycle of a GPU instance by searching for offers, provisioning the instance, running remote commands via SSH, and cleaning up resources to stop billing.

Inputs & outputs

You give it
GPU requirements and training job parameters
You get back
Model checkpoints and logs

When to use vastai-core-workflow-a

  • →Searching for GPU instances via command line
  • →Provisioning cloud GPUs for training jobs
  • →Running inference tasks on remote GPU instances
  • →Automating instance teardown after job completion

About this skill

Checkpointed Vast.ai Training Run

Overview

Treat a training run as a recoverable state machine, not an SSH session. Bind code and image identity, select an offer through policy, persist checkpoints outside the disposable root disk, export final evidence, and destroy.

Prerequisites

  • Immutable image and code revision with deterministic training command
  • GPU, VRAM, disk, reliability, geography, and price policy
  • Checkpoint destination, resume test, runtime deadline, and cleanup owner

Instructions

Step 1: Freeze the run manifest

Record code revision, image digest, dataset version, command, seed, expected checkpoint cadence, budget, and output destination.

Step 2: Select and create

Search only verified rentable offers that meet the manifest, record price components, and create one labeled instance. Persist new_contract immediately.

Step 3: Reach readiness safely

Poll structured instance state with a deadline and terminal branches. Confirm image identity, disk headroom, GPU model, and CUDA visibility.

Step 4: Run with external checkpoints

Start the workload so checkpoints are uploaded or copied to durable storage at the declared cadence. A local checkpoint alone is not recovery evidence.

Step 5: Verify completion or resume

Validate artifact checksums and run metadata. For an interruption, provision a replacement from policy and prove resume from the last durable checkpoint.

Step 6: Close the cost boundary

Copy final logs and checksums, destroy the instance, confirm removal, and reconcile actual spend against the manifest.

Authentication

Use a scoped control-plane key for search and instance operations. Give the workload only the storage credential needed for its checkpoint prefix, with no Vast.ai billing or team authority.

Tool Discipline

Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.

Output

  • Frozen run and offer-selection manifest
  • State, checkpoint, resume, artifact, and spend evidence
  • Confirmed instance destruction and discrepancy report

Return revision, image digest, offer and instance IDs, checkpoint URI/checksum, terminal result, actual spend, and cleanup confirmation.

Examples

A fine-tuning job checkpoints every ten minutes to a run-specific object prefix; after a simulated interruption, a replacement instance resumes from the last checksum and the original contract is destroyed.

Error Handling

FailureResponse
No compliant offer existsPause the run and report the binding constraint; do not silently weaken reliability or price policy.
Checkpoint upload failsStop training before the recovery window is exceeded and repair storage access.
Host goes offlineUse the last external checkpoint on a different host and preserve the affected instance ID for support.
Artifact checksum failsDo not mark the run complete; retain evidence and rerun from the last verified checkpoint.

Resources

Prerequisites

vastai-install-auth setupDocker image published to registrySSH key uploaded to Vast.aiTraining data accessible

Limitations

  • →Instance may be preempted if using spot instances
  • →Large Docker images can cause loading delays

How it compares

This workflow automates the entire lifecycle from search to teardown, preventing manual errors that lead to unnecessary billing.

Compared to similar skills

vastai-core-workflow-a side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
vastai-core-workflow-a (this skill)02moReviewIntermediate
modal59moReviewIntermediate
machine-learning-ops-ml-pipeline45moNo flagsAdvanced
hugging-face-jobs18moCautionAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore →

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

modal

davila7

Run Python code in the cloud with serverless containers, GPUs, and autoscaling. Use when deploying ML models, running batch processing jobs, scheduling compute-intensive tasks, or serving APIs that require GPU acceleration or dynamic scaling.

587

machine-learning-ops-ml-pipeline

sickn33

Design and implement a complete ML pipeline for: $ARGUMENTS

436

hugging-face-jobs

patchy631

This skill should be used when users want to run any workload on Hugging Face Jobs infrastructure. Covers UV scripts, Docker-based jobs, hardware selection, cost estimation, authentication with tokens, secrets management, timeout configuration, and result persistence. Designed for general-purpose compute workloads including data processing, inference, experiments, batch jobs, and any Python-based tasks. Should be invoked for tasks involving cloud compute, GPU workloads, or when users mention running jobs on Hugging Face infrastructure without local setup.

14

senior-computer-vision

davila7

World-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems. Expertise in PyTorch, OpenCV, YOLO, SAM, diffusion models, and vision transformers. Includes 3D vision, video analysis, real-time processing, and production deployment. Use when building vision AI systems, implementing object detection, training custom vision models, or optimizing inference pipelines.

1256

senior-prompt-engineer

davila7

World-class prompt engineering skill for LLM optimization, prompt patterns, structured outputs, and AI product development. Expertise in Claude, GPT-4, prompt design patterns, few-shot learning, chain-of-thought, and AI evaluation. Includes RAG optimization, agent design, and LLM system architecture. Use when building AI products, optimizing LLM performance, designing agentic systems, or implementing advanced prompting techniques.

743

senior-ml-engineer

davila7

World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. Expertise in PyTorch, TensorFlow, model deployment, feature stores, model monitoring, and ML infrastructure. Includes LLM integration, fine-tuning, RAG systems, and agentic AI. Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs into production systems.

634

Search skills

Search the agent skills registry