VA

vastai-common-errors

Diagnose Vast.ai errors using standard status codes and troubleshooting guides for GPU instance management.

Install

mkdir -p .claude/skills/vastai-common-errors && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8787" && unzip -o skill.zip -d .claude/skills/vastai-common-errors && rm skill.zip

Installs to .claude/skills/vastai-common-errors

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Diagnose and fix Vast.ai common errors and exceptions.
54 charsno explicit “when” trigger
Intermediate

Key capabilities

  • →Map HTTP status codes to fixes
  • →Diagnose instance status errors
  • →Troubleshoot SSH connection issues
  • →Verify GPU and CUDA availability
  • →Check platform connection status

How it works

The tool maps specific API responses and instance states to known causes and corrective commands.

Inputs & outputs

You give it
Error code or instance status
You get back
Actionable resolution steps

When to use vastai-common-errors

  • →Debug Vast.ai API authentication errors
  • →Troubleshoot instance creation and docker loading failures
  • →Resolve GPU rental configuration issues
  • →Check Vast.ai platform status for connection errors

About this skill

Vast.ai Failure Classifier

Overview

Diagnose from provider state and error evidence instead of retrying every failure. Separate identity and permission errors, marketplace scarcity, transient startup, terminal host states, billing stops, SSH configuration, and workload exits.

Prerequisites

  • Exact command, non-secret arguments, exit status, timestamp, and redacted response
  • Affected account context, instance ID, image identity, and expected state
  • Authority to inspect but not automatically destroy or fund resources

Instructions

Step 1: Capture structured evidence

Run the failing command with --raw where supported and record CLI version. For request diagnosis, use --explain or --curl only after ensuring generated output cannot expose the key.

Step 2: Classify control-plane failure

Treat 401 as credential failure, 403 as missing scoped permission, 429 as endpoint/identity rate limiting, and insufficient credit or spend-rate errors as billing policy—not host failure.

Step 3: Classify instance state

Loading may reflect an image pull; scheduling after a stop may wait indefinitely for the original GPU; exited is a workload/container failure; unknown or offline indicates missing host heartbeat.

Step 4: Check SSH and network facts

Wait for running, resolve the current SSH URL, verify the registered public key and selected private key, and do not disable host verification as a blanket fix.

Step 5: Choose reversible recovery

Relax offer filters explicitly, repair permissions, change host, restore credit through the approved owner, or resume from checkpoint according to the class.

Step 6: Close or escalate

Preserve identifiers and redacted evidence, confirm any replacement or cleanup, and escalate host or billing cases without speculative retries.

Authentication

Do not paste API keys into diagnostic commands. A scoped read key is usually sufficient for user, instance, logs, offers, and audit evidence; request additional authority only for the selected recovery.

Tool Discipline

Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.

Output

  • Failure class and evidence timeline
  • Ranked recovery action with mutation boundary
  • Redacted resolution or escalation receipt

Return command, CLI version, resource ID, observed status/error class, chosen response, outcome, and remaining billing risk.

Examples

An instance stuck in scheduling after being stopped is classified as GPU reacquisition, not image failure; the operator copies recoverable data or creates a new instance instead of waiting without a deadline.

Error Handling

FailureResponse
Evidence contains a credentialStop, redact, rotate if exposed, and recollect safely.
Error shape is inconsistentPreserve HTTP status plus msg or message and classify conservatively.
Host is offlineDo not attempt repair on the host; preserve the instance ID and use external checkpoints.
Balance is zeroEscalate to the billing owner because resources and data may be at risk.

Resources

When not to use it

  • →Debugging non-Vast.ai infrastructure
  • →Ignoring platform status alerts

Prerequisites

Vast.ai CLI installedAPI key configured

Limitations

  • →Instance may have been destroyed
  • →Host machine went down
  • →Very large Docker image

How it compares

It provides a centralized diagnostic reference for Vast.ai-specific errors instead of generic troubleshooting.

Compared to similar skills

vastai-common-errors side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
vastai-common-errors (this skill)02moReviewIntermediate
network-info36moReviewBeginner
debug-cluster210moReviewIntermediate
doctor14moCautionBeginner

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore →

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

Search skills

Search the agent skills registry