vastai-common-errors
Diagnose Vast.ai errors using standard status codes and troubleshooting guides for GPU instance management.
Install
mkdir -p .claude/skills/vastai-common-errors && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8787" && unzip -o skill.zip -d .claude/skills/vastai-common-errors && rm skill.zipInstalls to .claude/skills/vastai-common-errors
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Diagnose and fix Vast.ai common errors and exceptions.Key capabilities
- →Map HTTP status codes to fixes
- →Diagnose instance status errors
- →Troubleshoot SSH connection issues
- →Verify GPU and CUDA availability
- →Check platform connection status
How it works
The tool maps specific API responses and instance states to known causes and corrective commands.
Inputs & outputs
When to use vastai-common-errors
- →Debug Vast.ai API authentication errors
- →Troubleshoot instance creation and docker loading failures
- →Resolve GPU rental configuration issues
- →Check Vast.ai platform status for connection errors
About this skill
Vast.ai Failure Classifier
Overview
Diagnose from provider state and error evidence instead of retrying every failure. Separate identity and permission errors, marketplace scarcity, transient startup, terminal host states, billing stops, SSH configuration, and workload exits.
Prerequisites
- Exact command, non-secret arguments, exit status, timestamp, and redacted response
- Affected account context, instance ID, image identity, and expected state
- Authority to inspect but not automatically destroy or fund resources
Instructions
Step 1: Capture structured evidence
Run the failing command with --raw where supported and record CLI version. For request diagnosis, use --explain or --curl only after ensuring generated output cannot expose the key.
Step 2: Classify control-plane failure
Treat 401 as credential failure, 403 as missing scoped permission, 429 as endpoint/identity rate limiting, and insufficient credit or spend-rate errors as billing policy—not host failure.
Step 3: Classify instance state
Loading may reflect an image pull; scheduling after a stop may wait indefinitely for the original GPU; exited is a workload/container failure; unknown or offline indicates missing host heartbeat.
Step 4: Check SSH and network facts
Wait for running, resolve the current SSH URL, verify the registered public key and selected private key, and do not disable host verification as a blanket fix.
Step 5: Choose reversible recovery
Relax offer filters explicitly, repair permissions, change host, restore credit through the approved owner, or resume from checkpoint according to the class.
Step 6: Close or escalate
Preserve identifiers and redacted evidence, confirm any replacement or cleanup, and escalate host or billing cases without speculative retries.
Authentication
Do not paste API keys into diagnostic commands. A scoped read key is usually sufficient for user, instance, logs, offers, and audit evidence; request additional authority only for the selected recovery.
Tool Discipline
Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.
Output
- Failure class and evidence timeline
- Ranked recovery action with mutation boundary
- Redacted resolution or escalation receipt
Return command, CLI version, resource ID, observed status/error class, chosen response, outcome, and remaining billing risk.
Examples
An instance stuck in scheduling after being stopped is classified as GPU reacquisition, not image failure; the operator copies recoverable data or creates a new instance instead of waiting without a deadline.
Error Handling
| Failure | Response |
|---|---|
| Evidence contains a credential | Stop, redact, rotate if exposed, and recollect safely. |
| Error shape is inconsistent | Preserve HTTP status plus msg or message and classify conservatively. |
| Host is offline | Do not attempt repair on the host; preserve the instance ID and use external checkpoints. |
| Balance is zero | Escalate to the billing owner because resources and data may be at risk. |
Resources
When not to use it
- →Debugging non-Vast.ai infrastructure
- →Ignoring platform status alerts
Prerequisites
Limitations
- →Instance may have been destroyed
- →Host machine went down
- →Very large Docker image
How it compares
It provides a centralized diagnostic reference for Vast.ai-specific errors instead of generic troubleshooting.
Compared to similar skills
vastai-common-errors side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| vastai-common-errors (this skill) | 0 | 2mo | Review | Intermediate |
| network-info | 3 | 6mo | Review | Beginner |
| debug-cluster | 2 | 10mo | Review | Intermediate |
| doctor | 1 | 4mo | Caution | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
network-info
UKGovernmentBEIS
Gather network configuration and connectivity information including interfaces, routes, and DNS
debug-cluster
openshift
Provides systematic debugging approaches for HyperShift hosted-cluster issues. Auto-applies when debugging cluster problems, investigating stuck deletions, or troubleshooting control plane issues.
doctor
Yeachan-Heo
Diagnose and fix oh-my-claudecode installation issues
replit-advanced-troubleshooting
jeremylongshore
Apply Replit advanced debugging techniques for hard-to-diagnose issues. Use when standard troubleshooting fails, investigating complex race conditions, or preparing evidence bundles for Replit support escalation. Trigger with phrases like "replit hard bug", "replit mystery error", "replit impossible to debug", "difficult replit issue", "replit deep debug".
devops-troubleshooter
sickn33
Expert DevOps troubleshooter specializing in rapid incident response, advanced debugging, and modern observability. Masters log analysis, distributed tracing, Kubernetes debugging, performance optimization, and root cause analysis. Handles production outages, system reliability, and preventive monitoring. Use PROACTIVELY for debugging, incident response, or system troubleshooting.
environment-triage
parcadei
Environment Triage