vastai-webhooks-events
Implements event polling for Vast.ai instances to track status changes and trigger workflows.
Install
mkdir -p .claude/skills/vastai-webhooks-events && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8867" && unzip -o skill.zip -d .claude/skills/vastai-webhooks-events && rm skill.zipInstalls to .claude/skills/vastai-webhooks-events
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Build event-driven workflows around Vast.ai instance lifecycle events.Key capabilities
- →Poll Vast.ai REST API for instance status transitions
- →Trigger custom handlers on instance lifecycle events
- →Implement auto-recovery for preempted GPU instances
- →Track cost-related status transitions for billing
- →Monitor instance states like loading, running, and exited
How it works
The poller periodically queries the Vast.ai REST API to compare current instance statuses against previous states. When a transition is detected, it executes registered callback functions.
Inputs & outputs
When to use vastai-webhooks-events
- →Monitoring GPU instance lifecycle states
- →Building auto-recovery for training jobs
- →Triggering notifications on instance errors
- →Tracking cost-related status transitions
About this skill
Vast.ai Webhooks & Events
Overview
Build event-driven workflows around Vast.ai GPU instance lifecycle. Vast.ai does not provide traditional webhooks, so event detection relies on polling the REST API at cloud.vast.ai/api/v0 and reacting to instance status transitions (loading, running, exited, error, offline).
Prerequisites
- Vast.ai CLI authenticated
- Understanding of instance lifecycle states
- Python 3.8+ for event loop implementation
Instructions
Step 1: Instance Status Poller
import time, json, subprocess
from typing import Callable, Dict, List
class InstanceEventPoller:
"""Poll Vast.ai API and emit events on status transitions."""
def __init__(self, api_key: str, poll_interval: int = 30):
self.api_key = api_key
self.poll_interval = poll_interval
self.previous_states: Dict[int, str] = {}
self.handlers: Dict[str, List[Callable]] = {}
def on(self, event: str, handler: Callable):
self.handlers.setdefault(event, []).append(handler)
def poll_once(self):
result = subprocess.run(
["vastai", "show", "instances", "--raw"],
capture_output=True, text=True)
instances = json.loads(result.stdout)
for inst in instances:
inst_id = inst["id"]
status = inst.get("actual_status", "unknown")
prev = self.previous_states.get(inst_id)
if prev and prev != status:
event = f"{prev}_to_{status}"
for handler in self.handlers.get(event, []):
handler(inst)
for handler in self.handlers.get("any_change", []):
handler(inst, prev, status)
self.previous_states[inst_id] = status
def run(self):
print(f"Polling every {self.poll_interval}s...")
while True:
self.poll_once()
time.sleep(self.poll_interval)
Step 2: Event Handlers
def on_instance_running(instance):
print(f"Instance {instance['id']} is RUNNING")
print(f" SSH: ssh -p {instance['ssh_port']} root@{instance['ssh_host']}")
# Trigger: start training job, send notification, etc.
def on_instance_exited(instance):
print(f"Instance {instance['id']} EXITED")
# Trigger: collect results, check for errors, notify team
def on_spot_preemption(instance, old_status, new_status):
if old_status == "running" and new_status in ("exited", "offline"):
print(f"ALERT: Instance {instance['id']} may have been preempted")
# Trigger: auto-recovery, provision replacement
# Wire up handlers
poller = InstanceEventPoller(api_key)
poller.on("loading_to_running", on_instance_running)
poller.on("running_to_exited", on_instance_exited)
poller.on("any_change", on_spot_preemption)
poller.run()
Step 3: Auto-Recovery on Preemption
def auto_recover(instance, old_status, new_status):
"""Automatically replace preempted instances."""
if old_status != "running" or new_status not in ("exited", "offline", "error"):
return
gpu_name = instance.get("gpu_name", "RTX_4090")
image = instance.get("image_uuid", "pytorch/pytorch:latest")
print(f"Auto-recovering {instance['id']} ({gpu_name})...")
# Search for replacement
offers = json.loads(subprocess.run(
["vastai", "search", "offers",
f"gpu_name={gpu_name} reliability>0.98 rentable=true",
"--order", "dph_total", "--raw", "--limit", "3"],
capture_output=True, text=True, check=True).stdout)
if offers:
new_id = json.loads(subprocess.run(
["vastai", "create", "instance", str(offers[0]["id"]),
"--image", image, "--disk", "50", "--raw"],
capture_output=True, text=True, check=True).stdout)["new_contract"]
print(f"Replacement instance: {new_id}")
Step 4: Cost Event Tracking
def track_costs(instance, old_status, new_status):
"""Log cost events for billing tracking."""
if new_status == "running":
print(f"BILLING START: Instance {instance['id']} "
f"at ${instance.get('dph_total', 0):.3f}/hr")
elif old_status == "running":
print(f"BILLING STOP: Instance {instance['id']}")
Output
- Polling-based event detection for instance status changes
- Event handlers for running, exited, preempted states
- Auto-recovery on spot preemption
- Cost tracking event logger
Error Handling
| Error | Cause | Solution |
|---|---|---|
| Missed status transition | Poll interval too long | Reduce to 15-30s for critical instances |
| False preemption alert | Instance restarted intentionally | Track expected state changes |
| Auto-recovery loops | Same host keeps failing | Exclude failed host IDs from search |
| API timeout during poll | Network or rate limiting | Retry with backoff; continue polling |
Resources
Next Steps
For performance optimization, see vastai-performance-tuning.
Examples
Slack notifications: Wire on_instance_running to send a Slack message with SSH connection details. Wire on_spot_preemption to alert the team.
Training monitor: Track running_to_exited events. If exit was expected (job complete), collect results. If unexpected, trigger auto-recovery with checkpoint resume.
When not to use it
- →Real-time event streaming as Vast.ai requires polling
Prerequisites
Limitations
- →Poll interval latency can cause missed status transitions
- →Auto-recovery loops may occur if failed host IDs are not excluded
- →API rate limiting may require backoff strategies
How it compares
This approach automates event detection through polling rather than relying on native webhooks which are not provided by the platform.
Compared to similar skills
vastai-webhooks-events side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| vastai-webhooks-events (this skill) | 0 | 27d | Review | Intermediate |
| netalertx-plugin-run-development | 1 | 6mo | Review | Intermediate |
| ads-network-killswitch | 0 | 4mo | No flags | Beginner |
| sms-inbox-manager | 0 | 1mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
netalertx-plugin-run-development
netalertx
Create and run NetAlertX plugins. Use this when asked to create plugin, run plugin, test plugin, plugin development, or execute plugin script.
ads-network-killswitch
engremran07
Network killswitch: per-network enable/disable. Use when: disabling problematic networks, emergency ad shutoff, toggling networks without deletion.
sms-inbox-manager
Little-King2022
Maintain a self-hosted realtime SMS inbox service. Use when inspecting received SMS messages, managing the web UI, debugging SMS forwarding, restarting or enabling the service, adjusting authentication, updating nginx routing, or locating data/log files.
ha-config-fetch
calexandre
>-
amc-run-sample-calibration
NVIDIA
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
telegram-bot-builder
davila7
Expert in building Telegram bots that solve real problems - from simple automation to complex AI-powered bots. Covers bot architecture, the Telegram Bot API, user experience, monetization strategies, and scaling bots to thousands of users. Use when: telegram bot, bot api, telegram automation, chat bot telegram, tg bot.