Deploys AI workloads to Runpod serverless.

Install

mkdir -p .claude/skills/flash-machinaio && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/16519" && unzip -o skill.zip -d .claude/skills/flash-machinaio && rm skill.zip

Installs to .claude/skills/flash-machinaio

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

runpod-flash SDK and CLI for deploying AI workloads on Runpod serverless GPUs/CPUs.
83 charsno explicit “when” trigger
Intermediate

Key capabilities

  • →Start a local development server with `flash run`
  • →Package artifacts for deployment with `flash build`
  • →Deploy applications to specified environments with `flash deploy`
  • →Manage deployment environments (list, create, get, delete)
  • →Define queue-based endpoints using a decorator
  • →Deploy external Docker images as endpoints

How it works

The skill allows local development and testing, then packages and deploys AI workloads to Runpod's serverless infrastructure. It supports defining endpoints through decorators for Python code or by specifying external Docker images.

Inputs & outputs

You give it
Python code or Docker image
You get back
Deployed AI workload on Runpod serverless GPUs/CPUs

When to use flash

  • →Deploy an AI model to Runpod
  • →Manage serverless deployment environments
  • →Provision GPU resources

About this skill

Runpod Flash

Use for Runpod Flash endpoint code and deployment tasks. Ordinary mxx Rust/CUDA pod benchmarks use .codex/skills/bench-on-runpod/SKILL.md instead.

Choose the Requested Operation

Read the relevant section of .codex/skills/flash/guides/sdk-and-cli.md: Setup/CLI for environment operations, Endpoint modes for decorators/routes/image clients, constructor and job sections for API usage, or Common Patterns/Gotchas for debugging. Check installed SDK signatures and CLI help before relying on version-dependent defaults or resource enums.

Preserve the endpoint mode, user-selected GPU/CPU, worker limits, image, volume, and deployment environment. Prefer an explicit image version or digest for reproducibility. Do not increase worker counts or broaden GPU choices merely to match an example.

flash run exposes a local development server, but invoking its endpoints can provision and execute remote work. Treat it and cloud-calling examples according to their actual side effects. Prepare and inspect code or build artifacts before requesting missing cloud-execution authorization; a code-only task does not require deployment.

For remote functions, keep required imports inside the function, declare remote dependencies, and await asynchronous endpoint/job calls. GPU and CPU resource selections are alternatives.

Completion

For code changes, perform the relevant local validation and state whether cloud execution was performed. For authorized deployment or execution, verify the target environment/endpoint and job result, preserve useful logs, and follow the agreed resource lifecycle. Do not automatically run deletion examples as cleanup. Stop retrying when further progress needs new authorization, exceeds the run budget, or lacks new failure evidence.

When not to use it

  • →When the payload size exceeds 10MB
  • →When `gpu` and `cpu` parameters are specified simultaneously for an endpoint
  • →When `runsync` timeout is expected to exceed 60 seconds for cold starts without explicit timeout

Prerequisites

Python >=3.10runpod-flashRUNPOD_API_KEY

Limitations

  • →Payload limit is 10MB
  • →Imports must be inside the decorated function for queue-based endpoints
  • →Endpoint `gpu` and `cpu` parameters are mutually exclusive

How it compares

This skill automates the provisioning and deployment of AI workloads to Runpod, abstracting away manual infrastructure setup and management.

Compared to similar skills

flash side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
flash (this skill)05moReviewIntermediate
azure-functions107moReviewIntermediate
ml-pipeline-workflow96moNo flagsAdvanced
machine-learning-ops-ml-pipeline45moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

Search skills

Search the agent skills registry