Deploys AI workloads to Runpod serverless.
Install
mkdir -p .claude/skills/flash-machinaio && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/16519" && unzip -o skill.zip -d .claude/skills/flash-machinaio && rm skill.zipInstalls to .claude/skills/flash-machinaio
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
runpod-flash SDK and CLI for deploying AI workloads on Runpod serverless GPUs/CPUs.Key capabilities
- →Start a local development server with `flash run`
- →Package artifacts for deployment with `flash build`
- →Deploy applications to specified environments with `flash deploy`
- →Manage deployment environments (list, create, get, delete)
- →Define queue-based endpoints using a decorator
- →Deploy external Docker images as endpoints
How it works
The skill allows local development and testing, then packages and deploys AI workloads to Runpod's serverless infrastructure. It supports defining endpoints through decorators for Python code or by specifying external Docker images.
Inputs & outputs
When to use flash
- →Deploy an AI model to Runpod
- →Manage serverless deployment environments
- →Provision GPU resources
About this skill
Runpod Flash
Use for Runpod Flash endpoint code and deployment tasks. Ordinary mxx Rust/CUDA pod benchmarks use .codex/skills/bench-on-runpod/SKILL.md instead.
Choose the Requested Operation
Read the relevant section of .codex/skills/flash/guides/sdk-and-cli.md: Setup/CLI for environment operations, Endpoint modes for decorators/routes/image clients, constructor and job sections for API usage, or Common Patterns/Gotchas for debugging. Check installed SDK signatures and CLI help before relying on version-dependent defaults or resource enums.
Preserve the endpoint mode, user-selected GPU/CPU, worker limits, image, volume, and deployment environment. Prefer an explicit image version or digest for reproducibility. Do not increase worker counts or broaden GPU choices merely to match an example.
flash run exposes a local development server, but invoking its endpoints can provision and execute remote work. Treat it and cloud-calling examples according to their actual side effects. Prepare and inspect code or build artifacts before requesting missing cloud-execution authorization; a code-only task does not require deployment.
For remote functions, keep required imports inside the function, declare remote dependencies, and await asynchronous endpoint/job calls. GPU and CPU resource selections are alternatives.
Completion
For code changes, perform the relevant local validation and state whether cloud execution was performed. For authorized deployment or execution, verify the target environment/endpoint and job result, preserve useful logs, and follow the agreed resource lifecycle. Do not automatically run deletion examples as cleanup. Stop retrying when further progress needs new authorization, exceeds the run budget, or lacks new failure evidence.
When not to use it
- →When the payload size exceeds 10MB
- →When `gpu` and `cpu` parameters are specified simultaneously for an endpoint
- →When `runsync` timeout is expected to exceed 60 seconds for cold starts without explicit timeout
Prerequisites
Limitations
- →Payload limit is 10MB
- →Imports must be inside the decorated function for queue-based endpoints
- →Endpoint `gpu` and `cpu` parameters are mutually exclusive
How it compares
This skill automates the provisioning and deployment of AI workloads to Runpod, abstracting away manual infrastructure setup and management.
Compared to similar skills
flash side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| flash (this skill) | 0 | 5mo | Review | Intermediate |
| azure-functions | 10 | 7mo | Review | Intermediate |
| ml-pipeline-workflow | 9 | 6mo | No flags | Advanced |
| machine-learning-ops-ml-pipeline | 4 | 5mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by MachinaIO
View all by MachinaIO →You might also like
azure-functions
aj-geddes
Create serverless functions on Azure with triggers, bindings, authentication, and monitoring. Use for event-driven computing without managing infrastructure.
ml-pipeline-workflow
wshobson
Build end-to-end MLOps pipelines from data preparation through model training, validation, and production deployment. Use when creating ML pipelines, implementing MLOps practices, or automating model training and deployment workflows.
machine-learning-ops-ml-pipeline
sickn33
Design and implement a complete ML pipeline for: $ARGUMENTS
senior-ml-engineer
davila7
World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. Expertise in PyTorch, TensorFlow, model deployment, feature stores, model monitoring, and ML infrastructure. Includes LLM integration, fine-tuning, RAG systems, and agentic AI. Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs into production systems.
mlflow
davila7
Track ML experiments, manage model registry with versioning, deploy models to production, and reproduce experiments with MLflow - framework-agnostic ML lifecycle platform
mlops-automation
fmind
Guide to refine MLOps projects with task automation, containerization, CI/CD pipelines, and robust experiment tracking.