bench-on-runpod
Automate remote execution and benchmarking of repository code on Runpod GPU nodes.
Install
mkdir -p .claude/skills/bench-on-runpod && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/16563" && unzip -o skill.zip -d .claude/skills/bench-on-runpod && rm skill.zipInstalls to .claude/skills/bench-on-runpod
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Run mxx repository programs, tests, or benchmarks on Runpod instances with GPU machine selection, network volume setup, SSH/scp configuration, branch push and remote checkout synchronization, durable log capture, log retrieval, instance cleanup, and failed-test fix/push/pull/rerun loops. Use when Codex is asked to execute mxx work on a Runpod machine or instance, including requests that mention Runpod GPUs such as RTX5090, RTX PRO 6000, H200, remote benchmarking, remote tests, or running a command on a Runpod pod.Key capabilities
- →Run mxx repository programs on Runpod instances
- →Select GPU machines for benchmarking
- →Set up network volumes and SSH/scp configuration
- →Synchronize local branches to remote checkouts
- →Capture and retrieve durable logs
- →Clean up Runpod instances after runs
How it works
This skill provisions Runpod instances based on user requirements, configures SSH and network volumes, synchronizes the local branch to the remote checkout, executes the specified command while capturing durable logs, and manages the instance lifecycle.
Inputs & outputs
When to use bench-on-runpod
- →Running benchmarks on RTX5090 GPUs
- →Executing remote tests
- →Syncing and rerunning failed jobs
About this skill
Bench on Runpod
Use for requested mxx execution on Runpod pods. For planning or editing this workflow, inspect relevant files without provisioning or running it. Use .codex/skills/runpodctl/SKILL.md only when CLI resource operations are needed.
Run Contract
Resolve the command, exact source, GPU type, GPUs per pod, pod count, storage, and any time/cost limit from the task and existing environment. Preserve explicit choices. Choose routine settings from existing mxx configuration; ask only for unresolved decisions that affect the requested result or resource authorization. Prefer an existing network volume named for mxx; creation or use of a non-mxx volume needs authorization already present in the task or a focused question.
Carry existing authorization through the run. A test-only request does not authorize changing cryptographic semantics or an unlimited fix/rerun loop. The repository's integration-test restriction still applies. If authorized resources are unavailable, report that constraint before substituting hardware or raising capacity.
Source and Setup
- Inspect branch, commit, and dirty state. Use the user's branch; when a new branch is needed and none is specified, choose a descriptive
codex/name. - For the normal Git workflow, commit only task-owned changes when committing is authorized, push when authorized, and record the exact commit. Prepare the scoped diff before asking about a missing commit/push authorization. Honor no-push requests; if transfer is authorized, a source archive with a manifest and hashes can identify the exact dirty source without publishing it.
- Inspect remote dirty/untracked state before synchronization. Use an isolated checkout or preserve existing artifacts; do not blindly reset or clean a reused volume. Verify the checked-out commit or transferred manifest, including required submodule revisions, before execution.
- Reuse a suitable existing pod when requested. Otherwise provision only the agreed configuration, with SSH and
scpsupport. Record pod IDs and the volume mount. - For a new mxx volume, inspect
.codex/skills/bench-on-runpod/scripts/setup.shbefore running it remotely. It installs system dependencies and tools and expects the standard Runpod workspace mount; do not run it locally. Reuse an initialized environment when suitable. In the standard layout, source the mounted workspace'senv.shin each remote shell and enter itsmxxcheckout; verify actual paths rather than assuming a reused pod has this layout.
Execution and Evidence
Run the requested command and parameters with a durable combined stdout/stderr log and explicit exit-status evidence. Use release builds and RUST_LOG=debug for benchmarks unless the task specifies otherwise. Record VRAM usage about every 3 seconds during GPU measurements.
Record pod identity, GPU model/count, pod count, source commit or manifest, timestamp/timezone, command, non-secret environment overrides, remote paths, and exit code. Use a recognizable log filename with a short command summary. Do not log credentials.
For foreground pipelines, enable pipefail when using tee. For jobs that may outlive SSH, use nohup or an existing job manager, capture the PID/job ID, and persist the exit code separately. Reconnect to the same job rather than starting another copy. Follow root AGENTS.md for efficient waiting and progress updates.
Copy logs and relevant artifacts to logs/ in the active local checkout, or the user's requested destination, and verify retrieval before cleanup. Missing exit status means the result is unknown, even when the log contains successful intermediate stages.
Failures and Cleanup
When the task includes fixes, diagnose from evidence, implement the scoped repair locally, validate as appropriate, synchronize its exact source under the same authorization, and rerun affected commands with new logs. Stop retrying when the same blocker persists without new evidence, the agreed budget is reached, or a fix requires a new semantic or resource decision. Preserve the failure evidence and report the next action; finish independent authorized work.
After retrieval, stop the pod promptly unless the user requested that it remain running or an authorized retry will use it imminently within the run budget. If blocked on user input, stop it unless the user explicitly requested that it remain running. Stopping and deleting are different actions: do not terminate/delete a pod or volume without explicit authorization. Verify the final state through Runpod, and report a cleanup failure rather than claiming the pod stopped.
Completion means the requested command's result is known, logs are available locally, and the agreed pod state is verified. Report these with the source identity and any unmet correctness or performance target. A successful estimate or compile does not establish GPU runtime correctness.
When not to use it
- →When the task is not executing mxx work on a Runpod machine
- →When uncommitted local changes are not authorized for committing
- →When a machine cannot receive files through `scp`
Limitations
- →Requires explicit user authorization for committing uncommitted local changes
- →Requires SSH to be configured with `scp` support on the Runpod instance
- →Does not proceed if the mounted workspace path is not `/workspace`
How it compares
This skill provides a structured workflow for remote execution on Runpod, handling instance provisioning, synchronization, and logging, offering a controlled environment compared to manual remote execution.
Compared to similar skills
bench-on-runpod side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| bench-on-runpod (this skill) | 0 | 5mo | Review | Advanced |
| deployment-pipeline-design | 6 | 4mo | Review | Advanced |
| cloudflare-deploy | 3 | 7mo | Review | Intermediate |
| deployment-engineer | 4 | 5mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by MachinaIO
View all by MachinaIO →You might also like
deployment-pipeline-design
wshobson
Design multi-stage CI/CD pipelines with approval gates, security checks, and deployment orchestration. Use when architecting deployment workflows, setting up continuous delivery, or implementing GitOps practices.
cloudflare-deploy
davila7
Deploy applications and infrastructure to Cloudflare using Workers, Pages, and related platform services. Use when the user asks to deploy, host, publish, or set up a project on Cloudflare.
deployment-engineer
sickn33
Expert deployment engineer specializing in modern CI/CD pipelines, GitOps workflows, and advanced deployment automation. Masters GitHub Actions, ArgoCD/Flux, progressive delivery, container security, and platform engineering. Handles zero-downtime deployments, security scanning, and developer experience optimization. Use PROACTIVELY for CI/CD design, GitOps implementation, or deployment automation.
devops
mrgoonie
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm). Use for serverless, containers, CI/CD, GitOps, security audit.
kcli-cluster-deployment
karmab
Guides deployment and management of Kubernetes clusters with kcli. Use when deploying OpenShift, k3s, kubeadm, or other Kubernetes distributions.
managing-deployment-rollbacks
jeremylongshore
Deploy use when you need to work with deployment and CI/CD. This skill provides deployment automation and orchestration with comprehensive guidance and automation. Trigger with phrases like "deploy application", "create pipeline", or "automate deployment".