BE

bench-on-runpod

Automate remote execution and benchmarking of repository code on Runpod GPU nodes.

Install

mkdir -p .claude/skills/bench-on-runpod && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/16563" && unzip -o skill.zip -d .claude/skills/bench-on-runpod && rm skill.zip

Installs to .claude/skills/bench-on-runpod

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Run mxx repository programs, tests, or benchmarks on Runpod instances with GPU machine selection, network volume setup, SSH/scp configuration, branch push and remote checkout synchronization, durable log capture, log retrieval, instance cleanup, and failed-test fix/push/pull/rerun loops. Use when Codex is asked to execute mxx work on a Runpod machine or instance, including requests that mention Runpod GPUs such as RTX5090, RTX PRO 6000, H200, remote benchmarking, remote tests, or running a command on a Runpod pod.
519 chars✓ has a “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Advanced

Key capabilities

  • →Run mxx repository programs on Runpod instances
  • →Select GPU machines for benchmarking
  • →Set up network volumes and SSH/scp configuration
  • →Synchronize local branches to remote checkouts
  • →Capture and retrieve durable logs
  • →Clean up Runpod instances after runs

How it works

This skill provisions Runpod instances based on user requirements, configures SSH and network volumes, synchronizes the local branch to the remote checkout, executes the specified command while capturing durable logs, and manages the instance lifecycle.

Inputs & outputs

You give it
A request to execute an mxx program, test, or benchmark on a Runpod instance, specifying machine type and command
You get back
Execution results and durable logs from the Runpod instance, with the instance cleaned up or kept running as specified

When to use bench-on-runpod

  • →Running benchmarks on RTX5090 GPUs
  • →Executing remote tests
  • →Syncing and rerunning failed jobs

About this skill

Bench on Runpod

Use for requested mxx execution on Runpod pods. For planning or editing this workflow, inspect relevant files without provisioning or running it. Use .codex/skills/runpodctl/SKILL.md only when CLI resource operations are needed.

Run Contract

Resolve the command, exact source, GPU type, GPUs per pod, pod count, storage, and any time/cost limit from the task and existing environment. Preserve explicit choices. Choose routine settings from existing mxx configuration; ask only for unresolved decisions that affect the requested result or resource authorization. Prefer an existing network volume named for mxx; creation or use of a non-mxx volume needs authorization already present in the task or a focused question.

Carry existing authorization through the run. A test-only request does not authorize changing cryptographic semantics or an unlimited fix/rerun loop. The repository's integration-test restriction still applies. If authorized resources are unavailable, report that constraint before substituting hardware or raising capacity.

Source and Setup

  • Inspect branch, commit, and dirty state. Use the user's branch; when a new branch is needed and none is specified, choose a descriptive codex/ name.
  • For the normal Git workflow, commit only task-owned changes when committing is authorized, push when authorized, and record the exact commit. Prepare the scoped diff before asking about a missing commit/push authorization. Honor no-push requests; if transfer is authorized, a source archive with a manifest and hashes can identify the exact dirty source without publishing it.
  • Inspect remote dirty/untracked state before synchronization. Use an isolated checkout or preserve existing artifacts; do not blindly reset or clean a reused volume. Verify the checked-out commit or transferred manifest, including required submodule revisions, before execution.
  • Reuse a suitable existing pod when requested. Otherwise provision only the agreed configuration, with SSH and scp support. Record pod IDs and the volume mount.
  • For a new mxx volume, inspect .codex/skills/bench-on-runpod/scripts/setup.sh before running it remotely. It installs system dependencies and tools and expects the standard Runpod workspace mount; do not run it locally. Reuse an initialized environment when suitable. In the standard layout, source the mounted workspace's env.sh in each remote shell and enter its mxx checkout; verify actual paths rather than assuming a reused pod has this layout.

Execution and Evidence

Run the requested command and parameters with a durable combined stdout/stderr log and explicit exit-status evidence. Use release builds and RUST_LOG=debug for benchmarks unless the task specifies otherwise. Record VRAM usage about every 3 seconds during GPU measurements.

Record pod identity, GPU model/count, pod count, source commit or manifest, timestamp/timezone, command, non-secret environment overrides, remote paths, and exit code. Use a recognizable log filename with a short command summary. Do not log credentials.

For foreground pipelines, enable pipefail when using tee. For jobs that may outlive SSH, use nohup or an existing job manager, capture the PID/job ID, and persist the exit code separately. Reconnect to the same job rather than starting another copy. Follow root AGENTS.md for efficient waiting and progress updates.

Copy logs and relevant artifacts to logs/ in the active local checkout, or the user's requested destination, and verify retrieval before cleanup. Missing exit status means the result is unknown, even when the log contains successful intermediate stages.

Failures and Cleanup

When the task includes fixes, diagnose from evidence, implement the scoped repair locally, validate as appropriate, synchronize its exact source under the same authorization, and rerun affected commands with new logs. Stop retrying when the same blocker persists without new evidence, the agreed budget is reached, or a fix requires a new semantic or resource decision. Preserve the failure evidence and report the next action; finish independent authorized work.

After retrieval, stop the pod promptly unless the user requested that it remain running or an authorized retry will use it imminently within the run budget. If blocked on user input, stop it unless the user explicitly requested that it remain running. Stopping and deleting are different actions: do not terminate/delete a pod or volume without explicit authorization. Verify the final state through Runpod, and report a cleanup failure rather than claiming the pod stopped.

Completion means the requested command's result is known, logs are available locally, and the agreed pod state is verified. Report these with the source identity and any unmet correctness or performance target. A successful estimate or compile does not establish GPU runtime correctness.

When not to use it

  • →When the task is not executing mxx work on a Runpod machine
  • →When uncommitted local changes are not authorized for committing
  • →When a machine cannot receive files through `scp`

Limitations

  • →Requires explicit user authorization for committing uncommitted local changes
  • →Requires SSH to be configured with `scp` support on the Runpod instance
  • →Does not proceed if the mounted workspace path is not `/workspace`

How it compares

This skill provides a structured workflow for remote execution on Runpod, handling instance provisioning, synchronization, and logging, offering a controlled environment compared to manual remote execution.

Compared to similar skills

bench-on-runpod side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
bench-on-runpod (this skill)05moReviewAdvanced
deployment-pipeline-design64moReviewAdvanced
cloudflare-deploy37moReviewIntermediate
deployment-engineer45moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

deployment-pipeline-design

wshobson

Design multi-stage CI/CD pipelines with approval gates, security checks, and deployment orchestration. Use when architecting deployment workflows, setting up continuous delivery, or implementing GitOps practices.

670

cloudflare-deploy

davila7

Deploy applications and infrastructure to Cloudflare using Workers, Pages, and related platform services. Use when the user asks to deploy, host, publish, or set up a project on Cloudflare.

342

deployment-engineer

sickn33

Expert deployment engineer specializing in modern CI/CD pipelines, GitOps workflows, and advanced deployment automation. Masters GitHub Actions, ArgoCD/Flux, progressive delivery, container security, and platform engineering. Handles zero-downtime deployments, security scanning, and developer experience optimization. Use PROACTIVELY for CI/CD design, GitOps implementation, or deployment automation.

418

devops

mrgoonie

Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm). Use for serverless, containers, CI/CD, GitOps, security audit.

216

kcli-cluster-deployment

karmab

Guides deployment and management of Kubernetes clusters with kcli. Use when deploying OpenShift, k3s, kubeadm, or other Kubernetes distributions.

22

managing-deployment-rollbacks

jeremylongshore

Deploy use when you need to work with deployment and CI/CD. This skill provides deployment automation and orchestration with comprehensive guidance and automation. Trigger with phrases like "deploy application", "create pipeline", or "automate deployment".

11

Search skills

Search the agent skills registry