agent-fresh-operate
Operates the lifecycle of agent VMs, including setup, destruction, and verification.
Install
mkdir -p .claude/skills/agent-fresh-operate && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/17057" && unzip -o skill.zip -d .claude/skills/agent-fresh-operate && rm skill.zipInstalls to .claude/skills/agent-fresh-operate
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Use when deploying, tearing down, or reproducing a fresh NPA agent VM from scratch — npa-driven destroy/fresh-setup, profile selection, tiered verify gates, and teardown failure recovery.Key capabilities
- →Deploy a fresh NPA agent VM
- →Teardown existing agent infrastructure
- →Reproduce agent environments from scratch
- →Verify `/api/models` and `/api/chat` after deployment
- →Debug destroy/fresh-setup failures
- →Run smoke, grounded, and live verification gates
How it works
This skill manages the lifecycle of NPA agent VMs by providing commands for fresh setup, teardown, and verification loops. It uses `npa` commands and scripts to deploy, bootstrap, and test agent functionality on a development machine.
Inputs & outputs
When to use agent-fresh-operate
- →Deploy a fresh agent VM
- →Destroy agent infrastructure
- →Reproduce agent environment
- →Verify deployment health
About this skill
Agent Fresh Operate
When To Use
Use this skill to operate a clean agent VM lifecycle on the operator/dev VM:
- First-time
fresh-setupon a project alias - Teardown → redeploy loops (“reproduce from scratch”)
- Validate
/api/modelsand/api/chatafter deploy - Debug destroy/fresh-setup failures (IAM, orphan VMs, ingress rules)
For chat UX, API shapes, and Rerun iframe behavior, use npa-agent. For
npa configure / object-storage provisioning, use nebius-infra.
Entry Points
npa/.venv/bin/npa agent fresh-setup— initialize project env + deploy + bootstrapnpa/.venv/bin/npa agent destroy— npa-driven teardown (ingress cleanup, TF destroy, orphan VM delete)- If the project stanza is already gone, resume only from the opaque receipt ID
printed before removal (
agent destroy --receipt <id> --name <name> --yes) or from exact--project-id/--instance-idprovider identity. Conflicting receipt, operation-journal, record, or exact identities stop before deletion; NPA never performs a display-name/prefix VM sweep. npa/scripts/agent_fresh_setup_loop.sh— destroy → fresh-setup → smoke chat (loop until success)- Exact-name retries after client transport loss adopt a healthy exact VM or
resume its first incomplete phase. Do not use
--replacesolely because the final Terraform/SSH response was lost; mismatched or unavailable evidence is indeterminate and resumable. - A completed service installer writes a private receipt before credentials are staged. Interrupted bootstrap retries reuse that install only when its rendered contents match exactly, then restage credentials and verify health. Changed source/settings or a missing receipt require installation again.
npa/scripts/agent_mature_verify_loop.sh— bootstrap-first mature loop (existing agents; not fresh deploy)
All npa agent … and nebius commands run on the operator/dev VM with
the selected NPA configuration root (NPA_CONFIG_DIR, default ~/.npa).
Use the authorized checkout and branch for live tests.
The committed recovery regression is
npa/tests/e2e/test_agent_recovery_live.py. Set
NPA_AGENT_RECOVERY_LIVE_CONFIG to an owner-only JSON file containing
deploy_args (the argument array beginning with agent, deploy, including an
unused --name, exact project, --agent-only, and ingress settings) and
evidence_dir (an owner-only directory outside the checkout). Use isolated
NPA_CONFIG_DIR and NPA_OPERATION_JOURNAL_DIR, run credential/capacity preflights,
then run that test with the checkout's own Python. It injects an SSH failure during
credential staging, requires one service install across both attempts, verifies
authenticated health on the same VM, and destroys that exact test agent.
When the project already has agent records, cleanup uses --keep-iam to preserve
their shared identity and still requires provider-verified absence of the test
agent's infrastructure.
After setting the private configuration and completing the preflights, enable the common E2E gate as well as the recovery-specific configuration:
NPA_INTEGRATION_E2E=1 npa/.venv/bin/python -m pytest \
npa/tests/e2e/test_agent_recovery_live.py -q
Procedure
For an explicitly authorized existing bucket in another project, save its exact
bucket, endpoint, owner_project_id, and provider bucket_id in the selected
project's terraform_state configuration before deploy. The agent verifies both
projects belong to the selected tenant and the bucket's ID, name, and actual
parent agree, then probes the exact Terraform state key. This binding authorizes
data-plane use only: it creates no bucket or storage IAM grant and preserves the
actual owner in backend evidence. Missing or mismatched bindings fail before
Terraform. Without an explicit binding, the backend must exist in the compute
project. Keep credential selection in the supported private credential store or
process environment; do not represent external storage as newly owned resources.
When reusing configured storage, agent bootstrap creates or verifies a custom
group inside the compute project and an editor permit on that exact project
for the attached npa-agent account. It does not add a tenant-wide editors
membership. Creation IDs are journaled; rollback and last-agent cleanup verify
exact ownership and dependencies before deleting only run-created bindings.
Existing broad grants are not removed automatically.
-
Preconditions (dev VM).
cd <authorized-checkout> export NPA_CONFIG_DIR=<private-runtime-config-directory> export NPA_OPERATION_JOURNAL_DIR=<private-operation-journal-directory> export NPA_NEBIUS_PROFILE=<verified-profile> export NPA_SSH_KEY="${NPA_SSH_KEY:-$HOME/.ssh/id_ed25519}"Bootstrap the authorized checkout's own virtualenv before operating it.
NPA_CONFIG_DIRselects local config, credentials, cluster state, agent auth, workbench Terraform directories, and the default Terraform plugin cache. Set it before importing NPA. Concurrent operators should use separate private directories with the selected project stanza and authorized credentials. Provider calls honor the selected profile per command; deploying an agent does not activate or rewrite the host's shared default profile. Keep explicit operation-journal and SkyPilot isolation and Fleetwork_rootsettings when using those runtimes. The operation journal has its own configuration override; settingNPA_CONFIG_DIRalone does not relocate it.NPA_SSH_KEYis the SSH private-key path used after provisioning. It is not cloud-init key content and must never be passed as--ssh-public-key-path. That option defaults to the matching~/.ssh/id_ed25519.puband accepts exactly one OpenSSH public-key record. For a non-default private key, pass its existing matching.pubfile. If it is absent, derive only the public record withssh-keygen -y -f "$NPA_SSH_KEY"into an owner-controlled.pubfile, verify the two fingerprints match, and pass that.pubpath; never log either key's contents. -
Teardown (npa-driven — no manual
nebius vpcedits).npa/.venv/bin/npa agent destroy --project <alias> --name agent -
Fresh deploy.
npa/.venv/bin/npa agent fresh-setup \ --project <alias> --name agent \ --project-id <project-id> \ --tenant-id <tenant-id> \ --region us-central1The default capacity gate reserves the canonical follow-on cluster as well as the VM. Use
--agent-onlywhen this lifecycle intentionally creates only the UI VM or when a separately managed cluster already satisfies the workload plan; the flag still checks the VM's instance, disk, and public-IP capacity. Expect compute PermissionDenied with VM SA attachment on some cross-project profiles; npa retries apply without attachedservice_account_idand now emits a loud WARNING when it does — a VM without an attached SA cannot self-mint IAM tokens and needs an alternative token source (grant the deploying identitycompute.admin/equivalent, or inject a token on the VM).Agent VM IAM auth = attached service account (not a copied operator token). The VM authenticates to Nebius IAM using its attached
npa-agentservice account:get_iam_token()self-mints fresh tokens from the metadata/token-file sources the SA populates. npa no longer copies the operator's short-lived IAM token onto the VM (noNEBIUS_IAM_TOKEN/TF_VAR_iam_tokenin/opt/npa-agent/nebius.env, no/root/.npa/nebius-token, noagent-bootstrapprofile) — that token went stale and forced re-bootstrap. S3 access keys and the service API keys (Token Factory / HF / NGC) are still staged: object storage is HMAC-based and cannot use an IAM bearer token, and the product keys are independent of the SA. They are staged only after VM creation through the verified SSH channel; they never enter Terraform/cloud-init/user-data. On-VM Terraform (npa cluster …) mints a fresh token at run time vianebius iam get-access-token. -
Smoke gate (default “done” for fresh deploy).
source "${NPA_CONFIG_DIR:-$HOME/.npa}/agents/<alias>/agent/auth.env" BASE="$(npa/.venv/bin/npa agent status --project <alias> --name agent --json \ | npa/.venv/bin/python -c 'import json,sys; print(json.load(sys.stdin).get("public_url","").rstrip("/"))')" curl -sk -u "${AGENT_USER}:${AGENT_PASSWORD}" "${BASE}/api/models" curl -sk -u "${AGENT_USER}:${AGENT_PASSWORD}" -H 'Content-Type: application/json' \ -d '{"messages":[{"role":"user","content":"Say hello in one short sentence."}]}' \ "${BASE}/api/chat" -
Optional full gates.
- Grounded chat: ask “what is the current sim2real status” →
"grounded": true - Live regression:
NPA_AGENT_CHAT_LIVE=1 npa/.venv/bin/npa agent verify-live --project <alias> --name agent - Mature loop:
bash npa/scripts/agent_mature_verify_loop.sh(bootstrap-first)
- Grounded chat: ask “what is the current sim2real status” →
-
One-command loop.
export NPA_AGENT_PROJECT=<alias> NPA_AGENT_NAME=agent export NPA_AGENT_PROJECT_ID=<project-id> NPA_AGENT_TENANT_ID=<tenant-id> export NPA_AGENT_REGION=us-central1 NPA_NEBIUS_PROFILE=npa-mk8s bash npa/scripts/agent_fresh_setup_loop.sh
Preflight Before You Spend
npa/.venv/bin/npa agent preflight --project <alias> --name <name> [--agent-only]
Pass the same --name you will deploy. Capacity depends on it: an existing
agent of that name already holds its public IP and needs no headroom, while a new
name needs a free one. Preflighting the default agent and then deploying
--name something-else is how a "capacity ready" report is followed immediately
by a public-IP shortfall.
--agent-only drops the reserved PAIDF cluster shape, so it reserves no cluster
nodes and does not inspect mk8s. Use it when the operator can
Content truncated.
When not to use it
- →For chat UX, API shapes, and Rerun iframe behavior
- →For `npa configure` / object-storage provisioning
Prerequisites
Limitations
- →Cross-project deploy may require specific IAM permissions.
- →Destroy operations might encounter 'disk/SG in use' if orphan VMs exist.
- →The skill does not cover `npa configure` or object-storage provisioning.
How it compares
This workflow offers a complete, automated lifecycle management for NPA agent VMs, including specific verification tiers and failure recovery mechanisms, providing a more reliable and reproducible environment setup than manual VM provisioning
Compared to similar skills
agent-fresh-operate side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| agent-fresh-operate (this skill) | 0 | 3mo | Caution | Advanced |
| azure-role-selector | 0 | 7mo | No flags | Intermediate |
| setup-role-ops | 0 | 6mo | No flags | Intermediate |
| ops | 0 | 3mo | Review | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
You might also like
azure-role-selector
Tyler-R-Kendrick
|
setup-role-ops
tachiiri-org
Reconcile an ops repository to the expected shared-guidance baseline.
ops
denniszielke
>
env-stabilize
FlexNetOS
How to keep the environment reproducible and drift-free — the detect/drift, doctor-diagnostic, and content-hashed lock discipline envctl uses, plus how the built-in agent-env engine (envctl agent) provisions and locks the agent config. Use when checking environment health, diagnosing drift, regenera
cloudflare-manager
qdhenry
Comprehensive Cloudflare account management for deploying Workers, KV Storage, R2, Pages, DNS, and Routes. Use when deploying cloudflare services, managing worker containers, configuring KV/R2 storage, or setting up DNS/routing. Requires CLOUDFLARE_API_KEY in .env and Bun runtime with dependencies installed.
storage-networking
pluginagentmarketplace
Master Kubernetes storage management and networking architecture. Learn persistent storage, network policies, service discovery, and ingress routing.