running-tend
Provides project-specific CI guidance for tend workflows on Worktrunk.
Install
mkdir -p .claude/skills/running-tend && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/12454" && unzip -o skill.zip -d .claude/skills/running-tend && rm skill.zipInstalls to .claude/skills/running-tend
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Worktrunk-specific guidance for tend CI workflows. Adds codecov polling, Rust test commands, labels, and review criteria on top of the generic tend-* skills. Use when operating in CI.Key capabilities
- →Poll Codecov patch status in CI
- →Investigate local Codecov failures using `task coverage` and `cargo llvm-cov`
- →Query Codecov API for patch-level summaries and uncovered lines
- →Download `code-coverage-report` artifacts from CI runs
- →Run specific Rust test suites like unit or integration tests
- →Classify CI failures as transient infrastructure, flaky test, or real regression
How it works
This skill provides specific commands and guidance for interacting with Codecov and Rust tests within Worktrunk's CI environment. It details how to poll Codecov status, query its API, and download coverage artifacts.
Inputs & outputs
When to use running-tend
- →Check codecov status
- →Investigate CI failures
- →Run local tests
About this skill
Worktrunk Tend CI
Project-specific guidance for tend workflows running on worktrunk (a Rust
CLI for managing git worktrees). The generic skills
(/tend-ci-runner:run-tend, /tend-ci-runner:review,
/tend-ci-runner:triage, etc.) provide the workflow framework; this skill
adds worktrunk conventions.
Filing issues in other repos
Standing exception granted: file directly in agent-equipped targets (per
Filing issues in the bundled /tend-ci-runner:act-in-other-repos skill)
without asking permission here first. The default rule (open an issue here asking
permission first) still applies when the target shows no agent signals.
Codecov Monitoring
After required CI checks pass, poll codecov/patch — it is mandatory despite
being marked non-required:
for i in $(seq 1 5); do
CODECOV=$(gh pr checks <number> 2>&1 | grep 'codecov/patch' || true)
if echo "$CODECOV" | grep -q 'pass'; then
echo "codecov/patch passed"; exit 0
elif echo "$CODECOV" | grep -q 'fail'; then
echo "codecov/patch FAILED"; exit 1
fi
sleep 60
done
If codecov fails locally, investigate with task coverage and
cargo llvm-cov report --show-missing-lines | grep <file>.
Investigating codecov failures in CI
task and cargo-llvm-cov are not installed in the tend-setup action.
Don't try to cargo install them in the sandbox — past attempts at
source-compiling installs cascaded into bash-tool interrupts that blocked
even pwd and echo. Instead, query Codecov directly, following
tests/AGENTS.md → Coverage Investigation for the endpoints and their
traps. The scratch paths there and below go to ${TMPDIR:-/tmp} — write new
ones the same way.
If the Codecov API markers aren't enough, download the code-coverage-report
artifact from the PR head's coverage workflow run — it contains a
cobertura.xml with per-line hit counts:
REPO=$(gh repo view --json nameWithOwner --jq '.nameWithOwner')
# Find the coverage run on the PR head SHA:
CI_RUN=$(gh api "repos/$REPO/commits/<sha>/check-runs" --jq '.check_runs[] | select(.name == "code-coverage") | .details_url | capture("runs/(?<id>[0-9]+)") | .id')
# List artifacts, then download the coverage one:
gh api "repos/$REPO/actions/runs/$CI_RUN/artifacts" --jq '.artifacts[] | {name, id}'
gh api "repos/$REPO/actions/artifacts/<id>/zip" > "${TMPDIR:-/tmp}/coverage.zip"
unzip -q "${TMPDIR:-/tmp}/coverage.zip" -d "${TMPDIR:-/tmp}/coverage"
Test Commands
cargo run -- hook pre-merge --yes # full suite + lints
cargo test --lib --bins # unit tests only
cargo test --test integration # integration tests only
CI runs on Linux, Windows, and macOS.
Rework a test that reaches for its environment
A test that leans on inherited state — the process CWD, an ambient env var —
sets up its own instead, via TestRepo::with_initial_commit() plus a tempdir,
the way most worktrunk tests already do. Guarding it with an early return
(Don't "fix" tests by adding skip guards in /tend-ci-runner:fix-a-bug)
drops the coverage rather than restoring it. This governs every workflow that
fixes a test here, not just issue triage.
Session Log Paths
The artifact directory is named after the agent's working directory, which
moves between tend releases — 0.2.5–0.2.13 used a per-run
tend-agent-workspace-*/checkout, and from 0.2.14 the agent works in the
runner's own checkout, so -home-runner-work-worktrunk-worktrunk/ is the
prefix again. Match on neither literal: use the bundled
find "$DEST" -name '*.jsonl' recipe, which holds across all of them — every
shape is one <session-id>.jsonl under a single slugified directory.
Labels
automated-fix— fix PRs from triage and ci-fix workflowsnightly-cleanup— nightly sweep issues and PRs
CI Fix: Prefer Rerun for Transient Infrastructure Failures
Before opening a fix/ci-* PR, classify the failure:
- Transient infrastructure (link-check timeouts, apt-get flakes, GitHub outages, runner disk issues, codecov upload blips) — do not create a PR. The maintainer will rerun CI. Comment on the run or exit silently; a permanent config change for a one-off timeout is churn the maintainer will close.
- Flaky test (known-flaky or first-seen PTY/shell test) — try to fix it.
- Real regression — proceed with a fix PR.
Non-required ≠ transient. A non-required job (e.g. collect affected coverage, affected tests (linux, advisory)) can fail from a real regression. The required/non-required distinction is about merge-blocking, not about how the failure is classified. If a deterministic build error (error[E...], "binary not found", "ambiguous candidates", missing target) repeats across consecutive runs of the same shape, it's a real regression even when the job is advisory. Reserve "transient" for non-deterministic causes: BrokenPipe, connection reset, runner disk full, GitHub API timeouts, host-availability blips.
Lychee link-check timeouts are always transient unless the same URL has
failed on at least two separate runs within the last few days. The check runs
as the link-check job in the nightly workflow. .config/lychee.toml
already sets max_retries = 6 and lists known-unreliable hosts; one timeout
is not enough evidence to extend that list. Signals you have a transient
failure, not a broken link:
- The previous run on the same or a nearby commit passed.
- Only
[TIMEOUT]is reported (not404/403/410). - The URL is reachable from a local
curl.
When in doubt, post a comment on the failed run summarizing the diagnosis and wait — don't open a PR.
Applying GitHub Suggestions
Apply the literal suggestion only — change the lines it covers, nothing more. If surrounding lines also need updating, note that in your reply.
PR Review: Don't Self-Dismiss Over Unrelated Test Flakes
If a clearly-unrelated test fails after you've already approved a PR, leave the approval in place and post a comment noting the flake. Do not dismiss your own approval to "gate" on a rerun.
GitHub blocks both gh run rerun --failed and per-job rerun
(POST /repos/{owner}/{repo}/actions/jobs/{id}/rerun) with HTTP 403 while
any job in the same workflow run is still in_progress. The non-required
benchmarks job routinely runs 80+ minutes after test (linux|macos|windows)
finish, so dismiss-then-wait-then-rerun cascades into a long session for no
benefit — the maintainer can rerun the failed job directly once benchmarks
clears, or merge regardless if the failure is clearly a flake.
The codecov-failure dismissal pattern is different and remains correct:
AGENTS.md requires explicit user approval before merging with failing
codecov/patch, so dismissing the approval until the coverage gap is
addressed is intentional.
Weigh the root-cause fix before shipping a workaround
When a mismatch, a false positive, or a stale value has an obvious non-code workaround (a template change, a config value, an alias, a comment recording the drift), don't stop there. First check whether the workaround is lossy or foot-gunny, and weigh a proportionate root-cause code fix before opening a PR that only records it. A "docs-only, no risk" framing is not the same as good guidance — a zero-code-risk change can still steer users toward a collision-prone or lossy config, and annotating a stale value leaves the duplication that made it stale. If you do recommend a workaround, surface its downsides in the PR body up front, not only when challenged.
This governs every workflow that opens a PR here, not just issue triage.
Issue Triage
When you need more information to diagnose a reported bug, the primary
ask is wt -vv <command>. Re-running the failing command with -vv
writes a diagnostic bundle — a single report containing wt/git/OS versions,
shell integration, wt config show, git worktree list --porcelain, and a
trace.log of every git invocation with its output. The -vv output prints
the bundle's exact absolute path (Diagnostics and performance profile saved @ …) followed by a ready-to-run gh gist create --web <path> line — point the user at those printed lines; don't hand them a
hardcoded path. In particular, never tell them to cat .git/wt/logs/diagnostic.md: inside a linked worktree .git is a gitdir
file, not a directory, so that path fails with Not a directory (os error 20) (ENOTDIR). The bundle actually lives under the git common dir —
"$(git rev-parse --git-common-dir)/wt/logs/diagnostic.md" resolves from any
worktree, and the printed absolute path already points there. One gist URL
pasted into the issue gives us most of what we'd otherwise ask for piecemeal,
so lead with this for unexplained failures rather than chaining
version/config/repro questions across multiple round-trips.
When the report is about a slow wt command, read its Performance profile
section first. It renders the same breakdown as wt config state logs profile
(subprocess time by command type, slowest calls, repeated (command, context)
pairs) directly from the bundled trace.log, so you can spot redundant git
calls and slow commands without parsing the raw trace by hand. The same
report run against a statusline capture is the weekly per-render cache
check — see Weekly Maintenance: Statusline Cache-Check.
Reach for narrower asks only when the diagnostic is overkill:
wt --version— when the only question is whether a fix has landed.wt config show— when the suspicion is purely config/shell-integration and you already have the command + repro.
Don't ship fixes you can't verify
When the bug or proposed fix turns on runtime state the bot can't observe from CI — plugin hooks firing inside an agent CLI (Claude Code, Codex, Gemini), shell-integration side effects, interactive prompt rendering, signal forwarding into a TTY — do not open a PR premised on the hypothesis. Signals to stop:
- The pr
Content truncated.
When not to use it
- →When operating outside of CI workflows
- →When generic tend-* skills are sufficient without Worktrunk-specific conventions
- →When a CI failure is due to transient infrastructure issues, as a rerun is preferred over a fix PR
Limitations
- →`task` and `cargo-llvm-cov` are not installed in the `claude-setup` action
- →Source-compiling `cargo install` commands for `task` and `cargo-llvm-cov` is not supported in the sandbox
- →Lychee link-check timeouts are considered transient unless the same URL fails on at least two separate runs within a few days
How it compares
This skill adds Worktrunk-specific conventions and troubleshooting steps for Codecov and Rust tests, unlike generic CI workflows that lack these tailored instructions.
Compared to similar skills
running-tend side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| running-tend (this skill) | 0 | 3mo | Review | Intermediate |
| nx-run-tasks | 1 | 8mo | No flags | Intermediate |
| cts-triage | 1 | 6mo | Review | Advanced |
| pre_commit | 0 | 7mo | Review | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by max-sixty
View all by max-sixty →You might also like
nx-run-tasks
nrwl
Helps with running tasks in an Nx workspace. USE WHEN the user wants to execute build, test, lint, serve, or run any other tasks defined in the workspace.
cts-triage
gfx-rs
Run CTS test suites and investigate failures
pre_commit
pums974
Standardized pre-commit workflow for linting, formatting, and local quality gates.
ci-cd
ZacharyLuz
Continuous Integration and Continuous Deployment best practices. Use when setting up automated build pipelines, test automation, deployment workflows, or improving release processes.
ci-pr-helper
lance-format
Run local test/style checks and open GitHub PRs for lance-context. Use when asked to run CI-equivalent checks (uv pytest, ruff/pyright, cargo fmt/clippy/test) and then create a PR with a proper title/body.
e2e-test-service-management
raphaelmansuy
Service management for E2E testing in EdgeQuake. Start, stop, and monitor PostgreSQL, backend API, and frontend services. Includes health checks and logging utilities for interactive testing workflows.