VA

Runs batch validation of agent checkpoints to identify and fix issues.

Install

mkdir -p .claude/skills/validate-run && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/11968" && unzip -o skill.zip -d .claude/skills/validate-run && rm skill.zip

Installs to .claude/skills/validate-run

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Validate all checkpoints from an agent run directory in parallel. Spawns test-validator agents for each checkpoint and summarizes results. Invoke with /validate-run <run_path> [problem].
186 charsno explicit “when” trigger
Advanced

Key capabilities

  • Discover checkpoints within an agent run directory
  • Launch test-validator agents in parallel for each checkpoint
  • Collect and extract verdict, summary, and action items from validator reports
  • Verify sub-agent results for correctness of bug classification and proposed fixes
  • Generate a consolidated summary report of run validation

How it works

The skill discovers checkpoints in a run directory, spawns parallel test-validator agents for each, collects their reports, and then critically verifies and summarizes the findings.

Inputs & outputs

You give it
A run directory path and an optional problem name (e.g., '/path/to/run/directory file_backup')
You get back
A consolidated summary report detailing overall results, verdicts by type, action items, and corrections to sub-agent reports

When to use validate-run

  • Validating agent run results
  • Analyzing problem solutions
  • Automating checkpoint testing

About this skill

Validate Run

Analyze all checkpoints from an agent's run directory, spawning test-validator agents in parallel to validate each checkpoint against its specification.

Usage:

  • /validate-run /path/to/run/directory - Validate all problems/checkpoints
  • /validate-run /path/to/run/directory file_backup - Validate specific problem only

IMPORTANT: You Must Verify Sub-Agent Results

Sub-agents (Sonnet) make mistakes. After collecting their reports, you MUST:

  1. Read the actual spec and test files yourself
  2. Verify each TEST_BUG classification is correct
  3. Check that proposed fixes are actually valid
  4. Correct any errors before presenting to the user

See Step 4 for detailed verification process.


Step 1: Discover Checkpoints

The run directory structure varies. Look for checkpoints in these patterns:

# Agent run output structure
{run_path}/submissions/{problem}/checkpoint_N/

# Or direct problem solutions
{run_path}/checkpoint_N/

# Or solutions directory
{run_path}/solutions/checkpoint_N/

Use this command to discover checkpoints:

find {run_path} -type d -name "checkpoint_*" | sort

For each checkpoint found, extract:

  • problem_name: The problem being tested (from path or user input)
  • checkpoint: The checkpoint number (e.g., checkpoint_1, checkpoint_2)
  • snapshot_path: Full path to the checkpoint directory

Step 2: Launch Validators in Parallel

CRITICAL: Launch ALL test-validator agents in a SINGLE message with multiple Task tool calls.

For each checkpoint discovered, spawn a test-validator agent:

Task tool call 1: test-validator for {problem} checkpoint_1
Task tool call 2: test-validator for {problem} checkpoint_2
Task tool call 3: test-validator for {problem} checkpoint_3
...

Prompt template for each agent:

Validate the following checkpoint:

Problem: {problem_name}
Checkpoint: checkpoint_{N}
Snapshot Path: {snapshot_path}

Run the evaluation, analyze all test results against the specification, and save your report to:
problems/{problem_name}/checkpoint_{N}_report.md

Focus on determining whether any test failures indicate:
1. Solution bugs (code doesn't match spec)
2. Test bugs (tests expect behavior not in spec)
3. Spec ambiguity (unclear requirements)

Step 3: Collect Results

After all validators complete, read each report:

problems/{problem}/checkpoint_1_report.md
problems/{problem}/checkpoint_2_report.md
...

Extract from each report:

  • VERDICT line (the overall conclusion)
  • SUMMARY table (counts)
  • ACTION ITEMS (what needs fixing)

Step 4: VERIFY SUB-AGENT RESULTS (CRITICAL)

WARNING: Sub-agents (Sonnet) can and do make mistakes. You MUST verify their findings.

DO NOT blindly trust sub-agent reports. For each finding:

  1. Read the spec quote - Does it actually support the sub-agent's conclusion?
  2. Read the test assertion - Is the sub-agent's interpretation correct?
  3. Check the classification - Does TEST_BUG vs SOLUTION_BUG make sense?
  4. Verify the proposed fix - Is it actually more lenient, or just different?

Common Sub-Agent Mistakes to Catch:

MistakeHow to Spot It
Wrong classificationSpec clearly requires X, but agent says TEST_BUG
Invented spec requirementsAgent quotes spec but adds interpretation not present
Missed spec textAgent says "spec silent" but spec does address it
Bad fix proposal"Fix" changes behavior instead of relaxing constraint
Inconsistent verdictsSUMMARY counts don't match FINDINGS

Verification Process:

For each report with TESTS_HAVE_BUGS or MIXED verdict:

1. Open the spec file: problems/{problem}/checkpoint_N.md
2. Open the test file: problems/{problem}/tests/test_checkpoint_N.py
3. For each TEST_BUG finding:
   a. Find the spec quote - is it accurate and complete?
   b. Find the test assertion - does it really expect what agent claims?
   c. Is the classification correct?
   d. Is the proposed fix actually valid?
4. Correct any errors before including in final summary

If You Find Sub-Agent Errors:

  • Override the classification in your summary
  • Note the correction so user knows agent was wrong
  • Provide correct fix if agent's fix was wrong
  • Adjust counts in the summary table

Example correction note:

### Corrections to Sub-Agent Reports

- checkpoint_2 FINDING-003: Reclassified from TEST_BUG → SOLUTION_BUG
  (Agent missed spec requirement on line 45: "MUST return exactly 1")

Step 5: Generate Summary Report

Present a consolidated summary to the user:

# Run Validation Summary

**Run Path**: {run_path}
**Date**: {date}

## Overall Results

| Problem | Checkpoint | Verdict | Failing | Solution Bugs | Test Bugs |
|---------|------------|---------|---------|---------------|-----------|
| file_backup | 1 | SOLUTION_CORRECT | 0 | 0 | 0 |
| file_backup | 2 | TESTS_HAVE_BUGS | 3 | 0 | 3 |
| file_backup | 3 | MIXED | 5 | 2 | 3 |

## Verdicts by Type

- SOLUTION_CORRECT: N checkpoints
- SOLUTION_HAS_BUGS: N checkpoints
- TESTS_HAVE_BUGS: N checkpoints
- MIXED: N checkpoints
- SPEC_AMBIGUOUS: N checkpoints

## Action Items Required

<!-- Fix preference: MAKE_LENIENT > REMOVE_TEST >> UPDATE_SPEC -->

### Test Fixes: MAKE_LENIENT (Preferred)
- [ ] file_backup checkpoint_2 `test_error_code`: Change `assert rc == 1` → `assert rc != 0`
- [ ] file_backup checkpoint_3 `test_output_order`: Change `assert x == [a,b]` → `assert set(x) == {a,b}`

### Test Fixes: REMOVE_TEST
- [ ] file_backup checkpoint_2 `test_undefined_edge`: DELETE (tests undefined behavior)

### Solution Fixes
- [ ] file_backup checkpoint_3: {specific fix from report}

### Spec Clarifications (LAST RESORT)
- [ ] {only if test cannot be made lenient or removed}

## Corrections to Sub-Agent Reports

<!-- Include this section if you found any errors in sub-agent analysis -->

- checkpoint_2 FINDING-003: Reclassified TEST_BUG → SOLUTION_BUG
  - Agent claimed: "spec is silent on error codes"
  - Actually: spec line 45 says "MUST return exit code 1"
  - Correct fix: Solution must return 1, not test change needed

## Detailed Reports

Reports saved to:
- problems/file_backup/checkpoint_1_report.md
- problems/file_backup/checkpoint_2_report.md
- problems/file_backup/checkpoint_3_report.md

Quick Verdicts Reference

VerdictMeaningAction
SOLUTION_CORRECTAll good, solution passesNone
SOLUTION_HAS_BUGSSolution needs fixesFix solution code
TESTS_HAVE_BUGSTests are wrongFix test assertions
MIXEDBoth have issuesFix both
SPEC_AMBIGUOUSSpec unclearClarify spec first

Fix Preference Hierarchy

When tests have bugs, prefer fixes in this order:

1. MAKE_LENIENT  ← STRONGLY PREFERRED (relax assertion)
2. REMOVE_TEST   ← If test is fundamentally flawed
3. UPDATE_SPEC   ← LAST RESORT ONLY
  • Making tests lenient keeps coverage while accepting valid implementations
  • Removing tests is better than keeping broken ones
  • Changing specs is dangerous and affects all documentation

Recommended Tools for Lenient Tests

Many strict tests can be fixed using these tools:

ToolUse CaseExample
deepdiffIgnore order, precision, extra fieldsDeepDiff(a, b, ignore_order=True)
jsonschemaValidate structure not exact valuesvalidate(output, schema)
normalize()Handle whitespace, case, key ordernormalize(actual) == normalize(expected)
# Common fix pattern:
from deepdiff import DeepDiff

# BEFORE (too strict):
assert actual == expected

# AFTER (lenient):
diff = DeepDiff(expected, actual, ignore_order=True, significant_digits=5)
assert not diff

Example

User: /validate-run problems/file_backup/solutions file_backup

Assistant: [discovers checkpoints, launches test-validator agents in parallel]

Assistant: All validators complete. Here is the summary:

# Run Validation Summary
...

When not to use it

  • When sub-agent results are blindly trusted without verification
  • When the problem or checkpoint path is not correctly specified
  • When the goal is to change specs as a primary fix for test bugs

Limitations

  • Sub-agents (Sonnet) make mistakes, requiring verification.
  • It requires the spec and test files to be available for verification.
  • Changing specs is considered a last resort for fixing test bugs.

How it compares

This skill automates the parallel validation of agent run checkpoints and includes a critical human verification step for sub-agent reports, unlike manual, sequential validation or unverified automated reports.

Compared to similar skills

validate-run side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
validate-run (this skill)07moReviewAdvanced
bats97moReviewIntermediate
browser-daemon59moReviewIntermediate
examples-auto-run23moReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

bats

OleksandrKucherenko

Bash Automated Testing System (BATS) for TDD-style testing of shell scripts. Use when: (1) Writing unit or integration tests for Bash scripts, (2) Testing CLI tools or shell functions, (3) Setting up test infrastructure with setup/teardown hooks, (4) Mocking external commands (curl, git, docker), (5) Generating JUnit reports for CI/CD, (6) Debugging test failures or flaky tests, (7) Implementing test-driven development for shell scripts.

991

browser-daemon

noiv

Persistent browser automation via Playwright daemon. Keep a browser window open and send it commands (navigate, execute JS, inspect console). Perfect for interactive debugging, development, and testing web applications. Use when you need to interact with a browser repeatedly without opening/closing it.

587

examples-auto-run

openai

Run python examples in auto mode with logging, rerun helpers, and background control.

23

workflow-test-fix-cycle

catlog22

End-to-end test-fix workflow generate test sessions with progressive layers (L0-L3), then execute iterative fix cycles until pass rate >= 95%. Combines test-fix-gen and test-cycle-execute into a unified pipeline. Triggers on "workflow:test-fix-cycle".

12

qoobee-t&f-skill

Piaoxuemoli

Test & Fix skill for ImmerseAI. Supports automated testing (Agent runs code audits and auto-fixes) and manual test report fixing (Agent parses human test reports and fixes failures). Trigger with: 自动化测试, 自动测试, auto test, 人工测试, manual test, 测试报告, test report, 运行测试, run tests.

00

gsd-audit-fix

xai-bookkeeping

Autonomous audit-to-fix pipeline — find issues, classify, fix, test, commit

00

Search skills

Search the agent skills registry