repo-stage-review-loop
Provides a formal review process for plans and implementation code, ensuring changes meet the agreed contract.
Install
mkdir -p .claude/skills/repo-stage-review-loop && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/13572" && unzip -o skill.zip -d .claude/skills/repo-stage-review-loop && rm skill.zipInstalls to .claude/skills/repo-stage-review-loop
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Use when a RepoPilot plan or implementation needs formal review, when final code changed after earlier review, or when external findings require evidence-based triage.Key capabilities
- →Review the plan contract before implementation.
- →Review the final implementation state against the approved contract after code changes.
- →Report severity-ordered findings with file/line evidence.
- →Classify external findings as fix, clarify, reject, or defer.
- →Perform a focused Stage Debt Sweep over changed paths.
How it works
This skill systematically reviews plans or code against an approved contract, identifying issues by severity and providing evidence, then triages external feedback.
Inputs & outputs
When to use repo-stage-review-loop
- →Review proposed architecture plans
- →Validate final code changes against specs
- →Triage external feedback into fix/clarify/reject
About this skill
Repo Stage Review Loop
Core Rule
Review the plan contract before implementation when requested, and review the final implementation state against the approved contract after code changes. Passing tests and completed tasks are inputs to review, not substitutes for it.
Plan Contract Review
For plan review, read the proposed plan or active OpenSpec artifacts, Harness
boundaries, relevant specs, and directly implicated runtime/docs. Report
severity-ordered findings against intent, scope, non-goals, test plan,
review gates, and roadmap truth. Medium/high plans require internal plan review
plus two independent plan-review slots before implementation. Each first-round
slot must use a distinct reviewer with no inherited implementation context or
other first-round conclusion; Codex may use an empty-context task or
fork_turns="none". Inherited or unknown context keeps the slot open.
Review Loop
- Read the active OpenSpec contract, changed files or plan contract, tests, allowed paths, and current review checklist.
- Confirm the review occurs after the latest runtime/test change.
- Review in layers: scope, business logic, architecture boundary, minimality, failure semantics, security/privacy, test adequacy, and maintainability. For user-facing summaries in Chinese, keep precise English terms and add a short Chinese explanation or concrete example when the term is non-obvious.
- Report severity-ordered findings with file/line evidence, trigger, consequence, and missing regression coverage. If there are no findings, state inspected areas and residual risk.
- Use
external-review-triagefor external findings. Classify each asfix,clarify,reject, ordefer; never accept it by authority alone. - After remediation, rerun affected verification and review changed behavior. A same-slot remediation re-review may reuse the original reviewer for finding lineage, but every required slot must refresh to the same final content-addressed baseline.
- Materialize
.harness/reviews/<stage-id>/<phase>/review-set.jsonand runpython scripts/validate_independent_review.py --project-root . --receipt-set <path> --expected-stage <stage-id> --expected-phase <plan|implementation> --required-slots <count>. Missing receipt evidence, skipped execution, or nonzero exit keeps the gate open. A zero exit is mechanical consistency only (gate_ready=false); verify host-native dispatch provenance and pre-change-authority activation sequence separately before counting slots. - For final implementation review, consume the canonical byte-stable reviewed-change manifest and bounded diff derived from the planning base. The review subject excludes exactly four metadata paths: the manifest, diff, final review set, and delivery binding. No other path is implicitly excluded. Every required slot must bind the same host-retained final packet hash and exactly every existing manifest subject plus the manifest and deterministic inventory tail; any omission, arbitrary inventory, fifth metadata path, or non-tail change after the packet reopens review. For replay-capable assets, include all pre-tail event/receipt projections, adapter contracts, dormant activation wording, v1 cohort compatibility, and v2 template/validator bytes in the normal reviewed subject. A new material event or replay write after packet freeze is a non-tail change and reopens verification/review/archive; it cannot become a third evidence-tail file.
- Perform a focused Stage Debt Sweep over changed paths and directly dependent older paths. Record inspected paths, concrete findings, dispositions, and residual debt.
- Block archive when tasks are unchecked, review evidence is stale, validation failed, blocking findings remain, or delta operations do not match long-term specs.
Review Priorities
- stage scope and user-visible behavior match the approved contract
- business logic, state transitions, and failure paths match intended semantics
- code remains in the correct architectural layer and reuses existing boundaries
- functions/classes stay minimal enough to avoid hidden behavior coupling
- public contract and state-transition correctness
- fail-closed permissions, approval, identity, path, and lifecycle checks
- interruption, retry, rollback, and reconciliation behavior
- tests that assert the intended contract rather than implementation details
- scope drift and accidental roadmap capability claims
- stale assumptions in directly dependent older paths
External review should seek independent counterexamples, especially for medium/high-risk stages. Repeating the task checklist is not useful diversity. Reviewer tasks/subagents are development workflow adapters, not RepoPilot runtime capabilities. Final implementation review uses the risk-contract required-slot count; the two-slot rule is specific to medium/high plan review.
Review receipts and packet hashes never establish live human authority. Archive
must additionally pass the shared stage-authority archive preflight; merge and
push must bind the same final manifest/review set and host-retained exact
candidate under the controller workflow.
Replay reports likewise prove only mechanical_consistency_only. Review must
reject any claim that repository bytes activate v2, that an in-flight v1 stage
may change cohort, that replay PASS authorizes/blocks a v1 mutation, or that an
activated-v2 governed action may bypass exact-frontier equality because it is
called unaffected. Real host CAS, restart, dispatch and activation facts remain
external prerequisites.
Evidence Boundary
Store gate evidence in .harness/review_checklist.md. Store durable unresolved
debt in docs/PROGRESS.md. Put only next-session blockers in
HANDOFF_TO_NEXT_CHAT.md.
Do not perform merge/push handoff here. Return to repo-stage-workflow, which
uses repo-stage-handoff after integration.
Evals
Use references/evals.md when changing routing or review gates.
When not to use it
- →When the task is not related to formal review of plans or code implementation.
- →When the task is to perform merge/push handoff.
Limitations
- →The skill does not perform merge/push handoff.
- →The skill requires the review to occur after the latest runtime/test change.
How it compares
This skill enforces a multi-layered review process with severity-ordered findings and external feedback triage, unlike a simple checklist review.
Compared to similar skills
repo-stage-review-loop side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| repo-stage-review-loop (this skill) | 0 | 3mo | No flags | Advanced |
| test-plan | 1 | 7mo | No flags | Intermediate |
| flow-next-prime | 0 | 3mo | Review | Intermediate |
| gate-review | 0 | 7mo | No flags | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
You might also like
test-plan
quran
Generates a comprehensive testing plan based on the current branch changes or a specific PR. Use when creating QA checklists, test plans, or verifying PR readiness.
flow-next-prime
gmickel
Comprehensive codebase assessment for agent and production readiness. Scans 8 pillars (48 criteria), verifies commands work, checks GitHub settings. Reports everything, fixes agent readiness only. Triggers on /flow-next:prime.
gate-review
live-input-vector-output-node
Verify planned or implemented changes against active rules, scope lock, and required checks.
fmea-analysis
ddunnock
Conduct Failure Mode and Effects Analysis (FMEA) for systematic identification and risk assessment of potential failures in designs, processes, or systems. Supports DFMEA (Design), PFMEA (Process), and FMEA-MSR (Monitoring & System Response). Uses AIAG-VDA 7-step methodology with Action Priority (AP
spec-driven-qa
bankielewicz
>
quality-gate
raddue
Iterative red-teaming of any artifact (design docs, plans, code, hypotheses, mockups). Loops until clean or stagnation. Invoked by artifact-producing skills or their parent orchestrator.