GS

Evaluates the effectiveness of previous AI-driven work and creates a plan for improvements.

Install

mkdir -p .claude/skills/gsd-eval-review && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/13061" && unzip -o skill.zip -d .claude/skills/gsd-eval-review && rm skill.zip

Installs to .claude/skills/gsd-eval-review

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Audit an executed AI phase's evaluation coverage and produce an EVAL-REVIEW.md remediation plan.
96 charsno explicit “when” trigger
Intermediate

Key capabilities

  • Conduct a retroactive evaluation coverage audit.
  • Check whether the evaluation strategy from AI-SPEC.md was implemented.
  • Produce EVAL-REVIEW.md with score, verdict, gaps, and remediation plan.
  • Gather context for the audit.
  • Preserve all workflow gates.
  • Execute end-to-end.

How it works

The skill conducts a retroactive evaluation coverage audit of a completed AI phase, checking implementation against AI-SPEC.md, and then produces an `EVAL-REVIEW.md` remediation plan.

Inputs & outputs

You give it
Optional phase argument, defaulting to the last completed phase.
You get back
An `EVAL-REVIEW.md` remediation plan with score, verdict, and identified gaps.

When to use gsd-eval-review

  • Review AI performance
  • Check test coverage for a completed task
  • Plan remediation for bugs

About this skill

<codex_skill_adapter>

A. Skill Invocation

  • This skill is invoked by mentioning $gsd-eval-review.
  • Treat all user text after $gsd-eval-review as {{GSD_ARGS}}.
  • If no arguments are present, treat {{GSD_ARGS}} as empty.

B. AskUserQuestion → request_user_input Mapping

GSD workflows use AskUserQuestion (Claude Code syntax). Translate to Codex request_user_input:

Parameter mapping:

  • headerheader
  • questionquestion
  • Options formatted as "Label" — description{label: "Label", description: "description"}
  • Generate id from header: lowercase, replace spaces with underscores

Batched calls:

  • AskUserQuestion([q1, q2]) → single request_user_input with multiple entries in questions[]

Multi-select workaround:

  • Codex has no multiSelect. Use sequential single-selects, or present a numbered freeform list asking the user to enter comma-separated numbers.

Execute mode fallback:

  • When request_user_input is rejected or unavailable, you MUST stop and present the questions as a plain-text numbered list, then wait for the user's reply. Do NOT pick a default and continue (#3018).
  • You may only proceed without a user answer when one of these is true: (a) the invocation included an explicit non-interactive flag (--auto or --all), (b) the user has explicitly approved a specific default for this question, or (c) the workflow's documented contract says defaults are safe (e.g. autonomous lifecycle paths).
  • Do NOT write workflow artifacts (CONTEXT.md, DISCUSSION-LOG.md, PLAN.md, checkpoint files) until the user has answered the plain-text questions or one of (a)-(c) above applies. Surfacing the questions and waiting is the correct response — silently defaulting and writing artifacts is the #3018 failure mode.

C. Task() → spawn_agent Mapping

GSD workflows use Task(...) (Claude Code syntax). Translate to Codex collaboration tools:

Direct mapping:

  • Task(subagent_type="X", prompt="Y")spawn_agent(agent_type="X", message="Y")
  • Task(model="...") → omit. spawn_agent has no inline model parameter; GSD embeds the resolved per-agent model directly into each agent's .toml at install time so model_overrides from .planning/config.json and ~/.gsd/defaults.json are honored automatically by Codex's agent router.
  • fork_context: false by default — GSD agents load their own context via <files_to_read> blocks
  • Task(isolation="worktree") / Agent(isolation="worktree") → no direct Codex mapping. Codex spawn_agent does not create or bind a git worktree automatically. Workflows that require this isolation must fail closed or use an explicit manual worktree protocol before spawning (#3360).

Spawn restriction:

  • Codex restricts spawn_agent to cases where the user has explicitly requested sub-agents. When automatic spawning is not permitted, do the work inline in the current agent rather than attempting to force a spawn.

Parallel fan-out:

  • Spawn multiple agents → collect agent IDs → wait(ids) for all to complete

Result parsing:

  • Look for structured markers in agent output: CHECKPOINT, PLAN COMPLETE, SUMMARY, etc.
  • close_agent(id) after collecting results from each agent </codex_skill_adapter>
<objective> Conduct a retroactive evaluation coverage audit of a completed AI phase. Checks whether the evaluation strategy from AI-SPEC.md was implemented. Produces EVAL-REVIEW.md with score, verdict, gaps, and remediation plan. </objective>

<execution_context> @/Users/lmarques/Dev/efx-mux/.codex/get-shit-done/workflows/eval-review.md @/Users/lmarques/Dev/efx-mux/.codex/get-shit-done/references/ai-evals.md </execution_context>

<context> Phase: {{GSD_ARGS}} — optional, defaults to last completed phase. </context> <process> Execute end-to-end. Preserve all workflow gates. </process>

When not to use it

  • When `request_user_input` is rejected or unavailable and the user has not approved a default.

Limitations

  • It requires user answers for plain-text questions if `request_user_input` is unavailable.
  • It restricts `spawn_agent` to cases where the user has explicitly requested sub-agents.
  • Codex `spawn_agent` does not create or bind a git worktree automatically.

How it compares

This skill automates the audit of AI phase evaluation coverage and generates a structured remediation plan, ensuring that evaluation strategies are consistently checked and gaps are documented, unlike a manual review process.

Compared to similar skills

gsd-eval-review side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
gsd-eval-review (this skill)03moNo flagsIntermediate
flow-next-prime02moReviewIntermediate
skill-creator1283moReviewAdvanced
ralph172moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

Search skills

Search the agent skills registry