gsd-eval-review
Evaluates the effectiveness of previous AI-driven work and creates a plan for improvements.
Install
mkdir -p .claude/skills/gsd-eval-review && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/13061" && unzip -o skill.zip -d .claude/skills/gsd-eval-review && rm skill.zipInstalls to .claude/skills/gsd-eval-review
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Audit an executed AI phase's evaluation coverage and produce an EVAL-REVIEW.md remediation plan.Key capabilities
- →Conduct a retroactive evaluation coverage audit.
- →Check whether the evaluation strategy from AI-SPEC.md was implemented.
- →Produce EVAL-REVIEW.md with score, verdict, gaps, and remediation plan.
- →Gather context for the audit.
- →Preserve all workflow gates.
- →Execute end-to-end.
How it works
The skill conducts a retroactive evaluation coverage audit of a completed AI phase, checking implementation against AI-SPEC.md, and then produces an `EVAL-REVIEW.md` remediation plan.
Inputs & outputs
When to use gsd-eval-review
- →Review AI performance
- →Check test coverage for a completed task
- →Plan remediation for bugs
About this skill
<codex_skill_adapter>
A. Skill Invocation
- This skill is invoked by mentioning
$gsd-eval-review. - Treat all user text after
$gsd-eval-reviewas{{GSD_ARGS}}. - If no arguments are present, treat
{{GSD_ARGS}}as empty.
B. AskUserQuestion → request_user_input Mapping
GSD workflows use AskUserQuestion (Claude Code syntax). Translate to Codex request_user_input:
Parameter mapping:
header→headerquestion→question- Options formatted as
"Label" — description→{label: "Label", description: "description"} - Generate
idfrom header: lowercase, replace spaces with underscores
Batched calls:
AskUserQuestion([q1, q2])→ singlerequest_user_inputwith multiple entries inquestions[]
Multi-select workaround:
- Codex has no
multiSelect. Use sequential single-selects, or present a numbered freeform list asking the user to enter comma-separated numbers.
Execute mode fallback:
- When
request_user_inputis rejected or unavailable, you MUST stop and present the questions as a plain-text numbered list, then wait for the user's reply. Do NOT pick a default and continue (#3018). - You may only proceed without a user answer when one of these is true:
(a) the invocation included an explicit non-interactive flag (
--autoor--all), (b) the user has explicitly approved a specific default for this question, or (c) the workflow's documented contract says defaults are safe (e.g. autonomous lifecycle paths). - Do NOT write workflow artifacts (CONTEXT.md, DISCUSSION-LOG.md, PLAN.md, checkpoint files) until the user has answered the plain-text questions or one of (a)-(c) above applies. Surfacing the questions and waiting is the correct response — silently defaulting and writing artifacts is the #3018 failure mode.
C. Task() → spawn_agent Mapping
GSD workflows use Task(...) (Claude Code syntax). Translate to Codex collaboration tools:
Direct mapping:
Task(subagent_type="X", prompt="Y")→spawn_agent(agent_type="X", message="Y")Task(model="...")→ omit.spawn_agenthas no inlinemodelparameter; GSD embeds the resolved per-agent model directly into each agent's.tomlat install time somodel_overridesfrom.planning/config.jsonand~/.gsd/defaults.jsonare honored automatically by Codex's agent router.fork_context: falseby default — GSD agents load their own context via<files_to_read>blocksTask(isolation="worktree")/Agent(isolation="worktree")→ no direct Codex mapping. Codexspawn_agentdoes not create or bind a git worktree automatically. Workflows that require this isolation must fail closed or use an explicit manual worktree protocol before spawning (#3360).
Spawn restriction:
- Codex restricts
spawn_agentto cases where the user has explicitly requested sub-agents. When automatic spawning is not permitted, do the work inline in the current agent rather than attempting to force a spawn.
Parallel fan-out:
- Spawn multiple agents → collect agent IDs →
wait(ids)for all to complete
Result parsing:
- Look for structured markers in agent output:
CHECKPOINT,PLAN COMPLETE,SUMMARY, etc. close_agent(id)after collecting results from each agent </codex_skill_adapter>
<execution_context> @/Users/lmarques/Dev/efx-mux/.codex/get-shit-done/workflows/eval-review.md @/Users/lmarques/Dev/efx-mux/.codex/get-shit-done/references/ai-evals.md </execution_context>
<context> Phase: {{GSD_ARGS}} — optional, defaults to last completed phase. </context> <process> Execute end-to-end. Preserve all workflow gates. </process>When not to use it
- →When `request_user_input` is rejected or unavailable and the user has not approved a default.
Limitations
- →It requires user answers for plain-text questions if `request_user_input` is unavailable.
- →It restricts `spawn_agent` to cases where the user has explicitly requested sub-agents.
- →Codex `spawn_agent` does not create or bind a git worktree automatically.
How it compares
This skill automates the audit of AI phase evaluation coverage and generates a structured remediation plan, ensuring that evaluation strategies are consistently checked and gaps are documented, unlike a manual review process.
Compared to similar skills
gsd-eval-review side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| gsd-eval-review (this skill) | 0 | 3mo | No flags | Intermediate |
| flow-next-prime | 0 | 2mo | Review | Intermediate |
| skill-creator | 128 | 3mo | Review | Advanced |
| ralph | 17 | 2mo | No flags | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by electroheadfx
View all by electroheadfx →You might also like
flow-next-prime
gmickel
Comprehensive codebase assessment for agent and production readiness. Scans 8 pillars (48 criteria), verifies commands work, checks GitHub settings. Reports everything, fixes agent readiness only. Triggers on /flow-next:prime.
skill-creator
anthropics
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Claude's capabilities with specialized knowledge, workflows, or tool integrations.
ralph
Yeachan-Heo
Self-referential loop until task completion with architect verification
requesting-code-review
obra
Use when completing tasks, implementing major features, or before merging to verify work meets requirements
develop-ai-functions-example
vercel
Develop examples for AI SDK functions. Use when creating, running, or modifying examples under examples/ai-functions/src to validate provider support, demonstrate features, or create test fixtures.
survey-sdk-audit
PostHog
Audit PostHog survey SDK features and version requirements