evaluation-anchor-checker
Ensures numeric performance claims are contextually accurate and reviewer-safe.
Install
mkdir -p .claude/skills/evaluation-anchor-checker && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/12502" && unzip -o skill.zip -d .claude/skills/evaluation-anchor-checker && rm skill.zipInstalls to .claude/skills/evaluation-anchor-checker
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Audit and rewrite evaluation/numeric claims to ensure they carry minimal protocol context (task + metric + constraint) and avoid underspecified model naming. **Trigger**: evaluation anchor checker, numeric claim hygiene, underspecified numbers, protocol context, 评测锚点检查, 数字断言, 指标上下文. **Use when**: before final merge/polish, or when reviewers would likely flag claims as underspecified (numbers without task/metric/budget), or `pipeline-auditor` warns about suspicious model naming. **Skip if**: evidence is too thin to justify numeric claims (route upstream to C3/C4), or you are pre-C2 (NO PROSE). **Network**: none. **Guardrail**: do not invent numbers; do not add/remove/move citation keys; if protocol context is missing, weaken/remove the numeric claim rather than guessing.Key capabilities
- →Audit numeric claims in technical surveys
- →Ensure protocol context for numeric statements
- →Weaken or remove underspecified numeric claims
- →Generate an evaluation anchor report
- →Validate citation keys
How it works
This skill reviews numeric claims in technical documents, ensuring each claim includes sufficient protocol context (task, metric, constraint) or is downgraded if context is missing.
Inputs & outputs
When to use evaluation-anchor-checker
- →Check evaluation claims
- →Audit performance report
- →Improve numeric hygiene
About this skill
Evaluation Anchor Checker (make numbers reviewer-safe)
Purpose: fix a reviewer-magnet failure mode in agent surveys:
- strong numeric/performance statements appear
- but the minimal evaluation context is missing
This skill treats numeric claims as contracts:
- if a number stays, the same sentence must contain enough protocol context to interpret it
- if that context is not in evidence, the claim must be downgraded (no guessing)
Inputs
Preferred (pre-merge, keeps anchoring intact):
- the affected
sections/*.mdfiles
Optional context (read-only; helps you avoid guessing):
outline/writer_context_packs.jsonl(look forevaluation_anchor_minimal,evaluation_protocol,anchor_facts)outline/evidence_drafts.jsonl/outline/anchor_sheet.jsonlcitations/ref.bib
Outputs
- Updated
sections/*.md(oroutput/DRAFT.mdif you are post-merge), with safer evaluation anchoring output/EVAL_ANCHOR_REPORT.md(always; short report with files checked / changed / weakened sentences)- Optional completion marker:
output/eval_anchors_checked.refined.ok
Recommended slot in the survey pipeline
Use this as the last section-level numeric hygiene sweep before merge:
- after
style-harmonizer,opener-variator,section-logic-polisher, andparagraph-curator - immediately before the final
argument-selfloopsnapshot and merge
Reason:
- earlier section-level rewrite passes can legitimately rephrase or fuse numeric sentences
- if you only wait for
pipeline-auditor, numeric-context issues are discovered too late in the merged draft - section-scoped fixes are cheaper and preserve citation anchoring better than post-merge patching
Read Order
Always read:
references/numeric_hygiene.md
Machine-readable asset:
assets/numeric_hygiene.json
The asset defines the keyword families and qualitative fallback templates. Keep the script deterministic and let the policy live in the asset/reference pair.
Role prompt: Reviewer-minded Editor (evaluation hygiene)
You are a reviewer-minded editor for evaluation claims in a technical survey.
Goal:
- make every numeric/performance claim interpretable and reviewer-safe
Hard constraints:
- do not invent numbers
- do not add/remove/move citation keys
- if protocol context is missing, weaken or remove the numeric claim
Minimum context to include when keeping a number:
- task / setting (what kind of task)
- metric (what is being measured)
- constraint (budget/cost/tool access/horizon/seed/logging) when relevant
Avoid:
- ambiguous model naming that looks hallucinated (e.g., “GPT-5”) unless the cited paper uses it verbatim
Workflow (explicit inputs)
- Use
outline/writer_context_packs.jsonlto locate the subsection's allowed citations and any extractedevaluation_protocol/anchor_facts. - Cross-check
outline/evidence_drafts.jsonlandoutline/anchor_sheet.jsonlfor task/metric/constraint context before touching numbers. - Validate every cited key against
citations/ref.bib(do not introduce new keys). - Write
output/EVAL_ANCHOR_REPORT.mdso the pipeline has an auditable completion artifact for this sweep.
What to enforce (the “minimum protocol trio”)
When a sentence contains digits (%, x, or numbers):
- Keep the number only if you can attach at least 2 of the following in the same sentence without guessing:
- task family / benchmark name
- metric definition
- constraint (budget, tool access, cost model, retries, horizon)
If you cannot, downgrade:
- remove the number and rewrite as qualitative (“often”, “can”, “may”) with the same citation
- or move the specificity into a verification target (“evaluations need to report …”) without adding new facts
Mini examples (paraphrase; do not copy)
Bad (underspecified):
Model X achieves 75% exact performance [@SomeBench].
Better (minimal context):
On <task/benchmark>, Model X reaches ~75% <metric>, under <constraint/budget/tool access> [@SomeBench].
Better (downgrade when context is missing):
Reported gains vary, but comparisons remain fragile when budgets and retry policies are not reported [@SomeBench].
Done checklist
-
output/EVAL_ANCHOR_REPORT.mdexists and reports a non-zero file count. - No numeric claim remains without minimal protocol context.
- No ambiguous model naming remains unless explicitly supported by citations.
- Citation keys are unchanged.
- If you removed/downgraded numbers, the paragraph still makes a defensible, evidence-bounded point.
Script
Quick Start
uv run python .codex/skills/evaluation-anchor-checker/scripts/run.py --workspace <workspace>
All Options
--workspace <dir>: workspace containingsections/*.mdor merged draft artifacts--unit-id <id>: optional harness metadata--inputs <semicolon-separated>: optional override fromUNITS.csv--outputs <semicolon-separated>: optional output override; default includesoutput/EVAL_ANCHOR_REPORT.md--checkpoint <C*>: optional harness metadata
Examples
- Run the numeric hygiene sweep before merge:
uv run python .codex/skills/evaluation-anchor-checker/scripts/run.py --workspace <workspace> --inputs 'sections/*.md;outline/writer_context_packs.jsonl;citations/ref.bib' --outputs 'sections/*.md;output/EVAL_ANCHOR_REPORT.md;output/eval_anchors_checked.refined.ok'
When not to use it
- →If evidence is too thin to justify numeric claims
- →If you are pre-C2 (no prose)
Limitations
- →Does not invent numbers
- →Does not add, remove, or move citation keys
- →Weakens or removes claims if protocol context is missing
How it compares
This workflow systematically enforces evaluation hygiene by requiring explicit context for numeric claims, unlike a manual review that might overlook underspecified numbers.
Compared to similar skills
evaluation-anchor-checker side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| evaluation-anchor-checker (this skill) | 0 | 2mo | Review | Intermediate |
| tech-debt-tracker | 1 | 2mo | Review | Intermediate |
| quality-engineering-zephyr-coverage-analysis | 0 | 2mo | No flags | Intermediate |
| quality | 0 | 5mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by WILLOSCAR
View all by WILLOSCAR →You might also like
tech-debt-tracker
alirezarezvani
Scan codebases for technical debt, score severity, track trends, and generate prioritized remediation plans. Use when users mention tech debt, code quality, refactoring priority, debt scoring, cleanup sprints, or code health assessment. Also use for legacy code modernization planning and maintenance cost estimation.
quality-engineering-zephyr-coverage-analysis
HoangNguyen0403
Audit test coverage health, gaps, and QE debt for Jira stories or epics. Produces coverage_analysis_report.md with AC-to-TC heatmap, risk scores, and prioritized action plan. Use when assessing coverage percentage, pre-release readiness, sprint readiness, or identifying missing test cases. Do NOT us
quality
dpaola2
Generate a code quality report for a pipeline project by reading quality frontmatter from progress.md and optionally running fresh analysis.
code-stats
0xDarkMatter
Analyze codebase with tokei (fast line counts by language) and difft (semantic AST-aware diffs). Get quick project overview without manual counting. Triggers on: how big is codebase, count lines of code, what languages, show semantic diff, compare files, code statistics.
dbt-transformation-patterns
wshobson
Master dbt (data build tool) for analytics engineering with model organization, testing, documentation, and incremental strategies. Use when building data transformations, creating data models, or implementing analytics engineering best practices.
analyse-issue
monarch-initiative
Analyze MONDO GitHub issues for validity, suggest improvements, and generate structured reports with duplication checks and identifier validation