docs-cleaner
Merges duplicate documentation files to improve clarity and reduce project clutter.
Install
mkdir -p .claude/skills/docs-cleaner && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2818" && unzip -o skill.zip -d .claude/skills/docs-cleaner && rm skill.zipInstalls to .claude/skills/docs-cleaner
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Consolidates redundant documentation while preserving all valuable content. This skill should be used when users want to clean up documentation bloat, merge redundant docs, reduce documentation sprawl, or consolidate multiple files covering the same topic. Triggers include "clean up docs", "consolidate documentation", "too many doc files", "merge these docs", or when documentation exceeds 500 lines across multiple files covering similar topics.Key capabilities
- →Identify documentation files covering a specific topic
- →Map content overlap between multiple documents
- →Create section-by-section value analysis tables
- →Propose target structure for consolidated documentation
- →Update references in CLAUDE.md and README files
- →Verify link integrity after file deletion
How it works
It performs a discovery phase to map overlaps, followed by a value analysis to categorize sections as keep, condense, or delete. Finally, it executes the consolidation and updates all internal references.
Inputs & outputs
When to use docs-cleaner
- →Merging duplicate readme files
- →Reducing documentation sprawl
- →Consolidating project wikis
About this skill
Docs Cleaner
Documentation goes wrong in recurring ways, and they need different work.
It rots. Something changed — a port, a path, a procedure, a decision — and the docs still describe the old world. The field's name for this is documentation rot or version drift: artifacts fall out of sync one at a time until people stop trusting the set.
It sprawls. The same topic accumulates three files, each partly right, none authoritative.
They trace to one root cause: a fact was written down in more places than it was defined. Everything below follows from fixing that.
Core principle (governs both modes)
Critical evaluation before deletion. Never blindly delete. Analyze each section's unique value before proposing removal. The goal is reduction without information loss.
This governs every deletion-shaped action in this skill, not just consolidation: archiving a document, dropping a derived value, and removing a cross-reference are all deletions and all owe the same analysis first. "It looked redundant" is not that analysis.
Which mode are you in?
| The user said / the situation | Mode |
|---|---|
| "I changed X" / a diff exists / a procedure was updated | 1 — Post-change governance |
| "which docs are now wrong", "check for stale commands/paths" | 1 — Post-change governance |
| "clean up the docs", "too many files on this", "merge these" | 2 — Consolidation |
| A topic is spread over several files with no clear owner | 2 — Consolidation |
They share the Drift Test below. When both apply, run Mode 1 first: consolidating docs that are also factually stale just produces one confidently wrong document.
A note on CLAUDE.md / AGENTS.md: both modes apply to them like any other document
— govern them after a change, consolidate them when they duplicate each other. The one job
that belongs elsewhere is restructuring a single oversized one — pushing low-frequency
detail down into references while keeping the top level lean. That is what
daymade-claude-code's claude-md-progressive-disclosurer skill is built for; prefer it
when the problem is size and layering rather than truth or duplication. If it is not
available, Mode 2 still handles the file — you just do the value analysis by hand.
The Drift Test (the core of both modes)
Before writing any value into a document — and when deciding whether an existing one earns its place — ask these three questions in order, and stop at the first hit:
1. Can this be computed from details already recorded here or in the authoritative source? → It is a derived value. Do not write it. Compute it when someone asks. Examples: a count ("6 sections", "N groups covered"), a total, a status summarizing rows in a table below it, "last updated on" when the log underneath already says.
When you remove one, repair the sentence rather than leaving a hole: "the four checks below" becomes "the checks below". Deleting the number and leaving dangling grammar is a different defect, not a fix.
Exception — generated navigation. A table of contents restating the headings is a derived value by this test, and hand-maintained ones should indeed go. But leave it alone when it is produced by the docs toolchain at build time, or when the platform requires the marker to render navigation: those are computed on demand, which is what question 1 asks for. Check for a generator before deleting one.
2. Not computable, but the same fact is authoritatively defined somewhere else? → It is a copy. Link to that definition. Do not restate the value.
3. Neither — it records what was true at a particular moment? → Now you may write it down. Price history rows, decision log entries, "as of 2026-03 the vendor required X" — these are the legitimate case, and deleting them destroys an audit trail.
Question 3 needs care, because every stale value was true at some moment — that is what made it stale. But note what question 3 does and does not decide.
It decides one thing: this passage is not a drift problem. It is not a derived value and not a copy, so none of the drift remedies apply to it — do not recompute it, do not link it away, do not "update" it to the current value. Whether the passage is worth keeping at all is a content-value question, and it is answered in Mode 2's value analysis with the document's owner, not here.
Three earlier drafts of this section tried to decide keep-vs-delete right here, and each one failed review. The reason is worth stating so nobody rebuilds it: every version needed the agent to know something it does not have — whether the content could be regenerated, at what cost, by whom, or whether the other copy someone believes exists actually does. A delete authorized on an unverifiable belief is the self-certifying shape this skill exists to remove, and it would be the only evidence-free deletion in the file.
So the test is just the classification:
Is this passage describing how things are now? Then a dated old value in it is a question 1 or 2 case wearing a date, and the date does not rescue it.
Or is it recording what was true at a moment — a changelog line, a decision log entry, a price history row, an incident timeline, a dated measurement, a postmortem? Then it is question 3. Leave the value alone.
Apply this to the passage, not the file — where passage means the unit that would be read as one thing: the entry, the table row, the dated note, the section under one heading. A how-to guide can contain one genuinely historical block, and a changelog can carry a current-state summary at the top; a file-level verdict gets both wrong. But do not run it below the passage either — a postmortem's timeline and its root-cause section are one record, and splitting them so the timeline can be deleted as "reproducible from logs" destroys the thing whose parts they are.
Question 3 takes precedence over questions 1 and 2 for these passages, even though it is asked later. A changelog line "2026-01-04: raised the upload limit to 50 MB" does hit question 2 — the upload limit is authoritatively defined in config — and stopping there would replace it with a link and destroy the record. Question 1 does the same damage by a different route: "2026-01-04: added eu-west-3, bringing us to 12 regions" contains a count, and a count is question 1's own example — but recomputing it yields today's region count and rewrites what that day recorded. The stop-at-first-hit rule is about efficiency, not about routing history into a link or into a recomputation.
When you cannot tell whether a passage is history or current-state, ask — do not resolve it by deleting. The asymmetry is the reason: deleting a record you mistook for a current-state restatement destroys the only copy, while keeping a restatement you mistook for a record costs a few lines. If you are running unattended and cannot ask, leave the passage untouched and report it.
Linking only solves case 2. Pointing a link at a derived value is a category error: there is no single authoritative cell to point at, so the link becomes scaffolding that goes stale on its own and creates a false impression that the two places are aligned.
Why this ordering matters: the mature answer to drift in API documentation is "generate it from the source, don't hand-copy it" (OpenAPI/Swagger exist for exactly this reason). Question 1 is that same instinct applied to prose — the surest way to keep a number correct is for it not to be written down twice.
Position gives no exemption
A fact is a fact wherever it sits. All of these are body text for the purposes of the test above, and all of them are where stale values actually survive audits:
frontmatter · YAML/JSON metadata · parenthetical asides · a description column in an index table · checkbox state · introductory framing before the real content · cross-document footnotes · an extra column added to a table · commit-message-shaped notes left in prose
What counts as a fact
Anything that changes over time and has an authoritative definition somewhere: numbers, status, ownership, relationships, classifications, rules, decision sources, dates, counts. Business rules count too — "all quotes route through the finance lead" is defined once and referenced, never restated in each downstream doc.
Mode 1 — Post-change governance
Run these in order. The order is the point: fixing derived docs before establishing what is authoritative just multiplies the work.
1. Scope it from the change, not from the repo. Enumerate the facts this change altered — this port, this path, this procedure name, this default. That list is what bounds the work. The list is bounded; the search for each item on it is not — a fact's quiet copies live precisely in files the change never touched, so each fact gets searched repo-wide in step 5. What is forbidden is the other thing: opening files to look for unrelated problems. An unbounded sweep is how a doc task turns into an unreviewable refactor.
2. Identify the authoritative source (SSOT) for each affected fact. Which file defines this port / path / procedure — as opposed to mentioning it? Update that first. Everything downstream either points at it or is derived from it.
The practical test for "defines" is where a change has to be made for reality to change: the file the deployment actually reads, the schema the code loads, the runbook a human follows step by step. A file that recites the value while explaining something else is mentioning it. When two files both look declarative, prefer, in order: the one a machine consumes over one only humans read; the one nearest the thing it describes; the one other docs already cite. Then say in your report which you picked and on which of those grounds — an arbitrary pick becomes permanent once step 2 points everything else at it, so it should be a stated decision rather
Content truncated.
When not to use it
- →When documentation is already a single source of truth
- →When files cover entirely unrelated topics
Prerequisites
Limitations
- →Requires manual review of the value analysis table
- →Does not automatically rewrite content, only consolidates
How it compares
Unlike manual merging, this skill uses a structured value analysis table to justify every deletion and ensure 100% of valuable content is preserved.
Compared to similar skills
docs-cleaner side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| docs-cleaner (this skill) | 1 | 3mo | No flags | Intermediate |
| ml-paper-writing | 48 | 8mo | Review | Advanced |
| docs-review | 10 | 9mo | No flags | Beginner |
| claude-md-improver | 21 | 8mo | Review | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by daymade
View all by daymade →You might also like
ml-paper-writing
davila7
Write publication-ready ML/AI papers for NeurIPS, ICML, ICLR, ACL, AAAI, COLM. Use when drafting papers from research repos, structuring arguments, verifying citations, or preparing camera-ready submissions. Includes LaTeX templates, reviewer guidelines, and citation verification workflows.
docs-review
metabase
Review documentation changes for compliance with the Metabase writing style guide. Use when reviewing pull requests, files, or diffs containing documentation markdown files.
claude-md-improver
anthropics
Audit and improve CLAUDE.md files in repositories. Use when user asks to check, audit, update, improve, or fix CLAUDE.md files. Scans for all CLAUDE.md files, evaluates quality against templates, outputs quality report, then makes targeted updates. Also use when the user mentions "CLAUDE.md maintenance" or "project memory optimization".
write-docs
tldraw
Writing SDK documentation for tldraw. Use when creating new documentation articles, updating existing docs, or when documentation writing guidance is needed. Applies to docs in apps/docs/content/.
update-docs
vercel
This skill should be used when the user asks to "update documentation for my changes", "check docs for this PR", "what docs need updating", "sync docs with code", "scaffold docs for this feature", "document this feature", "review docs completeness", "add docs for this change", "what documentation is affected", "docs impact", or mentions "docs/", "docs/01-app", "docs/02-pages", "MDX", "documentation update", "API reference", ".mdx files". Provides guided workflow for updating Next.js documentation based on code changes.
wiki-architect
microsoft
Analyzes code repositories and generates hierarchical documentation structures with onboarding guides. Use when the user wants to create a wiki, generate documentation, map a codebase structure, or understand a project's architecture at a high level.