transcript-fixer
Corrects speech-to-text errors in transcripts using AI and rule sets.
Install
mkdir -p .claude/skills/transcript-fixer && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4280" && unzip -o skill.zip -d .claude/skills/transcript-fixer && rm skill.zipInstalls to .claude/skills/transcript-fixer
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Corrects speech-to-text transcription errors using dictionary rules and AI-powered analysis. Builds personalized correction databases that learn from each fix, auto-loads person-name ASR variants from your people roster, and reads per-domain context files that prime the AI pass for context-dependent homophones. Triggers when working with ASR/STT output containing recognition errors, homophones, garbled technical terms, person-name errors, or Chinese/English mixed content. Also triggers on requests to clean up meeting notes, lecture transcripts, interview recordings, or any text produced by speech recognition. Use this skill even when the user just says "fix this transcript", "clean up these meeting notes", or mentions garbled names without invoking ASR specifically.Key capabilities
- →Apply dictionary-based transcription corrections
- →Build personalized correction databases
- →Extract uncertain ASR tokens for review
- →Load domain-specific correction presets
- →Generate diff reports for changes
How it works
It uses a two-phase pipeline: a deterministic dictionary filter for known errors followed by an AI-powered pass to resolve remaining context-dependent mistakes.
Inputs & outputs
When to use transcript-fixer
- →Fix errors in interview transcripts
- →Clean up meeting notes
- →Apply custom vocabulary corrections
About this skill
Transcript Fixer
Use a two-phase loop:
- Stage 1 applies deterministic, already-known corrections.
- Native AI Correction reads the complete transcript, fixes one-off errors, verifies uncertain entities, and compounds reusable fixes.
Native AI Correction is the default. Stage 1 alone is incomplete. Stage 3 API exists only for automation that has no Claude/Codex agent available.
Operating contract
- Finish Stage 1 → Native AI Correction → compound confirmed recurring fixes. Do not report a transcript clean after Stage 1 alone.
- Skip Native AI only when the human explicitly limits this run to the dictionary pass or a dated artifact proves Native AI already ran on this exact transcript.
- In Claude Code or Codex, do not run Stage 3. Use Stage 1 plus the native workflow.
- Never rewrite speech for fluency. A correction must explain a plausible ASR error and preserve who said what.
- Never infer or reassign speaker identities. Preserve speaker-label lines; human-confirmed labels and user verdicts are authoritative.
- When comparing two transcripts of the same session (a 1.3x-sped Feishu 妙记 vs a full-speed one, DJI vs 录音豆, two devices, or any duplicate/overlapping capture), do NOT let an alignment script's output stand in for reading. Script diffs (
difflib, normalize-and-compare, per-turn speaker-label counts) are locators for candidate divergent segments only — their similarity ratios, "identical turn counts", and "speaker-divergence counts" are never the conclusion. Read both transcripts turn by turn yourself and decide from that. Real case 2026-09-18: a per-line regex that missed continuation lines reported "150 speaker divergences, all empty" then flipped to "1" once multi-line bodies were parsed — the whole numeric range was parser noise, and a canonical swap was nearly made on it. Corroboration that is genuinely independent requires a different recognizer over the source audio (rung 7), not a re-segmentation of the same recognition. - Parsing a Feishu 妙记
.txt(or anyspeaker HH:MM:SS.mmmtranscript): the spoken text lives on the lines AFTER the speaker-label line, up to the next label — not on the label line itself. A single-line regex (group(3)of the label) reads every multi-line turn's body as empty. Counting speakers or turns: use UTF-8-safe tooling (Python), neversort | uniq -corawk '{print $1}'— both silently collapse distinct CJK speaker labels into one (2026-08-29sort|uniqfolded 5 speakers' 106 turns into "说话人A 106"; 2026-09-18awk $1reported "说话人A 373" for a file whose real distribution was 说话人A17/说话人B287/说话人C60/说话人D9). - Before correcting any person name, directly read both the configured global people roster and the owning project's explicit identity roster or alias ledger. Stage 1 auto-loads only global
ASR 变体entries; it does not load project rosters or expose suppressed, disabled, and unlisted entries. If an expected source is missing or the sources conflict, leave the name unchanged and enqueue or ask once. Never use occurrence frequency as identity evidence. Read references/dictionary_identity_and_context.md before settling the name. - Resolve doubts from available evidence before escalating. Audio download is one evidence channel, not a prerequisite for Native correction. When it is unavailable, follow evidence selection and escalation; do not require the user to change download permissions or treat every pending row as a question only they can answer.
- Leave genuinely unresolved text unchanged and enqueue it. A visible garble is safer than a fluent wrong guess; a pending row records uncertainty, not an automatic human handoff.
- Treat an unfamiliar token as unknown, not as an error. Exhaust the local evidence ladder first. For a load-bearing token that remains unresolved, use the clip-level cross-recognizer rung only when source audio and a permitted second engine are already available; otherwise continue with the available evidence under the escalation policy above. Agreement from a genuinely different recognizer family strongly corroborates the sound, but never chooses between homophonic spellings or overrides the person-name gate. Read native workflow step 4, rung 7 before using it. A backlog of exhausted pendings is the batch form of this rung: when a native pass leaves a queue of locally-unresolvable rows, do not hand the queue to the user wholesale — check it with
verify_queue_audio.py(one transcript, one source audio, one second engine, two clip windows per row). The script can attach acoustic evidence but never records a verdict; identical clips and readings that do not support the disputed word boundary remain pending for adjudication. The operator matrix and timestamp-mapping pitfalls live in advanced_correction_evidence.md § Batch pending adjudication. - Treat a single-line
asr_notevalue as correction provenance: it intentionally cites old forms and is excluded from matching. Multi-line YAML ledger values are not masked; keywords, titles, other ASR-derived metadata, and body text remain in correction scope. - Read references/native_ai_full_workflow.md in full before performing a native pass. Read the task-specific references named below before their corresponding action.
Run context
Run every entrypoint through uv run; entrypoints that need third-party Python packages declare them with PEP 723, while stdlib/internal-only utilities may omit the metadata block. Execute commands from the skill directory printed when this skill was invoked, or prefix every script path with that directory. Do not rely on $CLAUDE_SKILL_DIR; it is not available in every harness.
If the bundle location is genuinely unknown, use the installation-resolution procedure in references/installation_setup.md. Do not select the first result from a broad find: caches, backups, and old versions can coexist.
Quick start
# Initialize once
uv run scripts/fix_transcription.py --init
# Stage 1 for one project domain. --apply-domain trusts that explicitly
# selected, human-curated project domain at every risk level.
uv run scripts/fix_transcription.py \
--input meeting.md --stage 1 \
--domain myproject --apply-domain --json
# Several sibling domains may be loaded as one union.
uv run scripts/fix_transcription.py \
--input meeting.md --stage 1 \
--domain myproject,myproject-alt --apply-domain --json
# Preview without writing the Stage 1 output.
uv run scripts/fix_transcription.py \
--input meeting.md --stage 1 --domain myproject --dry-run
# Scan all documented context traps after the native read-through.
uv run scripts/fix_transcription.py --scan-traps \
--context-file ~/.transcript-fixer/contexts/myproject.md \
--input meeting.md
Safe mode is the Stage 1 default: low-risk rules apply; medium/high-risk matches defer to *_needs_review.md and the persistent review queue. Applied: 0 is a valid result, not proof that the transcript is clean.
The Stage 1 JSON contract is:
{
"applied": 0,
"deferred": 0,
"output_path": null,
"needs_review_path": null,
"input_unchanged": true,
"review_enqueued": 0,
"stage1_only_incomplete": true,
"stage2_total_chunks": 0,
"stage2_failed_chunks": 0,
"stage2_degraded": false,
"boundary_refused": 0
}
Read every field in the result. boundary_refused counts dictionary matches the word-boundary check refused this run — neither applied nor deferred, so a caller comparing runs can see why a deferral disappeared; --apply-all switches that check off. stage1_only_incomplete is additive to the original caller contract and must remain true for a Stage 1 script run; only the caller can close it by running Native AI, or by explicitly choosing the agent-less Stage 2/3 route. The stage2_* telemetry fields are always present: Stage 1 reports 0, 0, and false; Stage 2/3 replace them with the actual API outcome. Do not infer no-op or success from whether a sidecar exists.
For a native end-to-end example, read references/example_session_dji_minutes.md.
Choose the route
| Route | Use when | Required reading |
|---|---|---|
| Fast native | Short/plain transcript, known speakers, low stakes | This file + native_ai_full_workflow.md |
| Full native | Domain-heavy, unfamiliar entities, 3+ speakers, long or decision-bearing transcript | native_ai_full_workflow.md, plus queue and evidence references below |
| Caller integration | Another skill or ingest pipeline invokes Stage 1 | Cross-skill caller contract below |
| Review queue/dashboard | Any item is uncertain or needs audio | review_queue_dashboard.md |
| Agent-less API | CI/batch automation with no agent available | glm_api_setup.md and workflow_guide.md |
| Multi-file batch | Several related transcripts; especially 10+ files | advanced_correction_evidence.md |
Use vocabulary and stakes as the primary tier signals; use length only as a tiebreaker. A five-minute medical interview can require the full tier, while a long plain two-person memo can use the fast tier.
Native correction checklist
- Give the file its final name before Stage 1. Queue anchors store absolute paths. Use a human-readable project filename before any deferral can enqueue. When the input arrives as inline text with no file yet — a slash-command argument, a pasted block — write it to a file before anything else;
--inputand the queue anchors both need a path, and a scratch location is fine when nothing downstream wil
Content truncated.
When not to use it
- →Correcting non-ASR/STT generated text
- →Applying high-risk rules without review
Prerequisites
Limitations
- →Safe mode defers medium/high-risk rules to sidecar files
- →Audit function cannot definitively identify all false positives
How it compares
It maintains a persistent, learning correction database that improves accuracy over time, unlike manual find-and-replace methods.
Compared to similar skills
transcript-fixer side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| transcript-fixer (this skill) | 1 | 2mo | Review | Intermediate |
| docs-write | 22 | 7mo | No flags | Beginner |
| content-research-writer | 15 | 11mo | No flags | Beginner |
| doc-coauthoring | 16 | 10mo | No flags | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by daymade
View all by daymade →You might also like
docs-write
metabase
Write documentation following Metabase's conversational, clear, and user-focused style. Use when creating or editing documentation files (markdown, MDX, etc.).
content-research-writer
ComposioHQ
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section. Transforms your writing process from solo effort to collaborative partnership.
doc-coauthoring
anthropics
Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.
research-grants
davila7
Write competitive research proposals for NSF, NIH, DOE, and DARPA. Agency-specific formatting, review criteria, budget preparation, broader impacts, significance statements, innovation narratives, and compliance with submission requirements.
teams-channel-post-writer
daymade
Creates educational Teams channel posts for internal knowledge sharing about Claude Code features, tools, and best practices. Applies when writing posts, announcements, or documentation to teach colleagues effective Claude Code usage, announce new features, share productivity tips, or document lessons learned. Provides templates, writing guidelines, and structured approaches emphasizing concrete examples, underlying principles, and connections to best practices like context engineering. Activates for content involving Teams posts, channel announcements, feature documentation, or tip sharing.
write-docs
tldraw
Writing SDK documentation for tldraw. Use when creating new documentation articles, updating existing docs, or when documentation writing guidance is needed. Applies to docs in apps/docs/content/.