TR

transcript-fixer

Corrects speech-to-text errors in transcripts using AI and rule sets.

Install

mkdir -p .claude/skills/transcript-fixer && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4280" && unzip -o skill.zip -d .claude/skills/transcript-fixer && rm skill.zip

Installs to .claude/skills/transcript-fixer

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Corrects speech-to-text transcription errors using dictionary rules and AI-powered analysis. Builds personalized correction databases that learn from each fix, auto-loads person-name ASR variants from your people roster, and reads per-domain context files that prime the AI pass for context-dependent homophones. Triggers when working with ASR/STT output containing recognition errors, homophones, garbled technical terms, person-name errors, or Chinese/English mixed content. Also triggers on requests to clean up meeting notes, lecture transcripts, interview recordings, or any text produced by speech recognition. Use this skill even when the user just says "fix this transcript", "clean up these meeting notes", or mentions garbled names without invoking ASR specifically.
776 chars✓ has a “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Intermediate

Key capabilities

  • →Apply dictionary-based transcription corrections
  • →Build personalized correction databases
  • →Extract uncertain ASR tokens for review
  • →Load domain-specific correction presets
  • →Generate diff reports for changes

How it works

It uses a two-phase pipeline: a deterministic dictionary filter for known errors followed by an AI-powered pass to resolve remaining context-dependent mistakes.

Inputs & outputs

You give it
Raw transcript file
You get back
Corrected transcript and change report

When to use transcript-fixer

  • →Fix errors in interview transcripts
  • →Clean up meeting notes
  • →Apply custom vocabulary corrections

About this skill

Transcript Fixer

Use a two-phase loop:

  1. Stage 1 applies deterministic, already-known corrections.
  2. Native AI Correction reads the complete transcript, fixes one-off errors, verifies uncertain entities, and compounds reusable fixes.

Native AI Correction is the default. Stage 1 alone is incomplete. Stage 3 API exists only for automation that has no Claude/Codex agent available.

Operating contract

  • Finish Stage 1 → Native AI Correction → compound confirmed recurring fixes. Do not report a transcript clean after Stage 1 alone.
  • Skip Native AI only when the human explicitly limits this run to the dictionary pass or a dated artifact proves Native AI already ran on this exact transcript.
  • In Claude Code or Codex, do not run Stage 3. Use Stage 1 plus the native workflow.
  • Never rewrite speech for fluency. A correction must explain a plausible ASR error and preserve who said what.
  • Never infer or reassign speaker identities. Preserve speaker-label lines; human-confirmed labels and user verdicts are authoritative.
  • When comparing two transcripts of the same session (a 1.3x-sped Feishu 妙记 vs a full-speed one, DJI vs 录音豆, two devices, or any duplicate/overlapping capture), do NOT let an alignment script's output stand in for reading. Script diffs (difflib, normalize-and-compare, per-turn speaker-label counts) are locators for candidate divergent segments only — their similarity ratios, "identical turn counts", and "speaker-divergence counts" are never the conclusion. Read both transcripts turn by turn yourself and decide from that. Real case 2026-09-18: a per-line regex that missed continuation lines reported "150 speaker divergences, all empty" then flipped to "1" once multi-line bodies were parsed — the whole numeric range was parser noise, and a canonical swap was nearly made on it. Corroboration that is genuinely independent requires a different recognizer over the source audio (rung 7), not a re-segmentation of the same recognition.
  • Parsing a Feishu 妙记 .txt (or any speaker HH:MM:SS.mmm transcript): the spoken text lives on the lines AFTER the speaker-label line, up to the next label — not on the label line itself. A single-line regex (group(3) of the label) reads every multi-line turn's body as empty. Counting speakers or turns: use UTF-8-safe tooling (Python), never sort | uniq -c or awk '{print $1}' — both silently collapse distinct CJK speaker labels into one (2026-08-29 sort|uniq folded 5 speakers' 106 turns into "说话人A 106"; 2026-09-18 awk $1 reported "说话人A 373" for a file whose real distribution was 说话人A17/说话人B287/说话人C60/说话人D9).
  • Before correcting any person name, directly read both the configured global people roster and the owning project's explicit identity roster or alias ledger. Stage 1 auto-loads only global ASR 变体 entries; it does not load project rosters or expose suppressed, disabled, and unlisted entries. If an expected source is missing or the sources conflict, leave the name unchanged and enqueue or ask once. Never use occurrence frequency as identity evidence. Read references/dictionary_identity_and_context.md before settling the name.
  • Resolve doubts from available evidence before escalating. Audio download is one evidence channel, not a prerequisite for Native correction. When it is unavailable, follow evidence selection and escalation; do not require the user to change download permissions or treat every pending row as a question only they can answer.
  • Leave genuinely unresolved text unchanged and enqueue it. A visible garble is safer than a fluent wrong guess; a pending row records uncertainty, not an automatic human handoff.
  • Treat an unfamiliar token as unknown, not as an error. Exhaust the local evidence ladder first. For a load-bearing token that remains unresolved, use the clip-level cross-recognizer rung only when source audio and a permitted second engine are already available; otherwise continue with the available evidence under the escalation policy above. Agreement from a genuinely different recognizer family strongly corroborates the sound, but never chooses between homophonic spellings or overrides the person-name gate. Read native workflow step 4, rung 7 before using it. A backlog of exhausted pendings is the batch form of this rung: when a native pass leaves a queue of locally-unresolvable rows, do not hand the queue to the user wholesale — check it with verify_queue_audio.py (one transcript, one source audio, one second engine, two clip windows per row). The script can attach acoustic evidence but never records a verdict; identical clips and readings that do not support the disputed word boundary remain pending for adjudication. The operator matrix and timestamp-mapping pitfalls live in advanced_correction_evidence.md § Batch pending adjudication.
  • Treat a single-line asr_note value as correction provenance: it intentionally cites old forms and is excluded from matching. Multi-line YAML ledger values are not masked; keywords, titles, other ASR-derived metadata, and body text remain in correction scope.
  • Read references/native_ai_full_workflow.md in full before performing a native pass. Read the task-specific references named below before their corresponding action.

Run context

Run every entrypoint through uv run; entrypoints that need third-party Python packages declare them with PEP 723, while stdlib/internal-only utilities may omit the metadata block. Execute commands from the skill directory printed when this skill was invoked, or prefix every script path with that directory. Do not rely on $CLAUDE_SKILL_DIR; it is not available in every harness.

If the bundle location is genuinely unknown, use the installation-resolution procedure in references/installation_setup.md. Do not select the first result from a broad find: caches, backups, and old versions can coexist.

Quick start

# Initialize once
uv run scripts/fix_transcription.py --init

# Stage 1 for one project domain. --apply-domain trusts that explicitly
# selected, human-curated project domain at every risk level.
uv run scripts/fix_transcription.py \
  --input meeting.md --stage 1 \
  --domain myproject --apply-domain --json

# Several sibling domains may be loaded as one union.
uv run scripts/fix_transcription.py \
  --input meeting.md --stage 1 \
  --domain myproject,myproject-alt --apply-domain --json

# Preview without writing the Stage 1 output.
uv run scripts/fix_transcription.py \
  --input meeting.md --stage 1 --domain myproject --dry-run

# Scan all documented context traps after the native read-through.
uv run scripts/fix_transcription.py --scan-traps \
  --context-file ~/.transcript-fixer/contexts/myproject.md \
  --input meeting.md

Safe mode is the Stage 1 default: low-risk rules apply; medium/high-risk matches defer to *_needs_review.md and the persistent review queue. Applied: 0 is a valid result, not proof that the transcript is clean.

The Stage 1 JSON contract is:

{
  "applied": 0,
  "deferred": 0,
  "output_path": null,
  "needs_review_path": null,
  "input_unchanged": true,
  "review_enqueued": 0,
  "stage1_only_incomplete": true,
  "stage2_total_chunks": 0,
  "stage2_failed_chunks": 0,
  "stage2_degraded": false,
  "boundary_refused": 0
}

Read every field in the result. boundary_refused counts dictionary matches the word-boundary check refused this run — neither applied nor deferred, so a caller comparing runs can see why a deferral disappeared; --apply-all switches that check off. stage1_only_incomplete is additive to the original caller contract and must remain true for a Stage 1 script run; only the caller can close it by running Native AI, or by explicitly choosing the agent-less Stage 2/3 route. The stage2_* telemetry fields are always present: Stage 1 reports 0, 0, and false; Stage 2/3 replace them with the actual API outcome. Do not infer no-op or success from whether a sidecar exists.

For a native end-to-end example, read references/example_session_dji_minutes.md.

Choose the route

RouteUse whenRequired reading
Fast nativeShort/plain transcript, known speakers, low stakesThis file + native_ai_full_workflow.md
Full nativeDomain-heavy, unfamiliar entities, 3+ speakers, long or decision-bearing transcriptnative_ai_full_workflow.md, plus queue and evidence references below
Caller integrationAnother skill or ingest pipeline invokes Stage 1Cross-skill caller contract below
Review queue/dashboardAny item is uncertain or needs audioreview_queue_dashboard.md
Agent-less APICI/batch automation with no agent availableglm_api_setup.md and workflow_guide.md
Multi-file batchSeveral related transcripts; especially 10+ filesadvanced_correction_evidence.md

Use vocabulary and stakes as the primary tier signals; use length only as a tiebreaker. A five-minute medical interview can require the full tier, while a long plain two-person memo can use the fast tier.

Native correction checklist

  1. Give the file its final name before Stage 1. Queue anchors store absolute paths. Use a human-readable project filename before any deferral can enqueue. When the input arrives as inline text with no file yet — a slash-command argument, a pasted block — write it to a file before anything else; --input and the queue anchors both need a path, and a scratch location is fine when nothing downstream wil

Content truncated.

When not to use it

  • →Correcting non-ASR/STT generated text
  • →Applying high-risk rules without review

Prerequisites

uv

Limitations

  • →Safe mode defers medium/high-risk rules to sidecar files
  • →Audit function cannot definitively identify all false positives

How it compares

It maintains a persistent, learning correction database that improves accuracy over time, unlike manual find-and-replace methods.

Compared to similar skills

transcript-fixer side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
transcript-fixer (this skill)12moReviewIntermediate
docs-write227moNo flagsBeginner
content-research-writer1511moNo flagsBeginner
doc-coauthoring1610moNo flagsBeginner

Try saying

Example prompts that trigger this skill in your AI assistant.

ppt-creator

daymade

Create professional slide decks from topics or documents. Generates structured content with data-driven charts, speaker notes, and complete PPTX files. Applies persuasive storytelling principles (Pyramid Principle, assertion-evidence). Supports multiple formats (Marp, PowerPoint). Use for presentations, pitches, slide decks, or keynotes.

75110

macos-cleaner

daymade

Analyze and reclaim macOS disk space through intelligent cleanup recommendations. This skill should be used when users report disk space issues, need to clean up their Mac, or want to understand what's consuming storage. Focus on safe, interactive analysis with user confirmation before any deletions.

1631

qa-expert

daymade

This skill should be used when establishing comprehensive QA testing processes for any software project. Use when creating test strategies, writing test cases following Google Testing Standards, executing test plans, tracking bugs with P0-P4 classification, calculating quality metrics, or generating progress reports. Includes autonomous execution capability via master prompts and complete documentation templates for third-party QA team handoffs. Implements OWASP security testing and achieves 90% coverage targets.

1427

repomix-unmixer

daymade

Extracts files from repomix-packed repositories, restoring original directory structures from XML/Markdown/JSON formats. Activates when users need to unmix repomix files, extract packed repositories, restore file structures from repomix output, or reverse the repomix packing process.

524

teams-channel-post-writer

daymade

Creates educational Teams channel posts for internal knowledge sharing about Claude Code features, tools, and best practices. Applies when writing posts, announcements, or documentation to teach colleagues effective Claude Code usage, announce new features, share productivity tips, or document lessons learned. Provides templates, writing guidelines, and structured approaches emphasizing concrete examples, underlying principles, and connections to best practices like context engineering. Activates for content involving Teams posts, channel announcements, feature documentation, or tip sharing.

591

twitter-reader

daymade

Fetch Twitter/X post content by URL using jina.ai API to bypass JavaScript restrictions. Use when Claude needs to retrieve tweet content including author, timestamp, post text, images, and thread replies. Supports individual posts or batch fetching from x.com or twitter.com URLs.

552

You might also like

docs-write

metabase

Write documentation following Metabase's conversational, clear, and user-focused style. Use when creating or editing documentation files (markdown, MDX, etc.).

22139

content-research-writer

ComposioHQ

Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section. Transforms your writing process from solo effort to collaborative partnership.

15111

doc-coauthoring

anthropics

Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.

1686

research-grants

davila7

Write competitive research proposals for NSF, NIH, DOE, and DARPA. Agency-specific formatting, review criteria, budget preparation, broader impacts, significance statements, innovation narratives, and compliance with submission requirements.

694

teams-channel-post-writer

daymade

Creates educational Teams channel posts for internal knowledge sharing about Claude Code features, tools, and best practices. Applies when writing posts, announcements, or documentation to teach colleagues effective Claude Code usage, announce new features, share productivity tips, or document lessons learned. Provides templates, writing guidelines, and structured approaches emphasizing concrete examples, underlying principles, and connections to best practices like context engineering. Activates for content involving Teams posts, channel announcements, feature documentation, or tip sharing.

591

write-docs

tldraw

Writing SDK documentation for tldraw. Use when creating new documentation articles, updating existing docs, or when documentation writing guidance is needed. Applies to docs in apps/docs/content/.

665

Search skills

Search the agent skills registry