HI

historical-document-ocr

An accurate OCR tool for transcribing faded, handwritten, or legacy document scans into digital text.

Install

mkdir -p .claude/skills/historical-document-ocr && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/12312" && unzip -o skill.zip -d .claude/skills/historical-document-ocr && rm skill.zip

Installs to .claude/skills/historical-document-ocr

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Transcribe scanned/photographed historical documents (PDFs or images) to text with Gemini vision, using accuracy-first practices: per-page high-res rendering, faded-scan image enhancement, strict verbatim prompting, and an optional multi-model consensus pass that reconciles disagreements by re-reading the page. Use this whenever the user wants to OCR, transcribe, or extract the text of scanned letters, manuscripts, typescripts, carbon copies, ledgers, archival records, genealogy documents, old correspondence, or any image-only PDF that has no real text layer — especially when the material is handwritten, typewritten, faded, rotated, or hard to read and accuracy matters. Trigger even if the user just says "read this old scan", "what does this letter say", "digitize these archive pages", or "transcribe this PDF" and the PDF turns out to be scanned images. Do NOT use this for audio/podcast transcription (that is gemini-podcast-transcribe) or for born-digital PDFs that already contain selectable text (a plain pdftotext is enough there).
1048 chars✓ has a “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Intermediate

Key capabilities

  • →Transcribe scanned historical documents to text
  • →Render high-resolution pages for review
  • →Enhance faded-scan images
  • →Use context blocks to disambiguate ambiguous characters
  • →Perform multi-model consensus pass for accuracy
  • →Report results and review list for human verification

How it works

The skill transcribes scanned historical documents using Gemini vision, optimizing for accuracy on degraded material. It renders pages, enhances images, uses context for disambiguation, and can perform a multi-model consensus pass.

Inputs & outputs

You give it
Scanned images or PDFs of historical documents
You get back
Transcribed text, page-delimited markdown files, manifest.json, rendered PNGs (optional)

When to use historical-document-ocr

  • →Transcribing historical archives
  • →Converting scanned PDFs to text
  • →Digitizing ledger pages
  • →Reading handwritten manuscripts

About historical-document-ocr

Transcribes scanned images or PDFs using AI vision models. It is specifically optimized for high accuracy on degraded material like handwritten notes or faded manuscripts.

>-

When not to use it

  • →For audio/podcast transcription
  • →For born-digital PDFs that already contain selectable text
  • →When `pdftotext` already returns the real text

Prerequisites

GEMINI_API_KEY or GOOGLE_API_KEYpopplerimagemagick (optional)

Limitations

  • →Expects ~1–4% word error on faded/handwritten material
  • →Does not invent text; failure modes are skipped lines and word substitutions
  • →Requires human pass on flagged disagreement spans for archival-grade output

How it compares

This skill prioritizes accuracy on degraded historical documents through image enhancement, context-aware disambiguation, and multi-model consensus, unlike standard OCR tools that may struggle with such material.

Compared to similar skills

historical-document-ocr side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
historical-document-ocr (this skill)04moReviewIntermediate
scientific-paper05moNo flagsIntermediate
content-research-writer1511moNo flagsBeginner
research-grants69moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

Search skills

Search the agent skills registry