Automates the parsing, extraction, and summarization of PDF, Word, and PowerPoint documents.

Install

mkdir -p .claude/skills/document-pro && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/16419" && unzip -o skill.zip -d .claude/skills/document-pro && rm skill.zip

Installs to .claude/skills/document-pro

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

文档处理技能 - 让 AI 能够读取、解析、提取 PDF、DOCX、PPT 等文档的关键信息。当用户要求分析文档、提取内容、总结报告时触发此技能。
73 charsno explicit “when” trigger
Beginner

Key capabilities

  • Read and parse PDF documents.
  • Extract text from DOCX files.
  • Extract text from PPTX presentations.
  • Extract data from XLSX spreadsheets.
  • Identify document types and select appropriate tools.
  • Summarize document content and extract key points.

How it works

The skill identifies the document type, uses specific Python libraries to read and extract content, then analyzes and summarizes the information.

Inputs & outputs

You give it
A document file (PDF, DOCX, PPTX, XLSX, TXT, Markdown)
You get back
Extracted text, tables, key points, or a summary in Chinese

When to use document-pro

  • Extracting table data from PDF invoices
  • Summarizing long technical reports
  • Converting DOCX content to structured text
  • Parsing key insights from project presentations

About this skill

Document Pro - 文档处理技能

概述

赋予 AI 强大的文档处理能力:

  • PDF 读取与提取
  • Word 文档解析
  • PowerPoint 提取
  • Excel 数据提取
  • 文档格式转换

触发场景

  1. 用户发送文档并要求"分析"、"总结"
  2. 用户要求"提取文档内容"
  3. 用户要求"转换成 PDF"
  4. 用户询问文档中的具体信息
  5. 用户要求"从报告/论文中提取要点"

支持的格式

格式读取写入工具
PDFpdfplumber, PyPDF2
DOCXpython-docx
PPTXpython-pptx
XLSXopenpyxl
TXT内置
Markdown内置

工具使用

PDF 处理

# 提取文本
import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    for page in pdf.pages:
        text = page.extract_text()
        print(text)

# 提取表格
with pdfplumber.open("document.pdf") as pdf:
    table = pdf.pages[0].extract_tables()

Word 文档

from docx import Document

doc = Document("document.docx")
for para in doc.paragraphs:
    print(para.text)

# 提取表格
for table in doc.tables:
    for row in table.rows:
        print([cell.text for cell in row.cells])

PowerPoint

from pptx import Presentation

prs = Presentation("presentation.pptx")
for slide in prs.slides:
    for shape in slide.shapes:
        if shape.has_text_frame:
            print(shape.text)

工作流

1. 识别文档类型 → 选择正确的工具
2. 读取内容 → 提取文本、表格、图片
3. 分析信息 → 理解结构、提取要点
4. 总结呈现 → 用中文总结给用户

进阶功能

文档摘要

  • 提取文档主要观点
  • 生成简短摘要
  • 列出关键要点

表格处理

  • 识别表格结构
  • 提取表格数据
  • 转换为 CSV/Excel

关键词提取

  • 找出重要名词/术语
  • 识别主题
  • 提取关键信息

输出格式

向用户呈现文档时:

  • 文档类型和页数
  • 主要内容摘要
  • 关键要点(3-5条)
  • 建议的后续操作

限制

  • 扫描版 PDF 需要 OCR
  • 复杂格式可能丢失
  • 图片/图表无法完全理解

When not to use it

  • When the user does not require document analysis or content extraction.
  • When the user does not need a summary or report from a document.

Limitations

  • Scanned PDFs require OCR for text extraction.
  • Complex document formats may lead to content loss.
  • Images and charts cannot be fully understood.

How it compares

This workflow provides a unified interface for processing various document formats, automating content extraction and summarization, unlike manual parsing or using separate tools for each format.

Compared to similar skills

document-pro side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
document-pro (this skill)03moNo flagsBeginner
biorxiv-database79moReviewBeginner
local-deep-research-guide05moReviewIntermediate
office-docs16moReviewBeginner

Try saying

Example prompts that trigger this skill in your AI assistant.

Search skills

Search the agent skills registry