BA

baoyu-url-to-markdown

Fetches any URL via Chrome CDP and converts the page content into clean, readable markdown.

Install

mkdir -p .claude/skills/baoyu-url-to-markdown && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2005" && unzip -o skill.zip -d .claude/skills/baoyu-url-to-markdown && rm skill.zip

Installs to .claude/skills/baoyu-url-to-markdown

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.
309 chars✓ has a “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Intermediate

Key capabilities

  • Fetch web content using Chrome CDP
  • Convert web pages to markdown format
  • Handle login and CAPTCHA via interaction wait modes
  • Download images and videos to local storage
  • Support site-specific adapters for YouTube, X, and Hacker News
  • Output content as markdown or JSON

How it works

The tool uses the baoyu-fetch CLI to launch Chrome via CDP, applying site-specific adapters to scrape and parse content into markdown. It supports interaction modes to navigate login prompts or CAPTCHAs before extraction.

Inputs & outputs

You give it
Target URL and optional CLI flags
You get back
Markdown file or JSON object containing page content

When to use baoyu-url-to-markdown

  • Saving documentation for offline reading
  • Converting YouTube transcripts to text
  • Capturing Hacker News discussions
  • Converting web articles to markdown notes

About this skill

URL to Markdown

Fetches any URL via baoyu-fetch CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.

User Input Tools

When this skill prompts the user, follow this tool-selection rule (priority order):

  1. Prefer built-in user-input tools exposed by the current agent runtime — e.g., AskUserQuestion, request_user_input, clarify, ask_user, or any equivalent.
  2. Fallback: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.
  3. Batching: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.

Concrete AskUserQuestion references below are examples — substitute the local equivalent in other runtimes.

CLI Setup

Important: The CLI source is vendored in {baseDir}/scripts/lib. scripts/package.json installs only third-party runtime dependencies.

Agent Execution Instructions:

  1. Determine this SKILL.md file's directory path as {baseDir}
  2. Resolve ${BUN} runtime: if bun installed → bun; else suggest installing Bun
  3. If {baseDir}/scripts/node_modules does not exist, run ${BUN} install --cwd {baseDir}/scripts
  4. ${READER} = {baseDir}/scripts/baoyu-fetch
  5. Replace all ${READER} in this document with the resolved value

Preferences (EXTEND.md)

Check EXTEND.md in priority order — the first one found wins:

PriorityPathScope
1.baoyu-skills/baoyu-url-to-markdown/EXTEND.mdProject
2${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.mdXDG
3$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.mdUser home
ResultAction
FoundRead, parse, apply settings
Not foundMUST run first-time setup (see below) — do NOT silently create defaults

EXTEND.md supports: download media by default, default output directory.

First-Time Setup ⛔ BLOCKING

When EXTEND.md is not found, you MUST use AskUserQuestion to gather preferences before creating EXTEND.md. NEVER create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call:

  • Q1 — Media (header "Media"): "How to handle images and videos in pages?"
    • "Ask each time (Recommended)" — Prompt after each save
    • "Always download" — Download to local imgs/ and videos/
    • "Never download" — Keep remote URLs
  • Q2 — Output (header "Output"): "Default output directory?"
    • "url-to-markdown (Recommended)" — Save to ./url-to-markdown/{domain}/{slug}.md
    • User may pick "Other" and type a custom path
  • Q3 — Save (header "Save"): "Where to save preferences?"
    • "User (Recommended)" — ~/.baoyu-skills/ (all projects)
    • "Project" — .baoyu-skills/ (this project only)

After answers, write EXTEND.md, confirm "Preferences saved to [path]", then continue.

Full template: references/config/first-time-setup.md.

Supported Keys

KeyDefaultValuesDescription
download_mediaaskask / 1 / 0ask = prompt each time, 1 = always, 0 = never
default_output_diremptypath or emptyDefault output directory (empty = ./url-to-markdown/)

EXTEND.md → CLI mapping:

EXTEND.md keyCLI argumentNotes
download_media: 1--download-mediaRequires --output to be set
default_output_dir: ./posts/Agent constructs --output ./posts/{domain}/{slug}.mdAgent generates path, not a direct flag

Value priority: CLI arguments → EXTEND.md → skill defaults.

Usage

# Default: headless capture, markdown to stdout
${READER} <url>

# Save to file
${READER} <url> --output article.md

# Save with media download
${READER} <url> --output article.md --download-media

# Wait for interaction (login/CAPTCHA) — auto-detect and continue
${READER} <url> --wait-for interaction --output article.md

# Wait for interaction — manual control (Enter to continue)
${READER} <url> --wait-for force --output article.md

# JSON output
${READER} <url> --format json --output article.json

# Force specific adapter
${READER} <url> --adapter youtube --output transcript.md

Options

OptionDescription
<url>URL to fetch
--output <path>Output file path (default: stdout)
--format <type>Output format: markdown (default) or json
--jsonShorthand for --format json
--adapter <name>Force adapter: x, youtube, hn, or generic (default: auto-detect)
--headlessForce headless Chrome (no visible window)
--wait-for <mode>Interaction wait mode: none (default), interaction, or force
--wait-for-interactionAlias for --wait-for interaction
--wait-for-loginAlias for --wait-for interaction
--timeout <ms>Page load timeout (default: 30000)
--interaction-timeout <ms>Login/CAPTCHA wait timeout (default: 600000 = 10 min)
--interaction-poll-interval <ms>Poll interval for interaction checks (default: 1500)
--download-mediaDownload images/videos to local imgs/ and videos/, rewrite markdown links. Requires --output
--media-dir <dir>Base directory for downloaded media (default: same as --output directory)
--cdp-url <url>Reuse existing Chrome DevTools Protocol endpoint
--browser-path <path>Custom Chrome/Chromium binary path
--chrome-profile-dir <path>Chrome user data directory (default: BAOYU_CHROME_PROFILE_DIR env or ./baoyu-skills/chrome-profile)
--debug-dir <dir>Write debug artifacts (document.json, markdown.md, page.html, network.json)

Agent Quality Gate

CRITICAL: treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without failing the CLI.

After every headless run, inspect the saved markdown. See references/quality-gate.md for the full checklist, recovery workflow, and capture-mode table. Read it whenever a run looks suspicious or the user asks about login/CAPTCHA handling.

Output Path Generation

The agent must construct the output file path — baoyu-fetch does not auto-generate paths.

Algorithm:

  1. Determine base directory from EXTEND.md default_output_dir or default ./url-to-markdown/
  2. Extract domain from URL (e.g., example.com)
  3. Generate slug from URL path or page title (kebab-case, 2-6 words)
  4. Construct: {base_dir}/{domain}/{slug}/{slug}.md — each URL gets its own directory so media files stay isolated
  5. Conflict resolution: append timestamp {slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md

Pass the constructed path to --output. Media files (--download-media) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.

Adapters & Media

See references/adapters.md for the adapter catalog (X, YouTube, Hacker News, generic), per-adapter notes, the media download flow (ask / always / never), and the JSON output schema. Read it before answering adapter-specific questions or handling media prompts.

Environment Variables

VariableDescription
BAOYU_CHROME_PROFILE_DIRChrome user data directory (can also use --chrome-profile-dir)

Troubleshooting: Chrome not found → use --browser-path. Timeout → increase --timeout. Login/CAPTCHA → --wait-for interaction. Debug → --debug-dir to inspect captured HTML and network logs.

Extension Support

Custom configurations via EXTEND.md. See Preferences section above for paths and supported keys.

When not to use it

  • When the target URL is inaccessible to the Chrome browser
  • When the user requires a non-markdown or non-JSON output format

Prerequisites

bun

Limitations

  • Headless capture may result in low-quality content for some sites
  • Requires manual inspection of output to verify capture quality
  • Does not auto-generate output paths without agent logic

How it compares

Unlike manual copy-pasting, this tool automates the extraction of clean markdown and handles complex web interactions through programmatic browser control.

Compared to similar skills

baoyu-url-to-markdown side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
baoyu-url-to-markdown (this skill)93moReviewIntermediate
web-to-markdown76moNo flagsIntermediate
defuddle02moReviewBeginner
firecrawl-core-workflow-a027dReviewBeginner

Try saying

Example prompts that trigger this skill in your AI assistant.

baoyu-xhs-images

JimLiu

Generates Xiaohongshu (Little Red Book) infographic series with 10 visual styles and 8 layouts. Breaks content into 1-10 cartoon-style images optimized for XHS engagement. Use when user mentions "小红书图片", "XHS images", "RedNote infographics", "小红书种草", or wants social media infographics for Chinese platforms.

2051

baoyu-article-illustrator

JimLiu

Analyzes article structure, identifies positions requiring visual aids, generates illustrations with Type × Style two-dimension approach. Use when user asks to "illustrate article", "add images", "generate images for article", or "为文章配图".

1848

baoyu-comic

JimLiu

Knowledge comic creator supporting multiple art styles and tones. Creates original educational comics with detailed panel layouts and sequential image generation. Use when user asks to create "知识漫画", "教育漫画", "biography comic", "tutorial comic", or "Logicomix-style comic".

1329

baoyu-infographic

JimLiu

Generates professional infographics with 20 layout types and 17 visual styles. Analyzes content, recommends layout×style combinations, and generates publication-ready infographics. Use when user asks to create "infographic", "信息图", "visual summary", or "可视化".

1230

baoyu-compress-image

JimLiu

Compresses images to WebP (default) or PNG with automatic tool selection. Use when user asks to "compress image", "optimize image", "convert to webp", or reduce image file size.

1129

baoyu-cover-image

JimLiu

Generates article cover images with 5 dimensions (type, palette, rendering, text, mood) combining 9 color palettes and 6 rendering styles. Supports cinematic (2.35:1), widescreen (16:9), and square (1:1) aspects. Use when user asks to "generate cover image", "create article cover", or "make cover".

1020

You might also like

web-to-markdown

davila7

Use ONLY when the user explicitly says: 'use the skill web-to-markdown ...' (or 'use a skill web-to-markdown ...'). Converts webpage URLs to clean Markdown by calling the local web2md CLI (Puppeteer + Readability), suitable for JS-rendered pages.

734

defuddle

ludotype

Extract clean markdown content from web pages using Defuddle CLI, removing clutter and navigation to save tokens. Use instead of WebFetch when the user provides a URL to read or analyze, for online documentation, articles, blog posts, or any standard web page. Do NOT use for URLs ending in .md — tho

00

firecrawl-core-workflow-a

jeremylongshore

Execute FireCrawl primary workflow: Core Workflow A. Use when implementing primary use case, building main features, or core integration tasks. Trigger with phrases like "firecrawl main workflow", "primary task with firecrawl".

00

okf-to-web

yzfly

把符合 OKF v0.1 的知识包打包成单个自包含且经过压缩(minify)的 HTML 文件, 内嵌导航, Markdown 阅读器与概念关系图谱, 数据不出页面, 无需后端. 适用场景包括把 OKF bundle 变成一个可分享的单文件网页, 离线浏览知识库, 生成对标官方 viz.html 的可视化. 当用户提到 okf to web, OKF 转单文件网页, 可视化 OKF, minify 知识库, 或要一个自包含 HTML 时触发.

00

fetch-url-md

1naichii

Fetch web content with automatic markdown version detection using curl. Use when Claude needs to retrieve documentation from websites that offer both HTML and markdown formats. First checks if a .md version exists (more efficient and cleaner), then falls back to HTML if unavailable. Ideal for fetchi

00

pdf-to-markdown

aliceisjustplaying

Convert entire PDF documents to clean, structured Markdown for full context loading. Use this skill when the user wants to extract ALL text from a PDF into context (not grep/search), when discussing or analyzing PDF content in full, when the user mentions "load the whole PDF", "bring the PDF into context", "read the entire PDF", or when partial extraction/grepping would miss important context. This is the preferred method for PDF text extraction over page-by-page or grep approaches.

1,1752,667

Search skills

Search the agent skills registry