baoyu-url-to-markdown
Fetches any URL via Chrome CDP and converts the page content into clean, readable markdown.
Install
mkdir -p .claude/skills/baoyu-url-to-markdown && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2005" && unzip -o skill.zip -d .claude/skills/baoyu-url-to-markdown && rm skill.zipInstalls to .claude/skills/baoyu-url-to-markdown
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.Key capabilities
- →Fetch web content using Chrome CDP
- →Convert web pages to markdown format
- →Handle login and CAPTCHA via interaction wait modes
- →Download images and videos to local storage
- →Support site-specific adapters for YouTube, X, and Hacker News
- →Output content as markdown or JSON
How it works
The tool uses the baoyu-fetch CLI to launch Chrome via CDP, applying site-specific adapters to scrape and parse content into markdown. It supports interaction modes to navigate login prompts or CAPTCHAs before extraction.
Inputs & outputs
When to use baoyu-url-to-markdown
- →Saving documentation for offline reading
- →Converting YouTube transcripts to text
- →Capturing Hacker News discussions
- →Converting web articles to markdown notes
About this skill
URL to Markdown
Fetches any URL via baoyu-fetch CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.
User Input Tools
When this skill prompts the user, follow this tool-selection rule (priority order):
- Prefer built-in user-input tools exposed by the current agent runtime — e.g.,
AskUserQuestion,request_user_input,clarify,ask_user, or any equivalent. - Fallback: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.
- Batching: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.
Concrete AskUserQuestion references below are examples — substitute the local equivalent in other runtimes.
CLI Setup
Important: The CLI source is vendored in {baseDir}/scripts/lib. scripts/package.json installs only third-party runtime dependencies.
Agent Execution Instructions:
- Determine this SKILL.md file's directory path as
{baseDir} - Resolve
${BUN}runtime: ifbuninstalled →bun; else suggest installing Bun - If
{baseDir}/scripts/node_modulesdoes not exist, run${BUN} install --cwd {baseDir}/scripts ${READER}={baseDir}/scripts/baoyu-fetch- Replace all
${READER}in this document with the resolved value
Preferences (EXTEND.md)
Check EXTEND.md in priority order — the first one found wins:
| Priority | Path | Scope |
|---|---|---|
| 1 | .baoyu-skills/baoyu-url-to-markdown/EXTEND.md | Project |
| 2 | ${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md | XDG |
| 3 | $HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md | User home |
| Result | Action |
|---|---|
| Found | Read, parse, apply settings |
| Not found | MUST run first-time setup (see below) — do NOT silently create defaults |
EXTEND.md supports: download media by default, default output directory.
First-Time Setup ⛔ BLOCKING
When EXTEND.md is not found, you MUST use AskUserQuestion to gather preferences before creating EXTEND.md. NEVER create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call:
- Q1 — Media (header "Media"): "How to handle images and videos in pages?"
- "Ask each time (Recommended)" — Prompt after each save
- "Always download" — Download to local
imgs/andvideos/ - "Never download" — Keep remote URLs
- Q2 — Output (header "Output"): "Default output directory?"
- "url-to-markdown (Recommended)" — Save to
./url-to-markdown/{domain}/{slug}.md - User may pick "Other" and type a custom path
- "url-to-markdown (Recommended)" — Save to
- Q3 — Save (header "Save"): "Where to save preferences?"
- "User (Recommended)" —
~/.baoyu-skills/(all projects) - "Project" —
.baoyu-skills/(this project only)
- "User (Recommended)" —
After answers, write EXTEND.md, confirm "Preferences saved to [path]", then continue.
Full template: references/config/first-time-setup.md.
Supported Keys
| Key | Default | Values | Description |
|---|---|---|---|
download_media | ask | ask / 1 / 0 | ask = prompt each time, 1 = always, 0 = never |
default_output_dir | empty | path or empty | Default output directory (empty = ./url-to-markdown/) |
EXTEND.md → CLI mapping:
| EXTEND.md key | CLI argument | Notes |
|---|---|---|
download_media: 1 | --download-media | Requires --output to be set |
default_output_dir: ./posts/ | Agent constructs --output ./posts/{domain}/{slug}.md | Agent generates path, not a direct flag |
Value priority: CLI arguments → EXTEND.md → skill defaults.
Usage
# Default: headless capture, markdown to stdout
${READER} <url>
# Save to file
${READER} <url> --output article.md
# Save with media download
${READER} <url> --output article.md --download-media
# Wait for interaction (login/CAPTCHA) — auto-detect and continue
${READER} <url> --wait-for interaction --output article.md
# Wait for interaction — manual control (Enter to continue)
${READER} <url> --wait-for force --output article.md
# JSON output
${READER} <url> --format json --output article.json
# Force specific adapter
${READER} <url> --adapter youtube --output transcript.md
Options
| Option | Description |
|---|---|
<url> | URL to fetch |
--output <path> | Output file path (default: stdout) |
--format <type> | Output format: markdown (default) or json |
--json | Shorthand for --format json |
--adapter <name> | Force adapter: x, youtube, hn, or generic (default: auto-detect) |
--headless | Force headless Chrome (no visible window) |
--wait-for <mode> | Interaction wait mode: none (default), interaction, or force |
--wait-for-interaction | Alias for --wait-for interaction |
--wait-for-login | Alias for --wait-for interaction |
--timeout <ms> | Page load timeout (default: 30000) |
--interaction-timeout <ms> | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |
--interaction-poll-interval <ms> | Poll interval for interaction checks (default: 1500) |
--download-media | Download images/videos to local imgs/ and videos/, rewrite markdown links. Requires --output |
--media-dir <dir> | Base directory for downloaded media (default: same as --output directory) |
--cdp-url <url> | Reuse existing Chrome DevTools Protocol endpoint |
--browser-path <path> | Custom Chrome/Chromium binary path |
--chrome-profile-dir <path> | Chrome user data directory (default: BAOYU_CHROME_PROFILE_DIR env or ./baoyu-skills/chrome-profile) |
--debug-dir <dir> | Write debug artifacts (document.json, markdown.md, page.html, network.json) |
Agent Quality Gate
CRITICAL: treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without failing the CLI.
After every headless run, inspect the saved markdown. See references/quality-gate.md for the full checklist, recovery workflow, and capture-mode table. Read it whenever a run looks suspicious or the user asks about login/CAPTCHA handling.
Output Path Generation
The agent must construct the output file path — baoyu-fetch does not auto-generate paths.
Algorithm:
- Determine base directory from EXTEND.md
default_output_diror default./url-to-markdown/ - Extract domain from URL (e.g.,
example.com) - Generate slug from URL path or page title (kebab-case, 2-6 words)
- Construct:
{base_dir}/{domain}/{slug}/{slug}.md— each URL gets its own directory so media files stay isolated - Conflict resolution: append timestamp
{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md
Pass the constructed path to --output. Media files (--download-media) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.
Adapters & Media
See references/adapters.md for the adapter catalog (X, YouTube, Hacker News, generic), per-adapter notes, the media download flow (ask / always / never), and the JSON output schema. Read it before answering adapter-specific questions or handling media prompts.
Environment Variables
| Variable | Description |
|---|---|
BAOYU_CHROME_PROFILE_DIR | Chrome user data directory (can also use --chrome-profile-dir) |
Troubleshooting: Chrome not found → use --browser-path. Timeout → increase --timeout. Login/CAPTCHA → --wait-for interaction. Debug → --debug-dir to inspect captured HTML and network logs.
Extension Support
Custom configurations via EXTEND.md. See Preferences section above for paths and supported keys.
When not to use it
- →When the target URL is inaccessible to the Chrome browser
- →When the user requires a non-markdown or non-JSON output format
Prerequisites
Limitations
- →Headless capture may result in low-quality content for some sites
- →Requires manual inspection of output to verify capture quality
- →Does not auto-generate output paths without agent logic
How it compares
Unlike manual copy-pasting, this tool automates the extraction of clean markdown and handles complex web interactions through programmatic browser control.
Compared to similar skills
baoyu-url-to-markdown side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| baoyu-url-to-markdown (this skill) | 9 | 3mo | Review | Intermediate |
| web-to-markdown | 7 | 6mo | No flags | Intermediate |
| defuddle | 0 | 2mo | Review | Beginner |
| firecrawl-core-workflow-a | 0 | 27d | Review | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by JimLiu
View all by JimLiu →You might also like
web-to-markdown
davila7
Use ONLY when the user explicitly says: 'use the skill web-to-markdown ...' (or 'use a skill web-to-markdown ...'). Converts webpage URLs to clean Markdown by calling the local web2md CLI (Puppeteer + Readability), suitable for JS-rendered pages.
defuddle
ludotype
Extract clean markdown content from web pages using Defuddle CLI, removing clutter and navigation to save tokens. Use instead of WebFetch when the user provides a URL to read or analyze, for online documentation, articles, blog posts, or any standard web page. Do NOT use for URLs ending in .md — tho
firecrawl-core-workflow-a
jeremylongshore
Execute FireCrawl primary workflow: Core Workflow A. Use when implementing primary use case, building main features, or core integration tasks. Trigger with phrases like "firecrawl main workflow", "primary task with firecrawl".
okf-to-web
yzfly
把符合 OKF v0.1 的知识包打包成单个自包含且经过压缩(minify)的 HTML 文件, 内嵌导航, Markdown 阅读器与概念关系图谱, 数据不出页面, 无需后端. 适用场景包括把 OKF bundle 变成一个可分享的单文件网页, 离线浏览知识库, 生成对标官方 viz.html 的可视化. 当用户提到 okf to web, OKF 转单文件网页, 可视化 OKF, minify 知识库, 或要一个自包含 HTML 时触发.
fetch-url-md
1naichii
Fetch web content with automatic markdown version detection using curl. Use when Claude needs to retrieve documentation from websites that offer both HTML and markdown formats. First checks if a .md version exists (more efficient and cleaner), then falls back to HTML if unavailable. Ideal for fetchi
pdf-to-markdown
aliceisjustplaying
Convert entire PDF documents to clean, structured Markdown for full context loading. Use this skill when the user wants to extract ALL text from a PDF into context (not grep/search), when discussing or analyzing PDF content in full, when the user mentions "load the whole PDF", "bring the PDF into context", "read the entire PDF", or when partial extraction/grepping would miss important context. This is the preferred method for PDF text extraction over page-by-page or grep approaches.