Retrieves Reddit post and comment data directly via their JSON API to bypass network restrictions.
Install
mkdir -p .claude/skills/reddit-fetch && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/1150" && unzip -o skill.zip -d .claude/skills/reddit-fetch && rm skill.zipInstalls to .claude/skills/reddit-fetch
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Fetch content from Reddit using the curl JSON API. Use when accessing Reddit URLs, researching topics on Reddit, or when Reddit returns 403/blocked errors.Key capabilities
- →Accesses Reddit JSON API via curl and user-agent emulation
- →Filters post listings by category (hot, new, top)
- →Parses JSON threads into readable comment trees
- →Saves output to temporary files to avoid pipeline encoding errors
How it works
It queries the public Reddit JSON API using simulated browser headers to bypass blockages and parses the payload via jq.
Inputs & outputs
When to use reddit-fetch
- →Researching community discussions on specific topics
- →Fetching subreddit post data when standard web access fails
- →Parsing Reddit comment trees via CLI
About this skill
Reddit Fetch
Reddit's public JSON API works by appending .json to any Reddit URL — but Reddit now hard-blocks automated access. curl 403s essentially every time (host AND container, regardless of User-Agent), and even a cold Playwright navigation to reddit.com hits a "You've been blocked by network security" challenge page. The reliable method is to arrive at Reddit through a DuckDuckGo result redirect: that sets a Reddit session cookie which unlocks direct .json access for the rest of the browser session.
Primary method: the DuckDuckGo-hop unlock
Step 1 - the DDG hop. Do this once per session before any .json fetch.
mcp__playwright__browser_navigatetohttps://html.duckduckgo.com/html/?q=site:reddit.com/r/SUBREDDIT+YOUR+QUERY- Grab the first result's full href - it's a DDG redirect that includes a
ruttoken (https://duckduckgo.com/l/?uddg=...&rut=...). The token is required; navigating to the bare/l/?uddg=without it 400s.() => document.querySelector('.result__a')?.href browser_navigateto that full redirect href. It lands on a realwww.reddit.compage (title = the post/subreddit title, not "Blocked") and sets the session cookie. The result doesn't have to be the exact thread you want - landing on any real Reddit page sets the cookie.
If Playwright errors with Browser is already in use, a stale instance is holding the profile - pkill -f ms-playwright-mcp (or the profile dir named in the error) and retry.
Step 2 - direct .json now works. For the rest of the session, navigate Playwright straight to any .json URL and JSON.parse(document.body.innerText). Use www.reddit.com (not old.reddit.com) for browser navigation. Full recency sorting (sort=new&t=week) is available.
browser_navigateto e.g.https://www.reddit.com/r/SUBREDDIT/search.json?q=QUERY&restrict_sr=on&sort=new&t=week&limit=25browser_evaluate, always wrapped in try/catch (returndocument.body.innerText.slice(0,200)on failure so you can see a challenge page if the session lapsed - just re-do the hop):() => { try { const data = JSON.parse(document.body.innerText); return data.data.children.map(c => ({ t: c.data.title, s: c.data.score, n: c.data.num_comments, id: c.data.id })); } catch (e) { return document.body.innerText.slice(0, 200); } }- For a thread, navigate to
.../comments/POST_ID.json?limit=30&sort=topand parsedata[0](post) anddata[1].data.children(comments).
JSON shapes (same for browser and curl)
# Listing - swap hot for new/top/rising; for top add &t=day|week|month|year|all
/r/SUBREDDIT/hot.json?limit=15
# Post + comments - JSON array where [0]=post, [1]=comment tree
/r/SUBREDDIT/comments/POST_ID.json?limit=20
# Search within a subreddit
/r/SUBREDDIT/search.json?q=QUERY&restrict_sr=on&sort=new&limit=15
- Listings:
.data.children[].datahastitle,score,num_comments,author,id. - Threads:
[0].data.children[0].datais the post;[1].data.children[](filterkind == "t1") are comments withauthor,score,body, and nestedrepliesof the same shape. - Truncate long comment bodies (e.g.
body.slice(0, 300)in JS,.body[:300]in jq 1.7+) to keep output readable.
Fallbacks
- More comments per thread: load the rendered thread page (after the hop) and scrape the
shredditDOM - it returns more comments than.json?limit=:() => ({ title: document.querySelector('shreddit-post')?.getAttribute('post-title'), comments: [...document.querySelectorAll('shreddit-comment')].map(c => ({ author: c.getAttribute('author'), score: c.getAttribute('score'), text: c.querySelector('.md')?.innerText })) }) - No Playwright at all: use Claude for Chrome to open the thread /
.jsonURL and read it off the page. - curl (last resort, expect 403): direct curl with a browser User-Agent used to work and is faster when it does, but it now gets 403'd essentially always - changing the UA doesn't help. Only worth a single quick try if you're already shelling out and a browser isn't available; on 403, go straight to the DDG hop.
Fetch to a temp file (UA="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" curl -s -L -o /tmp/reddit_result.txt -w "%{http_code}" -H "User-Agent: $UA" \ 'https://old.reddit.com/r/SUBREDDIT/hot.json?limit=15' jq -r '.data.children[] | .data | "\(.title)\n \(.score) pts | \(.num_comments) comments | u/\(.author) | id: \(.id)\n"' /tmp/reddit_result.txt-o) then parse;-w "%{http_code}"prints the status;-Lfollows redirects; single-quote the URL so the shell doesn't eat&.
Rate limiting
Reddit rate-limits aggressively even once you're unblocked:
- Don't fire parallel requests - run them sequentially with
sleep 2/sleep 3(or brief pauses between navigations). Fetch one listing, parse it, then fetch threads one at a time. - Empty response (0 bytes): wait 3-5s and retry. HTTP 429: back off 10-15s. A challenge page mid-session means the cookie lapsed - re-do the DDG hop.
When not to use it
- →High-frequency scraping tasks violating TOS
- →Accessing private or NSFW subreddits
Prerequisites
Limitations
- →JSON API may rate limit frequent requests
- →Formatted output depends on standard JSON structure
How it compares
It bypasses standard WebFetch blocks by directly hitting the JSON API interface.
Compared to similar skills
reddit-fetch side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| reddit-fetch (this skill) | 9 | 2mo | Review | Beginner |
| brightdata-web-mcp | 9 | 6mo | Review | Intermediate |
| firecrawl-scrape | 5 | 7mo | Review | Beginner |
| tavily-web | 5 | 6mo | Review | Beginner |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by ykdojo
View all by ykdojo →You might also like
brightdata-web-mcp
patchy631
Search the web, scrape websites, extract structured data from URLs, and automate browsers using Bright Data's Web MCP. Use when fetching live web content, bypassing blocks/CAPTCHAs, getting product data from Amazon/eBay, social media posts, or when standard requests fail.
firecrawl-scrape
parcadei
Scrape web pages and extract content via Firecrawl MCP
tavily-web
davila7
Web search, content extraction, crawling, and research capabilities using Tavily API
webclaw
0xmassi
Web extraction engine with antibot bypass. Scrape, crawl, extract, summarize, search, map, diff, monitor, research, and analyze any URL — including Cloudflare-protected sites. Use when you need reliable web content, the built-in web_fetch fails, or you need structured data extraction from web pages.
batch-research
miantiao-me
批量数据采集技能,负责分批并发调度 researcher agent 抓取所有数据源。
firecrawl
hardjunior
|