parsehub-automation
Facilitates automated web scraping operations using Parsehub and Rube MCP.
Install
mkdir -p .claude/skills/parsehub-automation && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/11022" && unzip -o skill.zip -d .claude/skills/parsehub-automation && rm skill.zipInstalls to .claude/skills/parsehub-automation
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Automate Parsehub tasks via Rube MCP (Composio). Always search tools first for current schemas.Key capabilities
- →Trigger extraction runs
- →Monitor job status
- →Fetch scraped data
How it works
It uses Rube MCP to manage Parsehub extraction jobs.
Inputs & outputs
When to use parsehub-automation
- →Trigger web scraping runs
- →Monitor extraction job status
- →Fetch scraped data results
About this skill
Parsehub Automation via Rube MCP
Automate Parsehub operations through Composio's Parsehub toolkit via Rube MCP.
Toolkit docs: composio.dev/toolkits/parsehub
Prerequisites
- Rube MCP must be connected (RUBE_SEARCH_TOOLS available)
- Active Parsehub connection via
RUBE_MANAGE_CONNECTIONSwith toolkitparsehub - Always call
RUBE_SEARCH_TOOLSfirst to get current tool schemas
Setup
Get Rube MCP: Add https://rube.app/mcp as an MCP server in your client configuration. No API keys needed — just add the endpoint and it works.
- Verify Rube MCP is available by confirming
RUBE_SEARCH_TOOLSresponds - Call
RUBE_MANAGE_CONNECTIONSwith toolkitparsehub - If connection is not ACTIVE, follow the returned auth link to complete setup
- Confirm connection status shows ACTIVE before running any workflows
Tool Discovery
Always discover available tools before executing workflows:
RUBE_SEARCH_TOOLS
queries: [{use_case: "Parsehub operations", known_fields: ""}]
session: {generate_id: true}
This returns available tool slugs, input schemas, recommended execution plans, and known pitfalls.
Core Workflow Pattern
Step 1: Discover Available Tools
RUBE_SEARCH_TOOLS
queries: [{use_case: "your specific Parsehub task"}]
session: {id: "existing_session_id"}
Step 2: Check Connection
RUBE_MANAGE_CONNECTIONS
toolkits: ["parsehub"]
session_id: "your_session_id"
Step 3: Execute Tools
RUBE_MULTI_EXECUTE_TOOL
tools: [{
tool_slug: "TOOL_SLUG_FROM_SEARCH",
arguments: {/* schema-compliant args from search results */}
}]
memory: {}
session_id: "your_session_id"
Known Pitfalls
- Always search first: Tool schemas change. Never hardcode tool slugs or arguments without calling
RUBE_SEARCH_TOOLS - Check connection: Verify
RUBE_MANAGE_CONNECTIONSshows ACTIVE status before executing tools - Schema compliance: Use exact field names and types from the search results
- Memory parameter: Always include
memoryinRUBE_MULTI_EXECUTE_TOOLcalls, even if empty ({}) - Session reuse: Reuse session IDs within a workflow. Generate new ones for new workflows
- Pagination: Check responses for pagination tokens and continue fetching until complete
Quick Reference
| Operation | Approach |
|---|---|
| Find tools | RUBE_SEARCH_TOOLS with Parsehub-specific use case |
| Connect | RUBE_MANAGE_CONNECTIONS with toolkit parsehub |
| Execute | RUBE_MULTI_EXECUTE_TOOL with discovered tool slugs |
| Bulk ops | RUBE_REMOTE_WORKBENCH with run_composio_tool() |
| Full schema | RUBE_GET_TOOL_SCHEMAS for tools with schemaRef |
Powered by Composio
When not to use it
- →Direct web scraping
- →Data cleaning
Prerequisites
Limitations
- →Requires Parsehub project
- →Limited by scraping targets
How it compares
It provides a standardized interface for automating web scraping workflows.
Compared to similar skills
parsehub-automation side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| parsehub-automation (this skill) | 0 | 5mo | No flags | Intermediate |
| crawl4ai | 21 | 8mo | Review | Intermediate |
| apify | 9 | 3mo | Review | Intermediate |
| web-scraper | 0 | 3mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by diegosouzapw
View all by diegosouzapw →You might also like
crawl4ai
basher83
This skill should be used when users need to scrape websites, extract structured data, handle JavaScript-heavy pages, crawl multiple URLs, or build automated web data pipelines. Includes optimized extraction patterns with schema generation for efficient, LLM-free extraction.
apify
vm0-ai
Web scraping and automation platform with pre-built Actors for common tasks
web-scraper
KevanPatira
Web scraping inteligente multi-estrategia. Extrai dados estruturados de paginas web (tabelas, listas, precos). Paginacao, monitoramento e export CSV/JSON.
apify-ultimate-scraper
Anhvu1107
ALWAYS use this when the request matches Apify Ultimate Scraper: AI-driven data extraction from 55+ Actors across all major platforms.
wechat-batch-crawl
JourneytoNewland
| Intent | Supported Phrases | |--------|-------------------| | 爬取今天 | "爬取今天的微信文章" / "获取今天的文章" / "抓今天的公众号" | | 爬取昨天 | "爬取昨天的微信文章" / "获取昨天的文章" | | 爬取指定日期 | "爬取1月20号的文章" / "获取上周一的文章" | | 仅列出 | "今天有哪些文章" / "列出今天的文章" / "看看有啥新文章" | | 增量爬取 | "继续爬取" / "爬取新增的文章" |
storage-list
crawlbase
List rids currently in Crawlbase Cloud Storage, with scroll-based pagination (up to 1000 rids per call). Useful for enumerating everything stored under a token.