firecrawl-architecture-variants
Architecture patterns for Firecrawl implementations based on scale and volume.
Install
mkdir -p .claude/skills/firecrawl-architecture-variants && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/5437" && unzip -o skill.zip -d .claude/skills/firecrawl-architecture-variants && rm skill.zipInstalls to .claude/skills/firecrawl-architecture-variants
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Choose and implement Firecrawl architecture patterns for different scalesKey capabilities
- →Design on-demand scraping for single-page extraction.
- →Implement scheduled crawl pipelines for content indexing.
- →Build real-time ingestion pipelines for AI/RAG applications.
- →Select architecture based on volume and latency requirements.
- →Manage credit control for different scraping patterns.
How it works
This skill provides three architecture patterns: on-demand for single requests, scheduled for periodic crawls, and real-time for high-volume ingestion. Each pattern uses Firecrawl functions to scrape or crawl URLs and process the results.
Inputs & outputs
When to use firecrawl-architecture-variants
- →Designing a RAG ingestion pipeline
- →Setting up site monitoring
- →Optimizing scraping for high volume
About this skill
Firecrawl Architecture Selection
Overview
Turn workload, freshness, compliance, throughput, and recovery requirements into an explicit architecture decision. Prefer the smallest Firecrawl surface that meets the outcome.
Prerequisites
- The target repository or integration path and the requested operator outcome.
- The source authorization, data classification, and environment policy.
- Current Firecrawl documentation, credentials only when needed, and an owner for approvals.
Current Contract
Scrape is a synchronous single-resource primitive; crawl and batch scrape support asynchronous jobs, pagination, webhooks, and status retrieval; map discovers URLs without retrieving page content; search discovers external sources; parse handles local file bytes. The default self-hosted Compose stack does not provide every Cloud capability and is not a production security design.
Authentication
For authenticated Cloud operations, inject FIRECRAWL_API_KEY from an approved secret manager. REST requests use Authorization: Bearer with the key. Never print, commit, transmit, or place a key in a URL. Keyless access is suitable only where the current documentation explicitly allows it and the workload accepts its limits; production workflows should make identity and team ownership explicit.
Instructions
- Inventory the source domains, ownership or authorization basis, expected page volume, freshness target, output formats, data classification, and recovery objective.
- Choose scrape for bounded single pages, map plus selective scrape for curated URL sets, batch scrape for known URL collections, and crawl for recursive site discovery.
- Choose polling, WebSocket streaming, or signed webhooks for asynchronous delivery based on network topology and recovery needs.
- Decide Cloud versus self-hosting using required capabilities, data flows, operator staffing, availability target, and upgrade ownership. Record unsupported self-hosted features explicitly.
- Place a policy gateway before Firecrawl for domain authorization, request shaping, budgets, key selection, and audit receipts. Keep content storage and indexing downstream.
- Define queue limits, explicit crawl limits, retention/cache choices, idempotency keys, pagination ownership, and degraded modes.
- Validate the chosen variant with one approved canary and document failure, retry, cancellation, and rollback paths.
Tool Discipline
Use Read, Glob, and Grep to inspect code, configuration, tests, and evidence. Use Write/Edit only for approved implementation or documentation changes. Do not call Firecrawl, rotate keys, change account settings, scrape a target, or deploy merely because this skill was invoked.
Approval Boundaries
Require approval before selecting self-hosting, introducing a new external provider, permitting authenticated-page capture, enabling Cloud-only capabilities, or broadening domains and retention.
Output
Return an architecture decision record containing workload facts, chosen endpoints, sequence, trust boundaries, data flows, capacity assumptions, recovery path, alternatives rejected, validation evidence, and open approvals.
Error Handling
- Requirements conflict: surface the conflict and propose bounded variants instead of hiding it in implementation.
- Self-host feature is unsupported: choose Cloud or identify and validate the required external service.
- Recovery path is undefined: block launch until cancellation, replay, and deduplication behavior is owned.
Examples
- "Design a nightly documentation ingest" selects a bounded crawl or map-plus-batch variant with async recovery.
- "Keep all traffic inside our infrastructure" evaluates self-hosted capability gaps and operating cost before choosing it.
Resources
Read official Firecrawl evidence before relying on an endpoint, SDK method, plan limit, price, retention option, or self-hosted release.
When not to use it
- →When real-time, user-facing response is not needed and volume is less than 500 pages/day.
- →When volume is between 500-10K pages/day and real-time response is not required.
Limitations
- →On-demand scraping is best for less than 500 pages per day.
- →Scheduled pipelines are for 500-10K pages per day.
- →Real-time pipelines are for 10K+ pages per day.
How it compares
This skill offers validated architecture blueprints for Firecrawl, unlike manually designing a scraping solution from scratch.
Compared to similar skills
firecrawl-architecture-variants side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| firecrawl-architecture-variants (this skill) | 1 | 2mo | Review | Intermediate |
| dev-browser | 53 | 6mo | Review | Intermediate |
| openspec-onboard | 10 | 8mo | Review | Beginner |
| workflow-orchestration-patterns | 10 | 4mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
dev-browser
SawyerHood
Browser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.
openspec-onboard
studyzy
Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.
workflow-orchestration-patterns
wshobson
Design durable workflows with Temporal for distributed systems. Covers workflow vs activity separation, saga patterns, state management, and determinism constraints. Use when building long-running processes, distributed transactions, or microservice orchestration.
simple-fetch
yoloshii
Basic MCP skill demonstrating CLI-based execution pattern for fetching URL content
agent-browser
vercel-labs
Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction.
meta-automation-architect
comzine
Use when user wants to set up comprehensive automation for their project. Generates custom subagents, skills, commands, and hooks tailored to project needs. Creates a multi-agent system with robust communication protocol.