firecrawl-reference-architecture
Build production-ready content ingestion pipelines with Firecrawl using scrape, crawl, and extract endpoints.
Install
mkdir -p .claude/skills/firecrawl-reference-architecture && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4122" && unzip -o skill.zip -d .claude/skills/firecrawl-reference-architecture && rm skill.zipInstalls to .claude/skills/firecrawl-reference-architecture
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Implement Firecrawl reference architecture with scrape/crawl/map/extractKey capabilities
- →Design three-tier scraping and ingestion pipelines
- →Implement on-demand scraping and scheduled crawls
- →Perform site-wide URL discovery and mapping
- →Process and deduplicate scraped markdown content
- →Build RAG-ready knowledge bases
How it works
It defines a pipeline architecture using Firecrawl endpoints to map, crawl, and extract content, followed by a processing layer that cleans, hashes, and chunks data for storage.
Inputs & outputs
When to use firecrawl-reference-architecture
- →Design new Firecrawl integrations
- →Structure content ingestion for RAG
- →Establish architectural standards for web scraping
- →Build pipelines for AI knowledge bases
About this skill
Firecrawl Governed Ingestion Architecture
Overview
Create explicit trust boundaries from request intake through deletion. Keep source discovery, content retrieval, validation, and downstream publication independently retryable and auditable.
Prerequisites
- The target repository or integration path and the requested operator outcome.
- The source authorization, data classification, and environment policy.
- Current Firecrawl documentation, credentials only when needed, and an owner for approvals.
Current Contract
Firecrawl v2 provides scrape, crawl, map, search, batch scrape, parse, JSON extraction, agentic, and browser surfaces with different async, credit, retention, and availability contracts. The SDK can auto-wait and paginate, while explicit start/status methods support durable orchestration. Provider webhooks supplement but do not replace reconciliation.
Authentication
For authenticated Cloud operations, inject FIRECRAWL_API_KEY from an approved secret manager. REST requests use Authorization: Bearer with the key. Never print, commit, transmit, or place a key in a URL. Keyless access is suitable only where the current documentation explicitly allows it and the workload accepts its limits; production workflows should make identity and team ownership explicit.
Instructions
- Define tenants, source authorization, freshness, scale, formats, data classification, SLO/RPO/RTO, retention, deletion, and cost constraints.
- Place authentication, tenant isolation, URL canonicalization, policy versioning, endpoint/format allowlists, budgets, and idempotency at the intake gateway.
- Separate discovery into map/search, retrieval into scrape/batch/crawl/parse, and model extraction into a validated stage. Use durable job records for asynchronous work.
- Persist normalized provenance and job/page state before downstream writes. Complete pagination and deduplicate by tenant, canonical source, version, and content hash.
- Run schema, origin-status, content-quality, malware/active-content, and prompt-injection gates before storage or agent context.
- Use an outbox or equivalent for indexing and publication; make retries idempotent and support tombstones across raw, normalized, embedding, cache, and search stores.
- Observe queue, job, origin, validation, freshness, spend, webhook, and deletion SLIs; test restore, reconciliation, cancellation, and regional/provider failure.
Tool Discipline
Use Read, Glob, and Grep to inspect code, configuration, tests, and evidence. Use Write/Edit only for approved implementation or documentation changes. Do not call Firecrawl, rotate keys, change account settings, scrape a target, or deploy merely because this skill was invoked.
Approval Boundaries
Require architecture and security/data approval before adding a new endpoint class, authenticated source, model extraction, long-term store, cross-region flow, or self-hosted provider.
Output
Return components and trust boundaries, sequence and state model, contracts, capacity and cost assumptions, security/privacy controls, failure and reconciliation paths, SLOs, tests, and phased rollout.
Error Handling
- A stage lacks durable identity or idempotency: block asynchronous rollout.
- Provider completion and stored counts diverge: reconcile pagination and downstream receipts before publication.
- Deletion cannot propagate to derived stores: fail the data architecture review.
Examples
- "Build a docs RAG pipeline" produces policy, acquisition, validation, outbox, index, and deletion stages.
- "Call crawl directly from the browser" is replaced with a server-side governed gateway.
Resources
Read official Firecrawl evidence before relying on an endpoint, SDK method, plan limit, price, retention option, or self-hosted release.
When not to use it
- →When real-time, sub-millisecond data access is required
- →When site structure is too dynamic for map-based discovery
Prerequisites
Limitations
- →Crawl depth is limited by configuration
- →Content deduplication relies on hash comparison
How it compares
This architecture standardizes content ingestion for AI applications compared to ad-hoc, unmanaged scraping scripts.
Compared to similar skills
firecrawl-reference-architecture side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| firecrawl-reference-architecture (this skill) | 1 | 2mo | Review | Intermediate |
| mcp-builder | 136 | 5mo | Review | Advanced |
| architecture-patterns | 55 | 4mo | No flags | Advanced |
| nodejs-best-practices | 28 | 8mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
mcp-builder
anthropics
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
architecture-patterns
wshobson
Implement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use when architecting complex backend systems or refactoring existing applications for better maintainability.
nodejs-best-practices
davila7
Node.js development principles and decision-making. Framework selection, async patterns, security, and architecture. Teaches thinking, not copying.
senior-fullstack
davila7
Comprehensive fullstack development skill for building complete web applications with React, Next.js, Node.js, GraphQL, and PostgreSQL. Includes project scaffolding, code quality analysis, architecture patterns, and complete tech stack guidance. Use when building new projects, analyzing code quality, implementing design patterns, or setting up development workflows.
workflow-orchestration-patterns
wshobson
Design durable workflows with Temporal for distributed systems. Covers workflow vs activity separation, saga patterns, state management, and determinism constraints. Use when building long-running processes, distributed transactions, or microservice orchestration.
langchain-architecture
wshobson
Design LLM applications using the LangChain framework with agents, memory, and tool integration patterns. Use when building LangChain applications, implementing AI agents, or creating complex LLM workflows.