firecrawl-reliability-patterns
Patterns for building fault-tolerant Firecrawl scrapers with timeouts, retries, and content validation.
Install
mkdir -p .claude/skills/firecrawl-reliability-patterns && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2760" && unzip -o skill.zip -d .claude/skills/firecrawl-reliability-patterns && rm skill.zipInstalls to .claude/skills/firecrawl-reliability-patterns
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Implement Firecrawl reliability patterns: circuit breakers, crawl fallbacks,Key capabilities
- →Implement reliable crawl with timeout and exponential backoff
- →Validate scraped content quality based on length and error patterns
- →Implement crawl-to-scrape fallback for critical pages
- →Create a circuit breaker for Firecrawl operations
- →Track and enforce credit usage with a CreditGuard
- →Build a resilient pipeline with URL mapping and content validation
How it works
The skill provides code examples for implementing a reliable crawl with timeout and backoff, validating scraped content, and falling back to individual scrapes if a batch fails. It also includes a circuit breaker to prevent cascading failures and a CreditGuard to manage API credit usage.
Inputs & outputs
When to use firecrawl-reliability-patterns
- →Implement circuit breakers for scrapers
- →Add crawl timeout handling
- →Build scraping fallbacks
- →Validate scraped content quality
About this skill
Firecrawl Reliability Controls
Overview
Make every unit of work recoverable without duplicate collection or downstream writes. Provider retry, client retry, queue replay, webhook redelivery, and operator replay must share one idempotency model.
Prerequisites
- The target repository or integration path and the requested operator outcome.
- The source authorization, data classification, and environment policy.
- Current Firecrawl documentation, credentials only when needed, and an owner for approvals.
Current Contract
Only errors classified retryable by Firecrawl's current error catalog should enter automatic backoff. Async crawl, batch, agent, and related jobs require durable IDs and terminal-state handling; crawl and batch results may paginate. Webhooks retry delivery on their schedule, while status reconciliation remains necessary after missed or exhausted delivery.
Authentication
For authenticated Cloud operations, inject FIRECRAWL_API_KEY from an approved secret manager. REST requests use Authorization: Bearer with the key. Never print, commit, transmit, or place a key in a URL. Keyless access is suitable only where the current documentation explicitly allows it and the workload accepts its limits; production workflows should make identity and team ownership explicit.
Instructions
- Define a stable work identity from tenant, approved source, operation, policy version, and requested version; store attempts and downstream state durably.
- Separate submission from completion. Persist the provider job ID before acknowledging work and resume status/pagination from durable state.
- Classify validation, authentication, credits, restrictions, origin responses, rate/concurrency, timeouts, server failures, and downstream errors independently.
- Retry only documented retryable classes with jitter, Retry-After, maximum attempts, total deadline, and a circuit breaker. Never rotate keys to evade limits.
- Make page acceptance, storage, indexing, webhook handling, and tombstones idempotent. Deduplicate webhook deliveries by webhookId plus event context.
- Reconcile provider job state, all result pages, accepted/rejected totals, downstream receipts, and spend on a schedule; cancel abandoned work where safe.
- Exercise crash-after-submit, crash-after-write, duplicate webhook, partial pagination, provider outage, stale cache, and rollback in tests.
Tool Discipline
Use Read, Glob, and Grep to inspect code, configuration, tests, and evidence. Use Write/Edit only for approved implementation or documentation changes. Do not call Firecrawl, rotate keys, change account settings, scrape a target, or deploy merely because this skill was invoked.
Approval Boundaries
Require approval before expanding retry budgets, replaying production jobs, cancelling shared work, using stale data as a degraded mode, or bypassing a circuit breaker.
Output
Return the state machine, idempotency keys, retry matrix, circuit and backpressure policy, reconciliation algorithm, degraded modes, test results, and recovery/rollback receipt.
Error Handling
- Job ID was not persisted: search only through approved evidence; do not submit a duplicate blindly.
- Pagination is incomplete: mark the result partial and withhold completeness claims.
- Provider and downstream state cannot reconcile: stop automated replay and escalate with redacted evidence.
Examples
- "Make crawl jobs restartable" persists job and pagination state before acknowledgment.
- "Retry every failure forever" is replaced with documented classes, deadlines, and a circuit breaker.
Resources
Read official Firecrawl evidence before relying on an endpoint, SDK method, plan limit, price, retention option, or self-hosted release.
When not to use it
- →When the primary concern is not fault-tolerant scraping pipelines
- →When the integration does not require crawl-to-scrape fallback
- →When content quality gates are not needed for Firecrawl integrations
Prerequisites
Limitations
- →Crawl jobs may timeout for large sites or slow rendering
- →Scraped content may be empty due to bot detection or JavaScript failures
- →Credit overruns can occur without proper budget tracking
How it compares
This skill provides specific code implementations for Firecrawl reliability patterns like circuit breakers and crawl-to-scrape fallbacks, addressing challenges unique to Firecrawl's async and credit-based model, unlike general reliability p
Compared to similar skills
firecrawl-reliability-patterns side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| firecrawl-reliability-patterns (this skill) | 3 | 2mo | Review | Advanced |
| chrome-devtools | 41 | 8mo | Review | Intermediate |
| groq-performance-tuning | 1 | 2mo | No flags | Intermediate |
| openrouter-streaming-setup | 1 | 2mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
chrome-devtools
mrgoonie
Browser automation, debugging, and performance analysis using Puppeteer CLI scripts. Use for automating browsers, taking screenshots, analyzing performance, monitoring network traffic, web scraping, form automation, and JavaScript debugging.
groq-performance-tuning
jeremylongshore
Optimize Groq API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Groq integrations. Trigger with phrases like "groq performance", "optimize groq", "groq latency", "groq caching", "groq slow", "groq batch".
openrouter-streaming-setup
jeremylongshore
Implement streaming responses with OpenRouter. Use when building real-time chat interfaces or reducing time-to-first-token. Trigger with phrases like 'openrouter streaming', 'openrouter sse', 'stream response', 'real-time openrouter'.
ideogram-rate-limits
jeremylongshore
Implement Ideogram rate limiting, backoff, and idempotency patterns. Use when handling rate limit errors, implementing retry logic, or optimizing API request throughput for Ideogram. Trigger with phrases like "ideogram rate limit", "ideogram throttling", "ideogram 429", "ideogram retry", "ideogram backoff".
linear-rate-limits
jeremylongshore
Handle Linear API rate limiting and quotas effectively. Use when dealing with rate limit errors, implementing throttling, or optimizing API usage patterns. Trigger with phrases like "linear rate limit", "linear throttling", "linear API quota", "linear 429 error", "linear request limits".
generating-grpc-services
jeremylongshore
Generate gRPC service definitions, stubs, and implementations from Protocol Buffers. Use when creating high-performance gRPC services. Trigger with phrases like "generate gRPC service", "create gRPC API", or "build gRPC server".