firecrawl-core-workflow-b
Use this skill for LLM-powered data extraction, batch scraping of multiple URLs, and rapid site discovery with Firecrawl.
Install
mkdir -p .claude/skills/firecrawl-core-workflow-b && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/7840" && unzip -o skill.zip -d .claude/skills/firecrawl-core-workflow-b && rm skill.zipInstalls to .claude/skills/firecrawl-core-workflow-b
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Execute Firecrawl secondary workflow: LLM extraction, batch scraping,Key capabilities
- →Extract structured data using JSON schemas
- →Perform batch scraping of multiple URLs
- →Discover site structure using map endpoint
- →Execute async batch scraping for large URL sets
- →Filter and scrape site sections intelligently
How it works
It utilizes LLM-powered extraction to convert unstructured web content into typed JSON based on provided schemas. It also provides batching and mapping utilities to handle multiple pages and site discovery efficiently.
Inputs & outputs
When to use firecrawl-core-workflow-b
- →Extract structured product pricing from web pages
- →Batch process a list of known URLs
- →Map entire website structures
- →Convert unstructured HTML to typed JSON
About this skill
Firecrawl Discovery and Structured Extraction
Overview
Discover first, select deliberately, and retrieve only the sources required for the outcome. Separate URL discovery from content acquisition and validate all model-produced JSON.
Prerequisites
- The target repository or integration path and the requested operator outcome.
- The source authorization, data classification, and environment policy.
- Current Firecrawl documentation, credentials only when needed, and an owner for approvals.
Current Contract
Use map for site URLs, search for web discovery, batchScrape or startBatchScrape for known URL sets, and parse for local/non-public file bytes. In v2, structured extraction is a json format object with prompt and optional JSON Schema; the legacy extract format name is not the current scrape contract.
Authentication
For authenticated Cloud operations, inject FIRECRAWL_API_KEY from an approved secret manager. REST requests use Authorization: Bearer with the key. Never print, commit, transmit, or place a key in a URL. Keyless access is suitable only where the current documentation explicitly allows it and the workload accepts its limits; production workflows should make identity and team ownership explicit.
Instructions
- Define the research question, approved domains, recency, maximum results/pages, required fields, JSON Schema, and evidence-retention policy.
- Use map for an owned site or search for broader discovery. Store normalized URLs and source metadata, not content, until selection policy passes.
- Filter duplicates, unsupported schemes, disallowed domains, and unnecessary query variants. Require review for sensitive or authenticated sources.
- Choose batchScrape for known URLs and startBatchScrape when work must be asynchronous. Retrieve all required pages through getBatchScrapeStatus pagination.
- For local PDF, DOCX, XLSX, HTML, or other supported bytes, use parse rather than inventing a public URL. Apply file-size and classification gates first.
- For typed extraction, request the v2 json format with the narrowest schema and prompt. Treat output as untrusted model data and validate types, constraints, provenance, and completeness.
- Return discovery, selection, retrieval, validation, cost, and failure receipts as separate stages so partial results are auditable.
Tool Discipline
Use Read, Glob, and Grep to inspect code, configuration, tests, and evidence. Use Write/Edit only for approved implementation or documentation changes. Do not call Firecrawl, rotate keys, change account settings, scrape a target, or deploy merely because this skill was invoked.
Approval Boundaries
Require approval before broad web search, uploading non-public files, using LLM-backed extraction on sensitive content, increasing batch size, or retaining raw source content.
Output
Return the discovery set, selection rationale, chosen v2 operation, job/pagination state, schema validation results, source provenance, rejection reasons, and redacted cost/evidence receipt.
Error Handling
- Discovery returns too many URLs: tighten search/path policy before retrieval.
- JSON extraction is invalid or prompt-injection protection blocks it: quarantine and require review; never coerce silently.
- Parse input is unsupported or oversized: stop before upload and choose an approved preprocessing path.
Examples
- "Find and extract product pages" maps or searches, filters URLs, then runs a bounded typed batch.
- "Parse this private PDF" checks classification and approval before uploading bytes to /v2/parse.
Resources
Read official Firecrawl evidence before relying on an endpoint, SDK method, plan limit, price, retention option, or self-hosted release.
When not to use it
- →When raw markdown is sufficient and structured data is not needed
Prerequisites
Limitations
- →LLM extraction adds variable credit costs
- →Large batch sets require async polling
How it compares
This workflow shifts from simple document retrieval to structured data extraction, allowing for direct integration of web content into typed applications.
Compared to similar skills
firecrawl-core-workflow-b side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| firecrawl-core-workflow-b (this skill) | 1 | 2mo | Review | Intermediate |
| json-render-core | 3 | 3mo | No flags | Advanced |
| turborepo | 61 | 3mo | Review | Intermediate |
| senior-fullstack | 35 | 9mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
json-render-core
vercel-labs
Core package for defining schemas, catalogs, and AI prompt generation for json-render. Use when working with @json-render/core, defining schemas, creating catalogs, or building JSON specs for UI/video generation.
turborepo
vercel
Turborepo monorepo build system guidance. Triggers on: turbo.json, task pipelines, dependsOn, caching, remote cache, the "turbo" CLI, --filter, --affected, CI optimization, environment variables, internal packages, monorepo structure/best practices, and boundaries. Use when user: configures tasks/workflows/pipelines, creates packages, sets up monorepo, shares code between apps, runs changed/affected packages, debugs cache, or has apps/packages directories.
senior-fullstack
davila7
Comprehensive fullstack development skill for building complete web applications with React, Next.js, Node.js, GraphQL, and PostgreSQL. Includes project scaffolding, code quality analysis, architecture patterns, and complete tech stack guidance. Use when building new projects, analyzing code quality, implementing design patterns, or setting up development workflows.
oracle
openclaw
Best practices for using the oracle CLI (prompt + file bundling, engines, sessions, and file attachment patterns).
codex-skill
feiskyer
Use when user asks to leverage codex, gpt-5, or gpt-5.1 to implement something (usually implement a plan or feature designed by Claude). Provides non-interactive automation mode for hands-off task execution without approval prompts.
svelte-expert
Raudbjorn
Expert Svelte/SvelteKit development assistant for building components, utilities, and applications. Use when creating Svelte components, SvelteKit applications, implementing reactive patterns, handling state management, working with stores, transitions, animations, or any Svelte/SvelteKit development task. Includes comprehensive documentation access, code validation with svelte-autofixer, and playground link generation.