firecrawl-core-workflow-b
Use this skill for LLM-powered data extraction, batch scraping of multiple URLs, and rapid site discovery with Firecrawl.
Install
mkdir -p .claude/skills/firecrawl-core-workflow-b && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/7840" && unzip -o skill.zip -d .claude/skills/firecrawl-core-workflow-b && rm skill.zipInstalls to .claude/skills/firecrawl-core-workflow-b
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Execute Firecrawl secondary workflow: LLM extraction, batch scraping,Key capabilities
- →Extract structured data using JSON schemas
- →Perform batch scraping of multiple URLs
- →Discover site structure using map endpoint
- →Execute async batch scraping for large URL sets
- →Filter and scrape site sections intelligently
How it works
It utilizes LLM-powered extraction to convert unstructured web content into typed JSON based on provided schemas. It also provides batching and mapping utilities to handle multiple pages and site discovery efficiently.
Inputs & outputs
When to use firecrawl-core-workflow-b
- →Extract structured product pricing from web pages
- →Batch process a list of known URLs
- →Map entire website structures
- →Convert unstructured HTML to typed JSON
About this skill
Firecrawl Core Workflow B — Extract, Batch & Map
Overview
Secondary workflow complementing the scrape/crawl workflow. Covers LLM-powered structured data extraction with JSON schemas, batch scraping multiple known URLs, and rapid site map discovery. Use this when you need typed data rather than raw markdown.
Prerequisites
@mendable/firecrawl-jsinstalledFIRECRAWL_API_KEYenvironment variable set- Understanding of JSON Schema (for extract)
Instructions
Step 1: LLM Extract — Structured Data from Pages
import FirecrawlApp from "@mendable/firecrawl-js";
const firecrawl = new FirecrawlApp({
apiKey: process.env.FIRECRAWL_API_KEY!,
});
// Extract structured data using an LLM + JSON schema
const result = await firecrawl.scrapeUrl("https://firecrawl.dev/pricing", {
formats: ["extract"],
extract: {
schema: {
type: "object",
properties: {
plans: {
type: "array",
items: {
type: "object",
properties: {
name: { type: "string" },
price: { type: "string" },
credits_per_month: { type: "number" },
features: { type: "array", items: { type: "string" } },
},
required: ["name", "price"],
},
},
},
},
},
});
console.log("Extracted plans:", JSON.stringify(result.extract, null, 2));
Step 2: Extract with Prompt (No Schema)
// Use natural language prompt instead of rigid schema
const result = await firecrawl.scrapeUrl("https://news.ycombinator.com", {
formats: ["extract"],
extract: {
prompt: "Extract the top 5 stories with their title, URL, points, and comment count",
},
});
console.log(result.extract);
Step 3: Batch Scrape Known URLs
// Scrape multiple specific URLs at once — more efficient than individual calls
const batchResult = await firecrawl.batchScrapeUrls(
[
"https://docs.firecrawl.dev/features/scrape",
"https://docs.firecrawl.dev/features/crawl",
"https://docs.firecrawl.dev/features/extract",
"https://docs.firecrawl.dev/features/map",
],
{
formats: ["markdown"],
onlyMainContent: true,
}
);
for (const page of batchResult.data || []) {
console.log(`${page.metadata?.title}: ${page.markdown?.length} chars`);
}
Step 4: Async Batch Scrape (Large Sets)
// Start async batch scrape for many URLs — returns job ID
const job = await firecrawl.asyncBatchScrapeUrls(
urls, // array of 100+ URLs
{ formats: ["markdown"] }
);
// Poll for completion
let status = await firecrawl.checkBatchScrapeStatus(job.id);
while (status.status !== "completed") {
await new Promise(r => setTimeout(r, 5000));
status = await firecrawl.checkBatchScrapeStatus(job.id);
}
console.log(`Batch complete: ${status.data?.length} pages`);
Step 5: Map — Rapid URL Discovery
// Discover all URLs on a site in ~2-3 seconds
// Uses sitemap.xml + SERP + cached crawl data
const mapResult = await firecrawl.mapUrl("https://docs.firecrawl.dev");
const urls = mapResult.links || [];
console.log(`Discovered ${urls.length} URLs`);
// Categorize by section
const sections = {
docs: urls.filter(u => u.includes("/docs/")),
api: urls.filter(u => u.includes("/api-reference/")),
features: urls.filter(u => u.includes("/features/")),
other: urls.filter(u => !u.includes("/docs/") && !u.includes("/api-reference/")),
};
Object.entries(sections).forEach(([name, list]) => {
console.log(` ${name}: ${list.length} URLs`);
});
Step 6: Map + Selective Scrape Pipeline
// 1. Map to discover URLs, 2. Filter, 3. Batch scrape relevant ones
async function intelligentScrape(siteUrl: string, pathFilter: string) {
const map = await firecrawl.mapUrl(siteUrl);
const relevant = (map.links || []).filter(url => url.includes(pathFilter));
console.log(`Map found ${map.links?.length} URLs, ${relevant.length} match filter`);
if (relevant.length === 0) return [];
if (relevant.length <= 10) {
return firecrawl.batchScrapeUrls(relevant, { formats: ["markdown"] });
}
// For large sets, use async batch
const job = await firecrawl.asyncBatchScrapeUrls(relevant.slice(0, 100), {
formats: ["markdown"],
});
// ...poll for completion
return job;
}
await intelligentScrape("https://docs.firecrawl.dev", "/features/");
Output
- Typed JSON objects extracted from web pages
- Batch scrape results for multiple URLs
- Complete site URL map for discovery
- Filtered scrape pipeline combining map + batch
Error Handling
| Error | Cause | Solution |
|---|---|---|
Empty extract | Page content too complex for LLM | Simplify schema, shorten prompt |
| Inconsistent extraction | Prompt too long | Keep prompts short and focused |
| Batch scrape timeout | Too many URLs | Use async batch with polling |
| Map returns few URLs | Site has no sitemap.xml | Use crawlUrl for thorough discovery |
402 Payment Required | Credits exhausted | Reduce batch size, check balance |
Examples
Extract Products from E-Commerce
const products = await firecrawl.scrapeUrl("https://store.example.com/products", {
formats: ["extract"],
extract: {
schema: {
type: "object",
properties: {
products: {
type: "array",
items: {
type: "object",
properties: {
name: { type: "string" },
price: { type: "number" },
availability: { type: "string" },
},
required: ["name", "price"],
},
},
},
},
},
});
Resources
Next Steps
For common errors, see firecrawl-common-errors.
When not to use it
- →When raw markdown is sufficient and structured data is not needed
Prerequisites
Limitations
- →LLM extraction adds variable credit costs
- →Large batch sets require async polling
How it compares
This workflow shifts from simple document retrieval to structured data extraction, allowing for direct integration of web content into typed applications.
Compared to similar skills
firecrawl-core-workflow-b side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| firecrawl-core-workflow-b (this skill) | 1 | 27d | Review | Intermediate |
| json-render-core | 3 | 2mo | No flags | Advanced |
| turborepo | 61 | 2mo | Review | Intermediate |
| senior-fullstack | 35 | 7mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by jeremylongshore
View all by jeremylongshore →You might also like
json-render-core
vercel-labs
Core package for defining schemas, catalogs, and AI prompt generation for json-render. Use when working with @json-render/core, defining schemas, creating catalogs, or building JSON specs for UI/video generation.
turborepo
vercel
Turborepo monorepo build system guidance. Triggers on: turbo.json, task pipelines, dependsOn, caching, remote cache, the "turbo" CLI, --filter, --affected, CI optimization, environment variables, internal packages, monorepo structure/best practices, and boundaries. Use when user: configures tasks/workflows/pipelines, creates packages, sets up monorepo, shares code between apps, runs changed/affected packages, debugs cache, or has apps/packages directories.
senior-fullstack
davila7
Comprehensive fullstack development skill for building complete web applications with React, Next.js, Node.js, GraphQL, and PostgreSQL. Includes project scaffolding, code quality analysis, architecture patterns, and complete tech stack guidance. Use when building new projects, analyzing code quality, implementing design patterns, or setting up development workflows.
oracle
openclaw
Best practices for using the oracle CLI (prompt + file bundling, engines, sessions, and file attachment patterns).
codex-skill
feiskyer
Use when user asks to leverage codex, gpt-5, or gpt-5.1 to implement something (usually implement a plan or feature designed by Claude). Provides non-interactive automation mode for hands-off task execution without approval prompts.
svelte-expert
Raudbjorn
Expert Svelte/SvelteKit development assistant for building components, utilities, and applications. Use when creating Svelte components, SvelteKit applications, implementing reactive patterns, handling state management, working with stores, transitions, animations, or any Svelte/SvelteKit development task. Includes comprehensive documentation access, code validation with svelte-autofixer, and playground link generation.