Source profileQuality 83/100

ComposioHQ/awesome-claude-skills/composio-skills/firecrawl-automation/SKILL.md

Firecrawl Automation

Automate web crawling and data extraction with Firecrawl -- scrape pages, crawl sites, extract structured data, batch scrape URLs, and map website structures through the Composio Firecrawl integration.

Source repository stars
71,746
Declared platforms
0
Static risk flags
1
Last source update
2026-07-24
Source checked
2026-08-04

Decision brief

What it does—and where it fits

Run Firecrawl web crawling and extraction directly from Claude Code. Scrape individual pages, crawl entire sites, extract structured data with AI, batch process URL lists, and map website structures without leaving your terminal.

Best for

    Not for

    • Tasks that require unconfirmed production actions or broad system permissions.
    • Environments where the pinned source and install steps cannot be inspected.

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexNot declaredNo explicit evidencePortability before use
    Claude CodeNot declaredNo explicit evidencePortability before use
    CursorNot declaredNo explicit evidencePortability before use
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/ComposioHQ/awesome-claude-skills --skill "composio-skills/firecrawl-automation"
    Safe inspection promptEditorial

    Inspect the Agent Skill "Firecrawl Automation" from https://github.com/ComposioHQ/awesome-claude-skills/blob/be2a406907dbc61b73e6827ded415c96139d13a2/composio-skills/firecrawl-automation/SKILL.md at commit be2a406907dbc61b73e6827ded415c96139d13a2. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      Setup

      1. Add the Composio MCP server to your configuration:

      Add the Composio MCP server to your configuration:Connect your Firecrawl account when prompted. The agent will provide an authentication link.Be mindful of credit consumption -- scope your crawls tightly and test on small URL sets before scaling.
    2. 02

      Core Workflows

      Fetch content from a URL in multiple formats with optional browser actions for dynamic pages.

      url (required) -- fully qualified URL to scrapeformats -- output formats: markdown (default), html, rawHtml, links, screenshot, jsononlyMainContent (default true) -- extract main content only, excluding nav/footer/ads
    3. 03

      1. Scrape a Single Page

      Fetch content from a URL in multiple formats with optional browser actions for dynamic pages.

      url (required) -- fully qualified URL to scrapeformats -- output formats: markdown (default), html, rawHtml, links, screenshot, jsononlyMainContent (default true) -- extract main content only, excluding nav/footer/ads
    4. 04

      2. Crawl an Entire Site

      Discover and scrape multiple pages from a website with configurable depth, path filters, and concurrency.

      url (required) -- starting URL for the crawllimit (default 10) -- max pages to crawlmaxDiscoveryDepth -- depth limit from the root page
    5. 05

      3. Extract Structured Data

      Extract structured JSON data from web pages using AI with a natural language prompt or JSON schema.

      urls (required) -- array of URLs to extract from (max 10 in beta). Supports wildcards like https://example.com/blog/prompt -- natural language description of what to extractschema -- JSON Schema defining the desired output structure

    Permission review

    Static risk signals and limitations

    Network access

    medium · line 14

    The documentation includes network, browsing, or remote request actions.

    https://rube.app/mcp

    Network access

    medium · line 25

    The documentation includes network, browsing, or remote request actions.

    Fetch content from a URL in multiple formats with optional browser actions for dynamic pages.

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score83/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars71,746SourceRepository attention, not individual Skill quality
    Compatibility0 platformsSourceDeclared in the catalog source record
    Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    ComposioHQ/awesome-claude-skills
    Skill path
    composio-skills/firecrawl-automation/SKILL.md
    Commit
    be2a406907dbc61b73e6827ded415c96139d13a2
    License
    Not declared
    Collected
    2026-08-04
    Default branch
    master
    View the original SKILL.md

    Firecrawl Automation

    Run Firecrawl web crawling and extraction directly from Claude Code. Scrape individual pages, crawl entire sites, extract structured data with AI, batch process URL lists, and map website structures without leaving your terminal.

    Toolkit docs: composio.dev/toolkits/firecrawl


    Setup

    1. Add the Composio MCP server to your configuration:
      https://rube.app/mcp
      
    2. Connect your Firecrawl account when prompted. The agent will provide an authentication link.
    3. Be mindful of credit consumption -- scope your crawls tightly and test on small URL sets before scaling.

    Core Workflows

    1. Scrape a Single Page

    Fetch content from a URL in multiple formats with optional browser actions for dynamic pages.

    Tool: FIRECRAWL_SCRAPE

    Key parameters:

    • url (required) -- fully qualified URL to scrape
    • formats -- output formats: markdown (default), html, rawHtml, links, screenshot, json
    • onlyMainContent (default true) -- extract main content only, excluding nav/footer/ads
    • waitFor -- milliseconds to wait for JS rendering (default 0)
    • timeout -- max wait in ms (default 30000)
    • actions -- browser actions before scraping (click, write, wait, press, scroll)
    • includeTags / excludeTags -- filter by HTML tags
    • jsonOptions -- for structured extraction with schema and/or prompt

    Example prompt: "Scrape the main content from https://example.com/pricing as markdown"


    2. Crawl an Entire Site

    Discover and scrape multiple pages from a website with configurable depth, path filters, and concurrency.

    Tool: FIRECRAWL_CRAWL_V2

    Key parameters:

    • url (required) -- starting URL for the crawl
    • limit (default 10) -- max pages to crawl
    • maxDiscoveryDepth -- depth limit from the root page
    • includePaths / excludePaths -- regex patterns for URL paths
    • allowSubdomains -- include subdomains (default false)
    • crawlEntireDomain -- follow sibling/parent links, not just children (default false)
    • sitemap -- include (default), skip, or only
    • prompt -- natural language to auto-configure crawler settings
    • scrapeOptions_formats -- output format for each page
    • scrapeOptions_onlyMainContent -- main content extraction per page

    Example prompt: "Crawl the docs section of firecrawl.dev, max 50 pages, only paths matching docs"


    3. Extract Structured Data

    Extract structured JSON data from web pages using AI with a natural language prompt or JSON schema.

    Tool: FIRECRAWL_EXTRACT

    Key parameters:

    • urls (required) -- array of URLs to extract from (max 10 in beta). Supports wildcards like https://example.com/blog/*
    • prompt -- natural language description of what to extract
    • schema -- JSON Schema defining the desired output structure
    • enable_web_search -- allow crawling links outside initial domains (default false)

    At least one of prompt or schema must be provided.

    Check extraction status with FIRECRAWL_EXTRACT_GET using the returned job id.

    Example prompt: "Extract company name, pricing tiers, and feature lists from https://example.com/pricing"


    4. Batch Scrape Multiple URLs

    Scrape many URLs concurrently with shared configuration for efficient bulk data collection.

    Tool: FIRECRAWL_BATCH_SCRAPE

    Key parameters:

    • urls (required) -- array of URLs to scrape
    • formats -- output format for all pages (default markdown)
    • onlyMainContent (default true) -- main content extraction
    • maxConcurrency -- parallel scrape limit
    • ignoreInvalidURLs (default true) -- skip bad URLs instead of failing the batch
    • location -- geolocation settings with country code
    • actions -- browser actions applied to each page
    • blockAds (default true) -- block advertisements

    Example prompt: "Batch scrape these 20 product page URLs as markdown with ad blocking"


    5. Map Website Structure

    Discover all URLs on a website from a starting URL, useful for planning crawls or auditing site structure.

    Tool: FIRECRAWL_MAP_MULTIPLE_URLS_BASED_ON_OPTIONS

    Key parameters:

    • url (required) -- starting URL (must be https:// or http://)
    • search -- guide URL discovery toward specific page types
    • limit (default 5000, max 100000) -- max URLs to return
    • includeSubdomains (default true) -- include subdomains
    • ignoreQueryParameters (default true) -- dedupe URLs differing only by query params
    • sitemap -- include, skip, or only

    Example prompt: "Map all URLs on docs.example.com, focusing on API reference pages"


    6. Monitor and Manage Crawl Jobs

    Track crawl progress, retrieve results, and cancel runaway jobs.

    Tools: FIRECRAWL_CRAWL_GET, FIRECRAWL_GET_THE_STATUS_OF_A_CRAWL_JOB, FIRECRAWL_CANCEL_A_CRAWL_JOB

    • FIRECRAWL_CRAWL_GET -- get status, progress, credits used, and crawled page data
    • FIRECRAWL_CANCEL_A_CRAWL_JOB -- stop an active or queued crawl

    Both require the crawl job id (UUID) returned when the crawl was initiated.

    Example prompt: "Check the status of crawl job 019b0806-b7a1-7652-94c1-e865b5d2e89a"


    Known Pitfalls

    • Rate limiting: Firecrawl can trigger "Rate limit exceeded" errors (429). Prefer FIRECRAWL_BATCH_SCRAPE over many individual FIRECRAWL_SCRAPE calls, and implement backoff on 429/5xx responses.
    • Credit consumption: FIRECRAWL_EXTRACT can fail with "Insufficient credits." Scope tightly and avoid broad homepage URLs that yield sparse fields. Test on small URL sets first.
    • Nested error responses: Per-page failures may be nested in response.data.code (e.g., SCRAPE_DNS_RESOLUTION_ERROR) even when the outer API call succeeds. Always validate inner status/error fields.
    • JS-heavy pages: Non-rendered fetches may miss key content. Use waitFor (e.g., 1000-5000ms) for dynamic pages, or configure scrapeOptions_actions to interact with the page before scraping.
    • Extraction schema precision: Vague or shifting schemas/prompts produce noisy, inconsistent output. Freeze your schema and test on a small sample before scaling to many URLs.
    • Crawl jobs are async: FIRECRAWL_CRAWL_V2 returns immediately with a job ID. Use FIRECRAWL_CRAWL_GET to poll for results. Cancel stuck crawls with FIRECRAWL_CANCEL_A_CRAWL_JOB to avoid wasting credits.
    • Extract job polling: FIRECRAWL_EXTRACT is also async for larger jobs. Retrieve final output with FIRECRAWL_EXTRACT_GET.
    • URL batching for extract: Keep extract URL batches small (~10 URLs) to avoid 429 rate limit errors.
    • Deeply nested responses: Results are often nested under data.data or deeper. Inspect the returned shape rather than assuming flat keys.

    Quick Reference

    Tool SlugDescription
    FIRECRAWL_SCRAPEScrape a single URL with format/action options
    FIRECRAWL_CRAWL_V2Crawl a website with depth/path control
    FIRECRAWL_EXTRACTExtract structured data with AI prompt/schema
    FIRECRAWL_BATCH_SCRAPEBatch scrape multiple URLs concurrently
    FIRECRAWL_MAP_MULTIPLE_URLS_BASED_ON_OPTIONSDiscover/map all URLs on a site
    FIRECRAWL_CRAWL_GETGet crawl job status and results
    FIRECRAWL_GET_THE_STATUS_OF_A_CRAWL_JOBCheck crawl job progress
    FIRECRAWL_CANCEL_A_CRAWL_JOBCancel an active crawl job
    FIRECRAWL_EXTRACT_GETGet extraction job status and results
    FIRECRAWL_CRAWL_PARAMS_PREVIEWPreview crawl parameters before starting
    FIRECRAWL_SEARCHWeb search + scrape top results

    Powered by Composio

    Alternatives

    Compare before choosing

    Computed 10023,781

    alirezarezvani/claude-skills

    app-store-optimization

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist

    Computed 10014,225

    wanshuiyin/Auto-claude-code-research-in-sleep

    citation-audit

    Use it for operations and research tasks; the detail page covers purpose, installation, and practical steps.

    Computed 1004,922

    dotnet/skills

    migrate-vstest-to-mtp

    Migrates .NET test projects from VSTest to Microsoft.Testing.Platform (MTP). Use when user asks to "migrate to MTP", "switch from VSTest", "enable Microsoft.Testing.Platform", "use MTP runner", set OutputType=Exe only for test projects in Directory.Build.props, or mentions EnableMSTestRunner, EnableNUnitRunner, or UseMicrosoftTestingPlatformRunner. USE FOR: MTP behavioral differences vs VSTest (exit code 8, zero tests discovered, --ignore-exit-code, TESTINGPLATFORM_EXITCODE_IGNORE); centralizing

    Computed 1002,504

    aaron-he-zhu/aaron-marketing-skills

    social-selling-planner

    Use when the user asks to "set up my founder social-selling routine", "build a daily engagement block for target accounts", or "turn funding / hiring signals into selling plays"; produces the founder/seller daily operating block — a time-boxed engagement-block spec (substantive value-add comments on target-account posts, never a pitch), warm-touch-before-ask cadence rules, trigger-response plays consuming the social-pulse-monitor B2B trigger watchlist (funding / hiring / launch signals), and a q