> For the complete documentation index, see [llms.txt](https://docs.flowiseai.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.flowiseai.com/integrations/langchain/document-loaders/firecrawl.md).

# FireCrawl

Load data from URL using FireCrawl.

## FireCrawl

<figure><img src="https://823733684-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F00tYLwhz5RyR7fJEhrWy%2Fuploads%2Fgit-blob-0ba04feea2500474a753311588b8b55995462e4e%2Fup-004.png?alt=media" alt="" width="347"><figcaption><p>FireCrawl Node</p></figcaption></figure>

## FireCrawl Document Loader

[FireCrawl](https://www.firecrawl.dev) is a powerful web crawling and scraping service that provides advanced capabilities for extracting content from websites. This module enables loading and processing web content through the FireCrawl API.

This module provides a sophisticated web crawler that can:

* Scrape single web pages
* Crawl entire websites
* Extract structured data
* Handle JavaScript-rendered content
* Process content with text splitters
* Customize metadata extraction
* Support multiple operation modes

### Inputs

#### Required Parameters

* **URL**: The webpage or website URL to process
* **Connect Credential**: FireCrawl API credentials
* **Mode**: Choose between:
  * Scrape: Single page extraction
  * Crawl: Multi-page website crawling
  * Extract: Structured data extraction

#### Optional Parameters

* **Text Splitter**: A text splitter to process the extracted content
* **Scrape Options**:
  * Include Tags: HTML tags to include
  * Exclude Tags: HTML tags to exclude
  * Mobile: Use mobile user agent
  * Skip TLS Verification: Bypass SSL checks
  * Timeout: Request timeout
* **Additional Metadata**: JSON object with additional metadata
* **Omit Metadata Keys**: Comma-separated list of metadata keys to omit

### Outputs

* **Document**: Array of document objects containing metadata and pageContent
* **Text**: Concatenated string from pageContent of documents

### Features

* Multiple operation modes
* Advanced scraping options
* Structured data extraction
* JavaScript rendering
* Mobile device emulation
* Custom timeout settings
* Error handling

### Operation Modes

#### Scrape Mode

* Single page processing
* Main content extraction
* Format selection
* Custom tag filtering

#### Crawl Mode

* Multi-page crawling
* Subdomain handling
* Sitemap processing
* Link extraction

#### Extract Mode

* Structured data extraction
* Schema-based parsing
* LLM-powered extraction
* Custom extraction prompts

### Document Structure

Each document contains:

* **pageContent**: Extracted content in markdown format
* **metadata**:
  * title: Page title
  * description: Meta description
  * language: Content language
  * sourceURL: Original URL
  * Additional custom metadata

### Notes

* Requires a valid [FireCrawl API key](https://www.firecrawl.dev/app/api-keys)
* Supports multiple content formats
* Handles rate limiting
* Job status monitoring
* Error handling and retries
* Customizable request options
* Memory-efficient processing


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.flowiseai.com/integrations/langchain/document-loaders/firecrawl.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
