Get structured website data with just a prompt
Get structured website data with just a prompt
Hey everyone! Eric, Caleb, and Nick here from Firecrawl (YC S22). We’re excited to announce the release of /extract - an endpoint that turns entire websites into structured data with just a prompt. With /extract, you can retrieve any information from anywhere on a website without being limited by crawling/scraping roadblocks or the typical context constraints of LLMs. Here’s how our new /extract endpoint works: Users provide a prompt or desired output schema along with URLs. We leverage our existing index, /map, and /crawl endpoints to gather relevant context. For pre-indexed sites, we use a mixture of vector search, keyword search and a custom classifier to identify the most relevant pages. Some thoughtful prompting and re-ranking algorithms analyze user intent and score pages accordingly. Once identified, relevant pages are batch scraped to retrieve fresh data. For complex tasks, an AI agent determines the type of extraction needed and routes it to the appropriate pipeline. For example, extracting thousands of products dynamically creates a custom multi-entity schema, breaking the user’s schema into smaller parts that can be processed independently. This avoids relying on an LLMs small context window and enables efficient parallelization and merging of each independent extraction at the end. We integrate structured outputs from OpenAI, small LLMs, and task-specific models for classification and prompting. The entire process is parallelized using BullMQ and Kubernetes on GCP GKE, ensuring scalability, speed and leverages our existing Firecrawl scraping infrastructure. The result is a structured, intent-aligned response tailored to the user’s needs. Since starting Firecrawl, we knew it wouldn't just change web scraping. We realized that AI tools could process vastly more data than humans and traditional web search methods weren't designed for this ability to consume data at scale. This opened up a new paradigm for information retrieval - one that required quickly querying structured and unstructured datasets from across the web. We set out to make building web datasets at scale easy and /extract is a major step towards this future. If you want to try out: - Visit our landing page here: https://www.firecrawl.dev/extract - See Extract documentation: https://docs.firecrawl.dev/features/extract - Also, most of our work including /extract is open-source. Check it out here at https://github.com/mendableai/firecrawl That's all for now! Let us know any feedback on /extract.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Get structured web data with just a prompt
GPT-4 solves a crackme with just 1 prompt
Edit Photos with Just a Prompt
R2d2 – radare2 plugin for GPT-4 solves a crackme with just 1 prompt
From idea to live app — no code, just one prompt.
Build Figma plugins with just a prompt
Start selling online with just a prompt
Vibe Foto editor, no mouse, just prompt
Explain Any Thing With Just a Prompt
Any website. Structured data. No code.