Skip to main content
A crawl starts at one URL, follows the links you describe, and returns a result for every page it visits, up to 50 pages. Use it when you know where a set of pages lives but not every URL. If you already have the URLs, Batch scrape is simpler. You can run a crawl from the Crawl page in the dashboard, from the API or an SDK, or by asking an AI assistant connected to the MCP server. This page crawls the Mystery category of books.toscrape.com, a practice site built for scraping, and collects each book’s title and price. The Crawling Playground with the form on the left and the Results panel on the right

The form, control by control

The Crawl form with numbers marking each control from 1 to 9

1. Target URL

Where the crawl starts. Spidra loads this page first and follows links from it. Choose an index, category or listing page, because it links to the pages you want. For this example:

2. Crawl instruction

Describe in plain language which pages to visit.
Treat this as a strong hint and not a hard rule. Spidra uses it to choose links, but it can still visit a page you didn’t mean, such as the starting page itself. When you need a firm boundary, use the include and exclude paths below.

3. Transform instruction (optional)

Describe what to extract from each page, in the same way as a scrape prompt. Every page gets the same instruction. Leave it empty and Spidra returns each page as markdown without using AI.

4. Output schema (optional)

A JSON Schema that every page’s result must follow. Use it when you load the results into a database.

5. Max pages

How many pages to visit, from 1 to 50. The default is 5. Start with 3 to 5 so you can check your instructions, then raise it.

6. Max depth (optional)

How many links away from the starting page the crawl may go. 0 is the starting page only and 1 adds the pages it links to. Leave it empty for no limit.

7. Stealth Mode

Routes the crawl through a residential proxy, as in the Scrape playground. See Stealth Mode.

8. Advanced options

Configure opens path filters, crawl behavior and authentication. The Advanced Options dialog with include paths, exclude paths and crawl behavior settings
  • Include paths. Only URLs matching one of these patterns are crawled, for example /blog/*.
  • Exclude paths. URLs matching any of these are skipped, for example /tag/*. Include is applied first, then exclude.
  • Allow subdomains. Also follow links to subdomains, such as docs.example.com from example.com.
  • Crawl entire domain. Follow any link on the same root domain, whatever the starting path. Pair it with path filters.
  • Ignore query parameters. Treat URLs that differ only by query string as one page.
  • Authentication. Pass session cookies to crawl pages behind a login. See Authenticated scraping.
A path pattern matches anywhere in the URL, not only at the start. /books/* can match /catalogue/category/books/mystery_3/. Make patterns specific enough to avoid surprises.

9. Fast Mode

Fetches pages over plain HTTP instead of a browser, which is quicker on static sites. See Fast Mode in the Scrape playground.

What comes back

A crawl returns one entry per page. Here is a real crawl of 3 pages with a schema for title, price and in_stock:
Each entry also carries html and markdown for the page. Three things in this result are worth understanding, because you will see them in your own crawls. The starting page is a result too. The first entry is the category page, not a book. A schema describes one object per page, so it returned only the first book on a page that lists many. To pull a list from a listing page, make the schema an object with an array property, such as books, and describe each item inside it. Some pages return {}. The second page is a real book, yet its data is empty. The page itself loaded (status is success). Extraction came back with nothing, and a second attempt can fill it. You can retry individual pages without crawling again, as described below. Without a schema, the shape can vary. Running the same crawl with only the instruction Return the book's title, price and star rating gave a sentence of text for the first page and an object for the next two:
The null is because the star rating isn’t in the page’s text. If you need consistent fields, use a schema. The price there is a string with its currency symbol, while the schema run above returned a number.

What a crawl costs

Each page visited costs 1 credit, plus credits for the AI tokens used on that page, plus 5 for every CAPTCHA solved. A crawl with no transform instruction and no schema uses no tokens, so it is 1 credit per page. A failed crawl costs nothing. See How credits work.

Tips

  • Start with 3 to 5 pages. Check the instructions on a small crawl before scaling up.
  • Be specific. Blog posts published in 2024 works better than everything.
  • Use path filters for a firm boundary and the crawl instruction for the rest.
  • Choose a listing page as the start, not a single article.
  • Use a schema when the output feeds a program.

Crawl with code

The same crawl through the API:
crawl() waits for the job to finish. For long crawls, startCrawl() returns the jobId immediately and getCrawl() checks on it later. The crawl endpoint reference lists every parameter. The Code button at the top right of the playground generates this request for whatever you have filled in.

Webhook

Provide a webhook URL to receive a POST request each time a page finishes processing. This lets you stream results into your own pipeline as they come in, rather than waiting for the whole job to complete.

After the crawl

  • CAPTCHA solving. Supported CAPTCHAs are solved during the crawl, at 5 credits each.
  • Download results. After a crawl completes, download all extracted data as a ZIP file containing the extracted content, raw markdown, and original HTML snapshots for each page.
  • Retry failed pages. If extraction fails on specific pages, retry them individually without re-crawling the whole site.

Extract from Crawl

After a crawl finishes, you can run a new extraction on the same pages without visiting the site again. Spidra reuses the HTML and markdown saved during the original crawl and runs your new prompt against it. This is useful when:
  • You want to pull different fields from pages you already crawled
  • Your first extraction prompt wasn’t quite right
  • You need the same pages in two different formats
Send a POST to /crawl/{jobId}/extract with a new transformInstruction. You’ll get back a new jobId to poll.

Extract from Crawl API Reference

Full API reference for the extract endpoint

Structured output

Write schemas that keep every page’s result the same

Logs

Reopen any crawl and see its credits