Skip to main content
Spidra gets data off the web for you. You point it at a page, a list of pages, a whole site or a search query, say what you want back, and it returns clean markdown or JSON. It opens real browsers, handles CAPTCHAs and proxies, and runs the AI extraction, so you don’t build any of that yourself. The examples on this page are real runs against books.toscrape.com, a practice site built to be scraped, and example.com.

The four things Spidra does

The first question is whether you know the URL. If you don’t, use Search to find it. If you know one page, Scrape it. If you have a list, use Batch so each URL gets its own result. If you only know where to start, Crawl.

How a request becomes a result

Scrapes, batches and crawls take a few seconds to a few minutes, so they run as jobs. You submit the request and Spidra answers straight away with a job ID. Then you check on the job until it finishes.
  1. Submit. POST /scrape returns 202 with a jobId.
  2. Wait. Spidra loads the page in a browser, runs any actions, solves CAPTCHAs and runs the extraction.
  3. Check. GET /scrape/{jobId} returns the job’s status. Poll every few seconds until it says completed, failed or cancelled.
The SDKs do the polling for you: scrape() waits and returns the finished result, and startScrape() returns the job ID right away if you’d rather check yourself. Search works the same way, and it can also return the result in the same call with wait: true. A scrape of one normal page took about 10 seconds in these examples. Pages behind a CAPTCHA can take over a minute.

What comes back

The same page can come back in three shapes, depending on what you ask for. Without a prompt, you get the page as text. No AI runs, no tokens are used, and it costs 1 credit. This is a real run against example.com, trimmed to the first lines:
With a prompt, you get what you described. Spidra reads the page and answers your instruction. Here the prompt was “Return the title and price of every book on the page”:
That came back with 20 books in total. The shape is whatever the AI decided, which is fine for a one-off. With a schema, you get the shape you defined. A JSON Schema fixes the field names and types, so every result looks the same and you can load it straight into a database. On a single book page:
Notice rating is null. The star rating on that page is stored in the page’s HTML and not in its visible text, and Spidra extracts from the text. On this run, a field that wasn’t in the text came back as null. A finished scrape always has the same outer shape: result.content holds the answer, result.data holds one entry per URL with the raw markdown, and result.stats records how long it took and how many tokens the AI used.

Options that change how a page is fetched

  • Browser actions. Click, type, scroll, wait, or loop over every item before extraction, for pages that load content on demand.
  • Authenticated scraping. Pass your own session cookies to read pages behind a login.
  • Stealth Mode. Route requests through residential proxies in the country you choose.
  • CAPTCHA solving. Spidra solves supported challenges on its own and charges 5 credits for each one it solves.
  • Fast Mode. Fetches the page over plain HTTP instead of opening a browser. It suits static pages, and it turns off actions, screenshots and Stealth Mode. Spidra switches to a browser when a page needs one.
  • Documents. The same request works on PDFs, Word, Excel and other document URLs.

What it costs

You pay in credits. A scrape costs 1 credit per page, plus a few more if AI extraction runs, plus 5 if a CAPTCHA is solved. Search costs 1 credit per 10 results. A job that fails costs nothing. The book list above cost 3 credits: 1 for the page and 2 for the extraction. How credits work has the full rules.

Where your results go

Every scrape, search and crawl is saved in your logs with its inputs, output and the credits it used, so you can reopen a result later or fetch it again through the API.

Four ways to use it

  • The dashboard. The Scrape, Crawl and Search playgrounds run everything above with forms and no code.
  • The REST API. Start at the API introduction.
  • SDKs. Official libraries for Node.js, Python and eight more languages. See the SDK overview.
  • An AI assistant. The MCP server lets Claude, Cursor and others call Spidra for you.

Run your first scrape

Get a key and make a request in a couple of minutes

Limits and rate limits

How many requests you can send, and what happens past that