Crawl jobs
A job is a batch of URLs. Each URL becomes one independent task, so a slow or failing page never holds up the rest of the batch.
Create a job
Body
| Field | Type | Default | Notes |
|---|---|---|---|
urls | string[] | required | 1–1000 entries, http or https |
maxAttempts | integer | 3 | 1–10. A blocked page is retried on a different exit IP |
resultTtlSeconds | integer | 600 | 60–604800. How long the HTML stays downloadable |
options | object | — | See below |
Render options
| Option | Default | What it does |
|---|---|---|
geo | any | Which country the request exits from — see below |
enableJavaScript | true | Turn off only for static pages; it is much faster |
blockAssets | true | Drops images, video and SVG from the DOM. Cuts size a lot. Set false if image URLs matter to you |
navigationTimeoutMs | 120000 | How long to wait for the page to load |
postLoadDelayMs | 1000 | Pause after load, before reading the DOM |
scrollToBottom | false | Scroll to trigger lazy-loaded content |
scrollTimeoutMs | 45000 | Cap on the scrolling phase |
expectedText | — | Treat the render as failed unless this string appears. The strongest tool you have |
postLoadDelayMs is the option people get wrong.
Pages that navigate or hydrate with JavaScript after load need 3,000–5,000 ms.
With the 1,000 ms default you may capture the page mid-transition and get
Execution context was destroyed instead of content.
Choosing a country
Three accepted shapes, in increasing specificity:
"geo": "de" // country
"geo": {"country": "de", "city": "fra"} // country and city
"geo": {"profile": "de-fra-de2"} // one exact exit
You cannot supply an endpoint, key, DNS server or credential. Geo selection names a pre-provisioned exit; it is not a tunnel configuration.
Response
{"ok": true, "jobId": "0199…", "taskCount": 20, "deliveryDeferred": 0}
deliveryDeferred counts tasks whose queue delivery is being
retried from the durable outbox. A positive number is not an error — the
work is committed and will run; it just was not handed to the queue on the
first attempt.
Read a job
{"ok": true,
"job": {"jobId": "…", "status": "RUNNING", "totalTasks": 20,
"doneTasks": 14, "failedTasks": 0},
"tasks": [{"taskId": "…", "url": "…", "status": "DONE",
"attempt": 1, "httpStatus": 200, "htmlBytes": 1312004,
"vpnCountry": "de", "vpnExitIp": "185.102.219.57"}],
"limit": 100, "offset": 0}
Tasks are paginated. A job of 1,000 URLs needs ten pages at the maximum
limit of 100.
Task status
| Status | Meaning |
|---|---|
PENDING | Waiting for a worker |
RUNNING | Being rendered right now |
DONE | HTML stored and downloadable |
FAILED | Out of attempts; error says why |
CANCELLED | Stopped before completion |
Download the HTML
- Requires the task to be
DONE; otherwise 409. X-HTML-SHA256is the digest of the stored bytes.- Served as an attachment;
?inline=1renders it instead. - Available until
resultTtlSecondselapses, then deleted.
Batch, do not loop
Sending 20 URLs in one job is measurably faster than 20 jobs of one URL, and it costs one request against your rate limit instead of twenty. On our own measurements the batched form finished 20 pages in 344 seconds where the sequential form took roughly three times as long.