Browser Crawl Engine

Crawl jobs

A job is a batch of URLs. Each URL becomes one independent task, so a slow or failing page never holds up the rest of the batch.

Create a job

POST /v1/jobsjobs:create

Body

FieldTypeDefaultNotes
urlsstring[]required1–1000 entries, http or https
maxAttemptsinteger31–10. A blocked page is retried on a different exit IP
resultTtlSecondsinteger60060–604800. How long the HTML stays downloadable
optionsobject—See below

Render options

OptionDefaultWhat it does
geoanyWhich country the request exits from — see below
enableJavaScripttrueTurn off only for static pages; it is much faster
blockAssetstrueDrops images, video and SVG from the DOM. Cuts size a lot. Set false if image URLs matter to you
navigationTimeoutMs120000How long to wait for the page to load
postLoadDelayMs1000Pause after load, before reading the DOM
scrollToBottomfalseScroll to trigger lazy-loaded content
scrollTimeoutMs45000Cap on the scrolling phase
expectedText—Treat the render as failed unless this string appears. The strongest tool you have
postLoadDelayMs is the option people get wrong. Pages that navigate or hydrate with JavaScript after load need 3,000–5,000 ms. With the 1,000 ms default you may capture the page mid-transition and get Execution context was destroyed instead of content.

Choosing a country

Three accepted shapes, in increasing specificity:

"geo": "de"                          // country
"geo": {"country": "de", "city": "fra"}   // country and city
"geo": {"profile": "de-fra-de2"}     // one exact exit
Pinning a profile disables rotation. With a country, a blocked exit is replaced by another in the same country and the task retries. With a fixed profile there is nothing to rotate to, so a block means the task fails. Pin only when you must, and expect failures.

You cannot supply an endpoint, key, DNS server or credential. Geo selection names a pre-provisioned exit; it is not a tunnel configuration.

Response

{"ok": true, "jobId": "0199…", "taskCount": 20, "deliveryDeferred": 0}

deliveryDeferred counts tasks whose queue delivery is being retried from the durable outbox. A positive number is not an error — the work is committed and will run; it just was not handed to the queue on the first attempt.

Read a job

GET /v1/jobs/{jobId}?limit=100&offset=0jobs:read
{"ok": true,
 "job": {"jobId": "…", "status": "RUNNING", "totalTasks": 20,
         "doneTasks": 14, "failedTasks": 0},
 "tasks": [{"taskId": "…", "url": "…", "status": "DONE",
            "attempt": 1, "httpStatus": 200, "htmlBytes": 1312004,
            "vpnCountry": "de", "vpnExitIp": "185.102.219.57"}],
 "limit": 100, "offset": 0}

Tasks are paginated. A job of 1,000 URLs needs ten pages at the maximum limit of 100.

Task status

StatusMeaning
PENDINGWaiting for a worker
RUNNINGBeing rendered right now
DONEHTML stored and downloadable
FAILEDOut of attempts; error says why
CANCELLEDStopped before completion

Download the HTML

GET /v1/tasks/{taskId}/htmlresults:read

Batch, do not loop

Sending 20 URLs in one job is measurably faster than 20 jobs of one URL, and it costs one request against your rate limit instead of twenty. On our own measurements the batched form finished 20 pages in 344 seconds where the sequential form took roughly three times as long.

Rule of thumb. One job per batch of work, not one job per URL. Poll the job, not each task.