Skip to content
Start here

Crawl websites.

client.browserRendering.crawl.create(CrawlCreateParamsparams, RequestOptionsoptions?): CrawlCreateResponse
POST/accounts/{account_id}/browser-rendering/crawl

Starts a crawl job for the provided URL and its children. Check available options like gotoOptions and waitFor* to control page load behaviour.

Security
API Token

The preferred authorization scheme for interacting with the Cloudflare API. Create a token.

Example:Authorization: Bearer Sn3lZJTBX6kkg7OdcBUAxOO963GEIyGQqnFTOFYY
API Email + API Key

The previous authorization scheme for interacting with the Cloudflare API, used in conjunction with a Global API key.

Example:X-Auth-Email: user@example.com

The previous authorization scheme for interacting with the Cloudflare API. When possible, use API tokens instead of Global API keys.

Example:X-Auth-Key: 144c9defac04969c7bfad8efaa8ea194
Accepted Permissions (at least one required)
Browser Rendering Write
ParametersExpand Collapse
CrawlCreateParams = Variant0 | Variant1
CrawlCreateParamsBase { account_id, url, cacheTTL, 27 more }
account_id: string

Path param: Account ID.

url: string

Body param: URL to navigate to, eg. https://example.com.

formaturi
cacheTTL?: number

Query param: Cache TTL default is 5s. Set to 0 to disable.

maximum86400
minimum0
actionTimeout?: number

Body param: The maximum duration allowed for the browser action to complete after the page has loaded (such as taking screenshots, extracting content, or generating PDFs). If this time limit is exceeded, the action stops and returns a timeout error.

maximum120000
addScriptTag?: Array<AddScriptTag>

Body param: Adds a <script> tag into the page with the desired URL or content.

id?: string
content?: string
type?: string
url?: string
formaturi
addStyleTag?: Array<AddStyleTag>

Body param: Adds a <link rel="stylesheet"> tag into the page with the desired URL or a <style type="text/css"> tag with the content.

content?: string
url?: string
formaturi
allowRequestPattern?: Array<string>

Body param: Only allow requests that match the provided regex patterns, eg. ’/^.*.(css)’.

allowResourceTypes?: Array<"document" | "stylesheet" | "image" | 15 more>

Body param: Only allow requests that match the provided resource types, eg. ‘image’ or ‘script’.

One of the following:
"document"
"stylesheet"
"image"
"media"
"font"
"script"
"texttrack"
"xhr"
"fetch"
"prefetch"
"eventsource"
"websocket"
"manifest"
"signedexchange"
"ping"
"cspviolationreport"
"preflight"
"other"
authenticate?: Authenticate

Body param: Provide credentials for HTTP authentication.

password: string
minLength1
username: string
minLength1
bestAttempt?: boolean

Body param: Attempt to proceed when ‘awaited’ events fail or timeout.

cookies?: Array<Cookie>

Body param: Check options.

name: string

Cookie name.

value: string
domain?: string
expires?: number
httpOnly?: boolean
partitionKey?: string
path?: string
priority?: "Low" | "Medium" | "High"
One of the following:
"Low"
"Medium"
"High"
sameParty?: boolean
sameSite?: "Strict" | "Lax" | "None"
One of the following:
"Strict"
"Lax"
"None"
secure?: boolean
sourcePort?: number
sourceScheme?: "Unset" | "NonSecure" | "Secure"
One of the following:
"Unset"
"NonSecure"
"Secure"
url?: string
crawlPurposes?: Array<"search" | "ai-input" | "ai-train">

Body param: List of crawl purposes to respect Content-Signal directives in robots.txt. Allowed values: ‘search’, ‘ai-input’, ‘ai-train’. Learn more: https://contentsignals.org/. Default: [‘search’, ‘ai-input’, ‘ai-train’].

One of the following:
"search"
"ai-input"
"ai-train"
depth?: number

Body param: Maximum number of levels deep the crawler will traverse from the starting URL.

maximum100000
minimum1
emulateMediaType?: string

Body param

formats?: Array<"html" | "markdown" | "json">

Body param: Formats to return. Default is html.

One of the following:
"html"
"markdown"
"json"
gotoOptions?: GotoOptions

Body param: Check options.

referer?: string
referrerPolicy?: string
timeout?: number
maximum60000
waitUntil?: "load" | "domcontentloaded" | "networkidle0" | "networkidle2" | Array<"load" | "domcontentloaded" | "networkidle0" | "networkidle2">
One of the following:
"load" | "domcontentloaded" | "networkidle0" | "networkidle2"
"load"
"domcontentloaded"
"networkidle0"
"networkidle2"
Array<"load" | "domcontentloaded" | "networkidle0" | "networkidle2">
"load"
"domcontentloaded"
"networkidle0"
"networkidle2"
jsonOptions?: JsonOptions

Body param: Options for JSON extraction.

custom_ai?: Array<CustomAI>

Optional list of custom AI models to use for the request. The models will be tried in the order provided, and in case a model returns an error, the next one will be used as fallback.

model: string

AI model to use for the request. Must be formed as <provider>/<model_name>, e.g. workers-ai/@cf/meta/llama-3.3-70b-instruct-fp8-fast.

authorization?: string

Authorization token for the AI model: Bearer <token>. Not needed for workers-ai models.

prompt?: string
response_format?: ResponseFormat { type, json_schema }
type: string
json_schema?: Record<string, unknown> | null

Schema for the response format. More information here: https://developers.cloudflare.com/workers-ai/json-mode/

limit?: number

Body param: Maximum number of URLs to crawl.

maximum100000
minimum1
maxAge?: number

Body param: Maximum age of a resource that can be returned from cache in seconds. Default is 1 day.

maximum604800
minimum0
modifiedSince?: number

Body param: Unix timestamp (seconds since epoch) indicating to only crawl pages that were modified since this time. For sitemap URLs with a lastmod field, this is compared directly. For other URLs, the crawler will use If-Modified-Since header when fetching. URLs without modification information (no lastmod in sitemap and no Last-Modified header support) will be crawled. Note: This works in conjunction with maxAge - both filters must pass for a cached resource to be used. Must be within the last year and not in the future.

exclusiveMinimum
minimum0
options?: Options

Body param: Additional options for the crawler.

excludePatterns?: Array<string>

Exclude links matching the provided wildcard patterns in the crawl job. Example: ‘https://example.com/privacy/**’.

includePatterns?: Array<string>

Include only links matching the provided wildcard patterns in the crawl job. Include patterns are evaluated before exclude patterns. URLs that match any of the specified include patterns will be included in the crawl job. Example: ‘https://example.com/blog/**’.

includeSubdomains?: boolean

Include links to subdomains in the crawl job. This option is ignored if includeExternalLinks is true.

rejectRequestPattern?: Array<string>

Body param: Block undesired requests that match the provided regex patterns, eg. ’/^.*.(css)’.

rejectResourceTypes?: Array<"document" | "stylesheet" | "image" | 15 more>

Body param: Block undesired requests that match the provided resource types, eg. ‘image’ or ‘script’.

One of the following:
"document"
"stylesheet"
"image"
"media"
"font"
"script"
"texttrack"
"xhr"
"fetch"
"prefetch"
"eventsource"
"websocket"
"manifest"
"signedexchange"
"ping"
"cspviolationreport"
"preflight"
"other"
render?: true

Body param: Whether to render the page or fetch static content. True by default.

setExtraHTTPHeaders?: Record<string, string>

Body param

setJavaScriptEnabled?: boolean

Body param

source?: "sitemaps" | "links" | "all"

Body param: Source of links to crawl. ‘sitemaps’ - only crawl URLs from sitemaps, ‘links’ - only crawl URLs scraped from pages, ‘all’ - crawl both sitemap and scraped links (default).

One of the following:
"sitemaps"
"links"
"all"
viewport?: Viewport

Body param: Check options.

height: number
width: number
deviceScaleFactor?: number
hasTouch?: boolean
isLandscape?: boolean
isMobile?: boolean
waitForSelector?: WaitForSelector

Body param: Wait for the selector to appear in page. Check options.

selector: string
hidden?: true
timeout?: number
maximum120000
visible?: true
waitForTimeout?: number

Body param: Waits for a specified timeout before continuing.

maximum120000
Variant0 extends CrawlCreateParamsBase { account_id, url, cacheTTL, 27 more }
Variant1 extends CrawlCreateParamsBase { account_id, url, cacheTTL, 27 more }
ReturnsExpand Collapse
CrawlCreateResponse = string

Crawl job ID.

Crawl websites.

import Cloudflare from 'cloudflare';

const client = new Cloudflare({
  apiToken: process.env['CLOUDFLARE_API_TOKEN'], // This is the default and can be omitted
});

const crawl = await client.browserRendering.crawl.create({
  account_id: 'account_id',
  url: 'https://example.com',
});

console.log(crawl);
{
  "result": "result",
  "success": true,
  "errors": [
    {
      "code": 0,
      "message": "message"
    }
  ]
}
{
  "errors": [
    {
      "code": 2001,
      "message": "Rate limit exceeded"
    }
  ],
  "success": false
}
Returns Examples
{
  "result": "result",
  "success": true,
  "errors": [
    {
      "code": 0,
      "message": "message"
    }
  ]
}
{
  "errors": [
    {
      "code": 2001,
      "message": "Rate limit exceeded"
    }
  ],
  "success": false
}