# Browser API

Use managed browser sessions and browser-backed scrape endpoints for pages that need rendering.

Markdown URL: `/api-docs/browser.md`
OpenAPI YAML: `/api-docs/public-v1/openapi/browser`

## AI Agent Task

Use this Markdown as the implementation brief for an AI coding agent. The goal is not to memorize endpoint names; the goal is to build a working client that can call SpideyData, parse the public response objects, and handle failures without guessing.

Minimum environment:

```bash
export SPIDER_API_KEY="spider_live_..."
```

The API base URL is filled into every example from the current documentation origin.

Required headers for JSON routes:

```http
Authorization: Bearer <spider_api_key>
Content-Type: application/json
```

For creation operations that declare it in OpenAPI, also send a
caller-generated `Idempotency-Key` when the client may retry. Reuse the key
only for the exact same request body.

Client rules:

- Use the origin already embedded in each example and call only documented `/v1` paths.
- Send JSON bodies exactly as shown; unsupported fields return `400 unsupported_feature` or `400 validation_error`.
- Treat `run.id`, `crawl.id`, and pagination cursors as opaque strings.
- Store `error.request_id` with failed calls so the server-side trace can be found later.
- Prefer the OpenAPI YAML when generating typed clients: `/api-docs/public-v1/openapi`.


## Common Response Objects

`PublicRun` identifies the unit of work:

```json
{
  "id": "run_01j...",
  "type": "fetch",
  "status": "completed",
  "created_at": "2026-06-21T10:00:00Z",
  "completed_at": "2026-06-21T10:00:02Z",
  "cost_units": 1,
  "quota": {
    "resource": "scrape",
    "quantity_reserved": 1
  }
}
```

Public run types are `fetch`, `browser_scrape`, and `crawl`. Public statuses are `queued`, `running`, `paused`, `completed`, `failed`, `low_quality`, `budget_exhausted`, and `expired`.

`PublicPageResult` is the normalized page payload returned by Fetch, Browser scrape, and Crawl result pages:

```json
{
  "url": "https://example.com/docs",
  "final_url": "https://example.com/docs",
  "status": "completed",
  "title": "Example Docs",
  "content": {
    "markdown": "# Example Docs",
    "html": null,
    "metadata": {},
    "links": ["https://example.com/docs/api"]
  },
  "artifacts": [
    {
      "key": "markdown",
      "url": "/v1/runs/run_01j/artifacts/markdown",
      "content_type": "text/markdown; charset=utf-8",
      "expires_at": null
    }
  ],
  "audit": {
    "engine_winner": "light",
    "engines_attempted": ["light"],
    "fallback_reasons": [],
    "fallback_events": [],
    "retry_stats": {"total_attempts": 1},
    "quality_score": 0.92,
    "block_reason": null
  },
  "usage": {
    "cost_units": 1,
    "budget_exhausted": false
  },
  "diagnostics": {
    "adaptive_extraction": {
      "enabled": true,
      "used": true,
      "action": "reused",
      "confidence": 0.88
    },
    "browser_debug": {
      "enabled": false,
      "network_event_count": 0,
      "console_message_count": 0
    },
    "browser_failure": {
      "present": false,
      "category": "none",
      "stage": "none",
      "attempt_count": 0,
      "browser_lease_acquired": null,
      "error_code": null
    },
    "content_quality": {
      "score": 0.92,
      "confidence": "high"
    }
  }
}
```

`diagnostics` is a public-safe explainability object. Page results expose
adaptive extraction diagnostics at `result.diagnostics.adaptive_extraction` and
browser debug capture counts at `result.diagnostics.browser_debug`, browser
failure category/stage diagnostics at `result.diagnostics.browser_failure`,
and content quality diagnostics at `result.diagnostics.content_quality`. These
are sanitized summaries; they do not expose raw selector memory, raw selectors,
raw debug events, raw URLs, headers, cookies, bearer tokens, response bodies,
provider internals, profile/proxy/CDP details, identity feedback kinds, retry
policy reasons, raw browser error text, or raw audit attempts.

Important parse targets for an AI-built client:

- `content.markdown`: primary text output for LLM ingestion.
- `content.html`: sanitized or extracted HTML when requested and available.
- `artifacts`: downloadable outputs; only request artifact keys listed here.
- `audit`: execution evidence for debugging fallback, quality, and retry behavior.
- `usage`: cost and budget outcome for the run.
- `diagnostics`: sanitized aggregate page diagnostics for adaptive extraction, browser debug counts, browser failure category/stage, and content quality.

When browser capacity is exhausted, audit uses
`block_reason="browser_capacity_full"` and diagnostics reports
`result.diagnostics.browser_failure.category="capacity"`; this is a local
browser-resource condition, not target-site network unreachability.

Light response guards use `block_reason="body_size_limit"`,
`block_reason="redirect_limit"`, or `block_reason="blocked_target"`; all are
non-retryable, and `blocked_target` prevents Browser fallback. The current server
runtime default is 10 MiB (10485760 bytes) and is not caller-configurable in
public v1. Only the sanitized category is exposed: the
response body and internal transport error text are never exposed.


## Error Handling

All JSON errors use `PublicErrorResponse`:

```json
{
  "error": {
    "code": "quota_exceeded",
    "message": "Monthly crawl quota exceeded.",
    "type": "quota",
    "retryable": false,
    "request_id": "trace_01j",
    "details": {
      "resource": "crawl"
    }
  }
}
```

Retry policy:

- `400 validation_error`, `400 invalid_request`, and `400 unsupported_feature`: fix the request body; do not retry unchanged payloads.
- `401 authentication_required` and `403 permission_denied`: stop and ask for a valid key or access change.
- `404 run_not_found`, `404 crawl_not_found`, `404 artifact_not_found`, `404 session_not_found`, and `404 profile_not_found`: stop or refresh the resource list.
- `408 timeout`, `502 fetch_failed`, `502 render_failed`, `503 service_unavailable`, and `500 internal_error`: retry only when `error.retryable` is true, using exponential backoff.
- `409 idempotency_conflict`: generate a new key for the changed request body; do not retry unchanged.
- `409 idempotency_in_progress`: retry the same key and body after a short backoff.
- `409 run_not_ready`: wait and retry the read/download.
- `410 artifact_expired`: do not retry; rerun the source job if the artifact is needed.
- `429 rate_limited`: honor `Retry-After` when present; otherwise back off.
- `429 quota_exceeded`: do not retry until quota changes.


### Goal

Use Browser API when a page needs rendering, browser identity, screenshots, recordings, or a reusable managed browser session. It has three public surfaces: Browser scrape, Browser Sessions, and persistent Browser Profiles.

### Browser scrape request body

```json
{
  "url": "https://example.com/app",
  "browser": {
    "wait_for_selector": ".post-card",
    "wait_for_timeout_ms": 7000,
    "process_iframes": true,
    "flatten_shadow_dom": true,
    "virtual_scroll": {
      "container_selector": ".feed",
      "scroll_count": 5,
      "scroll_by": "page_height"
    }
  },
  "output": {
    "screenshot": {
      "enabled": true,
      "full_page": true
    },
    "recording": {
      "enabled": true,
      "format": "mp4"
    }
  },
  "extraction": {
    "type": "json_css",
    "schema": {
      "name": "blog_posts",
      "base_selector": ".post-card",
      "fields": [
        { "name": "title", "selector": "h2", "type": "text", "transform": "strip" },
        { "name": "path", "selector": "a", "type": "attribute", "attribute": "href" }
      ]
    },
    "computed_fields": [
      { "name": "url", "operation": "template", "template": "https://example.com{path}" }
    ]
  }
}
```

`browser` exposes only safe render behavior controls: `wait_for_selector`,
`process_iframes`, `flatten_shadow_dom`, and `virtual_scroll`. It does not
expose arbitrary JavaScript, cookies, storage state, proxy URLs, user agents, or
DICloak profile IDs.

```bash
curl -sS "https://kieapi.com/v1/browser/scrape" \
  -H "Authorization: Bearer $SPIDER_API_KEY" \
  -H "Idempotency-Key: browser-example-blog-01" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com/blog","output":{"screenshot":{"enabled":true,"full_page":true}},"extraction":{"type":"json_css","schema":{"name":"blog_posts","base_selector":".post-card","fields":[{"name":"title","selector":"h2","type":"text","transform":"strip"},{"name":"path","selector":"a","type":"attribute","attribute":"href"}]},"computed_fields":[{"name":"url","operation":"template","template":"https://example.com{path}"}]}}'
```

### Browser scrape successful response

`200 OK` returns `PageRunResponse`. The envelope matches Fetch, but `run.type` is `browser_scrape`, `audit.engine_winner` is usually `browser`, and media artifacts may be listed:

```json
{
  "run": {
    "id": "run_browser_01j",
    "type": "browser_scrape",
    "status": "completed",
    "quota": {
      "resource": "scrape",
      "quantity_reserved": 1
    }
  },
  "result": {
    "url": "https://example.com/app",
    "status": "completed",
    "content": {
      "markdown": "# Rendered App",
      "html": null,
      "metadata": {},
      "links": []
    },
    "structured": {
      "type": "json_css",
      "source": "crawl4ai",
      "schema_name": "blog_posts",
      "items": [
        { "title": "Example A", "path": "/blog/a", "url": "https://example.com/blog/a" }
      ],
      "item_count": 1,
      "error": null
    },
    "artifacts": [
      {
        "key": "screenshot_png",
        "url": "/v1/runs/run_browser_01j/artifacts/screenshot_png",
        "content_type": "image/png",
        "expires_at": "2026-06-28T10:00:00Z"
      }
    ],
    "audit": {
      "engine_winner": "browser",
      "engines_attempted": ["light", "browser"]
    },
    "usage": {
      "cost_units": 3,
      "budget_exhausted": false
    }
  }
}
```

### Browser sessions

Use Sessions when your client needs a Team-owned live browser. Creation requires
an `Idempotency-Key`, reserves Team Credits, and returns only after the
browser-level relay is ready. A Profile-backed Session exclusively locks the
Profile through stop and provider save; omit `profile_id` for a disposable
Session. Session responses never expose the selected runtime node or provider
CDP URL.

Browser runtime costs 120 Credits/hour ($0.12/hour). Actual active seconds are
accumulated and rounded up to 1 Credit per 30 seconds. A 300-second TTL reserves
10 Credits; stopping at 31 seconds charges 2 and releases 8. Idle Profiles are
free. Per-Team Profile / concurrent Session limits: Free 5/1, Starter 20/2,
Scale 100/4, Enterprise 200/4. Proxy usage has separate pricing.
Each Free Team without an assigned billing period receives 5,000 included Credits
per UTC calendar month. Legacy local plans use Free limits. Profile and Session
limits aggregate all historical execution scopes within the Team.

Session creation enforces the published meter and Team plan's Browser
concurrency entitlement. Profile creation and cloning enforce the published
Profile-count entitlement.

```bash
curl -sS "https://kieapi.com/v1/browser/sessions" \
  -H "Authorization: Bearer $SPIDER_API_KEY"

curl -sS -X POST "https://kieapi.com/v1/browser/sessions" \
  -H "Authorization: Bearer $SPIDER_API_KEY" \
  -H "Idempotency-Key: checkout-session-20260905-001" \
  -H "Content-Type: application/json" \
  -d '{"profile_id":"bprof_01k","ttl_seconds":300}'

curl -sS -X POST "https://kieapi.com/v1/browser/sessions/bses_01k/connect" \
  -H "Authorization: Bearer $SPIDER_API_KEY"

curl -sS -X DELETE "https://kieapi.com/v1/browser/sessions/bses_01k" \
  -H "Authorization: Bearer $SPIDER_API_KEY"
```

Connect response:

```json
{
  "session_id": "bses_01k",
  "websocket_url": "wss://spideydata.example.com/v1/browser/sessions/bses_01k/devtools/browser/bses_01k?ticket=ONE_TIME_SCOPED_TICKET",
  "expires_at": "2026-09-05T10:00:30Z"
}
```

`expires_at` is the ticket expiry, not the Session expiry. The ticket is
one-time, at most 60 seconds, bound to Session/Team/user/credential, and the
last ticket minted before connection wins. Only one controller may attach.
Do not persist or log the ticket. Stop, API-key revocation or Team suspension
invalidates its control access. Removing a member invalidates that member's
personal control access; Team Keys remain independent of their creator. A runtime-node
failure ends the live Session; there is no cross-node live migration.

### Browser network

Disposable Sessions and Profile create/update accept `proxy` with one mode:
`direct`, `system_dynamic` (published `country`, optional `sticky_minutes`),
`account_static` (`resource_id`) or `custom_proxy` (`resource_id`). Do not combine
`proxy` with legacy `proxy_id`. Starting a Profile inherits its full selection;
omit network overrides. For example:

```json
{"ttl_seconds":300,"proxy":{"mode":"system_dynamic","country":"US","sticky_minutes":10}}
```

Dynamic response traffic is measured independently of Live or code connections.
Only confirmed completed 2xx/404/410 network responses are charged; cached or
failed responses and unmeasured upload/header bytes are excluded. Observation
loss ends the Session. `proxy_usage` reports bytes, Credits, charge status and
verified exit evidence separately from Browser runtime `usage`. The startup
price is retained for idempotent settlement. Saved custom resources resolve
current credentials at Session start and cannot be deleted while bound.

### Browser profiles

During cloning, the source rejects edits, deletion, new Sessions and additional
clones with 409 until the clone is verified. The clone has its own browser-data
save timestamps.

Use Profiles to manage persistent Team browser identity. Create, clone, delete,
and Proxy changes reconcile asynchronously with the provider; rename is
synchronous Spider-only display metadata and never renames the provider Profile.
`PATCH` checks ready-state and `If-Match` atomically; same-version concurrent
updates have at most one winner. Session idempotency replay precedes new-resource
Profile/runtime/quota checks. Accepted startup survives HTTP disconnect with its
original 60-second deadline; late results are cleanup-only. Termination and
billing are capped at Session expiry using the original price snapshot. A failed
CDP handshake consumes its ticket but releases the controller; get a fresh ticket
to reconnect to a healthy Session. Do not
expect provider profile/group IDs, home nodes, proxy credentials, cookies,
storage state, filesystem paths, or raw launch flags. The former public
`profile-pools` route is removed.

```bash
curl -sS "https://kieapi.com/v1/browser/profiles" \
  -H "Authorization: Bearer $SPIDER_API_KEY"

curl -sS -X POST "https://kieapi.com/v1/browser/profiles" \
  -H "Authorization: Bearer $SPIDER_API_KEY" \
  -H "Idempotency-Key: checkout-profile-20260905-001" \
  -H "Content-Type: application/json" \
  -d '{"name":"iOS checkout","device_profile":"ios","proxy_id":"proxy_01k"}'

curl -sS -X PATCH "https://kieapi.com/v1/browser/profiles/bprof_01k" \
  -H "Authorization: Bearer $SPIDER_API_KEY" \
  -H 'If-Match: "3"' \
  -H "Content-Type: application/json" \
  -d '{"name":"iOS checkout primary","proxy_id":null}'

curl -sS -X POST "https://kieapi.com/v1/browser/profiles/bprof_01k/clone" \
  -H "Authorization: Bearer $SPIDER_API_KEY" \
  -H "Idempotency-Key: clone-checkout-20260905-001" \
  -H "Content-Type: application/json" \
  -d '{"name":"iOS checkout clone"}'

curl -sS -X DELETE "https://kieapi.com/v1/browser/profiles/bprof_01k" \
  -H "Authorization: Bearer $SPIDER_API_KEY"
```

### Implementation Checklist

1. Use `POST /v1/browser/scrape` for one rendered page, optional structured extraction, and media artifacts.
2. Use `POST /v1/browser/sessions` plus `POST /v1/browser/sessions/{session_id}/connect` only when your app needs a live browser websocket.
3. Bind each API key to one Team and grant only the Browser scopes it needs.
4. Use Profile resources only for persistent browser identity; never send provider-native profile payloads.
5. For structured extraction, DICloak owns browser access, profile, proxy, fingerprint, and CDP execution while Crawl4AI executes the allowlisted schema over rendered HTML.
6. For screenshot or recording output, inspect `result.artifacts` and download through Runs artifact routes.
7. Handle `402 insufficient_credits`, `409 idempotency_*`, `412 precondition_failed`, `428 precondition_required`, `429 browser_capacity_exhausted`, `504 browser_start_timeout`, and the common errors explicitly.


### Endpoint Reference

- `POST /v1/browser/scrape`
  - Operation ID: `browserScrapePage`
  - Request schema: `BrowserScrapeRequest`
  - Response statuses: `200` `PageRunResponse`, `400`, `401`, `403`, `408`, `409`, `429`, `502`, `503`
  - Summary: Browser-scrape one URL.
  - Description: Runs one product-safe browser-backed scrape for one URL. Optional structured extraction runs over the rendered HTML through the safe Crawl4AI extraction contract.
- `GET /v1/browser/sessions`
  - Operation ID: `listBrowserSessions`
  - Response statuses: `200` `BrowserSessionsResponse`, `400`, `401`, `403`, `422`, `429`, `503`
  - Summary: List Team browser sessions.
  - Description: Lists Team-scoped Browser Sessions. The default filter is active; use status=all for retained stopped history. Runtime node, provider profile ID, and raw CDP endpoints are never returned.
- `POST /v1/browser/sessions`
  - Operation ID: `createBrowserSession`
  - Request schema: `BrowserSessionCreateRequest`
  - Response statuses: `201` `BrowserSessionCreateResponse`, `400`, `402`, `401`, `403`, `404`, `409`, `422`, `429`, `504`, `503`
  - Summary: Create a browser session.
  - Description: Creates one Team-scoped Browser Session, waits until the browser-level relay is ready, and returns the active resource. The request is Team-billed at 120 Credits/hour, rounded up at 1 Credit per 30 actual active seconds. The full TTL is reserved before startup and unused Credits are released on stop. Per-Team Profile/concurrent Session limits are Free 5/1, Starter 20/2, Scale 100/4, Enterprise 200/4. Each Free Team without an assigned billing period receives 5,000 included Credits per UTC calendar month. Idle Profiles are free. Profile use is exclusive, and Idempotency-Key is mandatory. After authentication and request validation, replay precedes new-resource Profile, runtime, and quota checks. Accepted startup survives HTTP disconnect but retains its original 60-second deadline; a late provider result is cleaned up and cannot activate the Session.
- `GET /v1/browser/sessions/{session_id}`
  - Operation ID: `getBrowserSession`
  - Response statuses: `200` `BrowserSessionResponse`, `401`, `403`, `404`, `429`, `503`
  - Summary: Get a browser session.
- `DELETE /v1/browser/sessions/{session_id}`
  - Operation ID: `stopBrowserSession`
  - Response statuses: `202` `BrowserSessionStopResponse`, `200` `BrowserSessionStopResponse`, `401`, `403`, `404`, `429`, `503`
  - Summary: Stop a browser session.
- `POST /v1/browser/sessions/{session_id}/connect`
  - Operation ID: `connectBrowserSession`
  - Response statuses: `200` `BrowserSessionConnectResponse`, `401`, `403`, `404`, `409`, `429`, `503`
  - Summary: Mint a one-time browser controller ticket.
  - Description: Returns a no-store, short-lived, one-time, session/Team/user/credential-scoped central browser-level CDP relay URL for an active Session. Last-issued-wins before consumption; only one controller may be attached. The URL supports Playwright, Puppeteer, and raw CDP without exposing the runtime node or provider `cdp_url`.
- `GET /v1/browser/profiles`
  - Operation ID: `listBrowserProfiles`
  - Response statuses: `200` `BrowserProfilesResponse`, `401`, `403`, `429`, `503`
  - Summary: List managed browser profiles.
  - Description: Lists Team-scoped persistent Browser Profile resources. Provider IDs, DICloak groups, cookies, storage paths, and proxy credentials remain private.
- `POST /v1/browser/profiles`
  - Operation ID: `createBrowserProfile`
  - Request schema: `BrowserProfileCreateRequest`
  - Response statuses: `202` `BrowserProfileResponse`, `400`, `401`, `403`, `409`, `422`, `429`, `503`
  - Summary: Create a managed browser profile resource.
  - Description: Accepts creation of a persistent Team Browser Profile. The public resource starts in creating and is reconciled asynchronously with its pinned runtime home node.
- `GET /v1/browser/profiles/{profile_id}`
  - Operation ID: `getBrowserProfile`
  - Response statuses: `200` `BrowserProfileResponse`, `401`, `403`, `404`, `429`, `503`
  - Summary: Get a managed browser profile resource.
- `PATCH /v1/browser/profiles/{profile_id}`
  - Operation ID: `updateBrowserProfile`
  - Request schema: `BrowserProfileUpdateRequest`
  - Response statuses: `200` `BrowserProfileResponse`, `202` `BrowserProfileResponse`, `400`, `401`, `403`, `404`, `409`, `412`, `428`, `422`, `429`, `503`
  - Summary: Update a managed browser profile resource.
  - Description: The name is Spider-owned display metadata, updated atomically with If-Match and without renaming the provider Profile. A name-only update returns 200. Proxy changes register durable asynchronous work; the previous binding remains authoritative until verification. Concurrent requests with the same version have at most one winner; a stale version returns 412.
- `DELETE /v1/browser/profiles/{profile_id}`
  - Operation ID: `deleteBrowserProfile`
  - Response statuses: `202` `BrowserProfileDeleteResponse`, `200` `BrowserProfileDeleteResponse`, `401`, `403`, `404`, `409`, `429`, `503`
  - Summary: Delete a managed browser profile resource.
- `POST /v1/browser/profiles/{profile_id}/clone`
  - Operation ID: `cloneBrowserProfile`
  - Request schema: `BrowserProfileCloneRequest`
  - Response statuses: `202` `BrowserProfileResponse`, `400`, `401`, `403`, `404`, `409`, `422`, `429`, `503`
  - Summary: Clone a ready Browser Profile.
  - Description: Accepts an asynchronous clone of a ready, unused Profile on its pinned runtime home node. The source is locked against edits, deletion, Sessions, and additional clones until cloning is verified. Conflicting requests return 409. Provider state remains private.
