SpideyDataAPI Reference

Crawl API

Create asynchronous crawls, check job status, and page through crawl results. Page billing is 1 Credit per delivered page, shadow by default; failed pages are free.

Quick start

Create a crawl with 1 to 100 HTTP(S) seed URLs without URL userinfo. Public Crawl requires a materialized plan before quota reservation; then poll status and list results by crawl ID. scope defaults to same_hostname. same_hostname follows only the exact hostname. same_domain follows hosts sharing the same registrable domain, including subdomains. same_origin additionally requires the same scheme and effective port. all accepts any valid HTTP(S) URL.

API base https://kieapi.com
curl -X POST "https://kieapi.com/v1/crawl" \
  -H "Authorization: Bearer $SPIDER_API_KEY" \
  -H "Idempotency-Key: crawl-example-01" \
  -H "Content-Type: application/json" \
  -d '{
    "urls": ["https://example.com/blog"],
    "max_depth": 1,
    "max_pages": 10,
    "scope": "same_hostname",
    "browser": {
      "wait_for_selector": ".post-card",
      "process_iframes": true,
      "flatten_shadow_dom": true,
      "virtual_scroll": {
        "container_selector": ".feed",
        "scroll_count": 5
      }
    },
    "extraction": {
      "type": "json_css",
      "schema": {
        "name": "blog_posts",
        "base_selector": ".post-card",
        "fields": [
          { "name": "title", "selector": "h2", "type": "text", "transform": "strip" },
          { "name": "path", "selector": "a", "type": "attribute", "attribute": "href" }
        ]
      },
      "computed_fields": [
        { "name": "url", "operation": "template", "template": "https://example.com{path}" }
      ]
    }
  }'

Authentication

Use a bearer token.

Team administrators manage API keys in Team Settings. Each key keeps its Team and billing scope when its creator leaves or switches Teams. Keys cannot manage members, billing or other keys.

Response shape

Crawl job

Accepted job metadata with IDs used by status and results endpoints.

Safe browser controls

Public crawl may use the same browser wait, iframe, shadow DOM, and virtual_scroll behavior controls while remaining an asynchronous job API.

Status

Progress, terminal state, counters, and failure information.

Results

Paginated public page results created by the crawl.

Page artifacts

Follow items[].artifacts[].url, including its crawl_id query parameter. Downloads verify parent ownership and page membership, create no independent Run, and are free. Configured file retention returns 410 on expiry; missing or foreign output returns 404.

Items

GET /v1/crawl/{crawl_id}/items returns extracted records with source page provenance.

Diagnostics

Use crawl.diagnostics.block_reasons for crawl blocked diagnostics, crawl.diagnostics.retry.categories for retry taxonomy, items[].diagnostics.adaptive_extraction for selector memory / adaptive extraction diagnostics, items[].diagnostics.browser_debug for sanitized browser debug capture counts, and items[].diagnostics.browser_failure for sanitized browser failure category/stage diagnostics.

Reference

Endpoint reference

OpenAPI-backed details for crawl creation, status polling, result pagination, and errors.

OpenAPI endpoint details

Interactive schemas, request bodies, response objects, and public error contracts load here.