# Fetch API

Fetch one URL and receive normalized page content, metadata, and run audit fields.

Markdown URL: `/api-docs/fetch.md`
OpenAPI YAML: `/api-docs/public-v1/openapi/fetch`

## AI Agent Task

Use this Markdown as the implementation brief for an AI coding agent. The goal is not to memorize endpoint names; the goal is to build a working client that can call SpideyData, parse the public response objects, and handle failures without guessing.

Minimum environment:

```bash
export SPIDER_API_KEY="spider_live_..."
```

The API base URL is filled into every example from the current documentation origin.

Required headers for JSON routes:

```http
Authorization: Bearer <spider_api_key>
Content-Type: application/json
```

For creation operations that declare it in OpenAPI, also send a
caller-generated `Idempotency-Key` when the client may retry. Reuse the key
only for the exact same request body.

Client rules:

- Use the origin already embedded in each example and call only documented `/v1` paths.
- Send JSON bodies exactly as shown; unsupported fields return `400 unsupported_feature` or `400 validation_error`.
- Treat `run.id`, `crawl.id`, and pagination cursors as opaque strings.
- Store `error.request_id` with failed calls so the server-side trace can be found later.
- Prefer the OpenAPI YAML when generating typed clients: `/api-docs/public-v1/openapi`.


## Common Response Objects

`PublicRun` identifies the unit of work:

```json
{
  "id": "run_01j...",
  "type": "fetch",
  "status": "completed",
  "created_at": "2026-06-21T10:00:00Z",
  "completed_at": "2026-06-21T10:00:02Z",
  "cost_units": 1,
  "quota": {
    "resource": "scrape",
    "quantity_reserved": 1
  }
}
```

Public run types are `fetch`, `browser_scrape`, and `crawl`. Public statuses are `queued`, `running`, `paused`, `completed`, `failed`, `low_quality`, `budget_exhausted`, and `expired`.

`PublicPageResult` is the normalized page payload returned by Fetch, Browser scrape, and Crawl result pages:

```json
{
  "url": "https://example.com/docs",
  "final_url": "https://example.com/docs",
  "status": "completed",
  "title": "Example Docs",
  "content": {
    "markdown": "# Example Docs",
    "html": null,
    "metadata": {},
    "links": ["https://example.com/docs/api"]
  },
  "artifacts": [
    {
      "key": "markdown",
      "url": "/v1/runs/run_01j/artifacts/markdown",
      "content_type": "text/markdown; charset=utf-8",
      "expires_at": null
    }
  ],
  "audit": {
    "engine_winner": "light",
    "engines_attempted": ["light"],
    "fallback_reasons": [],
    "fallback_events": [],
    "retry_stats": {"total_attempts": 1},
    "quality_score": 0.92,
    "block_reason": null
  },
  "usage": {
    "cost_units": 1,
    "budget_exhausted": false
  },
  "diagnostics": {
    "adaptive_extraction": {
      "enabled": true,
      "used": true,
      "action": "reused",
      "confidence": 0.88
    },
    "browser_debug": {
      "enabled": false,
      "network_event_count": 0,
      "console_message_count": 0
    },
    "browser_failure": {
      "present": false,
      "category": "none",
      "stage": "none",
      "attempt_count": 0,
      "browser_lease_acquired": null,
      "error_code": null
    },
    "content_quality": {
      "score": 0.92,
      "confidence": "high"
    }
  }
}
```

`diagnostics` is a public-safe explainability object. Page results expose
adaptive extraction diagnostics at `result.diagnostics.adaptive_extraction` and
browser debug capture counts at `result.diagnostics.browser_debug`, browser
failure category/stage diagnostics at `result.diagnostics.browser_failure`,
and content quality diagnostics at `result.diagnostics.content_quality`. These
are sanitized summaries; they do not expose raw selector memory, raw selectors,
raw debug events, raw URLs, headers, cookies, bearer tokens, response bodies,
provider internals, profile/proxy/CDP details, identity feedback kinds, retry
policy reasons, raw browser error text, or raw audit attempts.

Important parse targets for an AI-built client:

- `content.markdown`: primary text output for LLM ingestion.
- `content.html`: sanitized or extracted HTML when requested and available.
- `artifacts`: downloadable outputs; only request artifact keys listed here.
- `audit`: execution evidence for debugging fallback, quality, and retry behavior.
- `usage`: cost and budget outcome for the run.
- `diagnostics`: sanitized aggregate page diagnostics for adaptive extraction, browser debug counts, browser failure category/stage, and content quality.

When browser capacity is exhausted, audit uses
`block_reason="browser_capacity_full"` and diagnostics reports
`result.diagnostics.browser_failure.category="capacity"`; this is a local
browser-resource condition, not target-site network unreachability.

Light response guards use `block_reason="body_size_limit"`,
`block_reason="redirect_limit"`, or `block_reason="blocked_target"`; all are
non-retryable, and `blocked_target` prevents Browser fallback. The current server
runtime default is 10 MiB (10485760 bytes) and is not caller-configurable in
public v1. Only the sanitized category is exposed: the
response body and internal transport error text are never exposed.


## Error Handling

All JSON errors use `PublicErrorResponse`:

```json
{
  "error": {
    "code": "quota_exceeded",
    "message": "Monthly crawl quota exceeded.",
    "type": "quota",
    "retryable": false,
    "request_id": "trace_01j",
    "details": {
      "resource": "crawl"
    }
  }
}
```

Retry policy:

- `400 validation_error`, `400 invalid_request`, and `400 unsupported_feature`: fix the request body; do not retry unchanged payloads.
- `401 authentication_required` and `403 permission_denied`: stop and ask for a valid key or access change.
- `404 run_not_found`, `404 crawl_not_found`, `404 artifact_not_found`, `404 session_not_found`, and `404 profile_not_found`: stop or refresh the resource list.
- `408 timeout`, `502 fetch_failed`, `502 render_failed`, `503 service_unavailable`, and `500 internal_error`: retry only when `error.retryable` is true, using exponential backoff.
- `409 idempotency_conflict`: generate a new key for the changed request body; do not retry unchanged.
- `409 idempotency_in_progress`: retry the same key and body after a short backoff.
- `409 run_not_ready`: wait and retry the read/download.
- `410 artifact_expired`: do not retry; rerun the source job if the artifact is needed.
- `429 rate_limited`: honor `Retry-After` when present; otherwise back off.
- `429 quota_exceeded`: do not retry until quota changes.


### Goal

Use Fetch when the target is one URL and you want normalized Markdown/HTML with lightweight retrieval first. Fetch starts with the configured runtime identity/proxy path and can retry direct/no-proxy Light retrieval after proxy or network transport failures. Choose Browser scrape instead when JavaScript rendering, browser identity, screenshot, or recording output is required.

### Request body

```json
{
  "url": "https://example.com/docs"
}
```

Minimum request:

```json
{
  "url": "https://example.com/docs"
}
```

### Quick Start

```bash
curl -sS "https://kieapi.com/v1/fetch" \
  -H "Authorization: Bearer $SPIDER_API_KEY" \
  -H "Idempotency-Key: fetch-example-docs-01" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com/docs"}'
```

### Successful response

`200 OK` returns `PageRunResponse`:

```json
{
  "run": {
    "id": "run_fetch_01j",
    "type": "fetch",
    "status": "completed",
    "created_at": "2026-06-21T10:00:00Z",
    "completed_at": "2026-06-21T10:00:01Z",
    "cost_units": 1,
    "quota": {
      "resource": "scrape",
      "quantity_reserved": 1
    }
  },
  "result": {
    "url": "https://example.com/docs",
    "final_url": "https://example.com/docs",
    "status": "completed",
    "title": "Example Docs",
    "content": {
      "markdown": "# Example Docs",
      "html": "<main>Example Docs</main>",
      "metadata": {},
      "links": []
    },
    "artifacts": [],
    "audit": {
      "engine_winner": "light",
      "engines_attempted": ["light"],
      "fallback_reasons": [],
      "fallback_events": [],
      "retry_stats": {"total_attempts": 1},
      "quality_score": 0.92,
      "block_reason": null
    },
    "usage": {
      "cost_units": 1,
      "budget_exhausted": false
    }
  }
}
```

### Implementation Checklist

1. Validate the caller supplied an absolute HTTP or HTTPS `url`.
2. Call `POST /v1/fetch` with bearer auth and JSON headers.
3. If HTTP status is not 2xx, parse `PublicErrorResponse` and apply the retry policy.
4. On success, read `result.content.markdown` first; fall back to `result.content.html` only if your app can handle HTML.
5. Inspect `result.audit.fallback_reasons` and `result.audit.fallback_events` to see whether proxy/network failure required direct fallback.
6. Persist `run.id`, `result.final_url`, `result.audit`, and `result.usage` for debugging and cost reporting.
7. Download only artifact keys returned in `result.artifacts`.


### Endpoint Reference

- `POST /v1/fetch`
  - Operation ID: `fetchPage`
  - Request schema: `FetchRequest`
  - Response statuses: `402`, `200` `PageRunResponse`, `400`, `401`, `403`, `408`, `409`, `429`, `502`, `503`
  - Summary: Fetch one URL.
  - Description: Fetches one URL and returns a public page result.
