# Provider reference

One page per provider, describing exactly what PuffinParse sends, what comes back, and how the two are
mapped onto the unified `OcrResponse`, checked against the implementation in
`crates/puffinparse-core/src/providers/`.

Verification differs by provider, and the **Status** column says which applies:

* **live-verified** (Reducto, Extend, LlamaParse): the `#[ignore]`d live tests pass against the real
  API with a key, and the fixtures include redacted live responses.
* **verified locally** (Tesseract, Docling): self-hosted engines; the live tests pass against a local
  install and the fixtures are real output from it.
* **docs-only** (everything else): implemented from the provider's API documentation and tested
  against fixture payloads built from it, but not yet run against the live API. Expect wire-format
  differences until they are verified; [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10)
  tracks this, and help with a key is welcome.

| Provider | Doc | Implementation | Env var | Models | Status |
|---|---|---|---|---|---|
| Reducto | [`reducto.md`](/docs/providers/reducto/index.md) | `providers/reducto.rs` | `REDUCTO_API_KEY` | `standard` *(default)*, `r-1`, `agentic` | live-verified |
| Extend | [`extend.md`](/docs/providers/extend/index.md) | `providers/extend.rs` | `EXTEND_API_KEY` | `parse_performance` *(default)*, `parse_light`, `parse_auto` | live-verified |
| LlamaParse | [`llamaparse.md`](/docs/providers/llamaparse/index.md) | `providers/llamaparse.rs` | `LLAMA_API_KEY` | `fast`, `cost_effective` *(default)*, `agentic`, `agentic_plus` | live-verified |
| Mistral | [`mistral.md`](/docs/providers/mistral/index.md) | `providers/mistral.rs` | `MISTRAL_API_KEY` | `ocr-latest` *(default)*, `ocr-4-1`, `ocr-4-0`, `ocr-2512` | docs-only |
| Azure AI Document Intelligence | [`azure.md`](/docs/providers/azure/index.md) | `providers/azure.rs` | `AZURE_DOCUMENT_INTELLIGENCE_KEY` + `AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT` | `read` *(ocr default)*, `layout` *(parse default)*, `invoice` *(extract default)*, `receipt`, `id_document`, `tax_us_w2`, `custom` | docs-only |
| AWS Textract | [`textract.md`](/docs/providers/textract/index.md) | `providers/textract.rs` | `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY` (+ `AWS_SESSION_TOKEN`, `AWS_REGION`) | `detect-text` *(ocr default)*, `layout`, `queries` *(extract default)*, `forms` | docs-only |
| Google Gemini | [`gemini.md`](/docs/providers/gemini/index.md) | `providers/gemini.rs` | `GEMINI_API_KEY` | `2.5-flash` *(default)*, `2.5-pro`, `2.5-flash-lite`, `3.5-flash`, `3.5-flash-lite`, `3.8-flash` | docs-only |
| OpenAI | [`openai.md`](/docs/providers/openai/index.md) | `providers/openai.rs` | `OPENAI_API_KEY` | `gpt-5.6-luna` *(default)*, `gpt-5.6-terra`, `gpt-5.6-sol`, `gpt-6-astra` | docs-only |
| Anthropic | [`anthropic.md`](/docs/providers/anthropic/index.md) | `providers/anthropic.rs` | `ANTHROPIC_API_KEY` | `claude-sonnet-5` *(default)*, `claude-haiku-4-5`, `claude-opus-5` | docs-only |
| Mathpix | [`mathpix.md`](/docs/providers/mathpix/index.md) | `providers/mathpix.rs` | `MATHPIX_APP_ID` + `MATHPIX_APP_KEY` | `pdf` *(default)*, `text` | docs-only |
| Datalab (Marker) | [`datalab.md`](/docs/providers/datalab/index.md) | `providers/datalab.rs` | `DATALAB_API_KEY` | `fast`, `balanced` *(default)*, `accurate` | docs-only |
| Unstructured | [`unstructured.md`](/docs/providers/unstructured/index.md) | `providers/unstructured.rs` | `UNSTRUCTURED_API_KEY` | `hi_res` *(default)*, `fast`, `auto` | docs-only |
| Upstage | [`upstage.md`](/docs/providers/upstage/index.md) | `providers/upstage.rs` | `UPSTAGE_API_KEY` | `document-parse` *(default)*, `document-parse-nightly` | docs-only |
| Landing AI (ADE) | [`landingai.md`](/docs/providers/landingai/index.md) | `providers/landingai.rs` | `LANDINGAI_API_KEY` | `dpt-2` *(default)* | docs-only |
| Google Document AI | [`google-documentai.md`](/docs/providers/google-documentai/index.md) | `providers/google_documentai.rs` | `GOOGLE_DOCUMENTAI_ACCESS_TOKEN` (+ `_PROJECT`, `_LOCATION`, `_PROCESSOR_ID`) | `ocr` *(default)*, `layout`, `form`, `prebuilt` | docs-only |
| Tesseract *(local)* | [`tesseract.md`](/docs/providers/tesseract/index.md) | `providers/tesseract.rs` | none (`TESSERACT_CMD`, `PDFTOPPM_CMD`) | `default` | verified locally |
| Docling *(self-hosted)* | [`docling.md`](/docs/providers/docling/index.md) | `providers/docling.rs` | none (`DOCLING_BASE_URL`; optional `DOCLING_API_KEY`) | `default` | verified locally |
| PaddleOCR *(self-hosted)* | [`paddleocr.md`](/docs/providers/paddleocr/index.md) | `providers/paddleocr.rs` | none (`PADDLEOCR_BASE_URL`, `PADDLEOCR_PARSE_BASE_URL`) | `default` | docs-only |

The three self-hosted engines need no API key and are priced at $0/page (`puffinparse providers` shows
`local` in the Key column); they are the open baselines in the benchmark. Their shared helpers are
in `providers/local.rs`.

Every page follows the same structure: summary → models → request flow → response mapping → errors and
limits → gotchas → `provider_options` examples → links.

## At a glance

| | Reducto | Extend | LlamaParse |
|---|---|---|---|
| Upload method | `POST /upload` (multipart) → `reducto://<uuid>` id; URLs passed through as `input` | `POST /files/upload` (multipart) → `file_…` id; URLs passed through as `{"url", "name"}` | one multipart call: `file` part, or `input_url` field |
| Sync / async | Sync `POST /parse` by default; async `POST /parse_async` + `GET /job/{id}` with `provider_options={"async": true}` | Always async: `POST /parse_runs` + `GET /parse_runs/{id}` | Always async: `POST /api/v1/parsing/upload` + `GET …/job/{id}` + `GET …/job/{id}/result/json` |
| Page info source | `blocks[].bbox.page` (chunks carry no page number); PuffinParse pins `chunk_mode: "page"` | `chunk.metadata.pageRange` + `block.metadata.page.number`; PuffinParse pins `chunkingStrategy: {"type":"page"}` | `pages[].page` (native per-page objects with `md` and `text`) |
| Page dimensions | not reported — `Page.width`/`height` are `None` | `block.metadata.page.width/height` (raster pixels at the run's dpi) | `pages[].width/height` |
| Bbox units on the wire | already normalised 0–1, top-left origin | absolute `left/top/right/bottom` in page units | absolute `bBox {x,y,w,h}` in page units |
| Bbox after normalisation | clamped as-is | divided by page width/height | divided by page width/height |
| Confidence type | `granular_confidence.parse_confidence` (0–1), else coarse `"high"`/`"low"` → 0.9 / 0.5 | `metadata.avgOcrConfidence` (0–1) per block | `items[].bBox.confidence` (0–1) per item |
| Credits reported? | yes — `usage.credits` (`null` on new per-product pricing) | yes — `usage.credits` (`null` for runs before 2025-10-07) | effectively no — `job_credits_usage` is 0 until billing settles, so PuffinParse reports `None` |
| Large-result indirection | `result.type == "url"` → presigned JSON fetched without auth | `outputUrl` (15-min presigned) when `responseType=url` | none — result endpoints return inline JSON |
| Job id surfaced as | `job_id` | run `id` (`pr_…`) | job `id` (UUID) |

`Usage.pages` comes from `usage.num_pages` (Reducto), `metrics.pageCount` (Extend) and
`job_metadata.job_pages` (LlamaParse), falling back to the number of reconstructed pages.
`cost_usd` is always `pages × per_page_usd` from `crates/puffinparse-core/src/pricing.json` — no provider
reports dollar cost directly.

## Shared behaviour

These are implemented once, in `crates/puffinparse-core/src/{http,error,provider}.rs`, and apply to all
three providers:

* **API key**: request `api_key` first, then the provider's env var; missing ⇒ `authentication` error.
* **Base URL**: request `base_url`, then `<PROVIDER>_BASE_URL`, then the built-in default.
* **Deadline**: `timeout_secs` (default 300) covers upload, submission, polling and result download, and
  also caps each individual HTTP request.
* **Retries**: `max_retries` (default 2) with exponential backoff and full jitter, only on rate-limit,
  network, and 500/502/503/504 errors. 4xx is never retried.
* **Error kinds**: 401/403 → `authentication`, 429 → `rate_limit`, other 4xx → `bad_request`,
  5xx / malformed payload / failed job → `provider`, deadline → `timeout`.
* **None of the three returns rate-limit headers** (`X-RateLimit-*`), so backoff is always blind.

## How to add a provider

See [`CONTRIBUTING.md`](/docs/project/contributing/index.md), section **"3. Adding a provider"** — one file under
`crates/puffinparse-core/src/providers/`, registered in `providers/mod.rs` and `model::PROVIDERS`, plus
`pricing.json` entries, a fixture-backed normalisation test, and a doc page here following the same
eight sections as the pages above.

**Promoting a docs-only provider:** run `cargo test -p puffinparse-core <provider> -- --ignored` with a key, fix any wire-format differences, replace the hand-built fixture with a redacted real response, and change the status here, in the provider page's banner and in the README model table. Comment on [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10) to claim one.
