PuffinParse docs
/ llms.txt GitHub

Provider reference#

Reducto

REDUCTO_API_KEY

agentic deep_extract extract r-1 standard

Vendor docs

Extend

EXTEND_API_KEY

extraction_light extraction_performance parse_auto parse_light parse_performance

Vendor docs

LlamaParse (LlamaCloud)

LLAMA_API_KEY

agentic agentic_plus cost_effective fast

Vendor docs

Mistral Document AI

MISTRAL_API_KEY

ocr-2512 ocr-4-0 ocr-4-1 ocr-latest

Vendor docs

Azure AI Document Intelligence

AZURE_DOCUMENT_INTELLIGENCE_KEY

custom id_document invoice layout read receipt tax_us_w2

Vendor docs

AWS Textract

AWS_ACCESS_KEY_ID

detect-text forms layout queries

Vendor docs

Google Gemini

GEMINI_API_KEY

2.5-flash 2.5-flash-lite 2.5-pro 3.5-flash 3.5-flash-lite 3.8-flash

Vendor docs

OpenAI

OPENAI_API_KEY

gpt-5.6-luna gpt-5.6-sol gpt-5.6-terra gpt-6-astra

Vendor docs

Anthropic (Claude)

ANTHROPIC_API_KEY

claude-haiku-4-5 claude-opus-5 claude-sonnet-5

Vendor docs

Mathpix

MATHPIX_APP_KEY

pdf text

Vendor docs

Datalab (Marker)

DATALAB_API_KEY

accurate balanced fast

Vendor docs

Unstructured

UNSTRUCTURED_API_KEY

auto fast hi_res

Vendor docs

Upstage Document Parse

UPSTAGE_API_KEY

document-parse document-parse-nightly

Vendor docs

Google Cloud Document AI

GOOGLE_DOCUMENTAI_ACCESS_TOKEN

form layout ocr prebuilt

Vendor docs

One page per provider, describing exactly what PuffinParse sends, what comes back, and how the two are mapped onto the unified OcrResponse, checked against the implementation in crates/puffinparse-core/src/providers/.

Verification differs by provider, and the Status column says which applies:

  • live-verified (Reducto, Extend, LlamaParse): the #[ignore]d live tests pass against the real API with a key, and the fixtures include redacted live responses.
  • verified locally (Tesseract, Docling): self-hosted engines; the live tests pass against a local install and the fixtures are real output from it.
  • docs-only (everything else): implemented from the provider's API documentation and tested against fixture payloads built from it, but not yet run against the live API. Expect wire-format differences until they are verified; issue #10 tracks this, and help with a key is welcome.
Provider Doc Implementation Env var Models Status
Reducto reducto.md providers/reducto.rs REDUCTO_API_KEY standard (default), r-1, agentic live-verified
Extend extend.md providers/extend.rs EXTEND_API_KEY parse_performance (default), parse_light, parse_auto live-verified
LlamaParse llamaparse.md providers/llamaparse.rs LLAMA_API_KEY fast, cost_effective (default), agentic, agentic_plus live-verified
Mistral mistral.md providers/mistral.rs MISTRAL_API_KEY ocr-latest (default), ocr-4-1, ocr-4-0, ocr-2512 docs-only
Azure AI Document Intelligence azure.md providers/azure.rs AZURE_DOCUMENT_INTELLIGENCE_KEY + AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT read (ocr default), layout (parse default), invoice (extract default), receipt, id_document, tax_us_w2, custom docs-only
AWS Textract textract.md providers/textract.rs AWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY (+ AWS_SESSION_TOKEN, AWS_REGION) detect-text (ocr default), layout, queries (extract default), forms docs-only
Google Gemini gemini.md providers/gemini.rs GEMINI_API_KEY 2.5-flash (default), 2.5-pro, 2.5-flash-lite, 3.5-flash, 3.5-flash-lite, 3.8-flash docs-only
OpenAI openai.md providers/openai.rs OPENAI_API_KEY gpt-5.6-luna (default), gpt-5.6-terra, gpt-5.6-sol, gpt-6-astra docs-only
Anthropic anthropic.md providers/anthropic.rs ANTHROPIC_API_KEY claude-sonnet-5 (default), claude-haiku-4-5, claude-opus-5 docs-only
Mathpix mathpix.md providers/mathpix.rs MATHPIX_APP_ID + MATHPIX_APP_KEY pdf (default), text docs-only
Datalab (Marker) datalab.md providers/datalab.rs DATALAB_API_KEY fast, balanced (default), accurate docs-only
Unstructured unstructured.md providers/unstructured.rs UNSTRUCTURED_API_KEY hi_res (default), fast, auto docs-only
Upstage upstage.md providers/upstage.rs UPSTAGE_API_KEY document-parse (default), document-parse-nightly docs-only
Landing AI (ADE) landingai.md providers/landingai.rs LANDINGAI_API_KEY dpt-2 (default) docs-only
Google Document AI google-documentai.md providers/google_documentai.rs GOOGLE_DOCUMENTAI_ACCESS_TOKEN (+ _PROJECT, _LOCATION, _PROCESSOR_ID) ocr (default), layout, form, prebuilt docs-only
Tesseract (local) tesseract.md providers/tesseract.rs none (TESSERACT_CMD, PDFTOPPM_CMD) default verified locally
Docling (self-hosted) docling.md providers/docling.rs none (DOCLING_BASE_URL; optional DOCLING_API_KEY) default verified locally
PaddleOCR (self-hosted) paddleocr.md providers/paddleocr.rs none (PADDLEOCR_BASE_URL, PADDLEOCR_PARSE_BASE_URL) default docs-only

The three self-hosted engines need no API key and are priced at $0/page (puffinparse providers shows local in the Key column); they are the open baselines in the benchmark. Their shared helpers are in providers/local.rs.

Every page follows the same structure: summary → models → request flow → response mapping → errors and limits → gotchas → provider_options examples → links.

At a glance#

Reducto Extend LlamaParse
Upload method POST /upload (multipart) → reducto://<uuid> id; URLs passed through as input POST /files/upload (multipart) → file_… id; URLs passed through as {"url", "name"} one multipart call: file part, or input_url field
Sync / async Sync POST /parse by default; async POST /parse_async + GET /job/{id} with provider_options={"async": true} Always async: POST /parse_runs + GET /parse_runs/{id} Always async: POST /api/v1/parsing/upload + GET …/job/{id} + GET …/job/{id}/result/json
Page info source blocks[].bbox.page (chunks carry no page number); PuffinParse pins chunk_mode: "page" chunk.metadata.pageRange + block.metadata.page.number; PuffinParse pins chunkingStrategy: {"type":"page"} pages[].page (native per-page objects with md and text)
Page dimensions not reported — Page.width/height are None block.metadata.page.width/height (raster pixels at the run's dpi) pages[].width/height
Bbox units on the wire already normalised 0–1, top-left origin absolute left/top/right/bottom in page units absolute bBox {x,y,w,h} in page units
Bbox after normalisation clamped as-is divided by page width/height divided by page width/height
Confidence type granular_confidence.parse_confidence (0–1), else coarse "high"/"low" → 0.9 / 0.5 metadata.avgOcrConfidence (0–1) per block items[].bBox.confidence (0–1) per item
Credits reported? yes — usage.credits (null on new per-product pricing) yes — usage.credits (null for runs before 2025-10-07) effectively no — job_credits_usage is 0 until billing settles, so PuffinParse reports None
Large-result indirection result.type == "url" → presigned JSON fetched without auth outputUrl (15-min presigned) when responseType=url none — result endpoints return inline JSON
Job id surfaced as job_id run id (pr_…) job id (UUID)

Usage.pages comes from usage.num_pages (Reducto), metrics.pageCount (Extend) and job_metadata.job_pages (LlamaParse), falling back to the number of reconstructed pages. cost_usd is always pages × per_page_usd from crates/puffinparse-core/src/pricing.json — no provider reports dollar cost directly.

Shared behaviour#

These are implemented once, in crates/puffinparse-core/src/{http,error,provider}.rs, and apply to all three providers:

  • API key: request api_key first, then the provider's env var; missing ⇒ authentication error.
  • Base URL: request base_url, then <PROVIDER>_BASE_URL, then the built-in default.
  • Deadline: timeout_secs (default 300) covers upload, submission, polling and result download, and also caps each individual HTTP request.
  • Retries: max_retries (default 2) with exponential backoff and full jitter, only on rate-limit, network, and 500/502/503/504 errors. 4xx is never retried.
  • Error kinds: 401/403 → authentication, 429 → rate_limit, other 4xx → bad_request, 5xx / malformed payload / failed job → provider, deadline → timeout.
  • None of the three returns rate-limit headers (X-RateLimit-*), so backoff is always blind.

How to add a provider#

See CONTRIBUTING.md, section "3. Adding a provider" — one file under crates/puffinparse-core/src/providers/, registered in providers/mod.rs and model::PROVIDERS, plus pricing.json entries, a fixture-backed normalisation test, and a doc page here following the same eight sections as the pages above.

Promoting a docs-only provider: run cargo test -p puffinparse-core <provider> -- --ignored with a key, fix any wire-format differences, replace the hand-built fixture with a redacted real response, and change the status here, in the provider page's banner and in the README model table. Comment on issue #10 to claim one.