# PuffinParse — full documentation > PuffinParse is one API for every OCR and document-parsing provider (Reducto, Extend, LlamaParse, Mistral, Azure, Textract, Tesseract, Docling, ...). A Rust core with typed Python and TypeScript SDKs, a CLI, a self-hosted gateway, and an open benchmark that ranks providers on accuracy, latency and cost. Switching provider is a one-string change. Source: https://github.com/ajinkyashejul/puffinparse --- # PuffinParse **One API for every document parser: parse, OCR and extract.** Rust core, Python and TypeScript SDKs, a CLI, a self-hosted gateway, and an open benchmark that ranks providers on accuracy, latency and cost. ```python import puffinparse doc = puffinparse.parse("invoice.pdf", model="reducto/standard") # or "extend/parse_performance", "llamaparse/agentic", ... print(doc.markdown) # unified markdown, every provider print(doc.pages[0].blocks[0].bbox, doc.usage.pages, doc.cost_usd) text = puffinparse.ocr("scan.png", model="llamaparse/fast") # plain text + line/word boxes print(text.text, text.pages[0].lines[0].bbox) ``` Switch providers by changing one string. Same request, same response shape, same errors. ## Why Every document-parsing vendor has its own upload flow, polling loop, JSON layout, block vocabulary, coordinate system and billing unit. PuffinParse hides all of that behind one call, tracks cost per call, retries and falls back across providers, and ships a reproducible benchmark so you can pick a provider on evidence instead of marketing. Two things make switching real rather than aspirational. **Modes**: every call names `parse`, `ocr` or `extract`, models declare the modes they serve, and a model that cannot serve the one you asked for fails before any network call — so a provider swap can never quietly change the shape of your answer. **Native-format compatibility**: if you are already integrated with Reducto, Extend or LlamaParse, `output_format="reducto"` (and friends) renders *any* provider's result into that vendor's own JSON, so you can re-point a request without touching your parsing code — [see below](#keep-your-reducto-extend-llamaparse-code). | | | |---|---| | **Providers (v0.1)** | 18 providers · 60 models · 3 modes. **Live-verified** against the real APIs: [Reducto](https://reducto.ai), [Extend](https://extend.ai), [LlamaParse](https://cloud.llamaindex.ai). **Verified locally**: the self-hosted Tesseract and Docling. **Docs-only** (implemented from the provider's API documentation and tested against fixture payloads, not yet run live): Mistral, Azure, Textract, Gemini, OpenAI, Anthropic, Mathpix, Datalab, Unstructured, Upstage, Landing AI, Google Document AI and the self-hosted PaddleOCR ([help verify them](https://github.com/ajinkyashejul/puffinparse/issues/10)). [Full table](#model-names) | | **Modes** | `parse` (markdown + blocks), `ocr` (plain text + boxes), `extract` (JSON from a schema) | | **Core** | Rust (`puffinparse-core`): `reqwest` + `tokio`, no vendor SDKs, `#![forbid(unsafe_code)]` | | **SDKs** | Python 3.9+ (sync + async, fully typed) and Node.js / TypeScript ([`js/`](https://github.com/ajinkyashejul/puffinparse/blob/main/js/README.md)), both on the same Rust core | | **Gateway** | `puffinparse serve`: one HTTP endpoint with aliases, fallbacks, virtual keys, budgets, rate limits, JSON logs and Prometheus metrics ([`docs/SERVER.md`](/docs/gateway/index.md)) | | **Long documents** | `submit` / `retrieve` jobs and provider webhooks instead of a blocking call ([below](#long-documents-jobs-and-webhooks)) | | **CLI** | `puffinparse parse`, `puffinparse ocr`, `puffinparse extract`, `puffinparse providers`, `puffinparse bench` | | **Reliability** | Retries with jittered backoff, whole-call deadlines, `Router` with ordered fallbacks / round-robin | | **Compatibility** | `output_format` renders any provider's result in Reducto's, Extend's or LlamaParse's own JSON, so an existing integration keeps its parser ([`docs/COMPAT.md`](/docs/project/compat/index.md)) | | **Cost** | Embedded, overridable price table → `cost_usd` on every response | | **Benchmark** | One harness over synthetic data and public benchmarks (ParseBench, olmOCR-bench, OmniDocBench, DP-Bench); deterministic metrics and rule checks, latency, $/1k pages; every output inspectable at [puffinparse.com/benchmark-results](https://puffinparse.com/benchmark-results/) | ## Install ```bash pip install puffinparse # Python SDK (abi3 wheels: Linux, macOS, Windows) npm install puffinparse # Node.js SDK (prebuilt for Linux x64/arm64 glibc, macOS, Windows x64) cargo install puffinparse-cli # CLI + gateway; or grab an archive from GitHub Releases docker pull ghcr.io/ajinkyashejul/puffinparse # gateway image ``` Set the keys for the providers you use: ```bash export REDUCTO_API_KEY=... export EXTEND_API_KEY=... export LLAMA_API_KEY=llx-... # LlamaCloud / LlamaParse ``` No key yet? The self-hosted engines work out of the box once installed: ```bash sudo apt-get install tesseract-ocr poppler-utils # or: brew install tesseract poppler puffinparse ocr scan.png -m tesseract # free, local, word boxes + confidences ``` Every provider reads its own variable: [`.env.example`](https://github.com/ajinkyashejul/puffinparse/blob/main/.env.example) lists all of them, the [model tables](#model-names) say which belongs to which provider, and `puffinparse providers` shows which ones are set in your shell. From source (Rust stable + Python 3.9+): ```bash git clone https://github.com/ajinkyashejul/puffinparse && cd puffinparse python -m venv .venv && . .venv/bin/activate pip install maturin && maturin develop --release # builds the extension into the venv cargo build --release -p puffinparse-cli # ./target/release/puffinparse ``` ## Modes Document AI vendors sell three different products, and they are not interchangeable: layout **parsing**, plain-text **OCR**, and schema-driven **extraction**. A call picks one mode, and the mode decides the response type: | Mode | Call | Returns | Use it for | |---|---|---|---| | `parse` | `puffinparse.parse(...)` | `ParseResponse` — `markdown`, `pages[].blocks[]` with types and boxes | RAG chunks, tables, document structure | | `ocr` | `puffinparse.ocr(...)` | `TextResponse` — `text`, `pages[].lines[]` / `words[]` with boxes | search indexes, redaction, overlays | | `extract` | `puffinparse.extract(..., schema)` | `ExtractResponse` — `data` shaped by your JSON Schema, plus per-field confidence and citations | invoices, forms, anything with fields | **Providers are swappable only within a mode.** Every model declares the modes it serves, so a model that cannot do what you asked raises `UnsupportedModelError` *before* any network call instead of silently returning the wrong shape. `puffinparse.list_models("ocr")` lists the candidates for a mode; `puffinparse providers --mode ocr` does the same on the command line. Providers without a native OCR endpoint serve `ocr` from their parse output, flagged as `resp.metadata["puffinparse_derived_from"] == "parse"`. > Which models serve which mode is a registry fact, not a guess — ask > `puffinparse.list_models("extract")` or `puffinparse providers --mode extract`. Most providers serve > `parse` and `ocr`; `extract` needs a model built for it (`reducto/extract`, > `extend/extraction_light`, `azure/invoice`, the vision-LLM models, ...), and calling it with a > parse-only model raises `UnsupportedModelError` naming the mode. ## Usage ### One call, any provider ```python import puffinparse doc = puffinparse.parse("contract.pdf", model="extend/parse_performance") doc = puffinparse.parse("https://cdn.reducto.ai/samples/fidelity-example.pdf", model="reducto/r-1", pages="1-2") doc = puffinparse.parse(open("scan.png", "rb").read(), filename="scan.png", model="llamaparse/agentic") doc.markdown # whole document doc.text # plain text doc.pages[0].markdown # per page for block in doc.pages[0].blocks: block.type # text | title | section_header | list | table | figure | header | footer | ... block.content # markdown block.bbox # BBox(x0, y0, x1, y1) normalised 0..1, origin top-left (or None) block.confidence # 0..1 when the provider reports one doc.usage.pages, doc.usage.credits, doc.cost_usd, doc.latency_ms ``` ### Plain text and boxes (`ocr`) ```python text = puffinparse.ocr("scan.png", model="reducto/standard") text.text # whole document, pages joined by a blank line page = text.pages[0] page.text # plain text in reading order for line in page.lines: # Line(text, bbox, confidence) x0, y0, x1, y1 = line.bbox.to_pixels(page.width, page.height) for word in page.words: # Word(text, bbox, confidence) ... ``` ### Structured extraction (`extract`) ```python schema = { "type": "object", "properties": {"invoice_number": {"type": "string"}, "total": {"type": "number"}}, "required": ["invoice_number", "total"], } result = puffinparse.extract("invoice.pdf", schema, model="...", citations=True) result.data # {"invoice_number": "INV-42", "total": 1280.5} result.fields["/total"].confidence # per-field confidence, keyed by JSON pointer result.citations("/total") # [Citation(page_number=2, bbox=..., text="Total due 1,280.50")] ``` ### Async ```python doc = await puffinparse.aparse("contract.pdf", model="llamaparse/cost_effective") text = await puffinparse.aocr("scan.png", model="llamaparse/fast") ``` Runs on the Rust runtime; the event loop is never blocked. ### Long documents: jobs and webhooks `parse` waits for the provider. For long documents, batches or webhook-driven pipelines, split it: ```python job = puffinparse.submit("annual-report.pdf", model="reducto/standard", webhook_url="https://example.com/hooks/puffinparse") # optional store(job) # a Job is plain data: serialisable, holds no key result = puffinparse.retrieve(job) # Job (still pending) or ParseResponse # ...or, in your web handler, turn the provider's webhook body into a result: result = puffinparse.handle_webhook(request.json(), model="reducto") # verify the signature first ``` Reducto, Extend and LlamaParse support jobs; `webhook_url` maps to each provider's per-job webhook where one exists (Extend only has workspace-level webhooks, so it is rejected there). See [`docs/SPEC.md`](/docs/project/spec/index.md) §15. ### TypeScript / Node.js The same core as a napi-rs addon, with camelCase typed responses: ```ts import { parse, Router } from "puffinparse"; const doc = await parse("invoice.pdf", { model: "reducto/standard", fallbacks: ["llamaparse/agentic"] }); console.log(doc.markdown, doc.usage.pages, doc.costUsd); ``` `npm install puffinparse` ships prebuilt binaries for Linux x64/arm64 (glibc), macOS and Windows x64; other platforms build from source (`cd js && npm ci && npm run build`), see [`js/README.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/js/README.md). ### Model names `"/"`, like LiteLLM. A bare provider name picks that provider's default model **for the mode you called** — marked `*` below. `puffinparse providers` and `puffinparse.list_models(mode)` print the live list; the tables here are the built-in registry (`crates/puffinparse-core/src/model.rs` and `pricing.json`). Prices are public pay-as-you-go **list prices, per page, per mode**, shown as `parse · ocr · extract` with `—` where a model does not serve that mode. Override one mode at a time with `puffinparse.set_pricing({"reducto/standard": 0.012}, "parse")`, and estimate with `puffinparse.estimate_cost("reducto/standard", pages=1000, mode="ocr")`. The vision-LLM providers (Gemini, OpenAI, Anthropic) bill tokens rather than pages, so their per-page numbers are **estimates** — see the source lines in `pricing.json`. Each provider carries one of three verification labels: - **live-verified**: the live tests pass against the real API with a key, and the fixtures include redacted live responses (Reducto, Extend, LlamaParse). - **verified locally**: a self-hosted engine whose live tests pass against a local install (Tesseract, Docling). - **docs-only**: implemented from the provider's API documentation and tested against fixture payloads built from it, but not yet run against the live API. Expect wire-format differences until it is verified; [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10) tracks this, and a run with your own key is a welcome contribution. [`docs/providers/README.md`](/docs/providers/index.md) tracks the state and links one reference page per provider. **Reducto** · `REDUCTO_API_KEY` · live-verified — Layout parsing plus a schema extractor with citations; the default parse target. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `reducto/standard` | parse `*`, ocr `*` | $0.015 · $0.015 · — | Reducto Parse with account-default model (legacy standard) | | `reducto/r-1` | parse, ocr | $0.01 · $0.01 · — | Reducto Parse with settings.model=r-1 (newest model, cheaper) | | `reducto/agentic` | parse, ocr | $0.03 · $0.03 · — | Reducto Parse with agentic text+table enhancement (highest accuracy, 2x cost) | | `reducto/extract` | extract `*` | — · — · $0.035 | Reducto Extract: POST /extract with a JSON schema, citations on request | | `reducto/deep_extract` | extract | — · — · $0.055 | Reducto Deep Extract (settings.deep_extract = true) for long or complex documents | **Extend** · `EXTEND_API_KEY` · live-verified — Parse engines that trade accuracy for cost per page, and two extraction processors. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `extend/parse_performance` | parse `*`, ocr `*` | $0.025 · $0.025 · — | Extend engine=parse_performance (highest accuracy) | | `extend/parse_light` | parse, ocr | $0.00625 · $0.00625 · — | Extend engine=parse_light (fast, cheap, digital-native docs) | | `extend/parse_auto` | parse, ocr | $0.025 · $0.025 · — | Extend engine=parse_auto (picks light or performance per page) | | `extend/extraction_performance` | extract `*` | — · — · $0.0625 | Extend Extract, baseProcessor=extraction_performance (runs parse_performance) | | `extend/extraction_light` | extract | — · — · $0.015 | Extend Extract, baseProcessor=extraction_light (runs parse_light) | **LlamaParse (LlamaCloud)** · `LLAMA_API_KEY` · live-verified — Credit-priced tiers from plain text extraction to agentic parsing; extract from `cost_effective` up. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `llamaparse/fast` | parse, ocr | $0.00125 · $0.00125 · — | LlamaParse tier=fast (text extraction, no OCR of images) | | `llamaparse/cost_effective` | parse `*`, ocr `*`, extract `*` | $0.00375 · $0.00375 · $0.01 | LlamaParse tier=cost_effective | | `llamaparse/agentic` | parse, ocr, extract | $0.0125 · $0.0125 · $0.03125 | LlamaParse tier=agentic | | `llamaparse/agentic_plus` | parse, ocr, extract | $0.05625 · $0.05625 · $0.11875 | LlamaParse tier=agentic_plus (highest accuracy) | **Mistral Document AI** · `MISTRAL_API_KEY` · docs-only — One OCR model with pinned versions; annotations give it an extract mode. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `mistral/ocr-latest` | parse `*`, ocr `*`, extract `*` | $0.004 · $0.004 · $0.005 | Mistral OCR, latest alias (mistral-ocr-latest; currently OCR 4.1) | | `mistral/ocr-4-1` | parse, ocr, extract | $0.004 · $0.004 · $0.005 | Mistral OCR 4.1 pinned (mistral-ocr-4-1; blocks + block confidence scores) | | `mistral/ocr-4-0` | parse, ocr, extract | $0.004 · $0.004 · $0.005 | Mistral OCR 4.0 pinned (mistral-ocr-4-0; paragraph blocks, no block confidence) | | `mistral/ocr-2512` | parse, ocr, extract | $0.002 · $0.002 · $0.003 | Mistral OCR 3 pinned (mistral-ocr-2512; cheaper, no paragraph blocks) | **Azure AI Document Intelligence** · `AZURE_DOCUMENT_INTELLIGENCE_KEY` (+ `AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT`) · docs-only — Prebuilt Document Intelligence models: `read` for OCR, `layout` for structure, fixed schemas for extraction. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `azure/read` | ocr `*` | — · $0.0015 · — | Azure prebuilt-read: native OCR, words/lines with confidence (cheapest) | | `azure/layout` | parse `*`, ocr | $0.01 · $0.01 · — | Azure prebuilt-layout: markdown, paragraphs with roles, tables, polygons | | `azure/invoice` | extract `*` | — · — · $0.01 | Azure prebuilt-invoice: fixed invoice schema | | `azure/receipt` | extract | — · — · $0.01 | Azure prebuilt-receipt: fixed receipt schema | | `azure/id_document` | extract | — · — · $0.01 | Azure prebuilt-idDocument: fixed ID document schema | | `azure/tax_us_w2` | extract | — · — · $0.01 | Azure prebuilt-tax.us.w2: fixed W-2 schema | | `azure/custom` | parse, ocr, extract | $0.03 · $0.03 · $0.03 | Azure custom model, id via provider_options.model_id | **AWS Textract** · `AWS_ACCESS_KEY_ID` (+ `AWS_SECRET_ACCESS_KEY`, `AWS_REGION`) · docs-only — One AWS API, four feature sets; SigV4-signed, no SDK. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `textract/detect-text` | ocr `*` | — · $0.0015 · — | Textract DetectDocumentText: raw OCR, lines + words with boxes (cheapest) | | `textract/layout` | parse `*`, ocr | $0.015 · $0.015 · — | Textract AnalyzeDocument LAYOUT + TABLES: reading-order markdown + tables | | `textract/queries` | extract `*` | — · — · $0.015 | Textract AnalyzeDocument QUERIES: one natural-language query per schema field | | `textract/forms` | extract | — · — · $0.05 | Textract AnalyzeDocument FORMS: key-value pairs matched to schema fields | **Google Gemini** · `GEMINI_API_KEY` · docs-only — Vision-LLM transcription. Prices are per-page **estimates** (~1500 in + ~800 out tokens); `cost_usd` uses them until the response reports real usage. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `gemini/2.5-flash` | parse `*`, ocr `*`, extract `*` | $0.00245 · $0.00245 · $0.0012 | Gemini 2.5 Flash: vision-LLM transcription, the price/quality default | | `gemini/2.5-pro` | parse, ocr, extract | $0.009875 · $0.009875 · $0.004875 | Gemini 2.5 Pro: highest accuracy, ~4x the cost of Flash | | `gemini/2.5-flash-lite` | parse, ocr, extract | $0.00047 · $0.00047 · $0.00027 | Gemini 2.5 Flash-Lite: cheapest and fastest, clean documents | | `gemini/3.5-flash` | parse, ocr, extract | $0.00945 · $0.00945 · $0.00495 | Gemini 3.5 Flash: frontier Flash generation (GA 2026-05-19) | | `gemini/3.5-flash-lite` | parse, ocr, extract | $0.00245 · $0.00245 · $0.0012 | Gemini 3.5 Flash-Lite: low-latency 3.x tier (GA 2026-07-21) | | `gemini/3.8-flash` | parse, ocr, extract | $0.004125 · $0.004125 · $0.00225 | Gemini 3.8 Flash: newest Flash model (GA 2026-09-02, introductory pricing) | **OpenAI** · `OPENAI_API_KEY` · docs-only — Vision transcription through the Responses API. Per-page prices are estimates, as for Gemini. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `openai/gpt-5.6-luna` | parse `*`, ocr `*`, extract `*` | $0.00114 · $0.00114 · $0.00114 | OpenAI Responses API, gpt-5.6-luna (cheapest current vision model) | | `openai/gpt-5.6-terra` | parse, ocr, extract | $0.0114 · $0.0114 · $0.0114 | OpenAI Responses API, gpt-5.6-terra (balanced capability/price) | | `openai/gpt-5.6-sol` | parse, ocr, extract | $0.02 · $0.02 · $0.02 | OpenAI Responses API, gpt-5.6-sol (flagship GPT-5.6) | | `openai/gpt-6-astra` | parse, ocr, extract | $0.05 · $0.05 · $0.05 | OpenAI Responses API, gpt-6-astra (most capable, most expensive) | **Anthropic (Claude)** · `ANTHROPIC_API_KEY` · docs-only — Vision transcription through the Messages API. Per-page prices are estimates. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `anthropic/claude-sonnet-5` | parse `*`, ocr `*`, extract `*` | $0.01 · $0.01 · $0.01 | Claude Messages API, claude-sonnet-5 (balanced vision transcription) | | `anthropic/claude-haiku-4-5` | parse, ocr, extract | $0.005 · $0.005 · $0.005 | Claude Messages API, claude-haiku-4-5 (cheapest, 200K context) | | `anthropic/claude-opus-5` | parse, ocr, extract | $0.025 · $0.025 · $0.025 | Claude Messages API, claude-opus-5 (highest accuracy) | **Mathpix** · `MATHPIX_APP_KEY` (+ `MATHPIX_APP_ID`) · docs-only — Maths-first OCR: Mathpix Markdown with LaTeX, line and word polygons. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `mathpix/pdf` | parse `*`, ocr `*` | $0.005 · $0.005 · — | Mathpix v3/pdf document OCR (image inputs auto-routed to v3/text); MMD + line polygons | | `mathpix/text` | parse, ocr | $0.002 · $0.002 · — | Mathpix v3/text single-image OCR (line + word polygons, per-image billing) | **Datalab (Marker)** · `DATALAB_API_KEY` · docs-only — Marker's hosted API: three quality modes at two price points. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `datalab/fast` | parse, ocr | $0.004 · $0.004 · — | Datalab Convert mode=fast (lowest latency, digital-native documents) | | `datalab/balanced` | parse `*`, ocr `*` | $0.004 · $0.004 · — | Datalab Convert mode=balanced (Datalab's recommended default) | | `datalab/accurate` | parse, ocr | $0.01 · $0.01 · — | Datalab Convert mode=accurate (scans, dense layouts, complex tables) | **Unstructured** · `UNSTRUCTURED_API_KEY` · docs-only — The partitioner behind many RAG pipelines; flat per-page price across strategies. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `unstructured/hi_res` | parse `*`, ocr `*` | $0.015 · $0.015 · — | Unstructured strategy=hi_res (layout model + OCR; coordinates, table HTML, confidence) | | `unstructured/fast` | parse, ocr | $0.015 · $0.015 · — | Unstructured strategy=fast (text-layer extraction, no OCR, rejects images) | | `unstructured/auto` | parse, ocr | $0.015 · $0.015 · — | Unstructured strategy=auto (routes each page to fast / hi_res / VLM) | **Upstage Document Parse** · `UPSTAGE_API_KEY` · docs-only — Korean/English layout parsing to HTML or markdown. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `upstage/document-parse` | parse `*`, ocr `*` | $0.01 · $0.01 · — | Upstage Document Parse (layout to HTML/Markdown; 100 pages sync, 1000 async) | | `upstage/document-parse-nightly` | parse, ocr | $0.01 · $0.01 · — | Upstage Document Parse nightly build (newest layout model, may change without notice) | **Landing AI (Agentic Document Extraction)** · `LANDINGAI_API_KEY` · docs-only — Agentic Document Extraction (DPT-2): parse, then extract against your schema. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `landingai/dpt-2` | parse `*`, ocr `*`, extract `*` | $0.03 · $0.03 · $0.04 | Landing AI ADE DPT-2; extract runs parse then /v1/ade/extract with the JSON schema | **Google Cloud Document AI** · `GOOGLE_DOCUMENTAI_ACCESS_TOKEN` (+ `GOOGLE_DOCUMENTAI_PROJECT`, `_LOCATION`, `_PROCESSOR_ID`) · docs-only — Processor-based: pick the processor id in `provider_options`; OCR, layout, forms and prebuilt extractors. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `google_documentai/ocr` | parse `*`, ocr `*` | $0.0015 · $0.0015 · — | Document OCR processor: native text + word boxes (processor id from config) | | `google_documentai/layout` | parse, ocr | $0.01 · $0.01 · — | Layout Parser processor: documentLayout blocks (headings, tables, lists) | | `google_documentai/form` | parse, ocr, extract `*` | $0.03 · $0.03 · $0.03 | Form Parser processor: paragraphs + tables, entities as extraction | | `google_documentai/prebuilt` | parse, ocr, extract | $0.03 · $0.03 · $0.03 | Prebuilt or custom extractor (invoice, W2, ...): entities as extraction | **Self-hosted engines** · no key · $0/page — out-of-process: a local binary or a server you run. Tesseract and Docling are verified locally; PaddleOCR is docs-only. | Model | Modes | List price / page (parse · ocr · extract) | Notes | |---|---|---|---| | `tesseract/default` | ocr `*` (native), parse `*` | $0 · $0 · — | local `tesseract` binary (`TESSERACT_CMD`), PDFs via `pdftoppm`; word/line boxes + confidences, no layout model; verified locally | | `docling/default` | parse `*`, ocr `*` | $0 · $0 · — | your docling-serve (`DOCLING_BASE_URL`): layout, tables, OCR; verified locally | | `paddleocr/default` | ocr `*` (native), parse `*` (PP-StructureV3) | $0 · $0 · — | your PaddleOCR serving (`PADDLEOCR_BASE_URL`); docs-only | ### Keep your Reducto / Extend / LlamaParse code Already integrated with a vendor? Ask for its shape and PuffinParse renders the response into that vendor's own JSON, whatever provider actually ran the call — switch the model string, keep your parser: ```python doc = puffinparse.parse("invoice.pdf", model="extend/parse_light", output_format="reducto") block = doc["result"]["chunks"][0]["blocks"][0] # Reducto's shape, Extend's engine draw(block["bbox"]["left"], block["bbox"]["top"], block["type"]) ``` `output_format` accepts `reducto`, `extend`, `llamaparse`, or `puffinparse` (the unified shape, and the default). It returns a plain `dict` instead of a dataclass, and works on `parse`, `extract`, their async variants, the `Router`, and the CLI (`--output-format ` with `--format json`). An unknown name raises `BadRequestError` before any network call. What is guaranteed is **structural fidelity** — key set and nesting, one chunk/page per page, the content strings, the vendor's own block vocabulary and coordinate units, the billed page count — not byte equality with what the vendor would have returned. Fields PuffinParse does not model (presigned URLs, studio links, billing breakdowns, OCR word layers) are `null` or empty, and a few block types are lossy. [`docs/COMPAT.md`](/docs/project/compat/index.md) enumerates all of it, per format; `examples/switch_provider_keep_format.py` is a runnable version of the above. ### Provider-specific options Anything the common request doesn't cover is passed through verbatim: ```python puffinparse.parse("doc.pdf", model="reducto/standard", provider_options={"settings": {"return_ocr_data": True}, "async": True}) puffinparse.parse("doc.pdf", model="extend/parse_performance", provider_options={"blockOptions": {"figures": {"enabled": False}}}) puffinparse.ocr("doc.pdf", model="llamaparse/agentic", provider_options={"take_screenshot": True, "version": "2026-08-19"}) ``` `include_raw=True` attaches the provider's original payload as `resp.raw`. ### Router: fallbacks and load balancing ```python router = puffinparse.Router( ["reducto/standard", "llamaparse/agentic", "extend/parse_light"], mode="parse", # the mode every model must serve; "ocr" and "extract" too strategy="ordered", # or "round_robin" ) resp = router.parse("doc.pdf") # falls back on provider / rate-limit / timeout / network errors router.stats() # per-model successes, failures, latency, cost, pages ``` A router is bound to one mode: `Router([...], mode="parse").ocr(...)` raises `InputError` rather than quietly changing the answer's shape, and every model is validated against the mode when the router is built. Auth, bad-request and input errors never trigger a fallback. ### Errors All errors derive from `puffinparse.PuffinParseError` and carry `provider`, `status_code`, `job_id` and `retryable`: `AuthenticationError`, `RateLimitError`, `BadRequestError`, `ProviderError`, `TimeoutError`, `UnsupportedModelError`, `InputError`, `NetworkError`. `UnsupportedModelError` also covers a model asked to do a mode it does not serve; the message names the mode and the models that do serve it. ### Callbacks and logging Callbacks fire for every mode, with that mode's response object: ```python puffinparse.success_callback.append(lambda r: print(r.model, r.usage.pages, r.cost_usd)) puffinparse.failure_callback.append(lambda e: print("failed:", e)) puffinparse.init_logging("debug") # Rust core tracing on stderr (or PUFFINPARSE_LOG=debug for the CLI) ``` ### CLI ```bash puffinparse providers # models, modes, prices, key status puffinparse providers --mode ocr # only models that serve a mode puffinparse parse invoice.pdf -m extend/parse_light # markdown to stdout puffinparse parse scan.png -m llamaparse/agentic -f json --raw # full unified JSON (+ provider payload) puffinparse parse doc.pdf -m extend/parse_light -f json \ --output-format reducto # ... in Reducto's response shape puffinparse parse doc.pdf -m reducto/r-1 --pages 1-3 -f text puffinparse ocr scan.png -m reducto/standard # plain text to stdout puffinparse ocr scan.png -f json # TextResponse: text + lines + words puffinparse extract invoice.pdf -s schema.json --citations # JSON object from a schema puffinparse extract invoice.pdf -s '{"type":"object"}' # inline schema also works puffinparse extract invoice.pdf -s schema.json --output-format extend # Extend's extract_run shape puffinparse providers --json | jq '.output_formats' # the vendor shapes this build renders ``` ### Gateway server Run PuffinParse as one HTTP endpoint so applications never hold provider keys: ```bash puffinparse serve --config puffinparse.toml # or: docker run ghcr.io/ajinkyashejul/puffinparse (docs/SERVER.md) curl -H "Authorization: Bearer $TEAM_KEY" -F file=@invoice.pdf -F model=invoices \ http://localhost:4000/v1/parse ``` `puffinparse.toml` defines aliases (`invoices = [reducto/standard, extend/parse_performance]` with ordered or round-robin fallback), provider keys as `env:` references, and virtual keys with model allow-lists, monthly USD budgets and per-minute limits. `/v1/models`, `/v1/usage`, `/health` and Prometheus `/metrics` are built in; request logs are JSON lines that never contain document content or secrets. Reference: [`docs/SERVER.md`](/docs/gateway/index.md), sample: [`examples/server/puffinparse.toml`](https://github.com/ajinkyashejul/puffinparse/blob/main/examples/server/puffinparse.toml). ### Rust ```rust use puffinparse_core::{extract, ocr, parse, DocumentRequest, ExtractRequest}; let doc = parse(DocumentRequest::from_path("invoice.pdf").model("reducto/standard")).await?; println!("{} pages, ${:.4}\n{}", doc.usage.pages, doc.cost_usd.unwrap_or(0.0), doc.markdown); let text = ocr(DocumentRequest::from_path("scan.png").model("llamaparse/fast")).await?; println!("{} lines on page 1", text.pages[0].lines.len()); let req = ExtractRequest::new(DocumentRequest::from_path("invoice.pdf").model("..."), schema); let data = extract(req.citations(true)).await?.data; ``` ## Benchmark PuffinParse ships an open, reproducible benchmark. Ground truth is exact by construction (the documents are rendered from the same source as the truth files), metrics are deterministic text comparisons, and every run records the dataset hash, model, latency and cost. ```bash python benchmark/generate_synthetic.py # regenerate the dataset (byte-reproducible) puffinparse bench run --dataset benchmark/datasets/synthetic-v1 \ --models reducto/standard extend/parse_performance llamaparse/cost_effective puffinparse bench report benchmark/results/*.json > benchmark/LEADERBOARD.md puffinparse bench score prediction.md truth.md # metrics for one pair, no network ``` Metrics (after NFKC + markdown stripping + whitespace collapsing, case-insensitive by default): - **Overall** = `100 × mean(char_similarity)`, where `char_similarity = 1 − levenshtein / max(len)` - **CER**, **WER**, **word F1** (bag-of-words precision / recall) - **Order**: Kendall-τ-style agreement of shared line order (reading order) - **Table**: character similarity restricted to markdown table rows - **Latency** p50 / p95 and ms per page; **$/1k pages** from the price table The current leaderboard is in [`benchmark/LEADERBOARD.md`](/docs/benchmark/leaderboard/index.md), and every document, output, diff and rule check is browsable at [puffinparse.com/benchmark-results](https://puffinparse.com/benchmark-results/). Datasets: `synthetic-v1` (exact truth by construction) and the headline `combined-v3`, which adds subsets of [ParseBench](https://github.com/run-llama/ParseBench), [olmOCR-bench](https://huggingface.co/datasets/allenai/olmOCR-bench), [OmniDocBench](https://github.com/opendatalab/OmniDocBench) and [DP-Bench](https://huggingface.co/datasets/upstage/dp-bench) converted by [`benchmark/adapters/`](/docs/benchmark/adapters/index.md) at pinned revisions. Older `combined-v1` and `combined-v2` runs are kept for comparison. Long runs are safe: `--dry-run` and `--max-cost` show and cap the spend before any call, and `--resume` continues an interrupted run. ## How it maps providers | | Reducto | Extend | LlamaParse | |---|---|---|---| | Upload | `POST /upload` → `reducto://` id (URLs passed directly) | `POST /files/upload` → `file_…` (URLs passed directly) | multipart `file` or `input_url` | | Parse | `POST /parse` (sync) or `/parse_async` + `GET /job/{id}` | `POST /parse_runs` + `GET /parse_runs/{id}` | `POST /api/v1/parsing/upload` + `GET …/job/{id}` + `…/result/json` | | Pages | `chunk_mode=page`, blocks grouped by `bbox.page` | page chunks, blocks by `metadata.page.number` | `pages[].md` / `items[]` | | Boxes | already normalised | divided by `metadata.page.width/height` | `bBox` divided by page `width/height` | | Usage | `usage.num_pages`, `usage.credits` | `metrics.pageCount`, `usage.credits` | `job_metadata.job_pages` | Full details, including the exact wire formats verified against live responses, are in [`docs/SPEC.md`](/docs/project/spec/index.md). ## Project layout ``` crates/puffinparse-core Rust library: types, providers, router, pricing, benchmark metrics crates/puffinparse-cli `puffinparse` binary crates/puffinparse-python PyO3 extension (puffinparse._core) crates/puffinparse-node napi-rs addon for the Node.js SDK crates/puffinparse-server HTTP gateway behind `puffinparse serve` python/puffinparse Python package (typed public API) js/ Node.js / TypeScript package benchmark/ dataset generator, adapters, datasets, results, leaderboard, viewer docs/SPEC.md specification ``` ## Roadmap Planned work is tracked in [GitHub Issues](https://github.com/ajinkyashejul/puffinparse/issues). Open items: - **Packaging**: publish 0.1.0 to PyPI, npm and crates.io with prebuilt wheels, addons and CLI binaries ([#9](https://github.com/ajinkyashejul/puffinparse/issues/9)). - **Providers**: live-verify the docs-only providers ([#10](https://github.com/ajinkyashejul/puffinparse/issues/10)), including PaddleOCR against a real server ([#17](https://github.com/ajinkyashejul/puffinparse/issues/17)). - **Benchmark**: add `tesseract/default` ([#11](https://github.com/ajinkyashejul/puffinparse/issues/11)) and `docling/default` with a hardware note ([#12](https://github.com/ajinkyashejul/puffinparse/issues/12)) to `combined-v3`; regenerate the leaderboard in one command ([#13](https://github.com/ajinkyashejul/puffinparse/issues/13)); score table-cell neighbour relations exactly ([#15](https://github.com/ajinkyashejul/puffinparse/issues/15)); a READoc long-document track ([#16](https://github.com/ajinkyashejul/puffinparse/issues/16)). - **SDK**: `output_format="mistral"` for code written against Mistral OCR responses ([#14](https://github.com/ajinkyashejul/puffinparse/issues/14)). - **Playground**: upload a document and compare up to three models side by side ([#18](https://github.com/ajinkyashejul/puffinparse/issues/18)). ## Support PuffinParse is built in the open by one person. If it saved you time, a star on GitHub helps other developers find it. Reports of a wrong score, a missing provider or a confusing doc help even more: [open an issue](https://github.com/ajinkyashejul/puffinparse/issues/new/choose). ## Contributing See [CONTRIBUTING.md](/docs/project/contributing/index.md). Adding a provider is one Rust file plus a fixture test; the checklist is in the [new provider issue template](https://github.com/ajinkyashejul/puffinparse/blob/main/.github/ISSUE_TEMPLATE/new_provider.md). Issues labelled [`good first issue`](https://github.com/ajinkyashejul/puffinparse/labels/good%20first%20issue) are a good place to start, and [AGENTS.md](https://github.com/ajinkyashejul/puffinparse/blob/main/AGENTS.md) summarises the build commands and conventions for coding agents. ## License MIT. See [LICENSE](https://github.com/ajinkyashejul/puffinparse/blob/main/LICENSE). --- # Getting started PuffinParse gives you one call for every OCR / document-parsing provider. Install it, set one key, and parse a document in under a minute. ## 1. Install ### Python ```bash pip install puffinparse ``` Python 3.9+. The wheel bundles the Rust core — there is no toolchain to install and no provider SDK to add. ### CLI The `puffinparse` binary is built from the Rust workspace: ```bash cargo install --git https://github.com/ajinkyashejul/puffinparse puffinparse-cli # or, from a clone: cargo build --release -p puffinparse-cli # ./target/release/puffinparse ``` ### Rust ```toml [dependencies] puffinparse-core = "0.1" tokio = { version = "1", features = ["rt-multi-thread", "macros"] } ``` ### From source ```bash git clone https://github.com/ajinkyashejul/puffinparse && cd puffinparse python -m venv .venv && . .venv/bin/activate pip install maturin && maturin develop --release # builds puffinparse._core into the venv cargo build --release -p puffinparse-cli ``` ## 2. Set a key Each provider reads its own environment variable. Set only the ones you use. | Provider | Environment variable | Get a key | |---|---|---| | Reducto | `REDUCTO_API_KEY` | [reducto.ai](https://reducto.ai) | | Extend | `EXTEND_API_KEY` | [extend.ai](https://extend.ai) | | LlamaParse | `LLAMA_API_KEY` (starts with `llx-`) | [cloud.llamaindex.ai](https://cloud.llamaindex.ai) | ```bash export REDUCTO_API_KEY=... export EXTEND_API_KEY=... export LLAMA_API_KEY=llx-... ``` Keys can also be passed per call (`api_key=...` / `--api-key`), and base URLs overridden with `REDUCTO_BASE_URL`, `EXTEND_BASE_URL`, `LLAMA_BASE_URL`. See [`.env.example`](https://github.com/ajinkyashejul/puffinparse/blob/main/.env.example). Check what is configured: ```bash puffinparse providers ``` ## 3. First call ### Python ```python import puffinparse resp = puffinparse.parse("invoice.pdf", model="reducto/standard") print(resp.markdown) # unified markdown, identical shape for every provider print(resp.usage.pages, resp.cost_usd, resp.latency_ms) print(resp.pages[0].blocks[0].type, resp.pages[0].blocks[0].bbox) ``` Async is the same call with `await`: ```python resp = await puffinparse.aparse("invoice.pdf", model="llamaparse/cost_effective") ``` ### CLI ```bash puffinparse parse invoice.pdf -m extend/parse_light # markdown on stdout puffinparse parse scan.png -m llamaparse/agentic -f json --raw # unified JSON + provider payload ``` ### Rust ```rust use puffinparse_core::{parse, DocumentRequest}; #[tokio::main] async fn main() -> puffinparse_core::Result<()> { let resp = parse(DocumentRequest::from_path("invoice.pdf").model("reducto/standard")).await?; println!("{} pages, ${:.4}", resp.usage.pages, resp.cost_usd.unwrap_or(0.0)); println!("{}", resp.markdown); Ok(()) } ``` ## 4. Pick a mode `parse` is one of three modes. Each has its own response type, and providers can only be swapped *within* a mode. | Mode | Python | CLI | You get | |---|---|---|---| | `parse` | `puffinparse.parse(...)` | `puffinparse parse doc.pdf` | Layout-aware markdown + typed blocks with boxes. | | `ocr` | `puffinparse.ocr(...)` | `puffinparse ocr scan.png` | Plain text with line and word boxes. | | `extract` | `puffinparse.extract(..., schema)` | `puffinparse extract doc.pdf -s schema.json` | A JSON object shaped by your schema, with citations. | ```python text = puffinparse.ocr("scan.png", model="reducto/r-1") text.text, text.pages[0].lines[0].bbox data = puffinparse.extract( "invoice.pdf", {"type": "object", "properties": {"total": {"type": "number"}}}, model="reducto/standard", citations=True, ) data.data["total"], data.citations("/total") ``` `puffinparse providers --mode extract` lists the models that serve a mode; asking a model for a mode it does not support raises `UnsupportedModelError` before any network call. ## 5. Switch providers The model string is the only thing that changes. `"/"`, like LiteLLM; a bare provider name selects its default model. ```python puffinparse.parse("doc.pdf", model="reducto/r-1") puffinparse.parse("doc.pdf", model="extend/parse_performance") puffinparse.parse("doc.pdf", model="llamaparse/agentic_plus") puffinparse.parse("doc.pdf", model="reducto") # -> reducto's default parse model ``` Unknown providers or models raise `UnsupportedModelError` before any network call. ## 6. Add fallbacks ```python router = puffinparse.Router( ["reducto/standard", "llamaparse/agentic", "extend/parse_light"], mode="parse", # a router is bound to one mode strategy="ordered", # or "round_robin" ) resp = router.parse("doc.pdf") router.stats() # successes, failures, latency, cost, pages per model ``` Provider, rate-limit, timeout and network errors fall through to the next model. Auth, bad-request and input errors never do — they would fail on every provider. ## Where to next - [Python SDK](/docs/python/index.md) — every function, dataclass and exception. - [CLI](/docs/cli/index.md) — every subcommand and flag. - [Rust](/docs/rust/index.md) — `puffinparse-core` crate usage. - [Providers](/docs/providers/index.md) — exactly what PuffinParse sends and how the response is mapped. - [Benchmark](/docs/benchmark/index.md) — how the leaderboard is produced, and its caveats. - [Specification](/docs/project/spec/index.md) — the contract the implementations follow. --- # PuffinParse for agents PuffinParse is one API for hosted and local OCR / document-parsing providers (Reducto, Extend, LlamaParse, Mistral, Azure, Textract, Gemini, OpenAI, Anthropic, Tesseract, Docling and more). A Rust core with Python and TypeScript SDKs, a CLI and a self-hosted gateway. Every provider returns the same response shape and the same typed errors, so switching provider is a one-string change. MIT licensed. Source: https://github.com/ajinkyashejul/puffinparse This page is also served as plain Markdown at `https://puffinparse.com/agents.md`. The full site map for agents is `https://puffinparse.com/llms.txt`. ## Install ```bash pip install puffinparse # Python 3.9+, wheel bundles the Rust core npm install puffinparse # Node.js 18+, native addon cargo install puffinparse-cli # the `puffinparse` CLI ``` Prebuilt CLI archives are attached to each GitHub release. If a package is not available for your platform yet, [Getting started](/docs/getting-started/index.md) shows how to build from source. ## Model strings Every call takes a model string `/`, for example `reducto/standard`, `llamaparse/cost_effective` or `mistral/ocr-latest`. A bare provider name (`reducto`) means that provider's default model for the requested mode. Three modes, each with its own response type: `parse` (markdown + typed blocks with boxes), `ocr` (text + line/word boxes) and `extract` (JSON shaped by your schema). List what exists instead of guessing: ```bash puffinparse providers # providers, env vars, models, which keys are set python -c "import puffinparse; print(puffinparse.list_models('parse'))" ``` ## Choosing a model Use the open benchmark, not vendor claims. The [leaderboard](/docs/benchmark/leaderboard/index.md) ranks every benchmarked model on accuracy, p50 latency and cost per 1,000 pages; read the [methodology](/docs/benchmark/index.md) before comparing scores across datasets. Machine-readable data: - `https://puffinparse.com/benchmark-results/data/index.json`: every run, with dataset info and a per-model summary (`overall`, text and table metrics, `latency_p50_ms`, `cost_per_1k_pages_usd`, `by_category`). - `https://puffinparse.com/benchmark-results/data/runs/.json`: one full run, with per-document scores for every model. Prefer the newest `combined-v*` run. If a model is not in the benchmark, say so rather than inferring its quality. ## Python ```python import puffinparse doc = puffinparse.parse("invoice.pdf", model="reducto/standard") print(doc.markdown, doc.usage.pages, doc.cost_usd) puffinparse.estimate_cost("reducto/standard", 1000) # USD for 1,000 pages, None if unpriced ``` ## TypeScript ```ts import { parse, estimateCost } from 'puffinparse' const doc = await parse('invoice.pdf', { model: 'llamaparse/cost_effective' }) console.log(doc.markdown, doc.costUsd) estimateCost('llamaparse/cost_effective', 1000) // USD, or null when unpriced ``` ## CLI ```bash puffinparse parse invoice.pdf -m extend/parse_light # markdown on stdout puffinparse parse scan.png -m mistral/ocr-latest -f json # unified JSON ``` ## API keys Each provider reads its own environment variable; set only the ones you use: `REDUCTO_API_KEY`, `EXTEND_API_KEY`, `LLAMA_API_KEY`, `MISTRAL_API_KEY`, `AZURE_DOCUMENT_INTELLIGENCE_KEY` + `AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT`, `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY` (Textract), `GEMINI_API_KEY`, `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `MATHPIX_APP_ID` + `MATHPIX_APP_KEY`, `DATALAB_API_KEY`, `UNSTRUCTURED_API_KEY`, `UPSTAGE_API_KEY`, `LANDINGAI_API_KEY`, `GOOGLE_DOCUMENTAI_ACCESS_TOKEN` (+ `_PROJECT`, `_LOCATION`, `_PROCESSOR_ID`). Tesseract, Docling and PaddleOCR run locally or self-hosted and need no key. Full list: [Providers](/docs/providers/index.md). ## Rules for agents 1. Never print, log, echo or commit an API key. Read keys from the environment; do not paste them into code, notebooks, shell history or issue text. 2. Check verification status before relying on a provider. **live-verified**: Reducto, Extend, LlamaParse. **verified locally**: Tesseract, Docling. Everything else is **docs-only**: implemented from the provider's documentation and tested against fixtures, not yet run live. Tell the user when you pick a docs-only provider. 3. Every call costs the user money at the provider's price. Call `estimate_cost` (Python) or `estimateCost` (TypeScript) before parsing many pages, and ask before large batches. 4. Catch the typed errors (`AuthenticationError`, `RateLimitError`, `UnsupportedModelError`, ...) instead of retrying blindly; the provider's message is preserved on the exception. 5. Keep modes apart: providers are interchangeable only within one mode. ## Onboarding prompt Paste this into your coding agent to set PuffinParse up in a project: ```text Add document parsing to this project with PuffinParse (https://puffinparse.com). 1. Read https://puffinparse.com/agents.md and https://puffinparse.com/llms.txt first. 2. Install it: `pip install puffinparse` (Python) or `npm install puffinparse` (Node.js). 3. Pick a model string `/` from the benchmark leaderboard (https://puffinparse.com/docs/benchmark/leaderboard/) for my documents and budget, and tell me why. Prefer live-verified providers (Reducto, Extend, LlamaParse) or local ones (Tesseract, Docling); say so if you choose a docs-only provider. 4. Read the API key from the provider's environment variable (for example REDUCTO_API_KEY). Never print, log or commit a key; add the variable name to .env.example only. 5. Before parsing many pages, call estimate_cost / estimateCost and show me the estimate. 6. Handle PuffinParse's typed errors and keep the model string in configuration so the provider can be switched without code changes. ``` More: [Getting started](/docs/getting-started/index.md), [Python SDK](/docs/python/index.md), [TypeScript](/docs/typescript/index.md), [CLI](/docs/cli/index.md), [Gateway](/docs/gateway/index.md). --- # Python SDK `pip install puffinparse` — Python 3.9+, fully typed (`py.typed`, `mypy --strict` clean). The package is a thin, typed wrapper over the Rust core (`puffinparse._core`); no provider logic lives in Python. Everything below is the public surface exported from `puffinparse.__all__`. ```python import puffinparse puffinparse.__version__ # version of the compiled core puffinparse.modes() # ["parse", "ocr", "extract"] ``` ## Modes Every call names a **mode**, and the mode decides the response type. Providers can be swapped freely *within* a mode; a model that does not serve the mode you asked for raises `UnsupportedModelError` before any network call. | Mode | Call | Async | Returns | Use it for | |---|---|---|---|---| | `parse` | `puffinparse.parse` | `puffinparse.aparse` | `ParseResponse` | Layout-aware markdown and typed blocks with boxes. | | `ocr` | `puffinparse.ocr` | `puffinparse.aocr` | `TextResponse` | Plain text with line and word boxes — search indexes, redaction, overlays. | | `extract` | `puffinparse.extract` | `puffinparse.aextract` | `ExtractResponse` | A JSON object shaped by your schema, with per-field citations. | `Mode` is `Literal["parse", "ocr", "extract"]` and `MODES` is the same tuple as a constant. ```python doc = puffinparse.parse("invoice.pdf", model="reducto/standard") # markdown + blocks text = puffinparse.ocr("scan.png", model="llamaparse/fast") # plain text + boxes data = puffinparse.extract("invoice.pdf", schema, model="reducto/extract") ``` ## `parse` ```python def parse( input: DocumentLike, model: str = "reducto", *, output_format: Optional[str] = None, filename: Optional[str] = None, pages: Optional[str] = None, language: Optional[str] = None, output: Literal["markdown", "text"] = "markdown", provider_options: Optional[dict[str, Any]] = None, include_raw: bool = False, timeout: float = 300.0, max_retries: int = 2, api_key: Optional[str] = None, base_url: Optional[str] = None, metadata: Optional[dict[str, Any]] = None, ) -> ParseResponse # or dict[str, Any] when output_format is given ``` `DocumentLike` is `str | os.PathLike[str] | bytes | bytearray | memoryview`. | Parameter | Meaning | |---|---| | `input` | A file path, an `http(s)://` URL, or raw bytes (then `filename` is required). | | `model` | `"/"`, e.g. `"reducto/standard"`. A bare provider name selects its default model *for this mode*. | | `filename` | Required when `input` is bytes; used to infer the document type. | | `pages` | 1-based page selection such as `"1-3,7"`, forwarded best-effort. | | `language` | Language hint (ISO 639-1) when the provider supports it. | | `output` | Preferred block content: `"markdown"` (default) or `"text"`. | | `output_format` | Return a vendor's own JSON `dict` instead of the dataclass: `"reducto"`, `"extend"`, `"llamaparse"`, or `"puffinparse"` / `None` for the unified shape. See [Native-format output](#native-format-output-output_format). | | `provider_options` | Provider-specific options merged verbatim into the provider request body. | | `include_raw` | Attach the provider's raw payload as `response.raw`. | | `timeout` | Whole-call deadline in seconds (upload + polling + download). | | `max_retries` | Retries on 429 / 5xx / network errors with exponential backoff and full jitter. | | `api_key` | Override the API key (otherwise read from `REDUCTO_API_KEY`, `EXTEND_API_KEY`, `LLAMA_API_KEY`). | | `base_url` | Override the provider base URL. | | `metadata` | Free-form dict echoed back in `response.metadata`. | Passing bytes without `filename` raises `InputError` before any network call. ```python resp = puffinparse.parse("contract.pdf", model="extend/parse_performance") resp = puffinparse.parse("https://example.com/doc.pdf", model="reducto/r-1", pages="1-2") resp = puffinparse.parse(open("scan.png", "rb").read(), filename="scan.png", model="llamaparse/agentic") ``` `aparse` is the same call with `await`. It runs on the Rust runtime, so the event loop is never blocked. ```python resp = await puffinparse.aparse("contract.pdf", model="llamaparse/cost_effective") ``` ## `ocr` ```python def ocr( input: DocumentLike, model: str = "reducto", *, filename: Optional[str] = None, pages: Optional[str] = None, language: Optional[str] = None, provider_options: Optional[dict[str, Any]] = None, include_raw: bool = False, timeout: float = 300.0, max_retries: int = 2, api_key: Optional[str] = None, base_url: Optional[str] = None, metadata: Optional[dict[str, Any]] = None, ) -> TextResponse ``` Same parameters as `parse` minus `output` (the mode implies plain text). Use it when you want text and geometry rather than document structure; for markdown, tables and block types use `parse`. Providers without a native OCR endpoint serve this mode from their parse output. The response then carries `metadata["puffinparse_derived_from"] == "parse"`, so you can tell the difference. ```python text = puffinparse.ocr("scan.png", model="reducto/r-1") text.text # whole document for line in text.pages[0].lines: line.text, line.bbox, line.confidence text = await puffinparse.aocr("scan.png") # async ``` ## `extract` ```python def extract( input: DocumentLike, schema: dict[str, Any], *, output_format: Optional[str] = None, model: str = "reducto", instructions: Optional[str] = None, citations: bool = False, filename: Optional[str] = None, pages: Optional[str] = None, language: Optional[str] = None, provider_options: Optional[dict[str, Any]] = None, include_raw: bool = False, timeout: float = 300.0, max_retries: int = 2, api_key: Optional[str] = None, base_url: Optional[str] = None, metadata: Optional[dict[str, Any]] = None, ) -> ExtractResponse # or dict[str, Any] when output_format is given ``` | Parameter | Meaning | |---|---| | `schema` | A JSON Schema **object** describing the fields you want. | | `model` | Must support `extract`; a parse-only model raises `UnsupportedModelError` before any network call. | | `instructions` | Optional natural-language guidance, forwarded to providers that accept it. | | `citations` | Ask for per-field citations (page, box, source text) where the provider supports them. | | `output_format` | As for `parse`, rendering the vendor's *extract* envelope. Best effort — see [`docs/COMPAT.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/docs/COMPAT.md) §7. | Everything else matches `parse`. `aextract` is the async variant. ```python schema = { "type": "object", "properties": { "vendor": {"type": "string"}, "total": {"type": "number"}, }, } resp = puffinparse.extract("invoice.pdf", schema, model="reducto/extract", instructions="Totals are inclusive of tax.", citations=True) resp.data["total"] # the object your schema asked for resp.citations("/total") # [Citation(page_number=1, bbox=..., text="...")] resp.field_info("/total").confidence ``` ## `Router` ```python class Router: def __init__( self, models: list[str], *, mode: Mode = "parse", strategy: Literal["ordered", "round_robin"] = "ordered", fallback_on: Optional[list[str]] = None, ) -> None ``` A router is bound to **one mode** at construction: every model must support it, and calling a method for another mode raises `InputError`. That keeps fallbacks honest — a parse-only model can never quietly answer an extraction. | Member | Type | Description | |---|---|---| | `models` | `list[str]` (property) | The canonicalised model list. | | `mode` | `Mode` (property) | The mode this router serves. | | `plan()` | `list[str]` | The order models would be tried for the next call. Advances the round-robin cursor. | | `stats()` | `dict[str, dict[str, Any]]` | Per model: `successes`, `failures`, `total_latency_ms`, `total_cost_usd`, `total_pages`. | | `parse(input, **kw)` / `aparse` | `ParseResponse` | Same keyword arguments as `puffinparse.parse` except `model`, `output_format` included. | | `ocr(input, **kw)` / `aocr` | `TextResponse` | Same, minus `output`. | | `extract(input, schema, *, output_format, instructions, citations, **kw)` / `aextract` | `ExtractResponse` | Same as `puffinparse.extract` except `model`. | `fallback_on` is a list of error-kind names; the default is `provider`, `rate_limit`, `timeout`, `network`. Authentication, bad-request, unsupported-model and input errors never trigger a fallback. Unknown keyword arguments raise `TypeError` and the message lists what is accepted. ```python router = puffinparse.Router(["reducto/standard", "llamaparse/agentic"], strategy="round_robin") resp = router.parse("doc.pdf", pages="1-5", timeout=120) router.stats()["reducto/standard"]["successes"] text_router = puffinparse.Router(["reducto/r-1", "extend/parse_light"], mode="ocr") text_router.ocr("scan.png").text ``` ## Native-format output (`output_format`) PuffinParse normalises every provider to one response shape, which is the right default — and a migration cost if you are already integrated with a vendor. `output_format` removes it: ask for a vendor's shape and the response is rendered into *that vendor's own JSON*, whatever provider actually produced it. ```python doc = puffinparse.parse("invoice.pdf", model="extend/parse_light", output_format="reducto") for chunk in doc["result"]["chunks"]: # Reducto's shape, Extend's engine for block in chunk["blocks"]: draw(block["bbox"]["left"], block["bbox"]["top"], block["type"]) ``` | Value | Shape | |---|---| | `None` (default), `"puffinparse"`, `"unified"` | PuffinParse's own response — a dataclass for `None`, the same JSON as a `dict` for `"puffinparse"`. | | `"reducto"` | Reducto `POST /parse` response (`response_type: "parse"`). | | `"extend"` | Extend `parse_run` object (`GET /parse_runs/{id}`). | | `"llamaparse"` (`"llama"`, `"llama_parse"`) | LlamaParse `…/result/json` payload. | - Available on `parse`, `aparse`, `extract`, `aextract`, `Router.parse` / `aparse` / `extract` / `aextract`, and the CLI (`--output-format`, with `--format json`). - Names are case-insensitive and `-`/`_` are interchangeable; `puffinparse.output_formats()` lists them. An unknown value raises `BadRequestError` **before any network call**. - The return type follows the argument: `None` gives the dataclass, a string gives a `dict`. Both are `@overload`-typed, so a type checker knows which one it is. - **Callbacks always receive the dataclass**, whatever the caller asked for — logging and cost tracking are unaffected. - `output_format` is independent of `output`: `output` picks markdown vs plain text *inside* block content, `output_format` picks the JSON envelope around it. What is guaranteed is **structural fidelity**, not semantic identity: the key set and nesting, one chunk/page per unified page, the content strings, the vendor's own block vocabulary and coordinate units, and the billed page count. Not guaranteed: byte equality with what the vendor would have returned, fields PuffinParse does not model (they are rendered as `null` / `[]`, never invented), or vendor-specific enrichments. Extract-mode rendering is explicitly best effort. [`docs/COMPAT.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/docs/COMPAT.md) lists every always-null field, the lossy block-type mappings and the coordinate conversions, per format. ```python puffinparse.output_formats() # ["puffinparse", "reducto", "extend", "llamaparse"] # same call, both shapes doc = puffinparse.parse("invoice.pdf", model="reducto/standard") # ParseResponse raw = puffinparse.parse("invoice.pdf", model="reducto/standard", output_format="extend") raw["object"], raw["status"], raw["metrics"]["pageCount"] router = puffinparse.Router(["reducto/standard", "extend/parse_light"]) router.parse("doc.pdf", output_format="llamaparse")["pages"][0]["items"] ``` ## Response types All response dataclasses mirror the Rust structs in `puffinparse-core` 1:1 and live in `puffinparse.types`. `Response` is the union `ParseResponse | TextResponse | ExtractResponse`. Every response carries the same envelope: `id`, `provider`, `model`, `provider_job_id`, `usage`, `cost_usd`, `latency_ms`, `created_at`, `metadata`, `raw`, plus `to_dict()` and a `from_dict()` classmethod. `cost_usd` is `pages × per_page_usd` for the mode from the embedded price table (no provider reports dollars directly), and `raw` is `None` unless the call passed `include_raw=True`. ### `ParseResponse` ```python @dataclass class ParseResponse: id: str provider: str model: str pages: list[Page] markdown: str text: str usage: Usage latency_ms: int created_at: str provider_job_id: Optional[str] = None cost_usd: Optional[float] = None metadata: dict[str, Any] = field(default_factory=dict) raw: Any = None ``` | Member | Description | |---|---| | `num_pages` (property) | `len(self.pages)`. | | `blocks` (property) | Every block from every page, in reading order. | | `tables` (property) | Blocks whose `type` is `"table"`. | | `__str__` | Returns `markdown`. | ### `Page` and `Block` ```python @dataclass class Page: page_number: int markdown: str text: str blocks: list[Block] = field(default_factory=list) width: Optional[float] = None height: Optional[float] = None @dataclass class Block: type: BlockType content: str page_number: int text: Optional[str] = None bbox: Optional[BBox] = None confidence: Optional[float] = None ``` `Page.blocks_of(*types)` filters by block type; `Page.tables` is shorthand for `blocks_of("table")`. `width` / `height` are `None` for providers that do not report page dimensions (Reducto). `BlockType` is a `Literal` of `"text"`, `"title"`, `"section_header"`, `"list"`, `"table"`, `"figure"`, `"header"`, `"footer"`, `"footnote"`, `"caption"`, `"formula"`, `"other"`. ### `TextResponse`, `TextPage`, `Line`, `Word` ```python @dataclass class TextResponse: id: str provider: str model: str pages: list[TextPage] text: str usage: Usage latency_ms: int created_at: str provider_job_id: Optional[str] = None cost_usd: Optional[float] = None metadata: dict[str, Any] = field(default_factory=dict) raw: Any = None @dataclass class TextPage: page_number: int text: str lines: list[Line] = field(default_factory=list) words: list[Word] = field(default_factory=list) width: Optional[float] = None height: Optional[float] = None ``` `Line` and `Word` are both `{ text: str, bbox: Optional[BBox], confidence: Optional[float] }`. `TextResponse.lines` and `.words` flatten across pages, `num_pages` is the page count, and `__str__` returns `text`. ### `ExtractResponse`, `FieldInfo`, `Citation` ```python @dataclass class ExtractResponse: id: str provider: str model: str data: Any usage: Usage latency_ms: int created_at: str provider_job_id: Optional[str] = None fields: dict[str, FieldInfo] = field(default_factory=dict) cost_usd: Optional[float] = None metadata: dict[str, Any] = field(default_factory=dict) raw: Any = None ``` `fields` is keyed by **JSON pointer** into `data` (`"/invoice/total"`). `field_info(pointer)` returns the `FieldInfo` or `None`; `citations(pointer)` returns its citation list, or an empty list when the provider reported none. `__str__` returns `str(data)`. ```python @dataclass class FieldInfo: confidence: Optional[float] = None citations: list[Citation] = field(default_factory=list) @dataclass class Citation: page_number: int # 1-based bbox: Optional[BBox] = None text: Optional[str] = None # source text the value was read from ``` ### `BBox` and `Usage` ```python @dataclass(frozen=True) class BBox: x0: float y0: float x1: float y1: float @dataclass class Usage: pages: int = 0 credits: Optional[float] = None provider_cost_usd: Optional[float] = None ``` Boxes are normalised to 0..1 relative to page size, origin top-left. `BBox.width` and `.height` are properties; `to_pixels(page_width, page_height)` returns absolute `(x0, y0, x1, y1)`. ### `Metrics` Returned by `puffinparse.score`. ```python @dataclass class Metrics: char_similarity: float cer: float wer: float word_recall: float word_precision: float word_f1: float pred_chars: int truth_chars: int order_score: Optional[float] = None table_score: Optional[float] = None ``` ## Exceptions Every error raised by PuffinParse derives from `PuffinParseError`. ``` PuffinParseError kind = "error" ├── AuthenticationError "authentication_error" 401/403, or no API key configured ├── RateLimitError "rate_limit_error" 429 after retries were exhausted ├── BadRequestError "bad_request_error" other 4xx — the request itself is wrong ├── ProviderError "provider_error" 5xx, malformed payload, failed job ├── TimeoutError "timeout_error" the whole-call deadline was exceeded ├── UnsupportedModelError "unsupported_model_error" unknown model, or one that cannot serve the mode ├── InputError "input_error" unreadable input, bytes without filename └── NetworkError "network_error" network / TLS / DNS failure after retries ``` Every instance carries: | Attribute | Type | Description | |---|---|---| | `message` | `str` | The provider's message, never swallowed. | | `kind` | `str` (class attribute) | Stable machine-readable kind, as above. | | `provider` | `Optional[str]` | Provider the call was routed to. | | `status_code` | `Optional[int]` | HTTP status when there was one. | | `job_id` | `Optional[str]` | Provider job / run id when there was one. | | `retryable` | `bool` | Whether a retry could plausibly succeed. | | `to_dict()` | `dict` | All of the above as a plain dict. | ```python try: resp = puffinparse.parse("doc.pdf", model="reducto/standard") except puffinparse.RateLimitError as e: print(e.provider, e.status_code, e.retryable) except puffinparse.PuffinParseError as e: print(e.to_dict()) ``` `puffinparse.exceptions.from_core(exc)` converts a raw `puffinparse._core.CoreError` into the typed exception; the SDK applies it for you. ## Callbacks Two module-level lists, called after every call in every mode, including `Router` calls. Callbacks may be plain functions or coroutines; exceptions raised inside a callback are logged to the `puffinparse` logger and never propagate to the caller. ```python puffinparse.success_callback: list[Callable[[Response], None | Awaitable[None]]] puffinparse.failure_callback: list[Callable[[PuffinParseError], None | Awaitable[None]]] ``` ```python puffinparse.success_callback.append(lambda r: print(r.model, r.usage.pages, r.cost_usd)) puffinparse.failure_callback.append(lambda e: print("failed:", e)) ``` ## Models and pricing ```python puffinparse.modes() -> list[str] puffinparse.output_formats() -> list[str] # vendor shapes output_format accepts puffinparse.list_models(mode: Optional[Mode] = None) -> list[str] puffinparse.providers() -> list[dict[str, Any]] # name, env var, base URL, docs, models + modes puffinparse.resolve_model(model: str, mode: Optional[Mode] = None) -> str puffinparse.pricing() -> dict[str, dict[str, Any]] # model -> per-mode $/page, source, updated puffinparse.set_pricing(prices: dict[str, float], mode: Mode = "parse") -> None puffinparse.reset_pricing() -> None puffinparse.estimate_cost(model: str, pages: int, mode: Mode = "parse") -> Optional[float] ``` ```python puffinparse.list_models("extract") # only models that serve extract puffinparse.resolve_model("reducto", "ocr") # provider default *for that mode* puffinparse.set_pricing({"reducto/standard": 0.012}) # your negotiated parse rate puffinparse.estimate_cost("reducto/standard", 1000) # 12.0 puffinparse.reset_pricing() ``` `resolve_model` raises `UnsupportedModelError` for an unknown string, or for a model that does not serve the requested mode — which makes it a cheap validator for user input. `estimate_cost` returns `None` when a model has no published price for that mode. ## Scoring helpers The benchmark metrics are exposed directly — deterministic, offline, no LLM judge. ```python puffinparse.score( prediction: str, truth: str, *, case_insensitive: bool = True, strip_markdown: bool = True, strip_punctuation: bool = False, ) -> Metrics puffinparse.normalize_text(text, *, case_insensitive=True, strip_markdown=True, strip_punctuation=False) -> str puffinparse.markdown_to_text(markdown: str) -> str ``` ```python m = puffinparse.score(resp.markdown, open("truth.md").read()) print(m.char_similarity, m.cer, m.wer, m.word_f1, m.order_score, m.table_score) ``` See [Benchmark](/docs/benchmark/index.md) for what each metric means. ## Logging ```python puffinparse.init_logging(level: str = "info") -> None ``` Enables the Rust core's `tracing` output on stderr; `"debug"` shows every HTTP step. The CLI uses the `PUFFINPARSE_LOG` environment variable for the same thing. Python-side messages (such as a raising callback) go to the standard `logging` logger named `puffinparse`. ## See also - [Getting started](/docs/getting-started/index.md) - [Specification](/docs/project/spec/index.md) — the unified request/response contract these types implement - [Providers](/docs/providers/index.md) — per-provider mapping, supported modes and `provider_options` --- # TypeScript / Node.js SDK The `puffinparse` npm package is the same Rust core as the Python SDK and the CLI, loaded into Node.js as a native addon (N-API, built with [napi-rs](https://napi.rs)). Every call returns a `Promise`; responses are plain camelCase objects typed by the bundled `index.d.ts`. No provider logic lives in JavaScript, so a model behaves identically from Python, Node, Rust and the CLI. Node.js 18+. Source: [`js/`](https://github.com/ajinkyashejul/puffinparse/blob/main/js/README.md) (wrapper and typings) and [`crates/puffinparse-node`](https://github.com/ajinkyashejul/puffinparse/blob/main/crates/puffinparse-node/src/lib.rs) (the addon). ## Install Prebuilt binaries are not published to npm yet, so build the addon from a clone (needs a Rust toolchain): ```bash git clone https://github.com/ajinkyashejul/puffinparse && cd puffinparse/js npm install npm run build # cargo build --release of crates/puffinparse-node -> puffinparse..node npm test # offline unit tests ``` Then depend on it by path (`npm install ../puffinparse/js`) or `npm link`. Keys come from the same environment variables as every other surface (`REDUCTO_API_KEY`, `LLAMA_API_KEY`, `EXTEND_API_KEY`, ...; see [Providers](/docs/providers/index.md)). ## First call ```ts import { parse, ocr, extract } from 'puffinparse' const doc = await parse('invoice.pdf', { model: 'llamaparse/cost_effective' }) console.log(doc.markdown, doc.costUsd, doc.pages[0].blocks[0].bbox) const text = await ocr('scan.png', { model: 'mistral/ocr-latest' }) console.log(text.pages[0].lines.length) const inv = await extract<{ total: number }>('invoice.pdf', { model: 'reducto/extract', schema: { type: 'object', properties: { total: { type: 'number' } } }, citations: true, }) console.log(inv.data.total, inv.fields['/total']?.citations) ``` CommonJS works the same: `const puffinparse = require('puffinparse')`. ## Modes As everywhere in PuffinParse, the mode decides the response type, and a model that does not serve the mode you asked for rejects with `UnsupportedModelError` before any network call. | Mode | Call | Resolves to | |---|---|---| | `parse` | `parse(doc, options?)` | `ParseResponse` — markdown + typed blocks with boxes | | `ocr` | `ocr(doc, options?)` | `TextResponse` — plain text + line/word boxes | | `extract` | `extract(doc, { schema, ... })` | `ExtractResponse` — your schema's object + citations | ## Documents `doc` is one of: | Form | Example | |---|---| | Path string | `'invoice.pdf'` | | URL string | `'https://example.com/invoice.pdf'` (passed to the provider as a remote URL) | | `URL` | `new URL('https://...')` or `new URL('file:///tmp/a.pdf')` | | Bytes | `Buffer`, `Uint8Array` or `ArrayBuffer` — pass `filename` too (it sets the document type) | | Object | `{ url }`, `{ path }` or `{ data, filename }` | Bytes are handed to the core as a `Buffer`, never base64-encoded through JSON. ## Options The options mirror SPEC §4 and the Python keywords, in camelCase: | Option | Default | Meaning | |---|---|---| | `model` | `'reducto'` | `"/"`; a bare provider picks its default model for the mode. | | `fallbacks` | `[]` | More models to try, in order, when `model` fails with a retryable error (a one-off ordered `Router`). | | `filename` | — | Required with bytes. | | `pages` | all | `'1-3,7'`, `2` or `[1, 2, 5]`, 1-based, forwarded best-effort. | | `language` | — | Language hint when the provider supports one. | | `output` | `'markdown'` | `parse` only: preferred block content, `'markdown'` or `'text'`. | | `outputFormat` | unified | `parse` / `extract`: `'reducto'`, `'extend'` or `'llamaparse'` returns that vendor's JSON shape, verbatim ([compatibility](/docs/project/compat/index.md)). `'puffinparse'` is the unified response. | | `providerOptions` | — | Provider-specific options merged verbatim into the provider request. | | `includeRaw` | `false` | Attach the provider payload as `response.raw`. | | `timeout` | `300` | Whole-call deadline in **seconds** (upload + polling + download), as in Python and the CLI. | | `maxRetries` | `2` | Retries on 429 / 5xx / network errors with exponential backoff. | | `apiKey`, `baseUrl` | env | Override the key or the provider base URL. | | `metadata` | `{}` | Echoed back in `response.metadata`. | | `schema`, `instructions`, `citations` | — | `extract` only; `schema` is required and must be a JSON Schema object. | Unknown option names are a `TypeError` that lists the accepted ones, so a snake_case typo such as `provider_options` fails loudly instead of being ignored. ## Responses The interfaces in `index.d.ts` follow [SPEC §5](/docs/project/spec/index.md) field for field, in camelCase. Optional values are `null` (never missing), lists are always present, and three things are returned exactly as received: `data` (your extraction), `metadata` (your keys plus `puffinparse_*` keys such as `puffinparse_fallback_index` and `puffinparse_derived_from`) and `raw`. `fields` keeps its JSON-pointer keys. ```ts interface ParseResponse { id: string; provider: string; model: string; providerJobId: string | null pages: Page[]; markdown: string; text: string usage: { pages: number; credits: number | null; providerCostUsd: number | null } costUsd: number | null; latencyMs: number; createdAt: string metadata: Record; raw: unknown } interface Page { pageNumber: number; width: number | null; height: number | null; markdown: string; text: string; blocks: Block[] } interface Block { type: BlockType; content: string; text: string | null; bbox: BBox | null; confidence: number | null; pageNumber: number } // TextResponse: same envelope + pages: TextPage[] ({ pageNumber, width, height, text, lines, words }) and text // ExtractResponse: same envelope + data: T and fields: Record ``` `BBox` is `{ x0, y0, x1, y1 }`, normalised 0..1 with the origin top-left. ## Router ```ts import { Router } from 'puffinparse' const router = new Router({ models: ['reducto/standard', 'llamaparse/agentic', 'extend/parse_light'], mode: 'parse', // default; every model must serve it strategy: 'ordered', // or 'round_robin' fallbackOn: ['provider', 'rate_limit', 'timeout', 'network'], // the default }) const doc = await router.parse('contract.pdf') router.stats() // { 'reducto/standard': { successes, failures, totalLatencyMs, totalCostUsd, totalPages, avgLatencyMs }, ... } router.plan() // the order the next call would try ``` `fallbackOn` also accepts class names (`'ProviderError'`) or the classes themselves. Calling a method for another mode (`router.ocr(...)` on a parse router) rejects with `InputError`. Per-call options are the module-level ones minus `model` and `fallbacks`. When a fallback served the call, `response.metadata.puffinparse_fallback_index` says which. ## Async jobs and webhooks `parse()` waits for the provider (polling job-queue providers for you). For long documents, batches or webhook-driven pipelines, split the call in two and own the waiting yourself. Jobs are `parse` mode only and need a provider with a job queue: `reducto`, `extend`, `llamaparse`. ```ts import { submit, retrieve, handleWebhook, type Job } from 'puffinparse' const job: Job = await submit('200-pages.pdf', { model: 'reducto/standard', webhookUrl: 'https://example.com/hooks/puffinparse', // optional, see below }) await queue.put(JSON.stringify(job)) // a Job never holds an API key // later, anywhere: const result = await retrieve(JSON.parse(stored)) // the Job again while running, else ParseResponse if ('jobId' in result) console.log('still running', result.jobId) else console.log(result.markdown) ``` `submit(doc, options)` takes `parse()`'s options except `fallbacks` and `outputFormat`, plus `webhookUrl`; `timeout` covers the upload and submission only. It resolves to a `Job`: ```ts interface Job { provider: string; model: string; jobId: string; submittedAt: string output: 'markdown' | 'text'; includeRaw: boolean; baseUrl: string | null providerState: Record | null // non-secret options retrieve needs (Extend workspace_id) metadata: Record } ``` `retrieve(job, { apiKey?, baseUrl?, timeout = 120, maxRetries = 2, outputFormat? })` asks the provider once. It resolves to the **same** `Job` object while the job is pending, or to the `ParseResponse` (normalised exactly like `parse()`, `latencyMs` counted from submission; a vendor shape with `outputFormat`) once it is done. A job the provider reports as failed rejects with the typed `PuffinParseError`, `jobId` set, exactly like a failed `parse()`. The key is read from the environment again unless you pass `apiKey`. `webhookUrl` maps to Reducto `async.webhook` (direct mode) and LlamaParse `webhook_url`; Extend has no per-job webhook (register an endpoint in its dashboard) and rejects it with `InputError`. In your web handler, verify the provider's signature or your own secret first, then: ```ts app.post('/hooks/puffinparse', async (req, res) => { const result = await handleWebhook(req.body, { model: 'reducto' }) // Job | ParseResponse res.sendStatus(204) }) ``` `handleWebhook(payload, { model = 'reducto', apiKey?, baseUrl?, timeout?, maxRetries?, outputFormat? })` accepts the parsed body, a JSON string or bytes. Bodies that carry the whole result (a LlamaParse `webhook_url` push) are normalised directly; bodies that only name a finished job (Reducto, Extend `parse_run.*`, LlamaCloud `parse.*` events) trigger one `retrieve()`; a pending event resolves to its `Job`; a failure event rejects with the typed error. See SPEC §15 for every provider's body shape. ## Errors Every provider or core failure rejects with a subclass of `PuffinParseError`, mapped from the core's `ErrorKind`; argument mistakes are plain `TypeError`s thrown before anything else runs. | Class | `kind` | When | |---|---|---| | `AuthenticationError` | `authentication` | 401/403, or no API key configured | | `RateLimitError` | `rate_limit` | 429 after retries | | `BadRequestError` | `bad_request` | other 4xx, or an unknown `outputFormat` | | `ProviderError` | `provider` | 5xx, malformed payload, failed job | | `TimeoutError` | `timeout` | the whole-call deadline passed | | `UnsupportedModelError` | `unsupported_model` | unknown model, or one that does not serve the mode | | `InputError` | `input` | unreadable file, bytes without `filename`, bad mode, router mode mismatch | | `NetworkError` | `network` | network / TLS / DNS failure after retries | Each carries `message` (the provider's own message, verbatim), `provider`, `statusCode`, `jobId` and `retryable`; `toJSON()` returns all of them. ```ts import { parse, AuthenticationError, PuffinParseError } from 'puffinparse' try { await parse('a.pdf', { model: 'reducto/standard' }) } catch (e) { if (e instanceof AuthenticationError) console.error('set REDUCTO_API_KEY') else if (e instanceof PuffinParseError) console.error(e.kind, e.statusCode, e.message) else throw e } ``` ## Models, pricing and scoring ```ts import * as puffinparse from 'puffinparse' puffinparse.listModels() // every "/" puffinparse.listModels('extract') // only models serving extract puffinparse.resolveModel('reducto') // 'reducto/standard' puffinparse.providers() // [{ name, displayName, envVar, baseUrl, docs, models }] puffinparse.estimateCost('reducto/standard', 100) // USD, or null when unpriced puffinparse.setPricing({ 'reducto/standard': 0.012 }) // per page, mode defaults to 'parse' puffinparse.resetPricing() puffinparse.outputFormats() // ['puffinparse', 'reducto', 'extend', 'llamaparse'] puffinparse.score(prediction, truth) // { charSimilarity, cer, wer, wordF1, ..., tableScore } puffinparse.normalizeText(text) // the normalisation applied before scoring puffinparse.initLogging('debug') // core tracing on stderr ``` ## Differences from the Python SDK - Async only: every mode returns a `Promise` (there is no blocking variant); `submit` / `retrieve` / `handleWebhook` match Python's `asubmit` / `aretrieve` / `ahandle_webhook`, and `retrieve` / `handleWebhook` also take `outputFormat`. - `fallbacks` on a single call is a JavaScript convenience over `Router`. - Success/failure callbacks are not mirrored; wrap the promise instead. - Responses are plain objects, so the Python conveniences (`.tables`, `.num_pages`, `.field_info()`) are one-liners over `pages` / `fields`. ## See also - [Python SDK](/docs/python/index.md) — the same surface in Python. - [Specification](/docs/project/spec/index.md) — the unified request, response and error model. - [`js/README.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/js/README.md) — building, testing and the release plan for prebuilt binaries. --- # CLI The `puffinparse` binary wraps the same Rust core as the SDK. One subcommand per mode — `parse`, `ocr`, `extract` — plus `providers` and `bench`. ```bash cargo build --release -p puffinparse-cli # ./target/release/puffinparse puffinparse --help puffinparse --version ``` Logging is controlled by the `PUFFINPARSE_LOG` environment variable (`error` by default; try `PUFFINPARSE_LOG=debug` to see every HTTP step). All diagnostics go to stderr, so stdout stays clean for piping. **Exit codes:** `0` success · `1` any error · `2` unsupported model or bad input. ## Shared options `parse`, `ocr` and `extract` all take the same document options. The first argument is a local file path or an `http(s)://` URL. | Flag | Default | Description | |---|---|---| | `-m`, `--model ` | `reducto` | `"/"`. It must support the subcommand's mode; a bare provider name picks that provider's default model **for the mode**. | | `-p`, `--pages ` | — | 1-based page selection, e.g. `1-3,7`. | | `-l`, `--language ` | — | Language hint (ISO 639-1), forwarded when supported. | | `--options ` | — | Provider-specific options as a JSON object, merged verbatim into the provider request. | | `--raw` | off | Include the provider's raw payload (JSON output only). | | `--timeout ` | `300` | Whole-call deadline: upload + polling + result download. | | `--max-retries ` | `2` | Retries on transient errors (429 / 5xx / network) with jittered backoff. | | `--api-key ` | `$PUFFINPARSE_API_KEY` | Override the key; otherwise the provider's own env var is used. Never echoed. | | `--base-url ` | — | Override the provider base URL. | After a non-JSON run, a one-line summary goes to **stderr**: ``` [reducto/standard] 3 page(s) in 2841 ms, est. $0.0450 ``` ## `puffinparse parse` `parse` mode: layout-aware markdown and typed blocks. ```bash puffinparse parse [OPTIONS] ``` | Flag | Default | Description | |---|---|---| | `-f`, `--format ` | `markdown` | One of `markdown`, `text`, `json`. | | `--output-format ` | — | Render the JSON in a provider's **own** response shape: `reducto`, `extend`, `llamaparse`, or `puffinparse` for the unified one. | `--format json` prints the whole `ParseResponse` as pretty JSON (and no summary line). `--output-format` is the native-format compatibility layer: whatever provider ran the call, the response is rendered into the named vendor's JSON, so a script that already parses Reducto's or Extend's output keeps working after a model swap. It only applies to `--format json` — with `markdown` or `text` output the flag is ignored and a warning goes to stderr. An unknown name fails before any network call (exit code `2`) and the message lists the valid values. [`docs/COMPAT.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/docs/COMPAT.md) documents exactly what is guaranteed (key set, counts, content, block vocabulary, coordinate units, billed pages) and what is always `null`. ```bash puffinparse parse invoice.pdf -m extend/parse_light puffinparse parse scan.png -m llamaparse/agentic -f json --raw > out.json puffinparse parse doc.pdf -m reducto/r-1 --pages 1-3 -f text puffinparse parse https://example.com/doc.pdf -m reducto/standard \ --options '{"settings": {"return_ocr_data": true}}' puffinparse parse big.pdf -m extend/parse_auto --timeout 900 --max-retries 4 # Extend's engine, Reducto's response shape puffinparse parse invoice.pdf -m extend/parse_light -f json --output-format reducto \ | jq '.result.chunks[0].blocks[0].bbox' ``` ## `puffinparse ocr` `ocr` mode: plain text with line and word boxes. ```bash puffinparse ocr [OPTIONS] ``` | Flag | Default | Description | |---|---|---| | `-f`, `--format ` | `text` | `text` for the plain text, or `json` for the full `TextResponse` (text + lines + words). | ```bash puffinparse ocr scan.png -m reducto/r-1 puffinparse ocr scan.png -m llamaparse/fast -f json | jq '.pages[0].words[:5]' ``` Providers without a native OCR endpoint serve this mode from their parse output; the JSON response then carries `metadata.puffinparse_derived_from == "parse"`. ## `puffinparse extract` `extract` mode: pull a JSON object out of a document with a schema. Output is always pretty JSON. ```bash puffinparse extract --schema [OPTIONS] ``` | Flag | Default | Description | |---|---|---| | `-s`, `--schema ` | *required* | JSON Schema for the object to extract: a path to a `.json` file, or inline JSON (anything starting with `{`). | | `--instructions ` | — | Extra natural-language guidance for the extractor. | | `--citations` | off | Ask for per-field citations (page, box, source text) where the provider supports them. | | `--output-format ` | — | Render the extract envelope in `reducto`, `extend` or `llamaparse` shape instead of the unified one (best effort — see `docs/COMPAT.md` §7). `extract` always prints JSON, so it always applies. | ```bash puffinparse extract invoice.pdf -s invoice.schema.json -m reducto/extract --citations puffinparse extract invoice.pdf -s '{"type":"object","properties":{"total":{"type":"number"}}}' \ --instructions 'Totals are inclusive of tax.' | jq '.data.total' puffinparse extract invoice.pdf -s invoice.schema.json --output-format reducto | jq '.result' ``` A model that does not serve `extract` fails before any network call (exit code `2`); use `puffinparse providers --mode extract` to see the candidates. ## `puffinparse providers` List providers, models, modes, pricing, and whether an API key is configured. | Flag | Default | Description | |---|---|---| | `--mode ` | — | Only show models that serve this mode (`parse`, `ocr`, `extract`). | | `--json` | off | Emit JSON instead of a table. | ```bash puffinparse providers puffinparse providers --mode extract ``` ``` ┌───────────────────┬────────────┬─────────┬─────┬───────────────┬──────────────────────────┐ │ Model │ Modes │ Default │ Key │ $/page │ Description │ ... Modes: parse (markdown + blocks), ocr (plain text + boxes), extract (JSON schema). Keys are read from: REDUCTO_API_KEY, EXTEND_API_KEY, LLAMA_API_KEY, ... Native output formats (--output-format, json only): puffinparse | reducto | extend | llamaparse ``` `Default` marks each provider's default model, `Key` shows `✓` / `✗` for a non-empty environment variable, and `$/page` lists the price of every mode the model serves (one number when `--mode` is given). `--json` emits an object with two keys: - `providers` — one entry per provider with `name`, `display_name`, `env_var`, `key_configured`, `base_url`, `docs` and a `models` array of `{model, default, description, modes, per_page_usd}`, where `per_page_usd` is keyed by mode; - `output_formats` — the vendor shapes this build can render (`--output-format`, and `output_format=` in the SDK). Handy for scripting, or for an agent picking a model. ```bash puffinparse providers --json | jq -r '.providers[].models[] | select(.per_page_usd.parse < 0.005) | .model' puffinparse providers --mode ocr --json | jq -r '.providers[].models[].model' puffinparse providers --json | jq -r '.output_formats[]' ``` ## `puffinparse bench` Run and report the open OCR benchmark (it uses `parse` mode). Three subcommands: `run`, `report`, `score`. ### `puffinparse bench run` Run models over a dataset and write a result JSON. | Flag | Default | Description | |---|---|---| | `-d`, `--dataset ` | *required* | Dataset directory containing `manifest.json`. | | `-m`, `--models ...` | *required* | Models to evaluate. Repeatable / space-separated. | | `-o`, `--out ` | `benchmark/results/-.json` | Output JSON path. | | `-c`, `--concurrency ` | `4` | Concurrent requests per model. | | `--filter ` | — | Only run documents whose id contains this substring. | | `--limit ` | — | Limit the number of documents. | | `--timeout ` | `300` | Per-call timeout. | | `--save-outputs ` | — | Save each model's raw markdown per document, at `//.md`. | | `--case-sensitive` | off | Score case-sensitively (normalisation lowercases by default). | | `--allow-cache` | off | Allow provider-side result caches. Off by default so latency reflects real work. | Every model string is validated before any network call. Progress is drawn on stderr per model, followed by a summary line and any per-document failures; the result JSON is written to `--out` and the leaderboard table is printed to stdout. The result file records the `run_id`, `created_at`, the PuffinParse version, the dataset name, version, document count and **SHA-256 of the manifest plus every input and truth file**, the normalisation options, and for each model every document's metrics, latency, cost and error plus a summary (accuracy, `latency_p50_ms`, `latency_p95_ms`, `latency_per_page_ms`, `total_pages`, `total_cost_usd`, `cost_per_1k_pages_usd`, `by_category`). ```bash puffinparse bench run \ --dataset benchmark/datasets/synthetic-v1 \ --models reducto/standard reducto/r-1 extend/parse_performance extend/parse_light \ llamaparse/fast llamaparse/cost_effective llamaparse/agentic \ --concurrency 4 --save-outputs benchmark/runs/outputs puffinparse bench run -d benchmark/datasets/synthetic-v1 -m reducto/r-1 --filter table --limit 5 ``` ### `puffinparse bench report` Render one or more result JSON files as a leaderboard. | Argument / flag | Default | Description | |---|---|---| | `...` | *required* | Result JSON files. Globs are expanded by your shell. | | `-f`, `--format ` | `markdown` | `markdown` (alias `md`) or `json`. | Rows from every file are pooled and sorted by **Overall** descending. A per-category breakdown is appended when the results span more than one category. An unknown format exits with an error. ```bash puffinparse bench report benchmark/results/*.json > benchmark/LEADERBOARD.md puffinparse bench report benchmark/results/*.json -f json | jq '.[].summary.overall' ``` ### `puffinparse bench score` Score a single prediction file against a truth file. No network access, fully deterministic. ```bash puffinparse bench score ``` Prints the `Metrics` object as pretty JSON: `char_similarity`, `cer`, `wer`, `word_recall`, `word_precision`, `word_f1`, `order_score`, `table_score`, `pred_chars`, `truth_chars`. ```bash puffinparse parse doc.pdf -m reducto/r-1 > pred.md puffinparse bench score pred.md truth.md ``` The same function is available from Python as `puffinparse.score(prediction, truth)`. ## See also - [Getting started](/docs/getting-started/index.md) - [Benchmark](/docs/benchmark/index.md) — methodology, metrics and caveats - [Leaderboard](/docs/benchmark/leaderboard/index.md) — the current results - [Python SDK](/docs/python/index.md) --- # Gateway server A small HTTP server in front of every provider PuffinParse supports, in the spirit of the LiteLLM proxy: clients call one endpoint with one API shape and a gateway-issued key; the gateway holds the provider keys, routes aliases to provider models with fallbacks, enforces per-key model lists, monthly budgets and rate limits, and emits JSON-lines request logs and Prometheus metrics. It is a thin layer over `puffinparse-core` (crate `crates/puffinparse-server`, axum + tower-http). It adds no provider logic: every call is `puffinparse_core::parse / ocr / extract` with the request fields below, so responses are exactly the unified types of [SPEC §5](/docs/project/spec/index.md#5-unified-responses), or a vendor shape via `output_format` ([COMPAT.md](/docs/project/compat/index.md)). Long documents can go through the async jobs API instead (`POST /v1/jobs` + `GET /v1/jobs/{id}`, core `submit_parse` / `retrieve_parse`, [SPEC §15](/docs/project/spec/index.md#15-asynchronous-jobs-and-webhooks)), so no connection is held open while the provider works. ## Run it ```bash export PUFFINPARSE_MASTER_KEY=sk-master-change-me export PUFFINPARSE_KEY_BILLING=sk-billing-change-me PUFFINPARSE_KEY_RESEARCH=sk-research-change-me export REDUCTO_API_KEY=... EXTEND_API_KEY=... LLAMA_API_KEY=... puffinparse serve --config examples/server/puffinparse.toml # --host / --port override the file ``` Without `--config` it reads `./puffinparse.toml` if present (or `$PUFFINPARSE_CONFIG`), else runs with defaults: `127.0.0.1:4000`, no aliases, **no auth** (a warning is printed). Anything reachable beyond localhost should have a `master_key` or `[[keys]]`. Docker (multi-stage build, distroless runtime, runs as non-root). Releases publish a linux/amd64 image to `ghcr.io/ajinkyashejul/puffinparse` (tags `latest`, `0.1.0`, `0.1`); or build it yourself with `docker build -t puffinparse .`: ```bash docker run --rm -p 4000:4000 -v $PWD/examples/server/puffinparse.toml:/etc/puffinparse/puffinparse.toml:ro \ -e PUFFINPARSE_MASTER_KEY -e PUFFINPARSE_KEY_BILLING -e PUFFINPARSE_KEY_RESEARCH \ -e REDUCTO_API_KEY -e EXTEND_API_KEY -e LLAMA_API_KEY ghcr.io/ajinkyashejul/puffinparse ``` The image's default command is `serve --host 0.0.0.0 --config /etc/puffinparse/puffinparse.toml`; it exits if no config is mounted rather than serving an open gateway. ## Configuration (`puffinparse.toml`) Every secret may be written as `"env:VAR"`, resolved once at startup. A virtual key or master key whose reference resolves to nothing is a startup error (fail closed); a provider key that resolves to nothing prints a warning and the core falls back to the provider's default env var. ```toml master_key = "env:PUFFINPARSE_MASTER_KEY" # full access: all models, no budget, no rate limit [server] host = "127.0.0.1" port = 4000 log_stdout = true # JSON-lines request log on stdout log_file = "requests.jsonl" # optional append-only copy state_file = "usage.json" # optional: per-key monthly spend survives restarts max_body_mb = 50 # multipart upload / base64 JSON body cap max_timeout_secs = 300 # cap (and default) for a request's `timeout` max_retries = 2 # per provider call, when the client sends none allow_direct_models = true # false: only the aliases below may be requested job_retention_hours = 168 # how long POST /v1/jobs handles stay readable [webhooks] # optional provider webhook receiver (off by default) enabled = false secret = "env:PUFFINPARSE_WEBHOOK_SECRET" # required when enabled; sent as ?token= or a header [providers.reducto] # one table per provider (name as in `puffinparse providers`) api_key = "env:REDUCTO_API_KEY" # base_url = "https://eu.platform.reducto.ai" [[models]] # an alias clients can send as "model" name = "invoices" targets = ["reducto/standard", { model = "extend/parse_performance", api_key = "env:EXTEND_KEY_2" }] strategy = "ordered" # or "round_robin" fallback_on = ["provider", "rate_limit", "timeout", "network"] # the default [[keys]] id = "billing-team" # appears in logs, metrics, usage; never the secret key = "env:PUFFINPARSE_KEY_BILLING" # the bearer token the client sends models = ["invoices", "llamaparse/*"] # aliases, provider/model, provider/*, or "*"; empty = all monthly_budget_usd = 50.0 # calendar month, UTC rpm = 60 # requests per minute, sliding 60 s window ``` A complete sample is [`examples/server/puffinparse.toml`](https://github.com/ajinkyashejul/puffinparse/blob/main/examples/server/puffinparse.toml). **Routing.** `model` is looked up as an alias first, then (if `allow_direct_models`) as a registry model (`reducto/standard`, or a bare provider for its default model in that mode). An alias expands to its targets: `ordered` always starts at the first, `round_robin` rotates the start per request, and both fall back in order when the error kind is in `fallback_on`. A request's own `fallbacks` list is appended after that (aliases expand in order). Every target is checked against the endpoint's mode before any provider is called. Target-level `api_key` / `base_url` override the `[providers.*]` ones, so two deployments of the same provider (two accounts, two regions) can sit behind one alias. When a fallback served the call, the response metadata carries `puffinparse_fallback_index` and `puffinparse_fallback_from_error`, as with the SDK `Router`. **Budgets.** Spend is the response's `cost_usd` (provider-reported cost when available, otherwise the list-price estimate from `pricing.json`), summed per key per UTC calendar month. The check runs before the call: a key whose spend has reached its budget gets `402`. One in-flight request can take a key past its budget; nothing is charged for failed attempts (a provider may still bill a failed job). An async job is charged to the key that submitted it, once, when it is first observed succeeded (by `GET /v1/jobs/{id}` or a webhook); submitting and polling are free. State is in memory; with `state_file` it (spend, counts and submitted jobs) is rewritten (temp file + rename) after every change and reloaded on startup. Rate-limit windows are not persisted. ## API Authenticate with `Authorization: Bearer ` or `x-api-key: `. Send `x-request-id` to choose the request id (echoed in the `x-request-id` response header and the log), otherwise a UUID is generated. ### `POST /v1/parse`, `POST /v1/ocr`, `POST /v1/extract` Body is JSON or `multipart/form-data` with the same field names (multipart fields are text; the JSON-valued ones — `provider_options`, `schema`, `metadata`, `fallbacks` — are JSON strings, and `fallbacks` may also be comma-separated). | Field | Type | Notes | |---|---|---| | `model` | string | **Required.** Alias or `provider/model`. | | `document_url` | string | Public http(s) URL, passed to the provider. | | `document` | string | Base64 bytes (a `data:…;base64,` prefix is accepted). Needs `filename`. | | `file` | multipart file part | The upload; its filename sets the type (or send `filename`). | | `filename` | string | Sets the MIME type for `document` / overrides the part's filename. | | `pages`, `language` | string | As in SPEC §4.1. | | `output` | `"markdown"` \| `"text"` | Block content format (parse). | | `output_format` | `"reducto"` \| `"extend"` \| `"llamaparse"` \| `"puffinparse"` | Vendor-native response shape (parse, extract). | | `provider_options` | object | Merged into the provider request verbatim. | | `include_raw` | bool | Attach the provider payload as `raw`. | | `timeout` | number (s) | Capped at `server.max_timeout_secs`. | | `max_retries` | int | Per provider call (max 10). | | `fallbacks` | string[] | Extra aliases/models tried after `model`'s own targets. | | `metadata` | object | Echoed back in the response. | | `schema`, `instructions`, `citations` | | `/v1/extract` only; `schema` required there. | Exactly one of `document_url`, `document`, `file`. `api_key` and `base_url` are **rejected**: a client must never be able to point the gateway's provider credentials at another host. Local file paths are not accepted either. The response is the unified `ParseResponse` / `TextResponse` / `ExtractResponse` JSON (or the vendor shape), with headers `x-puffinparse-model` (served model), `x-puffinparse-cost-usd`, `x-request-id`. ### `POST /v1/jobs`, `GET /v1/jobs/{id}` (async parse) `POST /v1/jobs` takes the `/v1/parse` body (JSON or multipart, same fields) plus an optional `webhook_url`, uploads the document, starts the provider job and answers **202** at once: ```json {"id": "job_4f0c…", "object": "job", "status": "pending", "job": {"provider": "reducto", "model": "reducto/standard", "job_id": "c1e2…", "submitted_at": "2026-09-24T20:23:52.140Z", "output": "markdown", "include_raw": false}} ``` `job` is the core `JobHandle` (SPEC §15) minus the operator's `base_url`. Authentication, the key's model allow-list, aliases, the budget pre-check and `rpm` apply exactly as for `/v1/parse`. Only providers with a job queue can take jobs (`reducto`, `extend`, `llamaparse`; others give `400 unsupported_model_error`). **Fallback does not apply to jobs**: an alias submits to its first target (in `round_robin`, the next one in rotation) and a request `fallbacks` list is rejected, because a provider failure only shows up later, on retrieve. `webhook_url` is forwarded to the provider's per-job webhook (Reducto `async.webhook`, LlamaParse `webhook_url`; Extend rejects it) and is refused by the synchronous endpoints. `GET /v1/jobs/{id}` asks the provider once and returns ```json {"id": "job_4f0c…", "object": "job", "status": "pending" | "succeeded" | "failed", "model": "reducto/standard", "provider": "reducto", "provider_job_id": "c1e2…", "submitted_at": "…", "result": {…ParseResponse…}, "error": {…error object…}} ``` with `result` only when `succeeded` (the unified `ParseResponse`, or the vendor shape for the `output_format` sent at submit time; `?output_format=reducto` overrides it per call) and `error` only when `failed` (the same object as an error body's `error`, with the provider's message and `job_id`). A failed *job* is still HTTP 200; a failed *status check* (provider credentials, network) is an error response as usual. Job ids are opaque and bound to the key that created the job: any other key gets `404 not_found`, exactly as for an unknown id (the master key can read all jobs). Polls count toward the key's `rpm` but are not budget-gated. The job's cost is charged to its owner once, the first time it is seen succeeded. Credentials are resolved from the config on every poll (the alias target's own key, else `[providers.*]`), never stored; handles are kept for `job_retention_hours` (and in `state_file` when set). ### `POST /v1/webhooks/{provider}` (optional) Off by default (404). With `[webhooks] enabled = true` and a `secret`, the gateway accepts the body a provider POSTs when a job changes state — Reducto direct webhooks, Extend `parse_run.*` events, LlamaCloud `parse.*` events — authenticated by `?token=` or the `x-puffinparse-webhook-secret` header (compared in constant time). The body is read with core `parse_webhook`; when it only says the job finished, the gateway makes one status check. The provider's job id must match a job submitted through this gateway (else 404); a success is charged to the job's owner (once, shared with `GET`), and the answer is an acknowledgement `{"ids": [...], "status"}` — clients still collect the result with `GET /v1/jobs/{id}`. `ids` can hold several gateway jobs: LlamaParse returns the same (cached) job id for an identical upload, so two submissions of the same file share one provider job, and each is settled and charged to its own key. Point a provider's webhook at `https:///v1/webhooks/reducto?token=` (per job via `webhook_url`, or a workspace-level endpoint for Extend / LlamaCloud). This is a shared secret, not the vendors' HMAC signatures: keep the URL private and serve the gateway over TLS. A LlamaParse `webhook_url` result push that names no job id cannot be attributed and is rejected with 400. ### Errors Every error has the same body: ```json {"error": {"type": "provider_error", "message": "…the provider's own message…", "provider": "reducto", "provider_status": 503, "job_id": null, "request_id": "5c0f…"}} ``` | HTTP | `type` | Cause | |---|---|---| | 400 | `input_error`, `bad_request_error`, `unsupported_model_error` | Malformed request, provider 4xx, unknown model or model without this mode | | 401 | `unauthorized` | Missing or unknown gateway key | | 402 | `budget_exceeded` | Key's monthly budget spent | | 403 | `model_not_allowed` | Model (or a fallback) not in the key's `models` | | 404 | `not_found` | Unknown job id, a job another key owns, or webhooks disabled | | 413 | `payload_too_large` | Body over `max_body_mb` | | 429 | `key_rate_limited` | Key's `rpm` reached (`Retry-After` set) | | 429 | `rate_limit_error` | Provider rate limit after retries and fallbacks | | 502 | `provider_error`, `network_error`, `authentication_error` | Provider 5xx / failed job, network failure, provider rejected the gateway's credentials | | 504 | `timeout_error` | `timeout` exceeded | Provider credential failures are 502, not 401: the caller's key was fine, the operator's was not. ### `GET /v1/models` Aliases (targets, strategy, the modes every target supports) and, with `allow_direct_models`, every registry model with its modes, per-page list price per mode and whether a provider key is configured. Filtered to what the caller's key may use. ### `GET /v1/usage` This month's spend, requests, pages, remaining budget and limits: the caller's own key, or all keys for the master key. ### `GET /health`, `GET /metrics` `/health` → `{"status": "ok", "version": …}`. `/metrics` is Prometheus text, unauthenticated (it carries model names, counts and costs, no secrets or content; firewall it if that matters): | Metric | Labels | |---|---| | `puffinparse_requests_total` | `mode`, `model` (served model, or requested alias on failure, `-` if it never resolved), `status` | | `puffinparse_errors_total` | `type` (the error `type` above) | | `puffinparse_request_duration_seconds` (histogram, 0.25 s – 300 s buckets) | `mode` | | `puffinparse_pages_total` | `model` | | `puffinparse_cost_usd_total` | `model` | | `puffinparse_fallbacks_total` | — | | `puffinparse_jobs_total` | `event`: `submitted`, and `succeeded` / `failed` the first time a job is seen terminal | `mode` is `parse` / `ocr` / `extract` for the synchronous endpoints and `job_submit`, `job_retrieve`, `webhook` for the jobs API, so job latencies do not mix with blocking calls. A job's pages and cost are counted once, when it is first seen succeeded. ## Request log One JSON object per request, on stdout and/or `log_file`: ```json {"ts":"2026-09-24T20:23:52.140Z","request_id":"bf397168-…","key_id":"demo","method":"POST", "path":"/v1/parse","mode":"parse","model":"cheap","served_model":"llamaparse/cost_effective", "provider":"llamaparse","fallback_index":0,"pages":1,"cost_usd":0.00375,"latency_ms":10384, "status":200,"error_type":null,"provider_status":null} ``` Jobs API lines add `job_id` (the gateway id) and `job_status` (`pending` / `succeeded` / `failed`, as observed by that request); `method` is `GET` for status checks, and `key_id` is `webhook` for provider webhooks. `pages` and `cost_usd` are set only on the request that charged the job. The record has no field for document bytes, URLs, extracted content, provider error text (some providers echo document text in errors), provider keys or gateway key secrets. ## curl examples ```bash GW=http://127.0.0.1:4000; KEY=sk-billing-change-me # Upload a file (multipart) to an alias curl -s $GW/v1/parse -H "Authorization: Bearer $KEY" -F model=invoices -F file=@invoice.pdf | jq .markdown # URL input, vendor-native shape, a request-level fallback curl -s $GW/v1/parse -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{ "model": "invoices", "document_url": "https://example.com/invoice.pdf", "output_format": "reducto", "fallbacks": ["llamaparse/cost_effective"]}' # Base64 input, OCR mode curl -s $GW/v1/ocr -H "x-api-key: $KEY" -H 'content-type: application/json' \ -d "{\"model\": \"reducto\", \"filename\": \"scan.png\", \"document\": \"$(base64 -w0 scan.png)\"}" | jq .text # Extract with a schema curl -s $GW/v1/extract -H "Authorization: Bearer $KEY" -F model=invoice-fields -F file=@invoice.pdf \ -F 'schema={"type":"object","properties":{"total":{"type":"number"}}}' # Async job: submit, then poll until status is succeeded or failed JOB=$(curl -s $GW/v1/jobs -H "Authorization: Bearer $KEY" -F model=invoices -F file=@big.pdf | jq -r .id) curl -s $GW/v1/jobs/$JOB -H "Authorization: Bearer $KEY" | jq .status curl -s "$GW/v1/jobs/$JOB?output_format=reducto" -H "Authorization: Bearer $KEY" | jq .result curl -s $GW/v1/models -H "Authorization: Bearer $KEY" | jq '.data[].id' curl -s $GW/v1/usage -H "Authorization: Bearer $KEY" curl -s $GW/metrics ``` ## Not in scope (yet) Streaming, jobs for `ocr` / `extract` (jobs are parse-only, like the core), fallback for jobs, the gateway registering its own webhook URL with providers automatically, vendor HMAC signature checks on webhooks, key management over HTTP (keys live in the config file; restart to change them), a database, response caching, per-key budgets by model, TLS termination (put it behind a reverse proxy), and metrics auth. --- # Rust `puffinparse-core` is the whole implementation: unified types, every provider, the router, pricing and the benchmark metrics. The Python SDK and the CLI are thin wrappers over it. `#![forbid(unsafe_code)]`, no vendor SDK crates — every provider is spoken to over plain HTTPS with `reqwest` and `tokio`. API reference: [docs.rs/puffinparse-core](https://docs.rs/puffinparse-core) *(published on the first crates.io release)*. Until then, `cargo doc -p puffinparse-core --open` from a clone. ## Install ```toml [dependencies] puffinparse-core = "0.1" tokio = { version = "1", features = ["rt-multi-thread", "macros"] } ``` From the repository while it is pre-release: ```toml puffinparse-core = { git = "https://github.com/ajinkyashejul/puffinparse" } ``` ## First call ```rust use puffinparse_core::{parse, DocumentRequest}; #[tokio::main] async fn main() -> puffinparse_core::Result<()> { let resp = parse(DocumentRequest::from_path("invoice.pdf").model("reducto/standard")).await?; println!("{} pages, ${:.4}", resp.usage.pages, resp.cost_usd.unwrap_or(0.0)); println!("{}", resp.markdown); Ok(()) } ``` Each entry point resolves the model, builds the provider, times the call, fills in `provider`, `model`, `latency_ms` and `cost_usd`, merges request metadata into the response, and drops `raw` unless `include_raw` was set. ## Modes A call has a **mode**, and each mode has its own request and response type. Providers can only be swapped within a mode — a model that does not support the requested mode is rejected before any network call. | Mode | Entry point | Request | Response | What you get | |---|---|---|---|---| | `Mode::Parse` | `parse` | `DocumentRequest` | `ParseResponse` | Layout-aware markdown and typed blocks with boxes. | | `Mode::Ocr` | `ocr` | `DocumentRequest` | `TextResponse` | Plain text with line and word boxes, no layout semantics. | | `Mode::Extract` | `extract` | `ExtractRequest` | `ExtractResponse` | A JSON object shaped by your schema, with per-field citations. | ```rust pub async fn parse(request: DocumentRequest) -> Result; pub async fn ocr(request: DocumentRequest) -> Result; pub async fn extract(request: ExtractRequest) -> Result; /// Blocking wrapper around `parse`; builds a small current-thread runtime per call. pub fn parse_blocking(request: DocumentRequest) -> Result; ``` `Mode` is `Parse | Ocr | Extract`, with `Mode::ALL`, `as_str()` and a `Display` impl. ## `DocumentRequest` builder Every setter takes `self` and returns `Self`, so requests chain. ```rust use puffinparse_core::{DocumentRequest, OutputFormat}; let req = DocumentRequest::from_path("doc.pdf") .model("extend/parse_performance") .pages("1-3,7") .language("en") .output(OutputFormat::Markdown) .provider_options(serde_json::json!({ "blockOptions": { "figures": { "enabled": false } } })) .include_raw(true) .timeout_secs(120.0) .max_retries(3) .api_key(std::env::var("EXTEND_API_KEY").unwrap()) .base_url("https://api.extend.ai"); ``` ### Constructors | Constructor | Input | |---|---| | `DocumentRequest::from_path(path)` | Local file. | | `DocumentRequest::from_bytes(data, filename)` | In-memory bytes; the filename infers the type. | | `DocumentRequest::from_url(url)` | Remote document, passed to the provider where supported. | | `DocumentRequest::from_str_input(s)` | URL if `s` starts with `http://` / `https://`, else a path. | | `DocumentRequest::new(DocumentInput)` | The general form. | `DocumentInput` is an enum of `Path { path }`, `Bytes { data, filename }` and `Url { url }`, with `filename()`, `mime_type()` and `describe()` helpers. ### Fields and defaults | Field | Type | Default | |---|---|---| | `input` | `DocumentInput` | — | | `model` | `String` | `"reducto"` | | `pages` | `Option` | `None` | | `language` | `Option` | `None` | | `output` | `OutputFormat` | `Markdown` | | `provider_options` | `Option` | `None` | | `include_raw` | `bool` | `false` | | `timeout_secs` | `f64` | `300.0` | | `max_retries` | `u32` | `2` | | `api_key` | `Option` | `None` (falls back to the provider's env var) | | `base_url` | `Option` | `None` (falls back to `_BASE_URL`, then the built-in) | | `metadata` | `BTreeMap` | empty | `req.option("key")` reads a single key back out of `provider_options`. ## `ParseResponse` ```rust pub struct ParseResponse { pub id: String, pub provider: String, pub model: String, pub pages: Vec, pub markdown: String, pub text: String, pub usage: Usage, pub latency_ms: u64, pub created_at: String, pub provider_job_id: Option, pub cost_usd: Option, pub metadata: BTreeMap, pub raw: Option, } ``` `Page { page_number, markdown, text, blocks, width, height }`, `Block { type, content, page_number, text, bbox, confidence }`, `BBox { x0, y0, x1, y1 }` normalised 0..1 with a top-left origin, and `Usage { pages, credits, provider_cost_usd }`. Everything derives `Serialize` / `Deserialize`, so a response round-trips through JSON unchanged — that is exactly what `puffinparse parse -f json` prints. Helpers in `types`: `ParseResponse::from_pages`, `page_count`, `join_pages`, `pages_from_blocks`, `strip_html_tags`, `markdown_to_text`. ## `TextResponse` (ocr mode) Same envelope — `id`, `provider`, `model`, `provider_job_id`, `usage`, `cost_usd`, `latency_ms`, `created_at`, `metadata`, `raw` — with `text` for the whole document and `pages: Vec`. ```rust pub struct TextPage { pub page_number: u32, pub text: String, pub lines: Vec, pub words: Vec, pub width: Option, pub height: Option, } ``` `Line` and `Word` are both `{ text, bbox: Option, confidence: Option }`. ## `ExtractRequest` / `ExtractResponse` (extract mode) ```rust use puffinparse_core::{extract, DocumentRequest, ExtractRequest}; let req = ExtractRequest::new( DocumentRequest::from_path("invoice.pdf").model("reducto/standard"), serde_json::json!({ "type": "object", "properties": { "total": { "type": "number" }, "vendor": { "type": "string" } }, }), ) .instructions("Totals are inclusive of tax.") .citations(true); let resp = extract(req).await?; println!("{}", resp.data); // the object your schema describes for (pointer, info) in &resp.fields { // "/total" -> confidence + citations println!("{pointer}: {:?} {:?}", info.confidence, info.citations); } ``` `ExtractRequest` flattens a `DocumentRequest` (so every builder setter above still applies) and adds `schema` (a JSON Schema 2020-12 subset), optional `instructions`, and `citations: bool`. `ExtractResponse` carries `data`, plus `fields: BTreeMap` keyed by JSON pointer, where `FieldInfo { confidence: Option, citations: Vec }` and `Citation { page_number, bbox: Option, text: Option }`. ## Router ```rust use puffinparse_core::{DocumentRequest, Mode, Router, RouterConfig, Strategy}; let router = Router::new( RouterConfig::new(vec!["reducto/standard".into(), "llamaparse/agentic".into()]) .mode(Mode::Parse), // every model must support this mode )?; let resp = router.parse(&DocumentRequest::from_path("doc.pdf")).await?; for (model, s) in router.stats() { println!("{model}: {} ok, {} failed, avg {:?} ms", s.successes, s.failures, s.avg_latency_ms()); } ``` `RouterConfig::new(models)` fills in `Mode::Parse`, `Strategy::Ordered` and the default `fallback_on`: `Provider`, `RateLimit`, `Timeout`, `Network`. Set `strategy` to `Strategy::RoundRobin` to rotate the starting model per call. The request's own `model` is ignored — the router chooses. `router.models()` returns the canonicalised list, `router.mode()` the mode it is pinned to, and `router.plan()` the order the next call will try (advancing the round-robin cursor). `Strategy` parses from a string (`"ordered"` / `"fallback"`, `"round_robin"` / `"roundrobin"`). The router has one method per mode: `parse`, `ocr` and `extract`. Calling one whose mode does not match the router's configured mode is an error. `ModelStats` carries `successes`, `failures`, `total_latency_ms`, `total_cost_usd`, `total_pages` and `avg_latency_ms()`. ## Errors ```rust pub enum ErrorKind { Authentication, RateLimit, BadRequest, Provider, Timeout, UnsupportedModel, Input, Network, } ``` `Error` carries the `kind`, the provider message, and optional `provider`, `status_code` and `job_id`. `type Result = std::result::Result`. The Python exception hierarchy is a 1:1 mirror of these kinds. ```rust match parse(req).await { Ok(resp) => println!("{}", resp.markdown), Err(e) if e.kind == puffinparse_core::ErrorKind::RateLimit => { /* back off */ } Err(e) => eprintln!("{e}"), } ``` ## Models and pricing ```rust use puffinparse_core::{list_models, list_models_for, model_info, Mode, ModelRef, PROVIDERS}; for p in PROVIDERS { // name, display_name, env_var, base_url, docs, models for m in p.models { println!("{} {:?} {}", m.qualified(), m.modes, if m.default { "*" } else { "" }); } } list_models(); // every "/" list_models_for(Mode::Extract); // only the models that serve this mode model_info("reducto", "r-1"); // -> Option<&ModelInfo> let ModelRef { provider, model } = ModelRef::parse("reducto")?; // -> reducto / standard let checked = ModelRef::parse_for("reducto", Mode::Ocr)?; // also checks the mode let usd = puffinparse_core::pricing::estimate_cost("reducto/standard", Mode::Parse, 12); let table = puffinparse_core::pricing::all_prices(); ``` `ModelInfo` is `{ provider, model, description, default, modes }`; `default` marks the provider's default model *within each mode it supports*. `ModelRef::parse` / `parse_for` are the validation gate: unknown providers, unknown models, or a model that does not serve the requested mode produce `ErrorKind::UnsupportedModel` before any network call. Aliases `llama`, `llama_parse` and `llamacloud` resolve to `llamaparse`. ## Benchmark module `puffinparse_core::bench` is the deterministic scoring used by `puffinparse bench` and `puffinparse.score` — no LLM judge, no network. ```rust use puffinparse_core::bench::{normalize, score, summarize, Metrics, NormalizeOptions, Summary}; let opts = NormalizeOptions { case_insensitive: true, strip_markdown: true, strip_punctuation: false, }; let m: Metrics = score(&prediction, &truth, opts); println!("{:.3} char sim, {:.3} CER, {:.3} WER", m.char_similarity, m.cer, m.wer); let s: Summary = summarize(&[Some(m), None]); // None counts as a failed document println!("overall {:.2} over {} docs ({} failed)", s.overall, s.documents, s.failed); ``` | Item | Description | |---|---| | `normalize(s, opts)` | NFKC, markdown stripped, quotes/dashes straightened, whitespace collapsed. | | `score(pred, truth, opts)` | `Metrics { char_similarity, cer, wer, word_recall, word_precision, word_f1, order_score, table_score, pred_chars, truth_chars }`. | | `summarize(&[Option])` | `Summary { documents, failed, char_similarity, cer, wer, word_f1, order_score, table_score, overall }`. | | `levenshtein(a, b)` | The generic edit distance used underneath. | See [Benchmark](/docs/benchmark/index.md) for the definition of each metric. ## Module map | Module | Contents | |---|---| | `types` | `Mode`, `DocumentRequest`, `ParseResponse`, `TextResponse`, `TextPage`, `Line`, `Word`, `ExtractRequest`, `ExtractResponse`, `Citation`, `FieldInfo`, `Page`, `Block`, `BBox`, `Usage`, `DocumentInput`, `OutputFormat`. | | `error` | `Error`, `ErrorKind`, `Result`. | | `model` | `ModelRef`, `ModelInfo`, `ProviderInfo`, `PROVIDERS`, `list_models`, `list_models_for`, `model_info`. | | `pricing` | Embedded, overridable per-page price table. | | `http` | Shared client: retry/backoff with jitter, deadline, polling helper. | | `provider` | `Provider` trait plus key / base-URL / multipart helpers. | | `providers` | `reducto`, `extend`, `llamaparse`, and `build()`. | | `router` | `Router` (one method per mode), `RouterConfig`, `Strategy`, `ModelStats`. | | `bench` | Normalisation, metrics, summaries. | | `util` | `deep_merge`, page-range parsing. | ## Adding a provider One file under `crates/puffinparse-core/src/providers/`, implementing `Provider`, registered in `providers/mod.rs` and `model::PROVIDERS`, plus `pricing.json` entries, a fixture-backed normalisation test with a real redacted payload, and a page under [Providers](/docs/providers/index.md). The full checklist is in [Contributing](/docs/project/contributing/index.md). ## See also - [Specification](/docs/project/spec/index.md) — the contract all three surfaces implement - [CLI](/docs/cli/index.md) · [Python SDK](/docs/python/index.md) --- # Providers One page per provider, describing exactly what PuffinParse sends, what comes back, and how the two are mapped onto the unified `OcrResponse`, checked against the implementation in `crates/puffinparse-core/src/providers/`. Verification differs by provider, and the **Status** column says which applies: * **live-verified** (Reducto, Extend, LlamaParse): the `#[ignore]`d live tests pass against the real API with a key, and the fixtures include redacted live responses. * **verified locally** (Tesseract, Docling): self-hosted engines; the live tests pass against a local install and the fixtures are real output from it. * **docs-only** (everything else): implemented from the provider's API documentation and tested against fixture payloads built from it, but not yet run against the live API. Expect wire-format differences until they are verified; [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10) tracks this, and help with a key is welcome. | Provider | Doc | Implementation | Env var | Models | Status | |---|---|---|---|---|---| | Reducto | [`reducto.md`](/docs/providers/reducto/index.md) | `providers/reducto.rs` | `REDUCTO_API_KEY` | `standard` *(default)*, `r-1`, `agentic` | live-verified | | Extend | [`extend.md`](/docs/providers/extend/index.md) | `providers/extend.rs` | `EXTEND_API_KEY` | `parse_performance` *(default)*, `parse_light`, `parse_auto` | live-verified | | LlamaParse | [`llamaparse.md`](/docs/providers/llamaparse/index.md) | `providers/llamaparse.rs` | `LLAMA_API_KEY` | `fast`, `cost_effective` *(default)*, `agentic`, `agentic_plus` | live-verified | | Mistral | [`mistral.md`](/docs/providers/mistral/index.md) | `providers/mistral.rs` | `MISTRAL_API_KEY` | `ocr-latest` *(default)*, `ocr-4-1`, `ocr-4-0`, `ocr-2512` | docs-only | | Azure AI Document Intelligence | [`azure.md`](/docs/providers/azure/index.md) | `providers/azure.rs` | `AZURE_DOCUMENT_INTELLIGENCE_KEY` + `AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT` | `read` *(ocr default)*, `layout` *(parse default)*, `invoice` *(extract default)*, `receipt`, `id_document`, `tax_us_w2`, `custom` | docs-only | | AWS Textract | [`textract.md`](/docs/providers/textract/index.md) | `providers/textract.rs` | `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY` (+ `AWS_SESSION_TOKEN`, `AWS_REGION`) | `detect-text` *(ocr default)*, `layout`, `queries` *(extract default)*, `forms` | docs-only | | Google Gemini | [`gemini.md`](/docs/providers/gemini/index.md) | `providers/gemini.rs` | `GEMINI_API_KEY` | `2.5-flash` *(default)*, `2.5-pro`, `2.5-flash-lite`, `3.5-flash`, `3.5-flash-lite`, `3.8-flash` | docs-only | | OpenAI | [`openai.md`](/docs/providers/openai/index.md) | `providers/openai.rs` | `OPENAI_API_KEY` | `gpt-5.6-luna` *(default)*, `gpt-5.6-terra`, `gpt-5.6-sol`, `gpt-6-astra` | docs-only | | Anthropic | [`anthropic.md`](/docs/providers/anthropic/index.md) | `providers/anthropic.rs` | `ANTHROPIC_API_KEY` | `claude-sonnet-5` *(default)*, `claude-haiku-4-5`, `claude-opus-5` | docs-only | | Mathpix | [`mathpix.md`](/docs/providers/mathpix/index.md) | `providers/mathpix.rs` | `MATHPIX_APP_ID` + `MATHPIX_APP_KEY` | `pdf` *(default)*, `text` | docs-only | | Datalab (Marker) | [`datalab.md`](/docs/providers/datalab/index.md) | `providers/datalab.rs` | `DATALAB_API_KEY` | `fast`, `balanced` *(default)*, `accurate` | docs-only | | Unstructured | [`unstructured.md`](/docs/providers/unstructured/index.md) | `providers/unstructured.rs` | `UNSTRUCTURED_API_KEY` | `hi_res` *(default)*, `fast`, `auto` | docs-only | | Upstage | [`upstage.md`](/docs/providers/upstage/index.md) | `providers/upstage.rs` | `UPSTAGE_API_KEY` | `document-parse` *(default)*, `document-parse-nightly` | docs-only | | Landing AI (ADE) | [`landingai.md`](/docs/providers/landingai/index.md) | `providers/landingai.rs` | `LANDINGAI_API_KEY` | `dpt-2` *(default)* | docs-only | | Google Document AI | [`google-documentai.md`](/docs/providers/google-documentai/index.md) | `providers/google_documentai.rs` | `GOOGLE_DOCUMENTAI_ACCESS_TOKEN` (+ `_PROJECT`, `_LOCATION`, `_PROCESSOR_ID`) | `ocr` *(default)*, `layout`, `form`, `prebuilt` | docs-only | | Tesseract *(local)* | [`tesseract.md`](/docs/providers/tesseract/index.md) | `providers/tesseract.rs` | none (`TESSERACT_CMD`, `PDFTOPPM_CMD`) | `default` | verified locally | | Docling *(self-hosted)* | [`docling.md`](/docs/providers/docling/index.md) | `providers/docling.rs` | none (`DOCLING_BASE_URL`; optional `DOCLING_API_KEY`) | `default` | verified locally | | PaddleOCR *(self-hosted)* | [`paddleocr.md`](/docs/providers/paddleocr/index.md) | `providers/paddleocr.rs` | none (`PADDLEOCR_BASE_URL`, `PADDLEOCR_PARSE_BASE_URL`) | `default` | docs-only | The three self-hosted engines need no API key and are priced at $0/page (`puffinparse providers` shows `local` in the Key column); they are the open baselines in the benchmark. Their shared helpers are in `providers/local.rs`. Every page follows the same structure: summary → models → request flow → response mapping → errors and limits → gotchas → `provider_options` examples → links. ## At a glance | | Reducto | Extend | LlamaParse | |---|---|---|---| | Upload method | `POST /upload` (multipart) → `reducto://` id; URLs passed through as `input` | `POST /files/upload` (multipart) → `file_…` id; URLs passed through as `{"url", "name"}` | one multipart call: `file` part, or `input_url` field | | Sync / async | Sync `POST /parse` by default; async `POST /parse_async` + `GET /job/{id}` with `provider_options={"async": true}` | Always async: `POST /parse_runs` + `GET /parse_runs/{id}` | Always async: `POST /api/v1/parsing/upload` + `GET …/job/{id}` + `GET …/job/{id}/result/json` | | Page info source | `blocks[].bbox.page` (chunks carry no page number); PuffinParse pins `chunk_mode: "page"` | `chunk.metadata.pageRange` + `block.metadata.page.number`; PuffinParse pins `chunkingStrategy: {"type":"page"}` | `pages[].page` (native per-page objects with `md` and `text`) | | Page dimensions | not reported — `Page.width`/`height` are `None` | `block.metadata.page.width/height` (raster pixels at the run's dpi) | `pages[].width/height` | | Bbox units on the wire | already normalised 0–1, top-left origin | absolute `left/top/right/bottom` in page units | absolute `bBox {x,y,w,h}` in page units | | Bbox after normalisation | clamped as-is | divided by page width/height | divided by page width/height | | Confidence type | `granular_confidence.parse_confidence` (0–1), else coarse `"high"`/`"low"` → 0.9 / 0.5 | `metadata.avgOcrConfidence` (0–1) per block | `items[].bBox.confidence` (0–1) per item | | Credits reported? | yes — `usage.credits` (`null` on new per-product pricing) | yes — `usage.credits` (`null` for runs before 2025-10-07) | effectively no — `job_credits_usage` is 0 until billing settles, so PuffinParse reports `None` | | Large-result indirection | `result.type == "url"` → presigned JSON fetched without auth | `outputUrl` (15-min presigned) when `responseType=url` | none — result endpoints return inline JSON | | Job id surfaced as | `job_id` | run `id` (`pr_…`) | job `id` (UUID) | `Usage.pages` comes from `usage.num_pages` (Reducto), `metrics.pageCount` (Extend) and `job_metadata.job_pages` (LlamaParse), falling back to the number of reconstructed pages. `cost_usd` is always `pages × per_page_usd` from `crates/puffinparse-core/src/pricing.json` — no provider reports dollar cost directly. ## Shared behaviour These are implemented once, in `crates/puffinparse-core/src/{http,error,provider}.rs`, and apply to all three providers: * **API key**: request `api_key` first, then the provider's env var; missing ⇒ `authentication` error. * **Base URL**: request `base_url`, then `_BASE_URL`, then the built-in default. * **Deadline**: `timeout_secs` (default 300) covers upload, submission, polling and result download, and also caps each individual HTTP request. * **Retries**: `max_retries` (default 2) with exponential backoff and full jitter, only on rate-limit, network, and 500/502/503/504 errors. 4xx is never retried. * **Error kinds**: 401/403 → `authentication`, 429 → `rate_limit`, other 4xx → `bad_request`, 5xx / malformed payload / failed job → `provider`, deadline → `timeout`. * **None of the three returns rate-limit headers** (`X-RateLimit-*`), so backoff is always blind. ## How to add a provider See [`CONTRIBUTING.md`](/docs/project/contributing/index.md), section **"3. Adding a provider"** — one file under `crates/puffinparse-core/src/providers/`, registered in `providers/mod.rs` and `model::PROVIDERS`, plus `pricing.json` entries, a fixture-backed normalisation test, and a doc page here following the same eight sections as the pages above. **Promoting a docs-only provider:** run `cargo test -p puffinparse-core -- --ignored` with a key, fix any wire-format differences, replace the hand-built fixture with a redacted real response, and change the status here, in the provider page's banner and in the README model table. Comment on [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10) to claim one. --- # Reducto > **Status: live-verified.** The `#[ignore]`d live tests pass against the real Reducto API, > and its fixtures under `crates/puffinparse-core/tests/fixtures/` include redacted live responses. ## 1. Summary | | | |---|---| | Provider name | `reducto` | | Base URL | `https://platform.reducto.ai` (override: `base_url` on the request, or `REDUCTO_BASE_URL`) | | API key | `REDUCTO_API_KEY` (or `api_key` on the request) — sent as `Authorization: Bearer ` | | Docs | (append `.md` to any docs path for raw markdown; index at `llms.txt`) | | API version | Unversioned — no version header. Server string observed in the published OpenAPI: `v1.12.12-114-g9808b824f237` | | Verified | 2026-09-11, live against the production host | | Implementation | `crates/puffinparse-core/src/providers/reducto.rs` | Reducto's Parse product returns Markdown chunks plus typed, bbox-carrying blocks. PuffinParse uses the current **v3** request schema (`input` + `enhance`/`retrieval`/`formatting`/`spreadsheet`/`settings`), not the legacy `document_url` schema. ## 2. Models exposed by PuffinParse | Model | Provider parameters PuffinParse sets | List price (`pricing.json`) | |---|---|---| | `reducto/standard` *(default)* | none beyond the shared body — the account's default model (legacy Parse for most accounts) | $0.015 / page | | `reducto/r-1` | `settings.model = "r-1"` | $0.010 / page | | `reducto/agentic` | `enhance.agentic = [{"scope":"text"},{"scope":"table"}]` | $0.030 / page | | `reducto/extract` *(default for `extract`)* | `POST /extract` with `instructions.schema` | $0.035 / page (extract $0.020 + the parse it runs) | | `reducto/deep_extract` | `settings.deep_extract = true` | $0.055 / page (deep extract $0.040 + parse) | The first three models serve `parse` and `ocr`; the last two serve `extract` only (§5). Prices are public pay-as-you-go list prices (source: ), used only to fill `OcrResponse.cost_usd = per_page_usd × usage.pages`. Reducto bills "complex" pages a surcharge credit on top of the base page credit, so the estimate is a floor for `standard`/`agentic`. ## 3. Request flow PuffinParse uses All requests carry `Authorization: Bearer $REDUCTO_API_KEY`. There is no version or workspace header. 1. **Upload (only for path / bytes input).** `POST {base}/upload`, `multipart/form-data`, single part `file` with the filename and guessed MIME type. Response: `{"file_id":"reducto://.pdf", "presigned_url":null}`; `file_id` is used verbatim as `input`. **URL inputs skip this step** — the URL string is passed straight through as `input`, and Reducto downloads it server-side. 2. **Create.** `POST {base}/parse` (default) with `Content-Type: application/json`. The body is built by `build_body()`: ```json { "input": "reducto://.pdf", "retrieval": { "chunking": { "chunk_mode": "page" } }, "formatting": { "table_output_format": "md" }, "settings": {} } ``` * `reducto/r-1` adds `"settings": {"model": "r-1"}`. * `reducto/agentic` adds `"enhance": {"agentic": [{"scope":"text"},{"scope":"table"}]}`. * `pages="1-3,7"` becomes `settings.page_range = [{"start":1,"end":3},{"start":7}]` (1-based, matching Reducto's own indexing). * `chunk_mode: "page"` is deliberate: Reducto's default is `"disabled"`, which returns the whole document as one chunk and gives PuffinParse no page structure. 3. **Async variant.** When `provider_options.async == true`, PuffinParse posts the *same* body to `POST {base}/parse_async` → `{"job_id":"..."}`, then polls `GET {base}/job/{job_id}` (Bearer header required) starting at 1 s and backing off ×1.5 to a maximum of 8 s. Terminal statuses are `Completed`, `Failed`, `Cancelled`; anything else (`Pending`, `Idle`, `InProgress`, `Completing`, …) keeps polling. A completed job's `result` is deserialised as the same `ParseResponse` the sync call returns — i.e. the payload lives at `job.result.result.chunks`. 4. **Result.** If `result.type == "url"` (large results, or `settings.force_url_result`), PuffinParse issues a plain `GET` on the presigned URL **without** the Authorization header and expects the *whole* `FullResult` object back (`{"type":"full","chunks":[…]}`), not a bare chunk array. A second `url` result is an error, and a storage error (e.g. an expired link: S3 answers `403` with an XML body) is surfaced with its message. Verified live on 2026-09-24 with `settings.force_url_result: true`; the captured envelope and URL body are the fixtures `reducto_parse_url.json` / `reducto_parse_url_result.json` (ids and signatures redacted), replayed end to end by the loopback tests in `providers::reducto::wire`. The link is valid for 12 h (`X-Amz-Expires=43200`) and the object is deleted after 24 h. **Where `provider_options` are merged:** `build_body()` clones `provider_options`, removes the PuffinParse-only key `async`, and deep-merges the rest into the body above (objects merge recursively; scalars and arrays replace). So `provider_options` keys are top-level Reducto request keys — `settings`, `retrieval`, `formatting`, `enhance`, `spreadsheet`, `queue_priority`, `async` (the Reducto object of that name is *not* forwarded; only the boolean switch is consumed). ### Jobs API and webhooks (`submit_parse` / `retrieve_parse`, SPEC §15) * **Submit** — the same upload step and body as `parse`, sent to `POST {base}/parse_async`; the returned `job_id` becomes `JobHandle.job_id`. A Reducto-native `provider_options.async` *object* (`{"metadata": …, "priority": …}`) is forwarded as the body's `async`; the boolean switch of the blocking path is not. * **`webhook_url`** → `async.webhook = {"mode": "direct", "url": ""}`. Reducto POSTs `{"status": "Completed" | "Failed", "job_id": "…", "metadata": {…}}` (retried up to 3 times). The body has no result and no failure reason, so `parse_webhook` reports `Finished` and `resolve_webhook` / Python `handle_webhook` fetch `GET /job/{id}`. Direct webhooks are unsigned: put a secret in `async.metadata` and check it before trusting the body. Svix-mode webhooks are not configured by PuffinParse (pass `provider_options={"async": {"webhook": {"mode": "svix", …}}}`). * **Retrieve** — one `GET {base}/job/{job_id}` (retried on 429/5xx). `Completed` → the `result` (a `ParseResponse`, `url`-typed results followed as above); `Failed` / `Cancelled` → `JobStatus::Failed` with `reason` (or `error.message`); `Pending`, `Idle`, `InProgress`, `Completing` → `Pending`. Job ids expire after 12 h. * Verified live 2026-09-24 (`tests/live_jobs.rs::reducto_submit_retrieve_live`, 1 page: 3 status checks, ~4.6 s). ## 4. Response mapping (`parse` / `ocr`) | Reducto field | PuffinParse unified field | Notes | |---|---|---| | `job_id` | `OcrResponse.provider_job_id` | | | `result.chunks[].content` | `Page.markdown` | Only when every block in the chunk is on one page; multiple such chunks on a page are joined with a blank line. | | `result.chunks[].blocks[]` | `Page.blocks[]` | Grouped by page, reading order preserved. | | `blocks[].type` | `Block.type` | See mapping below. | | `blocks[].content` | `Block.content` | Markdown; `output="text"` runs it through `markdown_to_text`. | | `blocks[].bbox.page` | `Block.page_number` | 1-based; missing bbox ⇒ page 1. `original_page` is not used. | | `blocks[].bbox.{left,top,width,height}` | `Block.bbox` = `{x0,y0,x1,y1}` | **Already normalised 0–1**, origin top-left. `from_normalized_ltwh` clamps and converts to `x1=left+width`, `y1=top+height` — no page dimensions needed. `Page.width`/`height` stay `None`. | | `blocks[].granular_confidence.parse_confidence` | `Block.confidence` | Preferred, 0–1. | | `blocks[].confidence` (`"high"`/`"low"`) | `Block.confidence` | Fallback only: `high → 0.9`, `low → 0.5`, anything else `None`. | | `usage.num_pages` | `Usage.pages` | Falls back to the number of reconstructed pages when 0/absent. | | `usage.credits` | `Usage.credits` | `null` on accounts migrated to per-product pricing → `None`. | | `duration` | `metadata.reducto_duration_s` | | | `studio_link` | `metadata.reducto_studio_link` | | | — | `Usage.provider_cost_usd` | Never set; `cost_usd` comes from `pricing.json`. | Block types: `Text`, `Key Value`, `Comment` → `text`; `Title` → `title`; `Section Header` → `section_header`; `List Item` → `list`; `Table` → `table`; `Figure` → `figure`; `Header` → `header`; `Footer` → `footer`; `Footnote` → `footnote`; `Caption` → `caption`; `Formula`/`Equation` → `formula`; everything else (including `Page Number`, `Signature`, `Checkbox`) → `other`. Trimmed real response (`crates/puffinparse-core/tests/fixtures/reducto_parse.json`, two of three blocks elided): ```json { "response_type": "parse", "job_id": "04fe7cf4-1f40-4a03-b1a0-460bd8d6a89a", "duration": 0.7815132141113281, "pdf_url": "https://example-storage.invalid/converted.pdf?X-Amz-Signature=REDACTED", "studio_link": "https://studio.reducto.ai/job/04fe7cf4-1f40-4a03-b1a0-460bd8d6a89a", "usage": { "num_pages": 1, "credits": 1.0, "credit_breakdown": { "page": 1.0 }, "page_billing_breakdown": { "1": ["page"] }, "non_empty_cell_count": null }, "result": { "type": "full", "chunks": [ { "content": "# Hello PuffinParse\n\nInvoice #1234\nTotal: $56.78\n\nAcme Corporation, 123 Main Street", "embed": "# Hello PuffinParse\n\nInvoice #1234\nTotal: $56.78\n\nAcme Corporation, 123 Main Street", "enriched": null, "enrichment_success": false, "blocks": [ { "type": "Title", "bbox": { "left": 0.119281045751634, "top": 0.06818181818181818, "width": 0.24754901960784315, "height": 0.021464646464646464, "page": 1, "original_page": 1 }, "content": "Hello PuffinParse", "image_url": null, "chart_data": null, "confidence": "high", "granular_confidence": { "extract_confidence": null, "parse_confidence": 0.9461241155862807 }, "extra": null } ] } ], "ocr": null, "custom": null }, "parse_mode": null, "document_properties": null } ``` ## 5. Extract mode (`extract`) `puffinparse.extract(...)` posts to `POST {base}/extract` (sync) or `POST {base}/extract_async` + `GET {base}/job/{job_id}` when `provider_options={"async": true}` — the same upload step, auth header and polling schedule as parse. Implementation: `Reducto::extract` in `crates/puffinparse-core/src/providers/reducto.rs`. Body built by `build_extract_body()`: ```json { "input": "reducto://.png", "instructions": { "schema": { "...the request's JSON Schema, verbatim..." }, "system_prompt": "…only when `instructions` is set…" }, "settings": { "citations": { "enabled": true }, "deep_extract": true, "page_range": [{ "start": 1, "end": 3 }] } } ``` * The **schema is passed through unchanged** — Reducto accepts standard JSON Schema, so nothing is rewritten (unlike Extend — see `extend.md` §5). * `ExtractRequest.instructions` → `instructions.system_prompt` (Reducto's default is `"Be precise and thorough."`). * `ExtractRequest.citations = true` → `settings.citations.enabled`. Without it Reducto returns bare values and no boxes. * `reducto/deep_extract` adds `settings.deep_extract` (agentic refinement loop; `usage.extract_mode` comes back as `"super_agent"`). * `pages` → `settings.page_range`, and `provider_options` are deep-merged exactly as on the parse path (`settings`, `parsing`, `queue_priority`, …; the PuffinParse-only `async` key is consumed). ### Response mapping The response shape **changes with citations**, which is the main trap: | Citations | `response_type` | `result` | |---|---|---| | off | `"extract"` | a **list** of objects (one per chunk; length 1 unless chunking is on), plus a top-level `"citations": null` | | on | `"v3_extract"` | an **object** whose every *leaf* is `{"value": …, "citations": [ParseBlock…]}` — recursively, inside nested objects and arrays too | | Reducto field | PuffinParse unified field | Notes | |---|---|---| | `result` | `ExtractResponse.data` | Citation wrappers are stripped so `data` matches the request schema. A single-element list is unwrapped to the object; a longer list is kept as an array (pointers then start `/0/…`). | | `result..citations[]` | `ExtractResponse.fields[""]` | Keys are RFC 6901 pointers: `/invoice_number`, `/line_items/0/amount`, `/vendor_address/city`. Only leaves get an entry. | | citation `bbox{left,top,width,height,page}` | `Citation.bbox` / `Citation.page_number` | Already normalised 0–1, top-left origin — converted with `from_normalized_ltwh`. | | citation `content` | `Citation.text` | The matched span (its `parentBlock` — the whole source block — is **not** surfaced; set `settings.citations.parent_block="bbox_only"` to shrink responses). | | citation `granular_confidence.extract_confidence` | `FieldInfo.confidence` | Max over a field's citations; falls back to `confidence` (`high → 0.9`, `low → 0.5`). | | `usage.num_pages` | `Usage.pages` | | | `usage.credits` | `Usage.credits` | `null` on accounts on per-product pricing. | | `usage.num_fields` / `usage.extract_mode` | `metadata.reducto_num_fields` / `metadata.reducto_extract_mode` | `extract_mode` ∈ `extract`, `super_agent`, `spreadsheet_agent`. | | `confidence` / `confidence_reason` | `metadata.reducto_confidence…` | Document-level Deep Extract labels, when present. | | `studio_link` | `metadata.reducto_studio_link` | | | `job_id` | `ExtractResponse.provider_job_id` | Optional in the schema, so it may be `None`. | `settings.force_url_result` (and large results) replace `result` with `{"type":"url","url":"https://…"}`; PuffinParse fetches that URL **without** the Authorization header and uses the body, which is the *bare* result value — not a wrapper object like the parse path's. Trimmed real response (`crates/puffinparse-core/tests/fixtures/reducto_extract.json`, one citation shown): ```json { "response_type": "v3_extract", "job_id": "33e0fac4-a1bc-42ed-9aab-9c9faab42e31", "usage": { "num_pages": 1, "num_fields": 15, "credits": 3.333333, "extract_mode": "extract" }, "studio_link": "https://studio.reducto.ai/job/33e0fac4-a1bc-42ed-9aab-9c9faab42e31", "result": { "invoice_number": { "value": "INV-9865", "citations": [ { "type": "Key Value", "bbox": { "left": 0.157, "top": 0.269, "width": 0.092, "height": 0.011, "page": 1, "original_page": 1 }, "content": "INV-9865", "confidence": "high", "granular_confidence": { "extract_confidence": 0.996, "parse_confidence": 0.821 }, "parentBlock": { "type": "Key Value", "bbox": { "…": "…" }, "content": "" } } ] }, "line_items": [ { "description": { "value": "Hydraulic fluid, 5 gal", "citations": [ … ] }, "amount": { "value": "$439.20", "citations": [ … ] } } ] } } ``` **Verified live** on 2026-09-11 with `benchmark/datasets/synthetic-v1/docs/invoice_001.png`: `{invoice_number: "INV-9865", total: "$14,667.43", date: "2024-08-03", vendor: "Cedar Ridge Supply"}`, `usage.num_pages = 1`, `credits = 3.333333`, every field cited on page 1 with `extract_confidence ≈ 0.995`. Test: `providers::reducto::tests::live_extract` (`#[ignore]`). ### Extract gotchas * **`parentBlock` is the only camelCase key in the API.** It embeds the entire source block, so table-heavy schemas repeat the same block many times; `settings.citations.parent_block="bbox_only"` blanks its `content`. * **Citations change the response type**, including the container type of `result` (list → object). Code that indexes `result[0]` breaks as soon as citations are on. * **Credits are not the extract price alone**: a 1-page extract of this invoice billed `3.333333` credits — the extract itself, plus the page parse, plus Reducto's "complex page" surcharge. The `pricing.json` figure is a list-price estimate, not the billed amount. * **Deep Extract ≠ 2× credits in practice** (`4.666667` vs `3.333333` on the same page), though the list price is 2×. * An invalid JSON Schema comes back as `422` with a Pydantic validation array. * Passing a `jobid://…` input from a previous `/parse` skips (and stops billing) the parse step; `parsing` options are then ignored. ## 6. Errors, status codes, rate limits, timeouts Every non-2xx body goes through `Error::from_http`, which pulls a message out of `message` / `detail` / `error` and classifies by status: | Status | Reducto body | PuffinParse `ErrorKind` | |---|---|---| | 400 | `{"error":{"code":400,"name":"INVALID_CONFIG",…},"detail":"…"}` — bad config, source download failure, bad page range | `bad_request` | | 401 | `{"error":{"code":401,"name":"AUTH_ERROR","message":"Invalid access token"},…}` | `authentication` | | 403 | **nginx HTML page** when the `Authorization` header is missing entirely; also expired/inaccessible presigned source | `authentication` | | 404 | File not found | `bad_request` | | 413 | Image over 50 MP / 15 000 px per axis | `bad_request` | | 415 | Conversion error / corrupt PDF | `bad_request` | | 422 | Pydantic validation array (`{"detail":[…]}`) | `bad_request` | | 429 | `{"message":"[CODE 1000] rate limit exceeded, retry with exponential backoff"}` | `rate_limit` (retried) | | 442 | Password-protected document — Reducto's own non-standard code | `bad_request` | | 500 | Content/citation extraction error — not retriable upstream | `provider` | | 502/503/504 | LLM error, overload, timeout | `provider` (retried) | `reducto::map_status()` spells 442 out explicitly; the live path reaches the same verdict through `Error::from_http`'s generic `400..=499 → bad_request` arm. An async job that ends `Failed`/`Cancelled` becomes a `provider` error carrying `reason` (or `error.message`) and the `job_id`. **Retries.** `max_retries` (default 2) with exponential backoff + full jitter, on rate-limit, network, and 500/502/503/504 errors only. 4xx is never retried. **Rate limits.** 1 000 req/s per key across all endpoints (`[CODE 1000]`), 200 req/s on `GET /job/{id}` (`[CODE 2000]`, "use webhooks instead of polling"). A separate concurrency throttle (200/350/500+ concurrent pages by plan) does **not** return 429 — it just queues, so the symptom is latency. **No rate-limit headers of any kind** are returned (no `X-RateLimit-*`, no `Retry-After`), so backoff is blind. **Timeouts.** `timeout_secs` (default 300) is a whole-call deadline covering upload, parse, polling and result download; it is also the per-request timeout, shrinking as the budget is spent. Reducto's own sync `/parse` ceiling is 900 s — use `provider_options={"async": true}` for anything longer. ## 7. Gotchas (verified) * **`input`, not `document_url`.** The body is a three-way server-side union: `ParseConfigNew`, legacy `ParseConfig` (`document_url` + `options`/`advanced_options`/`experimental_options`), and the current `SyncParseConfig` (`input` + option groups). Legacy still works on sync `/parse`, but mixing the two (`document_url` + `retrieval`) yields `400 INVALID_CONFIG`. Do not push `document_url` through `provider_options`. * **One bad field produces several validation errors**, most of them complaints about a schema you never used; the `SyncParseConfig`-scoped entry is the relevant one. * **Error text always names `document_url`** even when you sent `input` — never parse error strings. * **`result.type == "url"`** appears for large results (~6 MB inline limit). The fetched body is the full `{"type":"full","chunks":[…]}` object, not a bare array — the vendor's own snippet gets this wrong. PuffinParse handles both variants; force the URL path deterministically with `provider_options={"settings":{"force_url_result":true}}`. * **442 is a real status code** (password-protected document), not a typo for 422. Supply `settings.document_password`. * **No rate-limit headers**, no `Retry-After`, even on 429. * **Missing auth header → 403 with an HTML body**, invalid token → 401 with JSON. PuffinParse's message extraction falls back to the raw (truncated) body for the HTML case. * **Chunks have no page number.** Page identity comes only from `blocks[].bbox.page`; with `chunk_mode: "variable"` chunks straddle pages, which is why PuffinParse pins `chunk_mode: "page"` and still derives page numbers from blocks rather than array position. * **Undocumented keys on the wire**: `response_type`, `parse_mode`, `document_properties`, `usage.credit_breakdown`, `usage.page_billing_breakdown` (1-based page numbers as *string* keys), `usage.non_empty_cell_count`. PuffinParse ignores all but `credit_breakdown`, which it deserialises but does not surface. * **`usage.credits` can be `null`** on accounts on the new per-product pricing (they get `usage_breakdown` instead) — `Usage.credits` is then `None` and `cost_usd` still comes from the table. * **Uploaded files expire after 24 h**, and results expire after 24 h unless `settings.persist_results` is set. * **`settings.return_images: ["page"]` does not populate `block.image_url`** (it stays `null`); the URL shows up under `block.extra.page_image_url`. PuffinParse surfaces neither. ## 8. Useful `provider_options` passthrough ```python # 1. Long documents: submit asynchronously and poll (the `async` key is consumed by PuffinParse). puffinparse.ocr("200-page.pdf", model="reducto/standard", provider_options={"async": True}) # 2. Word- and line-level OCR boxes in the raw payload (+$2 / 1k pages). puffinparse.ocr("scan.pdf", model="reducto/standard", include_raw=True, provider_options={"settings": {"return_ocr_data": True}}) # 3. HTML tables instead of PuffinParse's markdown default, and merge tables split across pages. puffinparse.ocr("report.pdf", model="reducto/r-1", provider_options={"formatting": {"table_output_format": "html", "merge_tables": True}}) # 4. Cheap bulk queue: 12-hour completion guarantee, 20% usage discount. puffinparse.ocr("batch.pdf", model="reducto/r-1", provider_options={"async": True, "queue_priority": "batch"}) # 5. Password-protected PDF, and always take the presigned-URL result path. puffinparse.ocr("locked.pdf", model="reducto/standard", provider_options={"settings": {"document_password": "…", "force_url_result": True}}) ``` ## 9. Links * Docs home: · agent guide: · index: * Parse API reference: * Legacy parse schema: * Credit usage / pricing: · * Error codes: * Studio (per-job inspector): --- # Extend > **Status: live-verified.** The `#[ignore]`d live tests pass against the real Extend API, > and its fixtures under `crates/puffinparse-core/tests/fixtures/` include redacted live responses. ## 1. Summary | | | |---|---| | Provider name | `extend` | | Base URL | `https://api.extend.ai` (override: `base_url` on the request, or `EXTEND_BASE_URL`). Regional hosts: `https://api.us2.extend.app` (note `.app`), `https://api.eu1.extend.ai` | | API key | `EXTEND_API_KEY` (or `api_key` on the request) — sent as `Authorization: Bearer ` | | Docs | (append `.md` to any docs path for raw markdown; index at `llms.txt`) | | API version | **`2026-02-09`**, pinned by PuffinParse in the mandatory `x-extend-api-version` header (`extend::API_VERSION`) | | Verified | 2026-09-11, live against the production host | | Implementation | `crates/puffinparse-core/src/providers/extend.rs` | Extend returns page chunks of markdown plus typed blocks with polygons, bounding boxes and OCR confidence. PuffinParse always uses the **async** run API (`/parse_runs`), never the 5-minute sync `/parse`. ## 2. Models exposed by PuffinParse | Model | Provider parameters PuffinParse sets | List price (`pricing.json`) | |---|---|---| | `extend/parse_performance` *(default)* | `config.engine = "parse_performance"` | $0.025 / page (2 credits) | | `extend/parse_light` | `config.engine = "parse_light"` | $0.00625 / page (0.5 credits) | | `extend/parse_auto` | `config.engine = "parse_auto"` | $0.025 / page (per-page: light pages bill 0.5 credits) | | `extend/extraction_performance` *(default for `extract`)* | `config.baseProcessor = "extraction_performance"`, `config.parseConfig.engine = "parse_performance"` | $0.0625 / page (3 + 2 credits) | | `extend/extraction_light` | `config.baseProcessor = "extraction_light"`, `config.parseConfig.engine = "parse_light"` | $0.015 / page (0.7 + 0.5 credits) | The first three models serve `parse` and `ocr`; the last two serve `extract` only (§5). An extract run triggers its own parse run and is billed for both, which is why the extract prices are the sum of the two line items; re-extracting a file Extend has already parsed bills only the extraction. Credits are the billing unit; pay-as-you-go is $0.0125/credit (Scale: $0.01). Prices above are the PAYG list rate from , used for `OcrResponse.cost_usd`. `parse_auto` is priced at the performance rate, so its estimate is an upper bound. Surcharges PuffinParse does not model: agentic text/table correction +1 credit per triggered page, priority parsing ×2, advanced Excel parsing 3 credits / 1 000 non-empty cells. ## 3. Request flow PuffinParse uses Every request carries: ``` Authorization: Bearer $EXTEND_API_KEY x-extend-api-version: 2026-02-09 x-extend-workspace-id: # only when supplied ``` 1. **Upload (only for path / bytes input).** `POST {base}/files/upload`, `multipart/form-data`, single part `file`. Response: a file object whose `id` (`file_…`) is used as `{"id": …}`. **URL inputs skip this step** — PuffinParse sends `{"url": "", "name": ""}` and Extend downloads the file itself, creating a `file_…` record it returns in `run.file`. 2. **Create the run.** `POST {base}/parse_runs` with `Content-Type: application/json`. Body from `build_body()`: ```json { "file": { "id": "file_bJ2ZXxacw7o206UnS6eR0" }, "config": { "target": "markdown", "chunkingStrategy": { "type": "page" }, "engine": "parse_performance", "blockOptions": { "tables": { "targetFormat": "markdown" } } } } ``` * `pages="1-3,7"` adds `config.advancedOptions.pageRanges = [{"start":1,"end":3},{"start":7,"end":750}]` — Extend requires an `end`, and 750 is its documented maximum page-range end. * `targetFormat: "markdown"` is deliberate: Extend's default is `"html"`. * The response echoes the **resolved** config (including `engineVersion`, e.g. `"2.0.0"`), which is what PuffinParse reads back to label the response model. 3. **Poll.** `GET {base}/parse_runs/{id}` starting at 1 s, backing off ×1.5 to a maximum of 10 s, until `status` is `PROCESSED` or `FAILED`. (`PENDING`/`PROCESSING` keep polling.) If the create call already came back terminal, the poll loop is skipped. 4. **Result.** Normally `run.output` is inline. `responseType` is a **query parameter of `GET /parse_runs/{id}`** (not a body key): with `provider_options={"responseType": "url"}` PuffinParse consumes the key (it is never sent in the create body) and polls `GET /parse_runs/{id}?responseType=url`. The finished run then has `output: null` and an `outputUrl`, which PuffinParse fetches with a plain `GET` and **no** auth header; the payload has exactly the shape of `output` (`{chunks, metadata, ocr?}`). Presigned output URLs expire in 15 minutes. (Before 2026-09-24 PuffinParse forwarded the key into the body, so this path never triggered.) Verified live on 2026-09-24; the captured run and output are the fixtures `extend_parse_run_url.json` / `extend_parse_run_url_output.json` (ids and signatures redacted), replayed by the loopback tests in `providers::extend::wire`. **Where `provider_options` are merged:** `build_body()` removes `workspace_id` (it becomes a header) and `responseType` (a query parameter, step 4), then lifts any of `target`, `chunkingStrategy`, `engine`, `engineVersion`, `blockOptions`, `advancedOptions` into `config` (deep-merged with an explicit `config` object if you passed one), and deep-merges everything that remains at the **top level** of the body — which is how `metadata`, `dataRetention` and even `file` overrides get through. So both spellings work: `{"blockOptions": {...}}` and `{"config": {"blockOptions": {...}}}`. ### Jobs API and webhooks (`submit_parse` / `retrieve_parse`, SPEC §15) * **Submit** — the same file reference and `POST {base}/parse_runs` as `parse`; the run id (`pr_…`) becomes `JobHandle.job_id`. `workspace_id` and `responseType` are kept in `JobHandle.provider_state` so retrieve sends the same header and query parameter. * **Retrieve** — one `GET {base}/parse_runs/{id}` (retried on 429/5xx). `PROCESSED` → the output (inline or via `outputUrl`); `FAILED` / `CANCELLED` → `JobStatus::Failed` with `failureReason: failureMessage` (kinds as in §6); `PENDING` / `PROCESSING` → `Pending`. * **`webhook_url` is rejected** (`input` error, no request sent). Extend has no per-run webhook: webhooks are workspace **endpoints** (`POST /webhook_endpoints` or the dashboard) subscribed to events such as `parse_run.processed` / `parse_run.failed`. Their body is `{"eventId", "eventType", "payload": {"object": "parse_run_status", "id", "status", "failureReason", "failureMessage", "metadata"}}`; `parse_webhook` maps `FAILED` straight to `Failed` (the reason is in the body), `PROCESSED` to `Finished` (then one retrieve), anything else to `Pending`. Endpoints configured for *signed download URL* delivery send `payload: {"data": ""}` — download it and pass the JSON inside. Verify the `HMAC-SHA256(v0:{timestamp}:{body})` signature before trusting a body. Non-`parse_run.*` events are rejected. * Verified live 2026-09-24 (`tests/live_jobs.rs::extend_submit_retrieve_live`, `parse_light`, 1 page: 3 status checks, ~4.4 s). ## 4. Response mapping (`parse` / `ocr`) | Extend field | PuffinParse unified field | Notes | |---|---|---| | `id` (`pr_…`) | `OcrResponse.provider_job_id` | | | `config.engine` | `OcrResponse.model` | `extend/` read back from the resolved config; falls back to `parse_performance`. | | `output.chunks[].content` | `Page.markdown` | Only for chunks with `type == "page"` and `pageRange.start == pageRange.end`. | | `output.chunks[].blocks[]` | `Page.blocks[]` | Reading order preserved. | | `blocks[].type` | `Block.type` | See mapping below. | | `blocks[].content` | `Block.content` | Markdown; `output="text"` runs it through `markdown_to_text`. | | `blocks[].metadata.page.number` | `Block.page_number` | Falls back to the chunk's `pageRange.start`. | | `blocks[].metadata.page.{width,height}` | `Page.width` / `Page.height` | First value seen per page wins; these are the *raster* dimensions (e.g. 1241×1754 at 150 dpi), not PDF points. | | `blocks[].boundingBox.{left,top,right,bottom}` | `Block.bbox` = `{x0,y0,x1,y1}` | **Absolute, top-left origin, page units.** Normalised by the page's `width`/`height` via `from_xywh(left, top, right-left, bottom-top, w, h)` and clamped to 0–1. No page dims for that page ⇒ `bbox = None`. | | `blocks[].metadata.avgOcrConfidence` | `Block.confidence` | 0–1 float. `minOcrConfidence` and chunk-level confidences are ignored. | | `metrics.pageCount` | `Usage.pages` | Rounded; falls back to the number of reconstructed pages. | | `usage.credits` | `Usage.credits` | `null` for runs created before 2025-10-07. `totalCredits`/`breakdown` are not surfaced. | | `metrics.processingTimeMs` | `metadata.extend_processing_time_ms` | | | `blocks[].polygon`, `output.ocr.words[]`, `output.metadata.pages[]` | — | Not mapped; visible with `include_raw=True`. | Block types: `text`, `key_value` → `text`; `heading` → `title`; `section_heading` → `section_header`; `table`, `table_head`, `table_cell` → `table`; `figure` → `figure`; `formula` → `formula`; `header` → `header`; `footer` → `footer`; everything else (`page_number`, `barcode`, …) → `other`. Trimmed real response (`crates/puffinparse-core/tests/fixtures/extend_parse_run.json`, one chunk with one block shown; the file holds 2 pages, 5 blocks and 21 OCR words): ```json { "object": "parse_run", "id": "pr_xi5wEyAbYlYVBDDy8QDRg", "file": { "object": "file", "id": "file_bJ2ZXxacw7o206UnS6eR0", "name": "test_multi.pdf", "type": "PDF", "parentFileId": null, "metadata": { "pageCount": 2 }, "dataRetention": { "mode": "workspace_default", "status": "available" }, "createdAt": "2026-09-11T07:18:23.026Z", "updatedAt": "2026-09-11T07:18:37.514Z" }, "status": "PROCESSED", "failureReason": null, "failureMessage": null, "metadata": null, "dataRetention": { "mode": "workspace_default", "status": "available" }, "output": { "chunks": [ { "object": "chunk", "id": "chunk_1_iSMg4N", "type": "page", "content": "# Hello PuffinParse\n\nInvoice #1234\nTotal: $56.78\nDate: 2026-09-11\n\n| Item | Amount |\n| --- | --- |\n| Widget | $56.78 |", "metadata": { "pageRange": { "start": 1, "end": 1 }, "minOcrConfidence": 0.905, "avgOcrConfidence": 0.986 }, "blocks": [ { "object": "block", "id": "block_1_YFFbD5", "type": "table", "content": "| Item | Amount |\n| --- | --- |\n| Widget | $56.78 |", "details": { "type": "table_details", "rowCount": 2, "columnCount": 2 }, "metadata": { "page": { "number": 1, "width": 1241, "height": 1754 }, "minOcrConfidence": 0.905, "avgOcrConfidence": 0.973 }, "polygon": [ { "x": 90.365, "y": 382.357 }, { "x": 704.950, "y": 382.357 }, { "x": 704.950, "y": 585.043 }, { "x": 90.365, "y": 585.043 } ], "boundingBox": { "left": 90.365, "top": 382.357, "right": 704.950, "bottom": 585.043 } } ] } ], "metadata": { "originalMimeType": "application/pdf", "finalMimeType": "application/pdf", "pages": [ { "number": 1, "rotationApplied": 0, "originalPageWidth": 596, "originalPageHeight": 842, "dpi": 150 } ] } }, "outputUrl": null, "metrics": { "processingTimeMs": 4383, "pageCount": 2 }, "config": { "target": "markdown", "chunkingStrategy": { "type": "page" }, "engine": "parse_performance", "engineVersion": "2.0.0", "blockOptions": { "tables": { "targetFormat": "markdown" }, "figures": { "enabled": true } } }, "batchId": null, "usage": { "credits": 4, "totalCredits": 4, "breakdown": [ { "object": "parse_run", "id": "pr_xi5wEyAbYlYVBDDy8QDRg", "credits": 4, "charges": [ { "product": "parse_performance", "unit": "page", "quantity": 2, "credits": 4 } ] } ] } } ``` ## 5. Extract mode (`extract`) `puffinparse.extract(...)` reuses the file reference from §3 (URL passed through as `file.url`, everything else uploaded to `POST /files/upload`) and then runs the async pair `POST {base}/extract_runs` → poll `GET {base}/extract_runs/{id}`, with the same `x-extend-api-version: 2026-02-09` (and optional `x-extend-workspace-id`) headers. Terminal statuses are `PROCESSED`, `FAILED`, `CANCELLED`. The sync `POST /extract` is not used: Extend documents it as onboarding-only and caps it at 5 minutes. Implementation: `Extend::extract` in `crates/puffinparse-core/src/providers/extend.rs`. Body built by `build_extract_body()`: ```json { "file": { "id": "file_…" }, "config": { "baseProcessor": "extraction_performance", "schema": { "...the request schema, adapted (see below)..." }, "parseConfig": { "engine": "parse_performance" }, "extractionRules": "…only when `instructions` is set…", "advancedOptions": { "citationsEnabled": true, "pageRanges": [{ "start": 1, "end": 3 }] } } } ``` * `ExtractRequest.instructions` → `config.extractionRules` (natural-language guidance). * `ExtractRequest.citations = true` → `config.advancedOptions.citationsEnabled`. This also switches on `ocrConfidence`, which is **absent** otherwise, and adds latency (a separate citation model runs). Granularity knobs (`citationMode: line|word|block`, `arrayCitationStrategy: item|property`) are available through `provider_options`. * `pages` → `config.advancedOptions.pageRanges` (open-ended ranges end at Extend's 750-page cap). * `provider_options` keys `baseProcessor`, `baseVersion`, `extractionRules`, `schema`, `advancedOptions`, `parseConfig` are merged into `config`; anything else (`metadata`, `dataRetention`, `extractor`) is merged at the top level — so a saved extractor is reachable with `provider_options={"extractor": {"id": "ex_…"}}`. ### Schema adaptation (mandatory) Extend rejects plain JSON Schema with `400 INVALID_REQUEST`. `adapt_schema()` rewrites the request schema before sending it: | PuffinParse input | Sent to Extend | Why | |---|---|---| | `{"type": "string"}` | `{"type": ["string", "null"]}` | *"Non-nullable primitive type "string" is not allowed."* Applies to `string`, `number`, `integer`, `boolean`. | | `{"type": "array", "items": {"type": "string"}}` | unchanged | Primitive **array items** must stay non-nullable — the documented exception. | | `{"enum": ["a", "b"]}` | `{"enum": ["a", "b", null]}` | Enums must offer a `null` option (a real JSON `null` here, unlike the `"null"` *string* used in a type union). | Everything else is passed through, including `required`, `description`, `extend:type`, `extend:name`. Unsupported constructs (`anyOf`/`oneOf`/`allOf`, `$ref`, `const`, regex/format validation, nesting deeper than 5) are *not* rewritten and will be rejected by the API. ### Response mapping `output` has two halves sharing the same field paths: `output.value` (the data) and `output.metadata` (per-field confidence + citations). | Extend field | PuffinParse unified field | Notes | |---|---|---| | `output.value` | `ExtractResponse.data` | Exactly the schema shape. | | `output.metadata["line_items[0].amount"]` | `ExtractResponse.fields["/line_items/0/amount"]` | Extend's path notation is converted to an RFC 6901 pointer; `~` and `/` inside names are escaped. Array *containers* (`line_items`, `line_items[0]`) get their own entries and are kept. | | `metadata[].ocrConfidence` | `FieldInfo.confidence` | Falls back to `logprobsConfidence`, which is being phased out (`null` on `extraction_light` and on `extraction_performance ≥ 4.6.0`). | | `metadata[].citations[].page.number` | `Citation.page_number` | 1-based. | | `metadata[].citations[].polygon[]` | `Citation.bbox` | The polygon's axis-aligned bounds, normalised by `page.width`/`page.height` (page pixels). No page dimensions ⇒ `bbox: None`. | | `metadata[].citations[].referenceText` | `Citation.text` | | | `usage.breakdown[].charges[]` where `unit == "page"` | `Usage.pages` | Max `quantity` over the charges; falls back to `file.metadata.pageCount`, then the highest cited page, then 1. | | `usage.totalCredits` (else `usage.credits`) | `Usage.credits` | `totalCredits` includes the parse run the extraction triggered. | | `parseRunId` / `dashboardUrl` / `reviewed` | `metadata.extend_parse_run_id` / `extend_dashboard_url` / `extend_reviewed` | | | `id` | `ExtractResponse.provider_job_id` | `exr_…` | | `failureReason` + `failureMessage` | error message | `OUT_OF_CREDITS → authentication`; `INTERNAL_ERROR`, `FAILED_TO_PROCESS_FILE`, `PARSING_ERROR`, `PRE_`/`POST_PROCESSING_FAILURE` → `provider`; everything else (`INVALID_CONFIGURATION`, `SCHEMA_GENERATION_FAILED`, …) → `bad_request`. | `reviewAgentScore` and `insights` (model reasoning) are not surfaced; enable them through `provider_options` and read `response.raw`. Trimmed real response (`crates/puffinparse-core/tests/fixtures/extend_extract_run.json`): ```json { "object": "extract_run", "id": "exr_TR4bUO18s2EjPzeLNB5vy", "status": "PROCESSED", "output": { "value": { "invoice_number": "INV-9865", "total": "$14,667.43", "vendor": "Cedar Ridge Supply", "line_items": [ { "description": "Hydraulic fluid, 5 gal", "amount": "$439.20" } ] }, "metadata": { "invoice_number": { "ocrConfidence": 0.929, "logprobsConfidence": null, "reviewAgentScore": null, "citations": [ { "fileId": "file_ffUqII9mSKKJsgQN1j1qz", "page": { "number": 1, "width": 1240, "height": 1754 }, "referenceText": "Invoice #: INV-9865", "polygon": [ { "x": 78, "y": 468 }, { "x": 310, "y": 467 }, { "x": 310, "y": 495 }, { "x": 78, "y": 496 } ] } ] }, "line_items[0].amount": { "…": "…" } } }, "parseRunId": "pr_q0b8az3CAWPoSbWGUk7RO", "usage": { "credits": 3, "totalCredits": 3, "breakdown": [ { "object": "extract_run", "credits": 3, "charges": [ { "product": "extraction_performance", "unit": "page", "quantity": 1, "credits": 3 } ] } ] } } ``` **Verified live** on 2026-09-11 with `benchmark/datasets/synthetic-v1/docs/invoice_001.png`: `{invoice_number: "INV-9865", total: "$14,667.43", date: "2024-08-03", vendor: "Cedar Ridge Supply"}` on both processors, every field cited on page 1 with `ocrConfidence` 0.93–0.99. A fresh `extraction_performance` run billed `totalCredits: 5` (3 extract + 2 parse); a second run on the already-parsed file billed only the extraction. Test: `providers::extend::tests::live_extract` (`#[ignore]`). ### Extract gotchas * **Plain JSON Schema is a `400`.** The error names one path at a time (*"Path: config.properties.invoice_number"*), so an unadapted schema fails field by field. Note the asymmetry PuffinParse handles: type unions use the **string** `"null"`, enums use a real JSON `null`. * **`metadata` keys are paths, not a tree** — `line_items[0].description`, with an entry for the array and for each item as well as each cell. * **`ocrConfidence` only exists with citations on**; without them a field's entry can be empty. * **Polygons are not rectangles** (four points, often slightly skewed) and are in page pixels, so the page dimensions in the same citation are required to normalise them. * **Extract implicitly bills a parse run** on a file Extend has not parsed yet; `usage.credits` alone understates the job (use `totalCredits`). * Omitting both `config` and `extractor` makes Extend infer a schema. PuffinParse always sends a schema — `extract` is schema-driven by definition — but `provider_options={"config": {"schema": null}}` is not a supported way around that; drop to `provider_options={"extractor": …}` instead. ## 6. Errors, status codes, rate limits, timeouts Extend's error body is `{"code","message","requestId","retryable"}` (sometimes `docUrl`). PuffinParse's `Error::from_http` uses `message` and classifies by HTTP status: | Status | Typical `code` | PuffinParse `ErrorKind` | |---|---|---| | 400 | `INVALID_REQUEST` (missing version header, bad enum with a `Path: config.target` suffix), `INVALID_CONFIG_OPTIONS`, `UNABLE_TO_DOWNLOAD_FILE`, `FILE_TYPE_NOT_SUPPORTED`, `FILE_SIZE_TOO_LARGE` | `bad_request` | | 401 | `UNAUTHORIZED` — "Invalid API key." | `authentication` | | 403 | workspace lacks permission | `authentication` | | 404 | `NOT_FOUND` — unknown file id, or an endpoint that does not exist | `bad_request` | | 410 | `ENDPOINT_REMOVED` — e.g. `GET /parser_runs/{id}` | `bad_request` | | 422 | corrupt / password-protected / conversion failure (sync `/parse` only) | `bad_request` | | 429 | `RATE_LIMIT_EXCEEDED` (`retryable: true`) | `rate_limit` (retried) | | 500 | `INTERNAL_ERROR`, OCR/chunking errors | `provider` (retried) | On `/parse_runs`, only 400/401/403 happen at creation; **processing failures return HTTP 200** with `status: "FAILED"`. PuffinParse maps `failureReason` itself: | `failureReason` | PuffinParse `ErrorKind` | |---|---| | `OCR_ERROR`, `INTERNAL_ERROR` | `provider` | | `OUT_OF_CREDITS` | `authentication` | | anything else (`CORRUPT_FILE`, `PASSWORD_PROTECTED_FILE`, `FILE_TYPE_NOT_SUPPORTED`, `CHUNKING_ERROR`, `FAILED_TO_CONVERT_TO_PDF`, …) | `bad_request` | The error message is `parse run failed: : ` and carries the run id as `job_id`. **Retries.** `max_retries` (default 2), exponential backoff with full jitter, on rate-limit, network and 500/502/503/504 only. **Rate limits.** Per organization, per category (independent GET / WRITE / RUN buckets), plus a throughput cap in files/minute: PAYG 10 req/s and 80 files/min; Scale 25+ and 120+; Enterprise 75+ and 300+. Free-form `429`s are retryable. **No `X-RateLimit-*` headers are returned at all** — reactive 429 handling only, though `Retry-After` is sometimes present. **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline (upload + create + polling + output download) and also caps each individual request. The async run itself has no server-side timeout; Extend's sync `/parse` (which PuffinParse does not use) has a hard 5-minute one. ## 7. Gotchas (verified) * **`x-extend-api-version` is mandatory** for any key created after 2025-04-21 — omitting it is a hard `400 INVALID_REQUEST`. Keys older than that silently fall back to the legacy `2024-12-23` behaviour. PuffinParse always pins `2026-02-09`; the response echoes the same header back. * **Endpoints from the old API are gone.** `POST /parse_async` → `404 NOT_FOUND`; `GET /parser_runs/{id}` → `410 ENDPOINT_REMOVED` ("Use GET /parse_runs/:id instead"). Async is `POST /parse_runs`. * **Enums are lowercase.** `"target": "MARKDOWN"` is a `400`; it must be `"markdown"` (or `"spatial"`). * **The run object *is* the response** on `2026-02-09` — there is no `parserRun`/`success` envelope, and chunks live at `output.chunks`, not `chunks`. * **`blockOptions.tables.targetFormat` defaults to `html`**, which surprises anyone expecting markdown end to end. PuffinParse sets `markdown` explicitly. * **`details` is often the literal empty object `{}`** (text, heading, footer blocks) even though the docs describe a tagged union. Anything reading `details.type` must tolerate its absence. * **`output.metadata.pages` is `null` for raw images** — only PDFs (and files converted to PDF) get page metadata. For images, block `metadata.page.width/height` are raw pixels, so PuffinParse still normalises boxes correctly. * **Undocumented-but-always-present keys**: `dataRetention` on both file and run objects, `chunk.id`, `block.id`. * **Coordinates are in the raster space at `dpi`**, not PDF points: a 596×842 pt page reports 1241×1754 at 150 dpi. Normalising by `metadata.page.width/height` (what PuffinParse does) is correct in either space; converting back to source points needs `originalPageWidth × dpi/72` and the inverse of `rotationApplied`. * **`file.metadata` is `{}` right after upload** even for PDFs; `pageCount` only appears once a run has processed the file. * **`usage` is `null` for runs created before 2025-10-07**, and `totalCredits`/`breakdown` are missing on runs persisted before 2026-05-14 — all three are optional. * **The docs site moved.** `https://docs.extend.ai/2025-04-21/developers/api-reference/…` URLs now 404; current docs live at the root. ## 8. Useful `provider_options` passthrough ```python # 1. Organization-scoped API keys need a workspace; PuffinParse turns this into a header. puffinparse.ocr("doc.pdf", model="extend/parse_performance", provider_options={"workspace_id": "ws_…"}) # 2. Word-level OCR boxes + confidence in the raw payload. puffinparse.ocr("scan.pdf", model="extend/parse_performance", include_raw=True, provider_options={"advancedOptions": {"returnOcr": {"words": True}}}) # 3. Skip figure processing, and let the table agent fix messy tables. puffinparse.ocr("report.pdf", model="extend/parse_performance", provider_options={"blockOptions": {"figures": {"enabled": False}, "tables": {"agentic": {"enabled": True}}}}) # 4. Password-protected PDF by URL, tagged for usage reporting, with no data retained. puffinparse.ocr("https://example.com/locked.pdf", model="extend/parse_light", provider_options={"file": {"settings": {"password": "…"}}, "metadata": {"extend:usage_tags": ["prod"]}, "dataRetention": {"mode": "zero"}}) # 5. Spreadsheets: advanced parsing, hidden content skipped. puffinparse.ocr("book.xlsx", model="extend/parse_performance", provider_options={"advancedOptions": {"excelParsingMode": "advanced", "excelSkipHiddenContent": True}}) ``` ## 9. Links * Docs home: · index: · compact platform context: * Credits and pricing: * API versions in use: `2026-02-09` (current), `2025-04-21`, `2024-12-23`, `2024-11-14`, `2024-07-30`, `2024-02-01` — pin one with `x-extend-api-version`. * Webhooks for `parse_run.processed` / `parse_run.failed` (HMAC-SHA256 over `v0:{timestamp}:{body}`) are the recommended alternative to polling at scale. --- # LlamaParse > **Status: live-verified.** The `#[ignore]`d live tests pass against the real LlamaParse API, > and its fixtures under `crates/puffinparse-core/tests/fixtures/` include redacted live responses. ## 1. Summary | | | |---|---| | Provider name | `llamaparse` (aliases accepted by `ModelRef`: `llama`, `llama_parse`, `llama-parse`, `llamacloud`, `llama_cloud`) | | Base URL | `https://api.cloud.llamaindex.ai` (override: `base_url` on the request, or `LLAMA_BASE_URL`). EU region: `https://api.cloud.eu.llamaindex.ai` | | API key | `LLAMA_API_KEY` (or `api_key` on the request) — sent as `Authorization: Bearer llx-…` | | Docs | (the old `docs.cloud.llamaindex.ai` host 308-redirects here) | | API version | Path-versioned: PuffinParse uses `/api/v1/parsing/*`. Quality is selected by `tier` + a dated `version` (PuffinParse sends `version=latest`) | | Verified | 2026-09-11, live against the North America host | | Implementation | `crates/puffinparse-core/src/providers/llamaparse.rs` | A key is bound to one region; using the wrong host returns `401 "Invalid API Key. Please check your region …"`. The OpenAPI spec is at `/api/openapi.json` (not `/openapi.json`); Swagger UI at `/docs`. ## 2. Models exposed by PuffinParse | Model | Provider parameters PuffinParse sets | List price (`pricing.json`) | |---|---|---| | `llamaparse/fast` | `tier=fast`, `version=latest` | $0.00125 / page (1 credit) | | `llamaparse/cost_effective` *(default)* | `tier=cost_effective`, `version=latest` | $0.00375 / page (3 credits) | | `llamaparse/agentic` | `tier=agentic`, `version=latest` | $0.0125 / page (10 credits) | | `llamaparse/agentic_plus` | `tier=agentic_plus`, `version=latest` | $0.05625 / page (45 credits) | `cost_effective`, `agentic` and `agentic_plus` also serve `extract` (§5); `fast` does not — it is a parse-only tier. Extract prices are per page, on top of the parse the extraction runs: | Model | Extract parameters | List price (`pricing.json`, `extract`) | |---|---|---| | `llamaparse/cost_effective` *(default for `extract`)* | `configuration.tier=cost_effective` | $0.01 / page (5 extract + 3 parse credits) | | `llamaparse/agentic` | `configuration.tier=agentic` | $0.03125 / page (15 + 10 credits) | | `llamaparse/agentic_plus` | `configuration.tier=agentic_plus` | $0.11875 / page (50 + 45 credits) | 1 000 credits = $1.25 ⇒ 1 credit = $0.00125. Prices come from and drive `OcrResponse.cost_usd`; add-ons PuffinParse does not model include `extract_layout` (+3 credits/page) and enriched forms (+10 credits per form page). The live per-tier version list is `GET /api/v2/parse/versions`. ## 3. Request flow PuffinParse uses 1. **Upload / create job — one call.** `POST {base}/api/v1/parsing/upload`, `multipart/form-data`, with `Authorization: Bearer …` and `accept: application/json`. The parts PuffinParse sends are: | Part | Value | |---|---| | `tier` | the model name (`fast` \| `cost_effective` \| `agentic` \| `agentic_plus`) | | `version` | `latest` | | `language` | the request's `language`, when set (single value; the field is repeatable for arrays) | | `target_pages` | from `pages`, **converted 1-based → 0-based**: `"1-3,5"` → `"0-2,4"`; an open range `"10-"` becomes `"9-100000"` | | `file` | the document bytes, with filename and guessed MIME type — **path / bytes input only** | | `input_url` | the URL string — **URL input only**, in place of `file` | Response: `{"id":"","status":"PENDING","error_code":null,"error_message":null}`. 2. **Poll.** `GET {base}/api/v1/parsing/job/{job_id}` starting at 1 s, backing off ×1.5 to a maximum of 5 s. Terminal statuses: `SUCCESS`, `PARTIAL_SUCCESS`, `ERROR`, `CANCELLED` (`PENDING` keeps polling). Both `SUCCESS` and `PARTIAL_SUCCESS` proceed to the result fetch. 3. **Result.** `GET {base}/api/v1/parsing/job/{job_id}/result/json` → `{"pages":[…],"job_metadata":{…}}`. PuffinParse always uses the JSON result (it carries per-page `md`, `text`, `items[]` and bboxes); `/result/markdown`, `/result/text` and the undocumented `/result/raw/markdown` are not used. **Where `provider_options` are merged:** LlamaParse has no JSON body — everything is a multipart form field — so `form_fields()` flattens `provider_options` (which must be a JSON object) into additional text parts. Strings pass through; booleans become `"true"`/`"false"`; numbers are stringified; `null` values are skipped; objects/arrays are serialised as JSON text. A key already present (e.g. `version`, `tier`, `language`, `target_pages`) is **replaced**, so provider options override PuffinParse's defaults. ### Jobs API and webhooks (`submit_parse` / `retrieve_parse`, SPEC §15) * **Submit** — the same `POST {base}/api/v1/parsing/upload`; the job `id` becomes `JobHandle.job_id`. `webhook_url` becomes the multipart field `webhook_url` (LlamaParse requires HTTPS, a domain name rather than an IP, and fewer than 200 characters). * **Retrieve** — one `GET {base}/api/v1/parsing/job/{id}` (retried on 429/5xx); `SUCCESS` / `PARTIAL_SUCCESS` then fetch `result/json` exactly like `parse`; `ERROR` / `CANCELLED` → `JobStatus::Failed` with `error_code error_message` (`INVALID*` → `bad_request`, else `provider`); `PENDING` → `Pending`. * **Webhook bodies** — two shapes reach a handler, and `parse_webhook` reads both: * the v1 `webhook_url` *result push*, `{"txt", "md", "json": [{"page", "text", "md", …}], "images"}`: this **is** the result (normalised through the `result/json` mapping, no geometry unless `items` are present) → `Succeeded`, but it carries no job id; * a LlamaCloud *event* (sent for jobs created with `webhook_configurations`, a v2 feature; that PuffinParse's v1 upload accepts it through `provider_options` is **not verified**), `{"event_id", "event_type": "parse.success", "timestamp", "data": {"job_id"}}`: `parse.pending` → `Pending`; `parse.success` / `partial_success` / `error` / `cancelled` → `Finished` (retrieve for the result or the error message). Events for other products (`extract.*`, …) are rejected. Verify the `LC-Signature` header when a signing secret is set. * Verified live 2026-09-24 (`tests/live_jobs.rs::llamaparse_submit_retrieve_live`, `fast`, 1 page: 3 status checks, ~4.4 s). ## 4. Response mapping (`parse` / `ocr`) | LlamaParse field | PuffinParse unified field | Notes | |---|---|---| | upload/poll `id` | `OcrResponse.provider_job_id` | | | `pages[].page` | `Page.page_number` | Already 1-based. | | `pages[].md` | `Page.markdown` | Trimmed. With `output="text"` the page text is used instead. | | `pages[].text` | `Page.text` | Trimmed; falls back to `markdown_to_text(md)` when blank. | | `pages[].width` / `height` | `Page.width` / `Page.height` | The coordinate space `items[].bBox` lives in. | | `pages[].items[]` | `Page.blocks[]` | In document order. | | `items[].type` (+ `lvl`) | `Block.type` | See mapping below. | | `items[].md` | `Block.content` | With `output="text"`, `items[].value` is preferred, else `markdown_to_text(md)`. | | `items[].value` (string) | `Block.text` | `null` for non-string values. | | `items[].bBox.{x,y,w,h}` | `Block.bbox` = `{x0,y0,x1,y1}` | **Absolute, in page units**; normalised by `pages[].width/height` and clamped to 0–1. No page dims ⇒ `bbox = None`. | | `items[].bBox.confidence` | `Block.confidence` | 0–1. `pages[].confidence` and `layoutAwareBbox[]` are not mapped. | | `job_metadata.job_pages` | `Usage.pages` | Falls back to the number of returned pages. | | `job_metadata.job_credits_usage` (else `credits_used`) | `Usage.credits` | **Only kept when > 0** — see gotchas. | | `job_metadata.job_is_cache_hit == true` | `metadata.llamaparse_cache_hit` | | | job `status == "PARTIAL_SUCCESS"` | `metadata.llamaparse_partial_success` | | | `pages[].images[]`, `charts`, `links`, `layout[]`, `printedPageNumber`, … | — | Not mapped; visible with `include_raw=True`. | Item types: `heading` → `title` when `lvl == 1`, otherwise `section_header`; `text` → `text`; `table` → `table`; `list`/`list_item` → `list`; `figure`/`image`/`chart` → `figure`; `formula`/`equation` → `formula`; `header` → `header`; `footer` → `footer`; everything else → `other`. Trimmed real response (`crates/puffinparse-core/tests/fixtures/llamaparse_result_json.json`; two pages, `images`/`layout`/`charts` and most page keys elided): ```json { "pages": [ { "page": 1, "text": "Hello PuffinParse\n\nInvoice #1234\nTotal: $56.78", "md": "# Hello PuffinParse\n\nInvoice #1234\n\nTotal: $56.78", "items": [ { "type": "heading", "md": "# Hello PuffinParse", "value": "Hello PuffinParse", "lvl": 1, "bBox": { "x": 60.178, "y": 56.553, "w": 308.672, "h": 40.835, "confidence": 0.84, "label": "doc_title" }, "layoutAwareBbox": [ { "x": 60.178, "y": 56.553, "w": 308.672, "h": 40.835, "startIndex": 0, "endIndex": 14, "confidence": 0.84, "label": "doc_title" } ] }, { "type": "text", "md": "Invoice #1234", "value": "Invoice #1234", "bBox": { "x": 59.731, "y": 138.867, "w": 217.678, "h": 28.367, "confidence": 0.75, "label": "paragraph_title" } } ], "status": "OK", "width": 1000, "height": 1300, "confidence": 0.875, "parsingMode": "accurate", "originalOrientationAngle": 0, "pageHeaderMarkdown": "", "pageFooterMarkdown": "", "printedPageNumber": "", "noTextContent": false, "costOptimized": false }, { "page": 2, "md": "## Line Items\n\n| Item | Qty | Price |\n| -------- | --- | ------ |\n| Widget | 2 | $10.00 |", "items": [ { "type": "heading", "md": "## Line Items", "value": "Line Items", "lvl": 2, "bBox": { "x": 60.159, "y": 55.671, "w": 235.113, "h": 44.238, "confidence": 0.84 } }, { "type": "table", "md": "| Item | Qty | Price |\n| -------- | --- | ------ |\n| Widget | 2 | $10.00 |", "html": "…
", "rows": [["Item","Qty","Price"],["Widget","2","$10.00"]], "isPerfectTable": true, "csv": "\"Item\",\"Qty\",\"Price\"\n\"Widget\",\"2\",\"$10.00\"", "bBox": { "x": 57.919, "y": 138.705, "w": 843.873, "h": 242.652, "confidence": 0.99, "label": "table" } } ], "width": 1000, "height": 1300 } ], "job_metadata": { "credits_used": 0, "job_credits_usage": 0, "job_pages": 2, "job_auto_mode_triggered_pages": 0, "job_is_cache_hit": false } } ``` ## 5. Extract mode (`extract`) — LlamaExtract Extraction is a different API surface from parsing: PuffinParse uses LlamaCloud's **v2 extract** endpoints, not the v1 extraction-agent ones, so no agent has to be created and the schema travels with the request. Implementation: `LlamaParse::extract` in `crates/puffinparse-core/src/providers/llamaparse.rs`. 1. **Upload.** `POST {base}/api/v1/beta/files`, `multipart/form-data` with parts `file` and `purpose=extract` → `201` `{"id": "", "name": …, "expires_at": …}` (files expire after 48 h). Unlike the parse path there is **no URL input**: an `https://…` document is downloaded by PuffinParse and re-uploaded. 2. **Create.** `POST {base}/api/v2/extract`: ```json { "file_input": "", "configuration": { "tier": "cost_effective", "data_schema": { "...the request's JSON Schema, verbatim..." }, "cite_sources": true, "confidence_scores": true, "system_prompt": "…only when `instructions` is set…", "target_pages": "1-3,5" } } ``` → `{"id": "ext-…", "status": "PENDING", "configuration": {…resolved defaults…}}`. 3. **Poll.** `GET {base}/api/v2/extract/{id}?expand=usage&expand=extract_metadata`, 1 s backing off ×1.5 to 5 s. **Both `expand` values matter**: without them `usage` and `extract_metadata` come back `null` even when citations were requested. Terminal statuses: `COMPLETED`, `FAILED`, `CANCELLED` (a failure carries `error_message`). Mapping of the unified request: * the schema is passed through **verbatim** (LlamaCloud accepts standard JSON Schema); * `instructions` → `configuration.system_prompt`; * `citations = true` → both `cite_sources` and `confidence_scores` (they are reported together under `extract_metadata`); * `pages` → `configuration.target_pages`, which is **1-based** here — the opposite of the 0-based `target_pages` form field on the parsing endpoint; * `provider_options` are merged into `configuration` key-by-key (`use_reasoning`, `extraction_target`, `parse_tier`, `disable_cache`, `max_pages`, `spreadsheet_mode`, …), and win over PuffinParse's defaults. ### Response mapping | LlamaCloud field | PuffinParse unified field | Notes | |---|---|---| | `extract_result` | `ExtractResponse.data` | The schema shape; no wrapper to strip. | | `extract_metadata.field_metadata.document_metadata` | `ExtractResponse.fields` | A **parallel tree** mirroring the data: objects keyed by field, arrays as lists indexed by position, and `{citation, confidence, extraction_confidence, parsing_confidence}` entries at the leaves. Walked into RFC 6901 pointers (`/sites/0/samples`). | | leaf `confidence` (else `extraction_confidence`) | `FieldInfo.confidence` | 0–1. `parsing_confidence` is not surfaced. | | leaf `citation[].page` | `Citation.page_number` | 1-based. | | leaf `citation[].bounding_boxes[]` (`{x,y,w,h}`) | `Citation.bbox` | Normalised by the same entry's `page_dimensions` (PDF points). One `Citation` per box; a citation with no boxes still yields one with `bbox: None` (the `turbo` tier is text-only). | | leaf `citation[].matching_text` | `Citation.text` | The matched **markdown** span, so it can contain `#`/`|`/`**`. | | `usage.credits` | `Usage.credits` | Total: `extract_credits` + `parse_credits`. | | — | `Usage.pages` | **Derived** — see below. | | `extract_metadata.parse_job_id` / `parse_tier` | `metadata.llamaparse_parse_job_id` / `llamaparse_parse_tier` | The parse defaults to the extract tier. | | `id` | `ExtractResponse.provider_job_id` | `ext-…` | **Page count is derived.** LlamaExtract reports no page count anywhere in the job. PuffinParse uses the highest page number seen in the citations; with citations off it divides `usage.extract_credits` by the tier's published per-page rate (5 / 15 / 50 credits for cost_effective / agentic / agentic_plus, 35 for turbo); failing both it reports 1. So `usage.pages` — and therefore `cost_usd` — is a best-effort figure, exact when citations are on and every page contributes a cited field. Trimmed real response (`crates/puffinparse-core/tests/fixtures/llamaparse_extract_job.json`, a 2-page PDF): ```json { "id": "ext-4mlkajo0l99nprnzarbjmt5x6143", "status": "COMPLETED", "extract_result": { "title": "A Short History of the Harbor", "sites": [ { "site": "Hazel Bend", "samples": 205 } ] }, "extract_metadata": { "field_metadata": { "document_metadata": { "title": { "citation": [ { "page": 1, "matching_text": "# A Short History of the Harbor", "bounding_boxes": [ { "x": 36.86, "y": 38.94, "w": 341.68, "h": 24.56 } ], "page_dimensions": { "width": 595.2, "height": 841.92 } } ], "confidence": 0.9425, "extraction_confidence": 0.9425, "parsing_confidence": 1.0 }, "sites": [ { "site": { "citation": [ { "page": 2, "…": "…" } ], "confidence": 0.9476 }, "samples": { "…": "…" } } ] }, "page_metadata": null, "row_metadata": null }, "parse_job_id": "pjb-12ue2qidroaofurlc4mfqgfgli1j", "parse_tier": "agentic" }, "usage": { "credits": 50.0, "extract_credits": 30.0, "parse_credits": 20.0 } } ``` **Verified live** on 2026-09-11: `invoice_001.png` on `cost_effective` returned `{invoice_number: "INV-9865", total: "$14,667.43", date: "2024-08-03", vendor: "Cedar Ridge Supply"}` with `usage.credits = 8` (5 extract + 3 parse) and every field cited on page 1; the 2-page `multipage_001.pdf` on `agentic` returned 7 table rows with cell-level citations on page 2 and `credits = 50` for 2 pages, matching 15 + 10 credits per page exactly. Tests: `providers::llamaparse::tests::live_extract` (`#[ignore]`) and `normalizes_extract_fixture`. ### Extract gotchas * **`expand` is not optional.** `GET /api/v2/extract/{id}` without `expand=usage&expand=extract_metadata` returns `usage: null` and `extract_metadata: null`, which looks exactly like "citations were not produced". * **`target_pages` flips base** between the two APIs: 0-based on `/api/v1/parsing/upload`, 1-based on the v2 extract configuration. * **No URL input and no page count** — both are handled by PuffinParse (download + re-upload, derived pages). * **`fast` is not an extract tier**; `turbo` exists on the API (35 credits/page, text-only citations) but is not registered as a PuffinParse model, so `llamaparse/turbo` resolves to an unsupported-model error even though the provider code accepts the tier. * **Results are cached**: an identical file + configuration returns in a couple of seconds and may not be billed again. Pass `provider_options={"disable_cache": true}` for benchmarking. * **`matching_text` is markdown**, taken from the parse output rather than the raw page text, and the boxes are the parse block's, not a tight box around the value. * Uploaded files carry `expires_at` (48 h) and `purpose=extract`; a file uploaded for parsing is not reusable here. ## 6. Errors, status codes, rate limits, timeouts Every error is FastAPI-shaped: either `{"detail": "message"}` or, on 422, `{"detail": [ValidationError…]}`. `Error::from_http` picks up `detail` (stringifying the array form) and classifies by status: | Status | Trigger | PuffinParse `ErrorKind` | |---|---|---| | 400 | `tier` without `version`; no input source; malformed job id | `bad_request` | | 401 | bad key (wrong region), or no `Authorization` header (`"Not authenticated"`) | `authentication` | | 404 | unknown job, or `Result for Parsing Job not found. Check job status…` | `bad_request` | | 422 | invalid enum/type in the form body (e.g. an unknown `language`) | `bad_request` | | 429 | rate limited | `rate_limit` (retried) | | 5xx | server error | `provider` (retried) | A job that reaches `ERROR` or `CANCELLED` is turned into an error whose message is `job : ` with the job id attached. The kind is `bad_request` when `error_code` starts with `INVALID` (e.g. `INVALID_TIER_VERSION_COMBINATION`), otherwise `provider`. **Retries.** `max_retries` (default 2), exponential backoff with full jitter, on rate-limit, network and 500/502/503/504 only. **Rate limits.** `POST /api/v1/parsing/upload`: 50 QPS over a 10-second window, per **organization**. `POST /api/v1/beta/files`: 50 QPS over 5 s, per project. Free-tier organizations: 20 requests/minute overall. **No `Retry-After` and no rate-limit headers on any endpoint**, so backoff is blind. **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline (upload + polling + result fetch) and caps each request. For reference, the official SDKs poll every 1 s with a 2 000 s ceiling; a 1-page image on `cost_effective` finished in ~4.3 s in testing. ## 7. Gotchas (verified) * **`tier` requires `version`.** Sending a tier alone is `400 "Must specify a version with a tier. Tier: cost_effective"`. PuffinParse always sends `version=latest`; pin a dated version through `provider_options` when you need reproducibility. * **`tier` is not validated at upload.** An unknown tier returns `200 PENDING` and only fails later with `status: "ERROR"`, `error_code: "INVALID_TIER_VERSION_COMBINATION"`. PuffinParse validates the tier client-side, but a tier overridden via `provider_options` bypasses that check. * **`PARTIAL_SUCCESS` is a real terminal status** (some pages failed within `page_error_tolerance`) and **results are retrievable**. PuffinParse treats it as success and flags `metadata.llamaparse_partial_success = true`. * **Credits are 0 until billing settles.** `credits_used` and `job_credits_usage` were `0` on every observed job — including `agentic` on 2 pages with `job_is_cache_hit: false` — because usage is recorded asynchronously. PuffinParse therefore drops non-positive values and leaves `Usage.credits` as `None`; `cost_usd` comes from the price table instead. For real billing numbers use `GET /api/v1/beta/usage-metrics`. * **Re-parsing the same file within 48 hours is a free cache hit**, which makes latency and cost benchmarks meaningless. `puffinparse bench` therefore injects `{"do_not_cache": true, "invalidate_cache": true}` for every `llamaparse/*` model (`crates/puffinparse-cli/src/bench.rs::cache_busting_options`) unless caching is explicitly allowed. `metadata.llamaparse_cache_hit` surfaces a hit when it happens. * **The docs moved** from `docs.cloud.llamaindex.ai` to `developers.llamaindex.ai` (308), and most old deep links 404. The OpenAPI spec is at `/api/openapi.json`. * **`result_type` is not an API parameter** — it is SDK-only and silently ignored by the server. The result flavour is chosen by which `/result/...` endpoint you call. * **`error_code` / `error_message` are omitted (not null)** on `GET /job/{id}` for successful jobs, but present-and-null on the upload response. * **Three coordinate spaces in one response**: `items[].bBox` and `images[]` are in page units (`pages[].width/height` — PuffinParse normalises against these), `pages[].layout[].bbox` is already normalised 0–1, and `images[].ocr[]` is in that image's own `original_width`×`original_height` pixels. * **Mixed casing.** `bBox`, `layoutAwareBbox`, `isPerfectTable`, `noTextContent`, `originalOrientationAngle` are camelCase while `job_metadata`, `original_width` are snake_case — no blanket rename rule works. * **`target_pages` is 0-based** while PuffinParse's `pages` is 1-based; the conversion happens in `form_fields()`. The document-level markdown is exactly `page_separator.join(pages[].md)` with a default separator of `"\n\n---\n\n"`. * **The `fast` tier degrades tables noticeably** (misaligned columns on a table `agentic` got right). Avoid it for anything structured. Also note `output_tables_as_HTML` (capital HTML) only affects the rendered markdown — `items[].html` is present either way. ## 8. Useful `provider_options` passthrough ```python # 1. Pin a dated parser version instead of `latest` (reproducible output). puffinparse.ocr("doc.pdf", model="llamaparse/cost_effective", provider_options={"version": "2026-08-19"}) # 2. Defeat the 48-hour result cache (what the benchmark does). puffinparse.ocr("doc.pdf", model="llamaparse/agentic", provider_options={"do_not_cache": True, "invalidate_cache": True}) # 3. Layout blocks and a full-page screenshot in the raw payload (+3 credits/page for layout). puffinparse.ocr("scan.png", model="llamaparse/agentic", include_raw=True, provider_options={"extract_layout": True, "take_screenshot": True}) # 4. Prompt steering and table tuning. puffinparse.ocr("statement.pdf", model="llamaparse/agentic_plus", provider_options={"parsing_instruction": "Preserve every table column.", "merge_tables_across_pages_in_markdown": True, "output_tables_as_HTML": True}) # 5. Skip OCR on a digital-native PDF, hide running headers/footers, tolerate bad pages. puffinparse.ocr("contract.pdf", model="llamaparse/fast", provider_options={"disable_ocr": True, "hide_headers": True, "hide_footers": True, "page_error_tolerance": 0.1, "replace_failed_page_mode": "raw_text"}) ``` ## 9. Links * Docs home: * Tiers: * Pricing: · * Rate limits: * Regions: * OpenAPI spec: · Swagger UI: * Live per-tier version list: `GET https://api.cloud.llamaindex.ai/api/v2/parse/versions` * Supported input extensions: `GET /api/v1/parsing/supported_file_extensions` (~130 extensions) --- # Mistral Document AI > **Status: docs-only.** Implemented from Mistral's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `mistral` | | Base URL | `https://api.mistral.ai` (override: `base_url` on the request, or `MISTRAL_BASE_URL`) | | API key | `MISTRAL_API_KEY` (or `api_key` on the request) — sent as `Authorization: Bearer ` | | Docs | · API reference | | API version | Unversioned path prefix `/v1`; no version header. Model version is pinned through the model id (`mistral-ocr-4-1`, …). | | Checked against | 2026-09-11, against the published docs and the machine-readable spec at (no live key in this environment — see §6) | | Implementation | `crates/puffinparse-core/src/providers/mistral.rs` | Mistral's Document AI OCR is a **single synchronous call**: `POST /v1/ocr` returns the whole document's markdown, per-page image boxes and (on OCR 4+) paragraph-level blocks in one response. There is no job queue, no polling and no result indirection, which makes it the simplest provider in PuffinParse. The same endpoint also does schema-driven extraction (`document_annotation_format`), so every `mistral` model serves all three modes — `parse`, `ocr` and `extract` — from the same call. ## 2. Models exposed by PuffinParse | Model | Mistral model id | List price (`pricing.json`) | |---|---|---| | `mistral/ocr-latest` *(default)* | `mistral-ocr-latest` (currently aliases OCR 4.1) | $0.004 / page · $0.005 / annotated page | | `mistral/ocr-4-1` | `mistral-ocr-4-1` (OCR 4.1, GA 2026-07-16) | $0.004 / page · $0.005 / annotated page | | `mistral/ocr-4-0` | `mistral-ocr-4-0` (OCR 4.0, GA 2026-06-23) | $0.004 / page · $0.005 / annotated page | | `mistral/ocr-2512` | `mistral-ocr-2512` (OCR 3, GA 2025-12-18) | $0.002 / page · $0.003 / annotated page | All four serve `parse`, `ocr` and `extract`. Prices are the public per-1 000-page list prices from the model cards (, ): $4 / 1 000 pages and $5 / 1 000 *annotated* pages for OCR 4.x, $2 / $3 for OCR 3. The "annotated page" rate is what `extract` mode bills, which is why the `extract` price in `pricing.json` is higher than `parse`/`ocr` for the same model. `mistral-ocr-2503` and `mistral-ocr-2505` (OCR 1 / OCR 2) are **retired** (2025-12-31 and 2026-05-31) and are deliberately not exposed. Feature availability differs by version and PuffinParse does not paper over it: | Feature | OCR 3 (`ocr-2512`) | OCR 4.0 | OCR 4.1 | |---|---|---|---| | Markdown + image boxes | yes | yes | yes | | `include_blocks` (paragraph blocks + labels) | accepted, returns empty | yes | yes | | `confidence_scores_granularity` block scores | no | no | yes | | `table_format`, `extract_header`, `extract_footer` | yes (OCR 2512+) | yes | yes | | Annotations (`extract`) | yes | yes | yes | ## 3. Request flow PuffinParse uses Every request carries `Authorization: Bearer $MISTRAL_API_KEY`. There is no workspace or version header. 1. **Input reference.** PuffinParse builds the `document` chunk: * **URL input** — passed straight through, and Mistral downloads it server-side. `{"type":"document_url","document_url":"…","document_name":"y.pdf"}`, or `{"type":"image_url","image_url":"…"}` when the URL's extension is an image type. * **Local image ≤ 10 MB** (path or bytes) — inlined as a data URL: `{"type":"image_url","image_url":"data:image/png;base64,…"}`. No upload round-trip, nothing left behind in the workspace's file storage. * **Everything else** (PDFs, DOCX, PPTX, large images) — `POST {base}/v1/files`, `multipart/form-data` with `purpose=ocr` and a `file` part → `{"id":"",…}`, then `GET {base}/v1/files/{id}/url?expiry=1` → `{"url":"https://…blob.core.windows.net/…?sig=…"}`, and that signed URL becomes `document_url`. The expiry is one hour — the shortest the API allows — and it only needs to outlive the OCR call. 2. **OCR.** `POST {base}/v1/ocr` with `Content-Type: application/json`: ```json { "model": "mistral-ocr-latest", "document": { "type": "document_url", "document_url": "…", "document_name": "invoice.pdf" }, "include_image_base64": false } ``` * `pages="1-3,7"` becomes `"pages": [0,1,2,6]` — **PuffinParse's page numbers are 1-based, Mistral's are 0-based**. Open-ended ranges (`"10-"`) are rejected with an `input` error, because the API has no way to express "to the end". * `include_image_base64` is pinned to `false` so responses stay small; set it through `provider_options` when you want the cropped images in `resp.raw`. * `language` is **ignored** — the OCR endpoint has no language parameter (the model is multilingual across 40+ languages and auto-detects). * PuffinParse does not send `include_blocks`; the API defaults it to `true`, so OCR 4.x returns blocks. 3. **Extract.** Same call, plus: ```json { "document_annotation_format": { "type": "json_schema", "json_schema": { "name": "document_annotation", "schema": { …your JSON Schema… }, "strict": true } }, "document_annotation_prompt": "…instructions, when given…" } ``` `strict` is `true` only when your schema already declares `"additionalProperties": false`; Mistral's strict mode requires a closed schema, so an open schema is sent with `strict: false` rather than being rejected by the API. `ExtractRequest.instructions` maps to `document_annotation_prompt`. 4. **Ocr mode** is derived from `parse` by `TextResponse::from_parse` — Mistral has no word- or line-level text endpoint, so `Line`s come from block content and `Word`s carry no geometry. **Where `provider_options` are merged:** the whole object is deep-merged into the request body after PuffinParse's own fields, so its keys are top-level `/v1/ocr` request keys — `include_image_base64`, `image_limit`, `image_min_size`, `table_format`, `extract_header`, `extract_footer`, `include_blocks`, `confidence_scores_granularity`, `bbox_annotation_format`, `document_annotation_prompt`, and even `document`/`pages` if you want to override the resolved input. Nested objects merge key-wise, so `{"document_annotation_format": {"json_schema": {"strict": true}}}` flips just that flag. ## 4. Response mapping | Mistral field | PuffinParse unified field | Notes | |---|---|---| | — | `OcrResponse.provider_job_id` | Never set: the call is synchronous and returns no job id. | | `pages[].index` | `Page.page_number` | **0-based on the wire**, `page_number = index + 1`. | | `pages[].markdown` | `Page.markdown` | Used verbatim; `output="text"` runs it through `markdown_to_text`. Whole-document `markdown` is the pages joined by a blank line. | | `pages[].dimensions.{width,height}` | `Page.width` / `Page.height` | Pixels of the page screenshot at `dimensions.dpi` (typically 200), not PDF points. | | `pages[].dimensions.dpi` | `metadata.mistral_dpi` | From the first page. | | `pages[].blocks[]` | `Page.blocks[]` | Present on OCR 4+ (`include_blocks` defaults to `true`), in reading order. | | `blocks[].type` | `Block.type` | See mapping below. | | `blocks[].content` | `Block.content` | Markdown (or HTML for tables when `table_format: "html"`). | | `blocks[].{top_left_x,top_left_y,bottom_right_x,bottom_right_y}` | `Block.bbox` | Absolute pixels → divided by `dimensions.width`/`height`, origin top-left. No dimensions ⇒ `None`. | | `blocks[].confidence_scores.average_content_confidence_score` | `Block.confidence` | Only populated when `confidence_scores_granularity: "block"` is requested (OCR 4.1). | | `pages[].images[]` | `Page.blocks[]` (`figure`) | **Fallback only**, when `blocks` is absent/empty: one `figure` block per image with `content` = the markdown placeholder `![img-0.jpeg](https://github.com/ajinkyashejul/puffinparse/blob/main/docs/providers/img-0.jpeg)` and the image's box. | | `pages[].markdown` | `Page.blocks[0]` (`text`) | **Fallback only**: a single `text` block per page carrying the page markdown, with no bbox. | | `usage_info.pages_processed` | `Usage.pages` | Falls back to the number of returned pages when 0/absent. | | `usage_info.doc_size_bytes` | `metadata.mistral_doc_size_bytes` | | | `model` | `metadata.mistral_model` | The concrete model that served the call (`mistral-ocr-4-1` even when you asked for `-latest`). | | `document_annotation` | `ExtractResponse.data` | A JSON **string**, parsed into `data`. Missing or unparseable ⇒ `provider` error. | | — | `ExtractResponse.fields` | Always empty: Mistral returns no per-field confidence or citations. | | — | `Usage.credits`, `Usage.provider_cost_usd` | Never set; `cost_usd` comes from `pricing.json`. | Block types: `text`, `aside_text` → `text`; `title` → `title`; `list` → `list`; `table` → `table`; `image` → `figure`; `equation` → `formula`; `caption` → `caption`; `header` → `header`; `footer` → `footer`; `code`, `references`, `signature` and anything unknown → `other`. Trimmed response (`crates/puffinparse-core/tests/fixtures/mistral_ocr.json`, second page only, base64 redacted): ```json { "pages": [ { "index": 1, "markdown": "![img-0.jpeg](img-0.jpeg)\n\nFigure 1: Quarterly revenue by segment.\n\nReference: ABC-9876", "images": [ { "id": "img-0.jpeg", "top_left_x": 292, "top_left_y": 217, "bottom_right_x": 1405, "bottom_right_y": 649, "image_base64": "data:image/jpeg;base64,REDACTED", "image_annotation": null } ], "tables": [], "hyperlinks": [], "header": null, "footer": null, "dimensions": { "dpi": 200, "height": 2200, "width": 1700 }, "confidence_scores": null, "blocks": null } ], "model": "mistral-ocr-latest", "document_annotation": null, "usage_info": { "pages_processed": 2, "doc_size_bytes": 30021 } } ``` `crates/puffinparse-core/tests/fixtures/mistral_ocr_blocks.json` covers the OCR 4.x `blocks` payload (labels, boxes, block confidence) and `mistral_annotation.json` the `document_annotation` string. ## 5. Errors, status codes, rate limits, timeouts Every non-2xx body goes through `Error::from_http`, which pulls a message out of `message` / `detail` / `error` and classifies by status: | Status | Mistral body | PuffinParse `ErrorKind` | |---|---|---| | 400 | `{"object":"error","message":"…","type":"invalid_request_error","param":null,"code":null}` — unreachable `document_url`, unsupported file type, bad page index | `bad_request` | | 401 | `{"message":"Unauthorized","request_id":"…"}` — missing/invalid key | `authentication` | | 403 | Key lacks access to the model (e.g. a Premier model on a free workspace) | `authentication` | | 404 | Unknown `file_id` on `/v1/files/{id}/url` | `bad_request` | | 422 | FastAPI validation array: `{"detail":[{"loc":["body","document"],"msg":"…","type":"…"}]}` — the whole array is kept as the message | `bad_request` | | 429 | Requests-per-second or tokens-per-minute limit | `rate_limit` (retried) | | 500 | `{"object":"error","message":"Internal Server Error",…}` | `provider` (retried) | | 502/503/504 | Gateway / capacity | `provider` (retried) | **Retries.** `max_retries` (default 2) with exponential backoff and full jitter, on rate-limit, network and 500/502/503/504 only. 4xx is never retried. The retry wraps each of the three calls (upload, signed URL, OCR) independently. **Rate limits.** Enforced per workspace as requests/second *and* tokens/minute, with a monthly token cap; the limits in force are shown at Admin ▸ API ▸ Limits. Mistral does return an `X-RateLimit-Remaining` header — the only provider in PuffinParse that does — but PuffinParse does not read it today, so backoff is still blind. Free-tier workspaces have the lowest limits. **Timeouts.** `timeout_secs` (default 300) is a whole-call deadline covering upload, signed-URL retrieval and the OCR call, and also caps each individual HTTP request. Because `/v1/ocr` is synchronous, a 1 000-page PDF is one long request — raise `timeout_secs` rather than expecting a job id, or use Mistral's Batch API directly for bulk work. ## 6. Gotchas * **No live-key verification.** Unlike the other provider pages, this one was written from the published docs and the OpenAPI spec (`https://docs.mistral.ai/openapi.yaml`), not from live traffic; the fixtures are built from the documented response examples. Treat exact field-by-field behaviour as "documented", not "observed", until the `#[ignore]`d live tests in `providers/mistral.rs` (`mistral_live_parse`, `mistral_live_extract`) are run with a key. * **`pages` is 0-based.** The request array, and `pages[].index` in the response. PuffinParse converts in both directions; if you pass `pages` through `provider_options` yourself, you own the conversion. * **The API reference's example response shows `"index": 1` for the first page**, contradicting the schema ("The page index in a pdf document starting from 0") and the `pages` parameter description. PuffinParse trusts the schema: `page_number = index + 1`. * **50 MB / 1 000 pages per document.** Larger files are rejected. The files API itself accepts up to 512 MB, so the OCR limit is what binds. * **`include_blocks` defaults to `true`** (per the OpenAPI spec) but OCR 3 and older *accept the parameter and return an empty array*. That is why PuffinParse keeps the "one text block + one figure block per image" fallback: with `mistral/ocr-2512` you get exactly that, with `ocr-4-x` you get real paragraph blocks. Block counts therefore differ sharply between models. * **Confidence is opt-in.** No `confidence_scores_granularity` ⇒ every `Block.confidence` is `None`. Ask for `"block"` (OCR 4.1) to populate it; `"word"` adds a large `word_confidence_scores` array that PuffinParse does not surface (visible in `resp.raw` with `include_raw=True`). * **Images and tables appear in the markdown as placeholders** — `![img-0.jpeg](https://github.com/ajinkyashejul/puffinparse/blob/main/docs/providers/img-0.jpeg)` and, when `table_format` is set, `[tbl-3.html](https://github.com/ajinkyashejul/puffinparse/blob/main/docs/providers/tbl-3.html)`. The bytes/HTML live in `pages[].images[]` and `pages[].tables[]`. PuffinParse leaves the placeholders in the markdown; with the default `table_format: null` tables are inlined as markdown and no placeholder appears, which is why PuffinParse does not set `table_format`. * **`document_annotation` is a JSON string, not an object** — double-encoded inside the response. It is `null` when no `document_annotation_format` was sent. * **Document annotation only sees the first eight image bounding boxes** (the OCR markdown plus those images is what the vision model is shown), so `extract` is best on text-heavy documents. There is no documented page cap, but the vendor's own examples restrict `pages` to the first eight. * **Strict schemas.** Mistral's `strict: true` follows the OpenAI convention: the schema must be closed (`additionalProperties: false`, every property `required`). PuffinParse only claims strictness when your schema already says so — otherwise the model is asked to follow the schema best-effort. * **`extract` bills the "annotated page" rate** ($5 / 1 000 pages on OCR 4.x, $3 on OCR 3), and it runs a vision LLM *after* OCR, so it is markedly slower than `parse` on the same document. * **No citations.** `ExtractResponse.fields` is always empty. Requesting `citations=True` sets `metadata.mistral_citations_unsupported = true` instead of failing. * **Signed URLs are public-ish.** `GET /v1/files/{id}/url` returns an unauthenticated blob URL; PuffinParse requests the minimum one-hour expiry. Uploaded files stay in the workspace until deleted — PuffinParse does not delete them (a `DELETE /v1/files/{id}` sweep would race with retries). * **`{"type":"file","file_id":""}` is also a valid `document`** (the spec's `FileChunk`), which would skip the signed-URL hop. PuffinParse uses the signed URL because that is the flow the guides document and exercise; pass a `FileChunk` yourself through `provider_options={"document": {...}}` if you already have a `file_id`. * **`bbox_annotation_format` is not wired into a PuffinParse mode.** It annotates individual figures rather than the document, so it only makes sense as a passthrough (see §7); the annotations come back in `pages[].images[].image_annotation` and are visible with `include_raw=True`. ## 7. Useful `provider_options` passthrough ```python # 1. Real paragraph blocks with per-block confidence (OCR 4.1 only). puffinparse.ocr("scan.pdf", model="mistral/ocr-4-1", provider_options={"confidence_scores_granularity": "block"}) # 2. Tables as separate HTML, and headers/footers split out of the body text. puffinparse.ocr("report.pdf", model="mistral/ocr-latest", include_raw=True, provider_options={"table_format": "html", "extract_header": True, "extract_footer": True}) # 3. Keep the cropped images (base64) in resp.raw, ignoring anything smaller than 200 px. puffinparse.ocr("figures.pdf", model="mistral/ocr-latest", include_raw=True, provider_options={"include_image_base64": True, "image_min_size": 200, "image_limit": 20}) # 4. Structured extraction with a prompt (instructions → document_annotation_prompt). puffinparse.extract("invoice.pdf", model="mistral/ocr-latest", schema=INVOICE_SCHEMA, instructions="Amounts are in EUR; ignore the shipping address.") # 5. Caption every figure while parsing (bbox annotations land in resp.raw). puffinparse.ocr("paper.pdf", model="mistral/ocr-latest", include_raw=True, provider_options={"include_image_base64": True, "bbox_annotation_format": { "type": "json_schema", "json_schema": {"name": "bbox_annotation", "strict": True, "schema": { "type": "object", "additionalProperties": False, "properties": {"image_type": {"type": "string"}, "summary": {"type": "string"}}, "required": ["image_type", "summary"]}}}}) ``` ## 8. Links * OCR processor guide: * Annotations guide: * API reference (`POST /v1/ocr`): · full spec: * Files API: (tag `files`) — `POST /v1/files`, `GET /v1/files/{id}/url` * Model cards: · · * Pricing: * Batch API (bulk OCR at 50% off): --- # Azure Document Intelligence > **Status: docs-only.** Implemented from Azure's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `azure` | | Base URL | **none built in** — Azure endpoints are per resource. `AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT` (or `base_url` on the request), e.g. `https://.cognitiveservices.azure.com` or `https://.api.cognitive.microsoft.com`. `base_url` *is* the endpoint. | | API key | `AZURE_DOCUMENT_INTELLIGENCE_KEY` (or `api_key` on the request) — sent as the `Ocp-Apim-Subscription-Key` header | | Docs | | | API version | **`2024-11-30`** (v4.0 GA), pinned by PuffinParse as the `api-version` query parameter (`azure::API_VERSION`) | | Checked against | 2026-09-11 — **against the published REST reference only**; no Azure resource was available when this provider was written, so the fixtures are built from the documented response schema, not captured from a live call. The `#[ignore]`d live tests in `providers/azure.rs` are the check to run once a key exists. | | Implementation | `crates/puffinparse-core/src/providers/azure.rs` | | Modes | `parse`, `ocr`, `extract` | Azure is the only provider so far that serves all three PuffinParse modes: `prebuilt-layout` returns markdown plus paragraphs/tables with polygons, `prebuilt-read` is a cheap native OCR endpoint with words, lines and confidences, and the `prebuilt-*` extraction models return typed fields with per-field confidence and bounding regions. Every model is reached through **one** endpoint shape, so the whole provider is a single request + poll loop. ## 2. Models exposed by PuffinParse | Model | Azure `modelId` | Modes | List price (`pricing.json`) | |---|---|---|---| | `azure/read` *(default for `ocr`)* | `prebuilt-read` | `ocr` | $0.0015 / page ($1.50 / 1 000) | | `azure/layout` *(default for `parse`)* | `prebuilt-layout` | `parse`, `ocr` | $0.01 / page ($10 / 1 000) | | `azure/invoice` *(default for `extract`)* | `prebuilt-invoice` | `extract` | $0.01 / page | | `azure/receipt` | `prebuilt-receipt` | `extract` | $0.01 / page | | `azure/id_document` | `prebuilt-idDocument` | `extract` | $0.01 / page | | `azure/tax_us_w2` | `prebuilt-tax.us.w2` | `extract` | $0.01 / page | | `azure/custom` | `provider_options.model_id` (required) | `parse`, `ocr`, `extract` | $0.03 / page ($30 / 1 000, custom extraction) | Prices are the public S0 pay-as-you-go rates from (the page renders its numbers client-side; the values above were cross-checked against Microsoft's retail-price listings in September 2026). Not modelled by `cost_usd`: * **Volume tiers** — Read drops to $0.60 / 1 000 pages above 1 M pages/month; custom extraction drops to $20 / 1 000. Commitment tiers go lower still. * **Add-on features** — `ocrHighResolution`, `formulas`, `barcodes`, `styleFont`, `languages` bill ~$6 / 1 000 pages *each* on top of the model, and `queryFields` ~$10 / 1 000. Turning them on through `provider_options` makes the real bill higher than `cost_usd`. * **Free tier (F0)**: 500 pages/month, but it analyses **only the first two pages** of any request — a silent truncation, not an error. `provider_options.model_id` overrides the `modelId` for *any* model name, which is also how you reach prebuilt models PuffinParse does not list (`prebuilt-document`, `prebuilt-tax.us.1098`, `prebuilt-healthInsuranceCard.us`, `prebuilt-contract`, …). Pricing then falls back to the price of the PuffinParse model you named, so pick `azure/custom` for custom models and `azure/invoice` for other prebuilt extraction models to keep the cost estimate honest. ## 3. Request flow PuffinParse uses Every request carries `Ocp-Apim-Subscription-Key: $AZURE_DOCUMENT_INTELLIGENCE_KEY`. 1. **Submit.** One call, no separate upload endpoint: ``` POST {endpoint}/documentintelligence/documentModels/{modelId}:analyze ?api-version=2024-11-30 &stringIndexType=unicodeCodePoint &outputContentFormat=markdown # parse only; ocr/extract use `text` [&pages=1-3,7][&locale=en-US][&features=…][&queryFields=…][&output=…] Content-Type: application/json {"urlSource": "https://…"} # URL input, Azure downloads it {"base64Source": "JVBERi0…"} # path / bytes input, inlined as base64 ``` * `pages="1-3,7,10-"` → `pages=1-3,7,10-2000`: Azure's grammar (`^(\d+(-\d+)?)(,\s*(\d+(-\d+)?))*$`) has no open-ended range, so PuffinParse closes it at the S0 page ceiling. * `language` → `locale`. * `stringIndexType=unicodeCodePoint` is deliberate: Azure's default is `textElements` (grapheme clusters), and PuffinParse slices `content` with Rust `char` indices, which *is* code points. * Query parameters are percent-encoded by hand — `reqwest`'s `query` feature is not enabled in this workspace. 2. **`202 Accepted`** with an empty body, an `Operation-Location` header (`{endpoint}/documentintelligence/documentModels/{modelId}/analyzeResults/{resultId}?api-version=…`) and usually `Retry-After: 1`. `{resultId}` becomes `provider_job_id`. 3. **Poll.** `GET {Operation-Location}` with the same key header. PuffinParse sleeps for `Retry-After` first, then polls starting at 2 s and backing off ×1.5 to a maximum of 10 s (Microsoft asks for at most one GET every 2 s per analyze request). `status` walks `notStarted` → `running` → `succeeded` | `failed`. Each poll is itself retried on 429 / 5xx / network errors, so a throttled poll does not fail the call. 4. **Result.** The succeeded payload carries the whole result inline under `analyzeResult`; there is no second fetch and no presigned URL indirection. **`provider_options`** map to query parameters (Azure's request body only holds the document): | Option | Effect | |---|---| | `model_id` | replaces the `modelId` path segment (required for `azure/custom`) | | `features` | `features=` — `ocrHighResolution`, `languages`, `barcodes`, `formulas`, `keyValuePairs`, `styleFont`, `queryFields` | | `query_fields` | `queryFields=` (and adds `queryFields` to `features` automatically) | | `output` | `output=` — `pdf` (searchable PDF) or `figures` (cropped figure images) | | `locale` | `locale=`, overriding `language` | | `output_content_format` | `markdown` or `text`, overriding the per-mode default | | `string_index_type` | overrides `unicodeCodePoint` — **only** do this if you also stop reading `Page.markdown` (see §6) | | `api_version` | pins a different `api-version` | Lists may be given as a JSON array or a single string. ## 4. Response mapping ### 4.1 `parse` (`prebuilt-layout`) | Azure field | PuffinParse unified field | Notes | |---|---|---| | `{resultId}` from `Operation-Location` | `ParseResponse.provider_job_id` | | | `analyzeResult.content` sliced by `pages[].spans` | `Page.markdown` | The page's slice of the document-level markdown, with the `` marker stripped and the edges trimmed. Offsets are code points. With no spans (or no content) the page markdown falls back to its blocks joined by blank lines. | | — | `Page.text` | Derived from `Page.markdown`: `` / `PageFooter` / `PageNumber` comments are unwrapped to their text, other HTML comments dropped, then `markdown_to_text` flattens headings and the HTML table. | | `pages[].pageNumber` | `Page.page_number` | 1-based, as Azure reports it. | | `pages[].width` / `height` | `Page.width` / `height` | **In `pages[].unit`: `inch` for PDFs, `pixel` for images.** PuffinParse passes the numbers through unchanged, so a PDF page is `8.5 × 11`, not `612 × 792`. | | `paragraphs[]` | `Page.blocks[]` | One block per paragraph, ordered by span offset within the page. | | `paragraphs[].role` | `Block.type` | `title`→`title`, `sectionHeading`→`section_header`, `pageHeader`→`header`, `pageFooter`→`footer`, `footnote`→`footnote`, `formulaBlock`→`formula`, `pageNumber`→`other`, no role→`text`. | | `paragraphs[].content` | `Block.content` | Markdown; `output="text"` runs it through `markdown_to_text`. | | `paragraphs[].boundingRegions[0]` | `Block.page_number`, `Block.bbox` | First region only. Multi-page paragraphs keep their first page. | | `pages[].words[].confidence` | `Block.confidence` | Mean confidence of the words whose span falls inside the block's spans; `None` when the input has no words (Office/HTML). | | `tables[]` | `Page.blocks[]` of type `table` | Cells are re-rendered as a **markdown** table (`| Item | Amount |` …) from `rowIndex`/`columnIndex`, with `columnHeader`/`stubHead` cells in row 0 becoming the header row and a `caption` prepended. Merged cells are placed at their origin cell; `rowSpan`/`columnSpan` are not replicated. | | `pages[].{angle, selectionMarks}`, `styles`, `languages`, `sections`, `figures`, `keyValuePairs` | — | Not mapped in `parse`; visible with `include_raw=True`. | | `pages.len()` | `Usage.pages` | Azure never reports credits or dollars, so `credits` and `provider_cost_usd` are `None` and `cost_usd` is `pages × per_page_usd`. | | `analyzeResult.contentFormat` | `metadata.azure_content_format` | | | resolved `modelId` | `metadata.azure_model_id` | Useful with `provider_options.model_id`. | **Bounding boxes.** Azure gives `polygon: [x1,y1,x2,y2,x3,y3,x4,y4]` — a *quadrilateral* that follows the text rotation. PuffinParse takes its axis-aligned hull (min/max of the x and y components) and divides by the page `width`/`height`, so `BBox` stays the unified 0–1, top-left-origin rectangle. The polygon unit and the page unit are always the same, so no unit conversion is needed. **Table paragraph de-duplication.** Azure emits table cell text *both* in `tables[].cells[]` and as ordinary `paragraphs[]`. PuffinParse drops paragraphs whose span starts inside a table's span, so cell text appears exactly once — in the table block. ### 4.2 `ocr` (`prebuilt-read`, also `prebuilt-layout`) `Provider::ocr` is overridden, so this is a native mapping and not derived from `parse`: | Azure field | PuffinParse unified field | |---|---| | `pages[].lines[].content` | `Line.text` | | `pages[].lines[].polygon` | `Line.bbox` (normalised) | | mean of the line's words' `confidence` | `Line.confidence` | | `pages[].words[].content` / `polygon` / `confidence` | `Word.text` / `bbox` / `confidence` | | lines joined by `\n` (or the page's `content` slice when there are no lines) | `TextPage.text` | | `pages.len()` | `Usage.pages` | `prebuilt-read` is 6–7× cheaper than layout and returns the same word geometry, which is why it is the `ocr` default; use `azure/layout` for `ocr` only when you want layout and text from one call. ### 4.3 `extract` (`prebuilt-*`, custom models) `documents[0].fields` is flattened into `ExtractResponse.data`: * Scalars come from `valueString` / `valueNumber` / `valueInteger` / `valueBoolean` / `valueDate` / `valueTime` / `valuePhoneNumber` / `valueCountryRegion` / `valueSelectionMark` / `valueSignature`. * Composite values are passed through as objects: `valueCurrency` → `{"amount": 56.78, "currencyCode": "USD", "currencySymbol": "$"}`, `valueAddress` → Azure's address object. * `valueArray` / `valueObject` recurse, so `Items[0].Description` is a normal nested JSON value. * A field with no typed value falls back to its `content` string; a field with neither is `null`. * `ExtractResponse.fields` is keyed by JSON pointer (`/InvoiceTotal`, `/Items/0/Amount`) and carries `confidence` plus `citations` built from `boundingRegions` (page number + normalised box) with the field's `content` as the citation text. * `documents[0].docType` → `metadata.azure_doc_type`; `documents[0].confidence` is not surfaced (it is the document-type confidence, not a field confidence) — read it from `raw`. * **Fallback**: if the model returned no `documents` but the response has `keyValuePairs` (a custom model, or `prebuilt-layout` with `features=keyValuePairs`), each pair becomes `data[key.content] = value.content` with the pair's confidence and the value's box. With neither, the call fails with a `provider` error. **The request schema does not change what Azure extracts.** Prebuilt models have a fixed field set (see the per-model field tables in the Azure docs) and custom models have the schema you trained. PuffinParse therefore uses `ExtractRequest.schema` only to **select and rename**: each key of `schema.properties` is matched against the returned field names ignoring case and non-alphanumeric characters (`invoice_total` ≡ `InvoiceTotal`), and matches are emitted under the *schema's* spelling, with `fields` pointers renamed to match. If nothing matches, the full Azure field set is returned unchanged. `metadata.azure_schema_selected_fields` says which of the two happened. `instructions` is ignored; to ask for a field Azure does not model, use `provider_options.query_fields` (billed separately). ### 4.4 Trimmed fixture `crates/puffinparse-core/tests/fixtures/azure_layout.json` (2 pages, 12 paragraphs, 1 table, 20 words; the three `azure_*.json` fixtures follow the documented `2024-11-30` schema exactly and drive the unit tests): ```json { "status": "succeeded", "createdDateTime": "2026-09-11T07:21:04Z", "lastUpdatedDateTime": "2026-09-11T07:21:09Z", "analyzeResult": { "apiVersion": "2024-11-30", "modelId": "prebuilt-layout", "stringIndexType": "unicodeCodePoint", "contentFormat": "markdown", "content": "\n\n# Hello PuffinParse\n\nInvoice #1234\n\n…\n\n\n\n\n
ItemAmount
Widget$56.78
\n\n\n\n\n\n\n## Page Two\n\nReference: ABC-9876", "pages": [ { "pageNumber": 1, "angle": 0.0, "width": 8.5, "height": 11.0, "unit": "inch", "words": [ { "content": "Hello", "polygon": [1.0, 1.0, 1.8, 1.0, 1.8, 1.5, 1.0, 1.5], "span": { "offset": 40, "length": 5 }, "confidence": 0.99 } ], "lines": [ { "content": "Hello PuffinParse", "polygon": [1.0, 1.0, 3.5, 1.0, 3.5, 1.3, 1.0, 1.3], "spans": [ { "offset": 40, "length": 13 } ] } ], "selectionMarks": [], "spans": [ { "offset": 0, "length": 249 } ] } ], "paragraphs": [ { "spans": [ { "offset": 17, "length": 14 } ], "boundingRegions": [ { "pageNumber": 1, "polygon": [1.0, 0.5, 3.5, 0.5, 3.5, 1.0, 1.0, 1.0] } ], "content": "PuffinParse sample", "role": "pageHeader" }, { "spans": [ { "offset": 40, "length": 13 } ], "boundingRegions": [ { "pageNumber": 1, "polygon": [1.0, 1.0, 3.5, 1.0, 3.5, 1.5, 1.0, 1.5] } ], "role": "title", "content": "Hello PuffinParse" } ], "tables": [ { "rowCount": 2, "columnCount": 2, "cells": [ { "kind": "columnHeader", "rowIndex": 0, "columnIndex": 0, "content": "Item", "boundingRegions": [ { "pageNumber": 1, "polygon": [1.0, 4.0, 2.8, 4.0, 2.8, 4.3, 1.0, 4.3] } ], "spans": [ { "offset": 104, "length": 4 } ] } ], "boundingRegions": [ { "pageNumber": 1, "polygon": [1.0, 3.9, 4.8, 3.9, 4.8, 4.8, 1.0, 4.8] } ], "spans": [ { "offset": 104, "length": 94 } ] } ], "figures": [], "sections": [ { "spans": [], "elements": ["/paragraphs/0"] } ], "styles": [] } } ``` ## 5. Errors, status codes, rate limits, timeouts Azure's error body is `{"error": {"code", "message", "target"?, "details"?, "innererror"?}}`, which `Error::from_http` reads out of the box (it picks up the nested `message`). | Status | Typical `error.code` | PuffinParse `ErrorKind` | |---|---|---| | 400 | `InvalidRequest` (+ `innererror.code` such as `InvalidContent`, `InvalidContentDimensions`, `InvalidArgument`), `NotSupportedApiVersion` | `bad_request` | | 401 | `401` / `Unauthorized` — "Access denied due to invalid subscription key or wrong API endpoint." | `authentication` | | 403 | `PermissionDenied`, disabled resource, network rules | `authentication` | | 404 | `ModelNotFound` — unknown `modelId`, or the wrong endpoint/region for that custom model | `bad_request` | | 405/415 | wrong verb or `Content-Type` | `bad_request` | | 429 | `429` — over the TPS quota (`Retry-After` is usually present) | `rate_limit` (retried) | | 500/503 | `InternalServerError`, `ServiceUnavailable` | `provider` (retried) | Failures **after** the 202 come back as `200 OK` with `status: "failed"` and the same error object. PuffinParse maps those itself: codes starting with `Invalid` / `Unsupported` / `NotSupported`, plus `ContentSourceNotAccessible`, `ContentSourceTimeout` and `ContentSourceSizeExceeded`, become `bad_request`; anything else is a `provider` error. The message is `analysis failed: : (: )`, and the error carries the `resultId` as `job_id`. **Missing configuration.** No key ⇒ `authentication` ("set `AZURE_DOCUMENT_INTELLIGENCE_KEY`"); no endpoint ⇒ `authentication` ("set `AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT` … or pass `base_url`") — both before any network call. `azure/custom` without `provider_options.model_id` ⇒ `unsupported_model`. **Rate limits (S0 defaults, adjustable by support ticket):** 15 analyze transactions/second, 50 GET operations/second, 5 model-management/second, 10 list/second. Free F0 is 1/second for each. There are no `X-RateLimit-*` headers; 429 carries `Retry-After` and is retried with PuffinParse's own jittered backoff. **Service limits:** 500 MB and 2 000 pages per document on S0 (4 MB / 2 pages on F0); max 500 MB of JSON response. PDF, JPEG/JPG, PNG, BMP, TIFF and HEIF work with every model; DOCX/PPTX/XLS(X) and HTML only with Read and Layout. **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline — submit, `Retry-After` sleep, every poll and the result download — and also caps each individual HTTP request. Azure imposes no server-side ceiling on how long an analysis may run; a 2 000-page PDF can take several minutes, so raise `timeout_secs` for large documents. ## 6. Gotchas (from the REST reference; re-verify the ⚠ ones against a live resource) * **The endpoint is not a constant.** Every Document Intelligence resource has its own host, so `base_url` is mandatory in the same sense an API key is. PuffinParse raises an authentication error rather than guessing a region. Keep the `/` -free form: `https://.cognitiveservices.azure.com`. * **Two auth schemes, one supported.** Azure also accepts Microsoft Entra ID bearer tokens (`Authorization: Bearer`, scope `https://cognitiveservices.azure.com/.default`). PuffinParse only sends `Ocp-Apim-Subscription-Key`; a managed identity / AAD setup is not reachable through `api_key`. * **Markdown tables are HTML.** In `2024-11-30` the markdown `content` renders tables as `
…` (to express merged cells), *not* as pipe tables, and selection marks as ☒ / ☐ rather than `:selected:`. `Page.markdown` therefore contains HTML fragments, while `Block.content` for a table block is a pipe table PuffinParse rendered from `tables[].cells`. `Page.text` flattens both. * **Page headers, footers and page numbers are HTML comments** in markdown (``), and pages are separated by ``. PuffinParse splits pages by `pages[].spans` (exact) and only strips the page-break marker; the comments stay in `Page.markdown`, while `Page.text` unwraps `PageHeader`/`PageFooter`/`PageNumber` to their text. * ⚠ **Offsets depend on `stringIndexType`.** PuffinParse pins `unicodeCodePoint` so `content` can be sliced with `char` indices. Overriding it with `textElements` (Azure's default) or `utf16CodeUnit` will mis-slice `Page.markdown` for documents containing emoji, combining marks or non-BMP script — everything else (blocks, boxes, fields) is unaffected. * **Page dimensions are in inches for PDFs.** `Page.width = 8.5` is not a bug; check `pages[].unit` in `raw` if you need to know which. Normalised `BBox` values are unaffected. * **Office and HTML inputs have no geometry.** For DOCX/PPTX/XLS(X)/HTML, v4.0 reports no `angle`, no `width`/`height`/`unit`, no polygons and no `lines`. Blocks then carry `bbox: None`, pages carry `width: None`, and `ocr` mode falls back to the page's slice of `content` with an empty `lines` list. Word/HTML pages are counted in blocks of 3 000 characters, XLSX per worksheet, PPTX per slide. * **Figures are not blocks.** `figures[]` (with `output=figures`, croppable via `/analyzeResults/{resultId}/figures/{figureId}`) is not mapped; the figure's caption usually also appears as a paragraph, so the text is not lost. Read `raw` for figure geometry. * **`prebuilt-layout` is the only model with roles.** `prebuilt-read` returns paragraphs without `role` (everything maps to `text`) and no `tables`, which is why `read` is registered for `ocr` only. * ⚠ **F0 silently truncates to 2 pages.** A 10-page document on the free tier returns `usage.pages = 2` and a 2-page result with no warning anywhere in the payload. * **Add-on features cost extra per page** and are off by default; `ocrHighResolution` also makes the analysis noticeably slower. `keyValuePairs` is the replacement for the retired `prebuilt-document` model. * **Analyze is always async**, even for a one-page PNG: there is no synchronous endpoint, so the minimum latency is one POST plus one GET. * **`Retry-After` is honoured once**, before the first poll; afterwards PuffinParse uses its own 2 s → 10 s backoff. A throttled poll (429) is retried inside the poll loop instead of failing the call. ## 7. Useful `provider_options` passthrough ```python # 1. Custom (trained) model — `model_id` is required for azure/custom and overrides any model name. puffinparse.extract("po.pdf", model="azure/custom", schema=schema, provider_options={"model_id": "purchase-orders-v3"}) # 2. Fine print / low-quality scans: high-resolution OCR (+$6 / 1k pages). puffinparse.parse("fine-print.pdf", model="azure/layout", provider_options={"features": ["ocrHighResolution"]}) # 3. Formulas as LaTeX and barcodes as markdown images, plus a searchable PDF of the result. puffinparse.parse("paper.pdf", model="azure/layout", provider_options={"features": ["formulas", "barcodes"], "output": ["pdf"]}) # 4. Ask a prebuilt model for fields it does not model (+$10 / 1k pages); `features` gets # `queryFields` added automatically. puffinparse.extract("receipt.png", model="azure/receipt", schema=schema, provider_options={"query_fields": ["StoreNumber", "CashierName"]}) # 5. General key/value pairs out of an unstructured form, via the layout model. puffinparse.extract("form.pdf", model="azure/layout", schema={}, provider_options={"model_id": "prebuilt-layout", "features": ["keyValuePairs"]}) # 6. A prebuilt model PuffinParse does not list, with a locale hint. puffinparse.extract("card.jpg", model="azure/invoice", schema=schema, provider_options={"model_id": "prebuilt-healthInsuranceCard.us", "locale": "en-US"}) ``` (Example 5 pairs `azure/layout` with `extract`; that combination is only reachable because `provider_options.model_id` bypasses the model→`modelId` table — the registry itself lists `layout` for `parse` and `ocr`.) ## 8. Links * Analyze Document REST reference (2024-11-30): * Get Analyze Result (the poll target): * Layout model and its JSON: * Markdown output elements: * Read model: * Prebuilt extraction models and their field lists: * Add-on capabilities: * Service quotas and limits: * Pricing: --- # AWS Textract > **Status: docs-only.** Implemented from AWS's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `textract` | | Base URL | `https://textract.{region}.amazonaws.com` (override: `base_url` on the request, or `TEXTRACT_BASE_URL`) | | API key | **A key pair, not a single token.** `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY`, optional `AWS_SESSION_TOKEN`, region from `AWS_REGION` (then `AWS_DEFAULT_REGION`, then `us-east-1`). `api_key` on the request overrides the **access key id only**. | | Auth header | `Authorization: AWS4-HMAC-SHA256 …` — SigV4, computed in `providers/textract.rs`; no AWS SDK crate is vendored | | Protocol | AWS JSON 1.1 RPC: always `POST /`, operation chosen by `X-Amz-Target`, `Content-Type: application/x-amz-json-1.1` | | API version | `textract-2018-06-27` (implicit in the target names; there is no version header) | | Checked against | 2026-09-11 — request/response shapes and quotas from the AWS API reference; pricing from the AWS Price List API (`AmazonTextract`, `us-east-1`, publication `2026-08-31`). **Response fixtures are built from the AWS documentation's own examples, not from a live capture** (see §6). | | Implementation | `crates/puffinparse-core/src/providers/textract.rs` | Textract is not a "parse a document to markdown" product: it returns a flat array of `Block` objects linked by ids. PuffinParse reassembles those into pages, blocks, markdown tables and extraction results. There is no upload endpoint and no remote-URL input — synchronous calls carry the document inline as base64, asynchronous ones read it from S3. ## 2. Models exposed by PuffinParse | Model | Modes | Textract operation + `FeatureTypes` | List price (`pricing.json`) | |---|---|---|---| | `textract/detect-text` *(default for `ocr`)* | `ocr` | `DetectDocumentText` (no features) | $0.0015 / page | | `textract/layout` | `parse`, `ocr` | `AnalyzeDocument`, `["LAYOUT", "TABLES"]` | $0.015 / page | | `textract/queries` *(default for `extract`)* | `extract` | `AnalyzeDocument`, `["QUERIES"]` + `QueriesConfig` | $0.015 / page | | `textract/forms` | `extract` | `AnalyzeDocument`, `["FORMS"]` | $0.050 / page | Prices are the us-east-1 pay-as-you-go list prices for the **first 1M pages/month**, taken from the AWS Price List API rather than the marketing page, and used only to fill `cost_usd = per_page_usd × usage.pages`: | Usage type (us-east-1) | $ / page (0–1M) | $ / page (1M+) | |---|---|---| | `USE1-SyncTextPagesProcessed` (DetectDocumentText) | 0.0015 | 0.0006 | | `USE1-SyncTablesPagesProcessed` (AnalyzeDocument TABLES) | 0.015 | 0.010 | | `USE1-SyncQueriesPagesProcessed` (AnalyzeDocument QUERIES) | 0.015 | 0.010 | | `USE1-SyncFormsPagesProcessed` (AnalyzeDocument FORMS) | 0.050 | 0.040 | | `USE1-SyncLayoutPagesProcessed` (AnalyzeDocument LAYOUT alone) | 0.004 | 0.003 | **Why `textract/layout` is priced at the TABLES rate and not TABLES + LAYOUT.** AWS bills a combined `AnalyzeDocument` call at a single combination rate, and there is no `Layout…Tables` usage type in the price list: *"Layout is available for free when used with the Tables feature."* So `["LAYOUT","TABLES"]` bills exactly like `["TABLES"]` — $0.015 / page. A LAYOUT-only call would be $0.004 / page, which you can get with `provider_options={"FeatureTypes": ["LAYOUT"]}` (the passthrough replaces the array), at the cost of losing markdown tables. Prices are region-dependent and PuffinParse's table is us-east-1 only, so `cost_usd` is an estimate for any other region. Async (`Start*`/`Get*`) pages cost the same as sync pages; the `USE1-Async…` usage types carry identical rates. ## 3. Request flow PuffinParse uses Every call is `POST {base}/` with these headers, all of them signed: ``` Content-Type: application/x-amz-json-1.1 X-Amz-Target: Textract. X-Amz-Date: 20260911T123456Z X-Amz-Content-Sha256: X-Amz-Security-Token: # only when set Authorization: AWS4-HMAC-SHA256 Credential=///textract/aws4_request, SignedHeaders=content-type;host;x-amz-content-sha256;x-amz-date; [x-amz-security-token;]x-amz-target, Signature= ``` SigV4 is implemented in-file with `hmac` + `sha2` (~60 lines: canonical request → string to sign → four chained HMACs for the signing key). It is re-signed on every retry, because the signature expires with `X-Amz-Date`. Unit tests assert AWS's own published `iam/ListUsers` example vectors (signing key `c4afb1cc…a4b9`, signature `5d672d79…b5d7`), so canonicalisation drift is caught without a network call. ### 3.1 Synchronous (default) 1. **Load bytes.** Path and bytes inputs are read locally. A **URL input is downloaded by PuffinParse** and sent inline — Textract cannot fetch URLs itself. 2. **Multi-page guard.** If the bytes are a PDF, `pdf_page_count()` counts `/Type /Page` objects (falling back to the page tree's `/Count`). More than one page ⇒ a `bad_request` error before any billable call (§5). 3. **Call.** `Textract.DetectDocumentText` or `Textract.AnalyzeDocument` with ```json { "Document": { "Bytes": "" }, "FeatureTypes": ["LAYOUT", "TABLES"], "QueriesConfig": { "Queries": [{ "Text": "What is the total?", "Alias": "total" }] } } ``` `FeatureTypes` is omitted for `detect-text`; `QueriesConfig` only for `textract/queries`. ### 3.2 Asynchronous (multi-page, `provider_options.s3_object`) Textract's async API reads **only from S3** — there is no way to hand it bytes. PuffinParse has no S3 client and deliberately does not grow one, so you upload the object yourself and name it: ```python provider_options={"s3_object": {"bucket": "my-bucket", "name": "invoices/2026-q3.pdf"}} # optional: "version": "" ``` Then PuffinParse: 1. `Textract.StartDocumentTextDetection` / `Textract.StartDocumentAnalysis` with `{"DocumentLocation": {"S3Object": {"Bucket": …, "Name": …}}, "FeatureTypes": […]}` → `{"JobId": …}`. 2. Polls `Textract.GetDocumentTextDetection` / `Textract.GetDocumentAnalysis` with `{"JobId": …, "MaxResults": 1000}` starting at 2 s, backing off ×1.5 to 10 s, until `JobStatus` leaves `IN_PROGRESS`. `SUCCEEDED` and `PARTIAL_SUCCESS` continue; anything else becomes a `provider` error carrying `StatusMessage` and the job id. 3. Follows `NextToken` until it is absent, concatenating every `Blocks` page (a 3 000-page document is many round trips — budget `timeout_secs` accordingly). `NotificationChannel` / SNS is not used: PuffinParse polls. You can still set it (and `OutputConfig`, `KMSKeyId`, `JobTag`, `ClientRequestToken`, `AdaptersConfig`) through `provider_options`, which is deep-merged into whichever body is actually sent. **Where `provider_options` are merged:** the keys `region` and `s3_object` are consumed by PuffinParse; everything else is deep-merged verbatim into the Textract request body, so the remaining keys are top-level Textract request members in `PascalCase` (`FeatureTypes`, `QueriesConfig`, `AdaptersConfig`, `HumanLoopConfig`, `OutputConfig`, `KMSKeyId`, `JobTag`, `ClientRequestToken`, `NotificationChannel`). ### 3.3 Page selection Textract has no page-range parameter. `pages="1-3,7"` is therefore applied **client-side**: blocks whose `Page` falls outside the selection are dropped after the response arrives. It reduces the output, never the bill. The one exception is `textract/queries` in async mode, where the ranges are also forwarded as `Query.Pages` (`["1-3", "7"]`), which Textract does honour. `language` is ignored — Textract auto-detects (English, French, German, Italian, Portuguese, Spanish) and never reports which language it found. ## 4. Response mapping Every operation returns the same envelope: `{"Blocks": [...], "DocumentMetadata": {"Pages": n}, "AnalyzeDocumentModelVersion": "1.0"}` (plus `JobStatus` / `NextToken` / `Warnings` for `Get*`). | Textract field | PuffinParse unified field | Notes | |---|---|---| | `DocumentMetadata.Pages` | `Usage.pages` | Falls back to the highest `Block.Page`, minimum 1. | | `JobId` (async only) | `provider_job_id` | `None` for synchronous calls — Textract returns no id there. | | `Geometry.BoundingBox.{Left,Top,Width,Height}` | `Block.bbox` / `Line.bbox` / `Word.bbox` | **Already normalised 0–1**, origin top-left; `from_normalized_ltwh` only clamps and converts to `{x0,y0,x1,y1}`. `Geometry.Polygon` and `RotationAngle` are ignored. | | `Confidence` (0–100) | `confidence` (0–1) | Divided by 100 and clamped. | | `Page` | `page_number` | 1-based. Always `1` for JPEG/PNG, even multi-page scans. | | — | `Page.width` / `height` | Always `None`: Textract never reports page dimensions. | | `AnalyzeDocumentModelVersion` | `metadata.textract_model_version` | | | resolved region | `metadata.textract_region` | | | `FeatureTypes` sent | `metadata.textract_feature_types` | | | `Warnings[]` | `metadata.textract_warnings` | Present only when non-empty (e.g. `INVALID_REQUEST_PARAMETERS` for an over-quota query page). | | — | `Usage.credits`, `Usage.provider_cost_usd` | Always `None`; `cost_usd` comes from `pricing.json`. | ### 4.1 `ocr` mode — `LINE` / `WORD` `LINE` blocks become `TextPage.lines`, `WORD` blocks become `TextPage.words`, both with box and confidence; `TextPage.text` is the lines joined by `\n` in response order. This is the native path for `textract/detect-text`, and `textract/layout` uses it too — `AnalyzeDocument` returns *all* lines and words regardless of `FeatureTypes`, so no second call is needed (and no extra charge). ### 4.2 `parse` mode — `LAYOUT_*` + `TABLE` Textract returns `LAYOUT_*` blocks **in implied reading order** (left to right, top to bottom; column by column on multi-column pages). PuffinParse walks them in array order, so `Page.markdown` is in reading order. A layout block's text is its descendant `LINE` blocks via `Relationships[CHILD]`. | `BlockType` | `Block.type` | `Block.content` | |---|---|---| | `LAYOUT_TITLE` | `title` | `# ` | | `LAYOUT_SECTION_HEADER` | `section_header` | `## ` | | `LAYOUT_HEADER` | `header` | plain text | | `LAYOUT_FOOTER` | `footer` | plain text | | `LAYOUT_TEXT`, `LAYOUT_KEY_VALUE` | `text` | plain text, one line per child `LINE` | | `LAYOUT_LIST` | `list` | `- ` per child `LAYOUT_TEXT` | | `LAYOUT_TABLE` | `table` | markdown table (below) | | `LAYOUT_FIGURE` | `figure` | the caption lines, if any | | `LAYOUT_PAGE_NUMBER` | `other` | plain text | `LAYOUT_LIST` points at `LAYOUT_TEXT` children, and those children *also* appear at the top level of `Blocks`; PuffinParse suppresses any layout block that is another layout block's child so list items are not emitted twice. **Tables.** A `LAYOUT_TABLE` is matched to its `TABLE` block by a direct `CHILD` reference, or — when it only points at lines — by the highest-overlap unclaimed `TABLE` on the same page (>10 % of the table's area). The `TABLE`'s `CHILD` `CELL` blocks are placed on a `RowIndex` × `ColumnIndex` grid (`MERGED_CELL` blocks are skipped: they repeat content already present in the individual cells) and rendered as markdown. Cell text is the `CHILD` `WORD` texts joined by spaces; a `SELECTION_ELEMENT` child becomes `[x]` / `[ ]`; `|` is escaped. The header row is the lowest `RowIndex` among cells with `EntityTypes: ["COLUMN_HEADER"]`, else row 1; rows above the header row and `TABLE_TITLE` / `TABLE_FOOTER` blocks are emitted as plain lines around the table. **Fallback.** If a response carries no `LAYOUT_*` blocks at all (LAYOUT disabled through `provider_options`, or nothing detected), PuffinParse emits every `TABLE` as a markdown table plus one `text` block per `LINE`, skipping lines whose words are all inside table cells so table content is not duplicated. `output="text"` renders every block through `markdown_to_text`. Trimmed fixture (`crates/puffinparse-core/tests/fixtures/textract_layout.json`, most blocks elided): ```json { "DocumentMetadata": { "Pages": 1 }, "AnalyzeDocumentModelVersion": "1.0", "Blocks": [ { "BlockType": "LAYOUT_TITLE", "Confidence": 98.12, "Id": "lay-title", "Page": 1, "Geometry": { "BoundingBox": { "Left": 0.12, "Top": 0.06, "Width": 0.26, "Height": 0.02 } }, "Relationships": [ { "Type": "CHILD", "Ids": ["ll1"] } ] }, { "BlockType": "LAYOUT_TABLE", "Confidence": 97.4, "Id": "lay-table", "Page": 1, "Relationships": [ { "Type": "CHILD", "Ids": ["tbl-1"] } ] }, { "BlockType": "TABLE", "Confidence": 99.21, "Id": "tbl-1", "Page": 1, "EntityTypes": ["STRUCTURED_TABLE"], "Relationships": [ { "Type": "CHILD", "Ids": ["c11","c12","c21","c22","c31","c32"] } ] }, { "BlockType": "CELL", "RowIndex": 1, "ColumnIndex": 1, "RowSpan": 1, "ColumnSpan": 1, "EntityTypes": ["COLUMN_HEADER"], "Id": "c11", "Page": 1, "Relationships": [ { "Type": "CHILD", "Ids": ["tw1"] } ] }, { "BlockType": "LINE", "Confidence": 98.9, "Text": "Quarterly Report", "Id": "ll1", "Page": 1, "Relationships": [ { "Type": "CHILD", "Ids": ["lw1","lw2"] } ] }, { "BlockType": "WORD", "Confidence": 99.0, "Text": "Region", "TextType": "PRINTED", "Id": "tw1", "Page": 1 } ] } ``` ### 4.3 `extract` mode — `textract/queries` Each **flat** property of the request schema becomes one Textract query: ``` {"Text": ?? ?? <key with _ and - as spaces>, "Alias": <sanitised key>} ``` The alias keeps `[A-Za-z0-9_.:-]` (everything else becomes `_`, duplicates get a numeric suffix) and maps back to the original property name, so `"invoice number"` in the schema is queried as `invoice_number` and returned under `"invoice number"`. Responses are `QUERY` blocks carrying `Query.Alias` and a `Relationships[{"Type":"ANSWER"}]` list of `QUERY_RESULT` ids. PuffinParse takes the highest-confidence non-empty `QUERY_RESULT`, coerces its `Text` to the schema's type (`number`/`integer` strip currency and separators; `boolean` understands yes/no/true/false/selected/`[x]`), and writes it to `data[<key>]`. Confidence lands in `fields["/<key>"].confidence`; with `citations=True` the answer's page, box and text land in `fields["/<key>"].citations`. A query with no answer yields `data[<key>] = null` and **no** entry in `fields` — the key is always present, so the shape of `data` matches the schema. > **Only flat string-like fields are supported.** Properties typed `object` or `array` (or carrying > `properties` / `items`) cannot be expressed as a Textract query; they are skipped, reported in > `metadata.textract_unsupported_fields`, and set to `null`. Textract Queries answers one question > with one span of text — there is no nesting and no repeated-row extraction. Flatten the schema > (`line_item_1_total`, …) or use a provider with native structured extraction. `instructions` on the request is ignored (Textract has no free-text guidance parameter); the fact is recorded in `metadata.textract_instructions_ignored`. Textract allows **15 queries per page synchronously and 30 asynchronously**. A schema with more flat properties than the applicable limit is rejected with an `input` error before any call is made. ### 4.4 `extract` mode — `textract/forms` (best effort) `FORMS` returns `KEY_VALUE_SET` blocks: a block with `EntityTypes: ["KEY"]` holds the label (its `CHILD` `WORD`s) and points at its value block through `Relationships[{"Type":"VALUE"}]`; the value block's `CHILD`ren are `WORD`s or a `SELECTION_ELEMENT`. PuffinParse matches each schema property to a detected key by **case- and punctuation-insensitive** comparison (both sides reduced to lowercase alphanumerics), trying the property name and its `title`, first for an exact match and then for "one name contains the other" (≥3 characters). Each detected key is consumed at most once. Values are coerced by the schema's type exactly as in §4.3, and a selected checkbox becomes `true`. > **This is a heuristic, not schema-driven extraction.** Textract decides what the keys are; PuffinParse > only tries to line them up with your field names. Properties with no match are set to `null` and > listed in `metadata.textract_unmatched_fields`; `metadata.textract_form_keys_found` reports how many > key-value pairs Textract actually detected, which is the first thing to look at when fields come > back empty. Nested (`object`/`array`) properties are never matched. At $0.050/page `textract/forms` > is also the most expensive model here — prefer `textract/queries` when you know what you want. ## 5. Errors, status codes, rate limits, timeouts Textract reports failures as `{"__type": "<Exception>", "message": "…"}`, almost always with **HTTP 400 — including throttling and server-side failures**, so the shared `Error::from_http` mapping (`4xx → bad_request`) is wrong for Textract. `textract::map_error` keys off `__type` instead (a fully-qualified `com.amazonaws.textract#ThrottlingException` is accepted too) and prefixes the exception name onto the message: | `__type` | HTTP | PuffinParse `ErrorKind` | Retried? | |---|---|---|---| | `AccessDeniedException` | 400 | `authentication` | no | | `UnrecognizedClientException`, `InvalidClientTokenId`, `ExpiredTokenException` | 400/403 | `authentication` | no | | `IncompleteSignature`, `InvalidSignatureException`, `MissingAuthenticationTokenException` | 400/403 | `authentication` | no | | `ThrottlingException` | 400/500 | `rate_limit` | **yes** | | `ProvisionedThroughputExceededException` | 400 | `rate_limit` | **yes** | | `LimitExceededException` (too many concurrent async jobs) | 400 | `rate_limit` | **yes** | | `InternalServerError`, `ServiceUnavailable`, `InternalFailure` | 500 | `provider` | **yes** | | `BadDocumentException`, `UnsupportedDocumentException`, `DocumentTooLargeException` | 400 | `bad_request` | no | | `InvalidParameterException`, `InvalidS3ObjectException`, `InvalidKMSKeyException`, `InvalidJobIdException` | 400 | `bad_request` | no | | `IdempotentParameterMismatchException`, `HumanLoopQuotaExceededException` | 400 | `bad_request` | no | | anything unrecognised | — | falls back to `Error::from_http` | per status | An async job that ends `FAILED` becomes a `provider` error carrying `StatusMessage` and the job id. `PARTIAL_SUCCESS` is accepted, logged, and its `Warnings` surfaced in metadata. **Multi-page PDFs.** Synchronous Textract accepts PDF and TIFF at **one page only**. PuffinParse raises a `bad_request` before the call: > textract: synchronous operations accept single-page PDF/TIFF only (this document has 2 pages). > Multi-page documents require Textract's asynchronous API, which reads the file from S3. Upload the > file yourself and pass `provider_options={"s3_object": {"bucket": "my-bucket", "name": "path/doc.pdf"}}`, > or split the PDF into single pages first. The page count is best effort — it reads uncompressed page objects and the `/Count` in the page tree, which covers most producers but not PDFs that hide their structure in object streams. When the count cannot be determined, the call goes out and Textract's own 4xx is rewritten into the same message, so the advice is identical either way. **Retries.** `max_retries` (default 2), exponential backoff with full jitter, only on the `rate_limit`, `network` and 5xx `provider` kinds above. **Textract returns no `Retry-After` and no rate-limit headers**, so backoff is blind. Default us-east-1 quotas worth knowing (all adjustable in Service Quotas, and lower in most other regions): synchronous `AnalyzeDocument` 10 TPS, `DetectDocumentText` 25 TPS; `StartDocumentAnalysis` 10 TPS, `StartDocumentTextDetection` 15 TPS; `GetDocumentAnalysis` 10 TPS, `GetDocumentTextDetection` 25 TPS; at most 600 asynchronous jobs existing simultaneously per account (exceeding that is `LimitExceededException`, which PuffinParse retries). **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline — download, signing, the call, job polling and `NextToken` pagination — and also caps each individual HTTP request, shrinking as the budget is spent. ## 6. Gotchas * **The "API key" is a key pair.** `api_key` on a PuffinParse request can only stand in for `AWS_ACCESS_KEY_ID`; `AWS_SECRET_ACCESS_KEY` must be in the environment, otherwise the request is refused with an `authentication` error that says so. Set `AWS_SESSION_TOKEN` as well for STS / assumed-role credentials — it is signed as `x-amz-security-token` and included in `SignedHeaders`. PuffinParse reads **only** those environment variables: it does not parse `~/.aws/credentials`, does not honour `AWS_PROFILE`, and does not call IMDS or the ECS credential endpoint. * **The region is part of the signature, not just the URL.** Signing with the wrong region gives a 400 that talks about the *credential scope*, not about the host. If you override `base_url` to a VPC endpoint or a mock, set `provider_options={"region": …}` to match. * **No fixtures were captured live.** The credentials available in this repository's build environment are proxy placeholders; a real `DetectDocumentText` call against `textract.us-east-1.amazonaws.com` returned `{"__type":"UnrecognizedClientException","message":"The security token included in the request is invalid."}`. That round trip does confirm the transport and the signature format (a malformed canonical request yields `IncompleteSignature`/`InvalidSignatureException`, not `UnrecognizedClientException`) and the error mapping (`ErrorKind::Authentication`), but every `textract_*.json` fixture is assembled from the shapes in the AWS API reference, with confidences, ids and boxes filled in to be realistic. **Re-capture them from a real account before trusting the numbers**, and run the `#[ignore]`d live tests at the bottom of `textract.rs`. * **Confidence is 0–100, not 0–1.** Every `Confidence` is a percentage; PuffinParse divides by 100. The `QUERY_RESULT` sample in the AWS docs shows `"Confidence": 1.0`, which is 1 %, not certainty. * **Boxes are already normalised.** Unlike most providers, `BoundingBox` is 0–1 relative to the page, so no page dimensions are needed — which is just as well, because Textract never reports them and `Page.width`/`height` are always `None`. * **`Page` is always 1 for JPEG/PNG**, even for a scanned image that visually contains several pages. Only PDF and TIFF produce `Page > 1`, and only through the async API. * **Synchronous PDFs are single-page, full stop** (10 MB in memory). Async accepts 500 MB / 3 000 pages but **only from S3** — there is no bytes variant of `StartDocument*`. This is the single biggest limitation of this provider and the reason `provider_options.s3_object` exists. * **Password-protected PDFs and XFA PDFs are rejected**; images must be ≤10 000 px per side. * **`AnalyzeDocument` always returns every `LINE` and `WORD`**, whatever `FeatureTypes` says. That is why `textract/layout` can serve `ocr` mode from the same response, and why a `FORMS`-only call still gives you the full text in `raw`. * **`LAYOUT_LIST` children are `LAYOUT_TEXT`, not `LINE`** — one level of indirection that also means those `LAYOUT_TEXT` blocks appear twice in `Blocks` (once nested, once at the top level). Naive iteration duplicates every list item. * **`MERGED_CELL` duplicates content.** A merged cell's `CHILD` ids are the individual `CELL`s, which are *also* children of the `TABLE`. Rendering both repeats the text; PuffinParse renders only the plain cells, so a row/column span shows its text in the first cell and blanks beside it. * **Queries are English-only** and capped at 15 per page synchronously / 30 asynchronously. Answers are capped at 128 characters. A query aimed at a page that does not exist comes back as an `INVALID_REQUEST_PARAMETERS` entry in `Warnings`, not as an error. * **Throttling arrives as HTTP 400.** Treating Textract's 400s as non-retryable (the usual rule) means giving up on `ThrottlingException` and `ProvisionedThroughputExceededException`; the mapping in §5 exists entirely for this. * **`Warnings` is silent data loss.** A `PARTIAL_SUCCESS` job returns blocks for the pages that worked and lists the failures in `Warnings` — check `metadata.textract_warnings` before trusting the page count. * **No `provider_job_id` for sync calls.** Textract returns the request id only in the `x-amzn-RequestId` response header, which PuffinParse does not surface today. * **Pagination is per 1 000 blocks, not per page.** A dense 50-page document can need dozens of `Get*` round trips; they all come out of `timeout_secs`. ## 7. Useful `provider_options` passthrough ```python # 1. Multi-page PDF: upload to S3 yourself, then let PuffinParse drive Start*/Get* + NextToken. puffinparse.parse("ignored-when-s3.pdf", model="textract/layout", provider_options={"s3_object": {"bucket": "my-bucket", "name": "reports/q3.pdf"}}, timeout=1200) # 2. Cheapest layout: drop TABLES to bill at the LAYOUT rate ($4 vs $15 per 1k pages). # Tables then come back as plain lines instead of markdown grids. puffinparse.parse("memo.png", model="textract/layout", provider_options={"FeatureTypes": ["LAYOUT"]}) # 3. A different region (also changes the signing scope, not just the host). puffinparse.ocr("scan.png", model="textract/detect-text", provider_options={"region": "eu-west-1"}) # 4. Signatures alongside layout, and the raw Block array for anything PuffinParse does not map. puffinparse.parse("contract.png", model="textract/layout", include_raw=True, provider_options={"FeatureTypes": ["LAYOUT", "TABLES", "SIGNATURES"]}) # 5. A trained Custom Queries adapter (adapters are Queries-only). puffinparse.extract("claim.png", model="textract/queries", schema=schema, provider_options={"AdaptersConfig": {"Adapters": [ {"AdapterId": "abc123", "Version": "1"}]}}) # 6. Async with your own output bucket, KMS key and an idempotency token. puffinparse.parse("big.pdf", model="textract/layout", timeout=1800, provider_options={"s3_object": {"bucket": "in", "name": "big.pdf"}, "OutputConfig": {"S3Bucket": "out", "S3Prefix": "textract/"}, "KMSKeyId": "alias/textract", "ClientRequestToken": "big-pdf-2026-09-11"}) ``` ## 8. Links * What is Amazon Textract: <https://docs.aws.amazon.com/textract/latest/dg/what-is.html> * `AnalyzeDocument`: <https://docs.aws.amazon.com/textract/latest/dg/API_AnalyzeDocument.html> · `DetectDocumentText`: <https://docs.aws.amazon.com/textract/latest/dg/API_DetectDocumentText.html> * `StartDocumentAnalysis`: <https://docs.aws.amazon.com/textract/latest/dg/API_StartDocumentAnalysis.html> · `GetDocumentAnalysis`: <https://docs.aws.amazon.com/textract/latest/dg/API_GetDocumentAnalysis.html> * `Block` reference: <https://docs.aws.amazon.com/textract/latest/dg/API_Block.html> * Layout response objects: <https://docs.aws.amazon.com/textract/latest/dg/layoutresponse.html> * Tables: <https://docs.aws.amazon.com/textract/latest/dg/how-it-works-tables.html> · Form data: <https://docs.aws.amazon.com/textract/latest/dg/how-it-works-kvp.html> · Queries: <https://docs.aws.amazon.com/textract/latest/dg/queryresponse.html> * Quotas: <https://docs.aws.amazon.com/textract/latest/dg/limits-document.html> · <https://docs.aws.amazon.com/general/latest/gr/textract.html> * Pricing: <https://aws.amazon.com/textract/pricing/> · machine-readable price list: <https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonTextract/current/index.json> * Signature Version 4: <https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_sigv4-signing-examples.html> --- # Google Gemini <!-- source: docs/providers/gemini.md | url: https://puffinparse.com/docs/providers/gemini/ --> > **Status: docs-only.** Implemented from Google's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `gemini` | | Base URL | `https://generativelanguage.googleapis.com` (override: `base_url` on the request, or `GEMINI_BASE_URL`) | | API key | `GEMINI_API_KEY` (or `api_key` on the request) — sent as the header `x-goog-api-key: AIza…` | | Docs | <https://ai.google.dev/gemini-api/docs> | | API version | Path-versioned; PuffinParse uses `/v1beta` (the Files API and the newest models live there). `/v1` serves the same `generateContent` method. | | Modes | `parse`, `ocr` (derived from `parse`), `extract` | | Checked against | 2026-09-11 against the documented wire format; **the live call was not exercised** — see §6, "Unverified against a live key" | | Implementation | `crates/puffinparse-core/src/providers/gemini.rs` | Gemini is not a document-AI product but a general vision LLM: PuffinParse sends the file plus a transcription prompt and pins the answer's shape with `generationConfig.response_schema` (structured output), so the model must return one entry per page instead of a single markdown blob. There is exactly **one** HTTP call per parse — no upload step, no polling — unless the document is larger than ~14 MB, in which case the resumable Files API is used first. The consequence of using an LLM is that **nothing geometric comes back**: no bounding boxes, no page dimensions, no per-block confidence. Anything that needs overlays or coordinates should use a layout provider (Reducto, Extend, LlamaParse, Azure, Textract) instead. ## 2. Models exposed by PuffinParse The PuffinParse model name is the Gemini model id minus the `gemini-` prefix. | Model | API model id | Token price (in / out, per 1M) | `pricing.json` per-page estimate | |---|---|---|---| | `gemini/2.5-flash` *(default)* | `gemini-2.5-flash` | $0.30 / $2.50 | $0.00245 | | `gemini/2.5-pro` | `gemini-2.5-pro` | $1.25 / $10.00 (>200k prompt tokens: $2.50 / $15.00) | $0.009875 | | `gemini/2.5-flash-lite` | `gemini-2.5-flash-lite` | $0.10 / $0.40 | $0.00047 | | `gemini/3.5-flash` | `gemini-3.5-flash` | $1.50 / $9.00 | $0.00945 | | `gemini/3.5-flash-lite` | `gemini-3.5-flash-lite` | $0.30 / $2.50 | $0.00245 | | `gemini/3.8-flash` | `gemini-3.8-flash` | $0.75 / $3.75 *(introductory, through 2026-12-31; $1.50 / $7.50 after)* | $0.004125 | All six support `parse`, `ocr` and `extract` — the mode is a prompt + schema, not an endpoint, so every model serves every mode. The 2.5 trio is the safe default; the 3.x entries come from the models page and the changelog (3.5 Flash GA 2026-05-19, 3.5 Flash-Lite GA 2026-07-21, 3.8 Flash GA 2026-09-02) and have not been exercised against a live key here. **Cost is computed from tokens, not from pages.** Each response's `usageMetadata` is priced exactly (`promptTokenCount × input + (candidatesTokenCount + thoughtsTokenCount) × output`) and lands in `usage.provider_cost_usd`, which `puffinparse-core` prefers over the price table. The per-page numbers in `pricing.json` are only a fallback for the case where a response carries no `usageMetadata` (and for `puffinparse providers`-style estimates). They were derived, not measured: ``` per_page_usd = (1500 × input_price_per_1M + N × output_price_per_1M) / 1e6 N = 800 output tokens for parse/ocr, 300 for extract ``` so `gemini/2.5-flash` parse = (1500 × 0.30 + 800 × 2.50) / 1e6 = $0.00245/page, and the `extract` column of `pricing.json` is the same formula with N = 300 ($0.0012/page). 1 500 input tokens/page is a deliberately conservative blend: a PDF page bills at a flat 258 tokens plus the rendered-image tokens, while a full-page scan image is ~1 000–1 600 tokens. Thinking tokens are not in the estimate; with a thinking budget left at its default they can double the output side. A model that is not in this table can be used without a registry change: `provider_options={"model": "gemini-3.1-pro-preview"}` replaces the API model id verbatim (PuffinParse then has no token prices for it and falls back to the registry model's per-page price). ## 3. Request flow PuffinParse uses One call: `POST {base}/v1beta/models/{api_model}:generateContent`, headers `x-goog-api-key` and `content-type: application/json`. ```json { "contents": [{"role": "user", "parts": [ {"inline_data": {"mime_type": "application/pdf", "data": "JVBERi0xLjQK…"}}, {"text": "Transcribe the attached document to GitHub-flavoured Markdown, page by page. …"} ]}], "generationConfig": { "temperature": 0, "response_mime_type": "application/json", "response_schema": { "type": "object", "properties": {"pages": {"type": "array", "items": { "type": "object", "properties": {"page_number": {"type": "integer"}, "markdown": {"type": "string"}}, "required": ["page_number", "markdown"], "propertyOrdering": ["page_number", "markdown"]}}}, "required": ["pages"] } } } ``` **Input handling.** | Input | What PuffinParse does | |---|---| | Path / bytes ≤ 14 MB | base64 into `inline_data` | | Path / bytes > 14 MB | resumable Files API upload, then `file_data: {file_uri, mime_type}` | | URL | **downloaded first** (Gemini cannot fetch URLs), then treated as bytes | The MIME type is sniffed from the magic bytes (`%PDF-`, PNG, JPEG, WebP, GIF) and only falls back to the extension guess, because Gemini rejects a mismatched `mime_type` outright. PDFs are native input (up to 1 000 pages / 50 MB); the accepted image types are `image/png`, `image/jpeg`, `image/webp`, `image/heic` and `image/heif` — **not GIF**, which PuffinParse still labels correctly so that Gemini's rejection names the real reason. The 14 MB inline cut-off exists because a `generateContent` request is capped at ~20 MB and base64 inflates the payload by a third. The Files API path is `POST {base}/upload/v1beta/files` with `X-Goog-Upload-Protocol: resumable` and `X-Goog-Upload-Command: start` → the `x-goog-upload-url` response header → a second request with `X-Goog-Upload-Command: upload, finalize` carrying the bytes → `{"file": {"uri", "name", "state"}}`; if `state` is not `ACTIVE` PuffinParse polls `GET {base}/v1beta/files/{id}` (0.5 s, ×1.5, max 5 s) until it is. Uploaded files expire after 48 hours. **Per mode.** * `parse` — the prompt above plus the `pages` schema. The model returns `{"pages": [{"page_number": 1, "markdown": "…"}, …]}`, so the page split is exact rather than guessed from a separator. * `ocr` — not implemented natively; the trait default derives a `TextResponse` from `parse` (`metadata.puffinparse_derived_from = "parse"`). Lines come from the page text, words carry no boxes. Passing `output="text"` additionally tells the model to skip Markdown syntax. * `extract` — the request schema is rewritten into Gemini's schema subset (see §4) and used as `response_schema`; the schema and `instructions` are also restated in the prompt, which measurably improves adherence. `ExtractResponse.data` is the parsed JSON. `citations=True` is accepted but cannot be honoured (`metadata.gemini_citations_unsupported = true`). **Page selection.** Gemini always reads the whole file, so `pages` is a *prompt instruction* ("Transcribe ONLY these pages of the PDF: 2-3 … keep the original page numbers"), applied only when the input is a PDF. It is best effort: the full document is still uploaded and still billed as input tokens, and a model may ignore the restriction. For a non-PDF input `pages` is dropped and `metadata.gemini_pages_ignored = true` is set. **Where `provider_options` are merged:** the whole object is deep-merged into the request body after PuffinParse builds it, so `generationConfig`, `safetySettings`, `systemInstruction`, `tools`, … can all be set or overridden. Three keys are consumed by PuffinParse and never reach Gemini: `model` (API model id), `prompt` (replaces the built-in instruction entirely) and `prompt_suffix` (appended to it). ## 4. Response mapping ```json { "candidates": [{ "content": {"parts": [{"text": "{\"pages\": [{\"page_number\": 1, \"markdown\": \"# …\"}]}"}], "role": "model"}, "finishReason": "STOP", "index": 0 }], "usageMetadata": { "promptTokenCount": 1809, "candidatesTokenCount": 1060, "totalTokenCount": 2869, "promptTokensDetails": [{"modality": "DOCUMENT", "tokenCount": 1548}, {"modality": "TEXT", "tokenCount": 261}] }, "modelVersion": "gemini-2.5-flash", "responseId": "HfPJaNXKO7OXm9IPk7P1kQc" } ``` (The full payloads are `crates/puffinparse-core/tests/fixtures/gemini_parse_multipage.json` and `gemini_extract_invoice.json`, which drive the normalisation tests.) | Gemini field | PuffinParse unified field | Notes | |---|---|---| | `responseId` | `ParseResponse.provider_job_id` | Gemini has no job concept; this is the only per-call id. | | `candidates[0].content.parts[].text` | — | All text parts are concatenated, then parsed as JSON (a stray ```` ```json ```` fence is stripped defensively). | | `pages[].page_number` | `Page.page_number` | 1-based; falls back to the array index + 1 if the model omits it. | | `pages[].markdown` | `Page.markdown` | Trimmed. With `output="text"` the markdown-to-text conversion is stored instead. | | derived | `Page.text` | Always `markdown_to_text(markdown)`. | | derived | `Page.blocks` | Exactly one `text` block per non-empty page, `content` = the page, `bbox: None`, `confidence: None`. | | — | `Page.width` / `height` | Always `None` — **Gemini reports no geometry**. | | number of returned pages | `Usage.pages` | For an inline PDF the `/Type /Page` count is compared against it (see below). | | `usageMetadata` | `Usage.provider_cost_usd` | `prompt × input_price + (candidates + thoughts) × output_price`. | | — | `Usage.credits` | Always `None`; Gemini bills tokens, not credits. | | `usageMetadata.*` | `metadata.gemini_tokens` | `{prompt, candidates, thoughts, cached, total}`. | | `modelVersion` | `metadata.gemini_model_version` | The exact served version behind an alias. | | the API model id sent | `metadata.gemini_api_model` | Useful when `provider_options.model` overrode it. | | `finishReason` ≠ `STOP` | `metadata.gemini_finish_reason` | | | `promptFeedback.blockReason` | → error | `bad_request`, message `gemini blocked the prompt: <reason>`. | For `extract`, the parsed JSON becomes `ExtractResponse.data` verbatim, `fields` stays empty (no citations, no per-field confidence) and `usage.pages` is the PDF page count, or `1` for an image. **PDF page-count sanity check.** For an inline PDF PuffinParse counts `/Type /Page` objects (ignoring `/Type /Pages`) in the raw bytes and records it as `metadata.gemini_pdf_page_count`. If it differs from the number of pages the model returned — a dropped or hallucinated page, or simply a page-selection request — `metadata.gemini_page_count_mismatch = true` is set and a warning is logged. `Usage.pages` still follows what the model returned. The count is best effort and **inline-only**: PDFs whose page objects live in compressed object streams return no count, and neither does a document that went through the Files API — in `extract` mode that means `usage.pages` falls back to `1`. **`sanitize_schema`** (public, unit-tested) rewrites a draft-2020-12 JSON Schema into the subset Gemini's `response_schema` accepts: * keeps only `type`, `format`, `title`, `description`, `nullable`, `enum`, `items`, `prefixItems`, `properties`, `required`, `propertyOrdering`, `minItems`, `maxItems`, `minimum`, `maximum`, `minLength`, `maxLength`, `pattern`, `anyOf` — everything else (`$schema`, `$id`, `additionalProperties`, `unevaluatedProperties`, `default`, `examples`, `allOf`, `not`, `if`/`then`, `patternProperties`, …) is dropped; * `type: ["string", "null"]` → `type: "string"` + `nullable: true`; * `oneOf` → `anyOf`; `const: X` → `enum: [X]` with the type inferred; * local `$ref`s into `$defs` / `definitions` are inlined (depth-limited to 12; a recursive schema degrades gracefully instead of expanding forever, and an unresolvable `$ref` becomes `{"type": "object"}`); * a node with `properties` but no `type` gets `type: "object"`, with `items` gets `type: "array"`. ## 5. Errors, status codes, rate limits, timeouts Errors are `{"error": {"code": 400, "message": "…", "status": "INVALID_ARGUMENT", "details": [...]}}`; `Error::from_http` lifts `error.message` and classifies by status: | Status | `status` field | Trigger | PuffinParse `ErrorKind` | |---|---|---|---| | 400 | `INVALID_ARGUMENT` | malformed body, unsupported `mime_type`, **invalid API key**, schema Gemini rejects | `bad_request` | | 400 | `FAILED_PRECONDITION` | free tier not available in the caller's country / billing required | `bad_request` | | 403 | `PERMISSION_DENIED` | key lacks access to the model, or key restrictions (IP/referrer) | `authentication` | | 404 | `NOT_FOUND` | unknown model id, or an expired Files API `file_uri` | `bad_request` | | 429 | `RESOURCE_EXHAUSTED` | per-minute request/token quota | `rate_limit` (retried) | | 500 / 503 | `INTERNAL` / `UNAVAILABLE` | server error, model overloaded | `provider` (retried) | | 504 | `DEADLINE_EXCEEDED` | prompt too large to finish in time | `provider` (retried) | Note the odd one: **an invalid API key is a 400, not a 401**, so it surfaces as `bad_request` rather than `authentication`. A *missing* key is caught before the call and is `authentication`. Failures that arrive with HTTP 200 are turned into errors too: * `promptFeedback.blockReason` set, or a candidate with `finishReason: "SAFETY"` and no text → `bad_request` / `provider` with the reason in the message. * `finishReason: "MAX_TOKENS"` → `provider`, "gemini hit the output token limit, so the JSON is truncated — raise `generationConfig.maxOutputTokens` … or parse fewer pages per call". * Text that is not the promised JSON → `provider`, with the first 200 characters in the message. **Retries.** `max_retries` (default 2) with exponential backoff and full jitter, on 429, 5xx and network errors only (the shared `http::with_retry`). Gemini's 429 body sometimes carries a `RetryInfo` detail with `retryDelay`; PuffinParse does not read it yet and backs off blind. **Rate limits.** Counted per project and per model in three dimensions — requests/minute, *input* tokens/minute and requests/day — and they depend on the usage tier, so the authoritative numbers are the ones shown in Google AI Studio rather than any number quoted here. A single large document can exhaust the token dimension in very few requests. **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline and caps every individual request, including the Files API upload and the URL download. ## 6. Gotchas * **Unverified against a live key.** Everything here follows the published wire format, and both the normalisation path (fixture tests) and the transport (`mod transport` — a loopback HTTP server that asserts the exact URL, headers, base64 `inline_data`, schema, 429 retry and 400 mapping) are covered by tests, but the `GEMINI_API_KEY` present in the development environment is rejected by Google with `400 API_KEY_INVALID` on every endpoint, so no call has ever reached the real service. The fixtures under `tests/fixtures/gemini_*.json` were therefore **constructed from the documented response schema**, not captured. Run `cargo test -p puffinparse-core gemini_live -- --ignored --nocapture` with a working key, then replace the fixtures with real payloads and fill in measured latency and cost here. * **No geometry, ever.** `Block.bbox`, `Page.width`/`height` and `Block.confidence` are always `None`. `ocr` mode returns words without boxes. This is a property of the model, not of PuffinParse. * **The page split comes from the model, not from the file.** Structured output makes it reliable in practice, but a model can still merge or drop a page. `metadata.gemini_page_count_mismatch` is the tripwire for PDFs; there is no equivalent check for multi-page TIFFs or images. * **`pages` is a suggestion.** Whole-file upload, prompt-level selection, full input billing. * **Thinking tokens are billed as output and are invisible in `candidatesTokenCount`.** On 2.5 models a page of dense tables can spend more on reasoning than on the transcript. Set `provider_options={"generationConfig": {"thinkingConfig": {"thinkingBudget": 0}}}` to turn thinking off on Flash / Flash-Lite (Pro cannot go below 128; Gemini 3 models use `thinkingLevel` instead). * **Long documents hit `MAX_TOKENS` before they hit the page limit.** The 1 000-page ceiling is theoretical: the transcript of ~40 dense pages already approaches the 64k output budget. Split long PDFs with `pages`, or raise `generationConfig.maxOutputTokens`. * **`temperature: 0` does not make the output deterministic.** Repeated runs differ slightly, which matters for benchmark reproducibility. * **Gemini cannot fetch a URL.** PuffinParse downloads it first; a URL that needs authentication has to be fetched by the caller and passed as bytes. * **`additionalProperties` is the classic 400.** Most schema generators (Pydantic, zod) emit it and Gemini rejects it; `sanitize_schema` strips it, along with `$schema`, `default` and `examples`. The current docs page lists `additionalProperties` as supported, but rejecting it has been the observed behaviour for long enough that stripping is the safe default. * **The API is versioned in the path and the newest surface is moving.** Google's *Interactions API* (`POST /v1beta/interactions`, GA June 2026) is now the documented default and uses a different shape (`input`, `steps`, `response_format`). `generateContent` remains supported and is still the recommended path for stable deployments — PuffinParse targets it deliberately; expect the docs links below to show the Interactions shape. * **Files API objects expire after 48 hours** and are scoped to the project, so a `file_uri` cannot be reused across keys. PuffinParse uploads per call and never reuses or deletes (files are free and capped at 20 GB per project). * **Image vs. PDF tokenisation differ.** A PDF page is a flat 258 tokens plus image tokens; a standalone image is tiled at 768×768 (≈258 tokens per tile). A scan sent as PNG can therefore cost several times what the same page costs inside a PDF. * **`x-goog-api-key` and `?key=` are equivalent**; PuffinParse uses the header so keys never land in request logs or URLs. ## 7. Useful `provider_options` passthrough ```python # 1. Turn thinking off for cheap, fast transcription (Flash / Flash-Lite only). puffinparse.parse("scan.png", model="gemini/2.5-flash", provider_options={"generationConfig": {"thinkingConfig": {"thinkingBudget": 0}}}) # 2. Raise the output budget for a long PDF. puffinparse.parse("report.pdf", model="gemini/2.5-pro", provider_options={"generationConfig": {"maxOutputTokens": 65536}}) # 3. Reach a model that is not in the registry. puffinparse.parse("doc.pdf", model="gemini/2.5-flash", provider_options={"model": "gemini-3.1-pro-preview"}) # 4. Steer the transcription without rewriting the whole prompt. puffinparse.parse("statement.pdf", model="gemini/2.5-flash", provider_options={"prompt_suffix": "Keep every stamp and handwritten note, " "and transcribe struck-through text as ~~text~~."}) # 5. Replace the instruction entirely (the pages schema still applies). puffinparse.parse("form.pdf", model="gemini/2.5-flash-lite", provider_options={"prompt": "Return each page of this form as a Markdown table of " "field name and value, one row per field."}) # 6. Loosen safety filters for documents that trip them (medical, legal, incident reports). puffinparse.parse("report.pdf", model="gemini/2.5-flash", provider_options={"safetySettings": [ {"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_NONE"}, {"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_NONE"}]}) # 7. A system instruction, e.g. to pin the output language. puffinparse.parse("brief.pdf", model="gemini/2.5-flash", provider_options={"systemInstruction": {"parts": [{"text": "Always answer in German."}]}}) ``` ## 8. Links * Docs home: <https://ai.google.dev/gemini-api/docs> * Document (PDF) understanding: <https://ai.google.dev/gemini-api/docs/document-processing> * Image understanding: <https://ai.google.dev/gemini-api/docs/image-understanding> * Structured output: <https://ai.google.dev/gemini-api/docs/structured-output> * Models: <https://ai.google.dev/gemini-api/docs/models> · live list: `GET /v1beta/models` * Pricing: <https://ai.google.dev/gemini-api/docs/pricing> * Rate limits: <https://ai.google.dev/gemini-api/docs/rate-limits> * Files API: <https://ai.google.dev/gemini-api/docs/files> * `generateContent` reference: <https://ai.google.dev/api/generate-content> * API versions: <https://ai.google.dev/gemini-api/docs/api-versions> * Interactions API (the newer surface PuffinParse does *not* use): <https://ai.google.dev/gemini-api/docs/migrate-to-interactions> --- # OpenAI <!-- source: docs/providers/openai.md | url: https://puffinparse.com/docs/providers/openai/ --> > **Status: docs-only.** Implemented from OpenAI's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `openai` | | Base URL | `https://api.openai.com` (override: `base_url` on the request, or `OPENAI_BASE_URL`) | | API key | `OPENAI_API_KEY` (or `api_key` on the request) — sent as `Authorization: Bearer sk-…` | | Docs | <https://developers.openai.com/api/docs> (the old `platform.openai.com/docs/*` links 301 here) | | Endpoint | `POST /v1/responses` — one synchronous call per document, no job id, no polling | | Modes | `parse`, `ocr` (derived from `parse`), `extract` | | Checked against | 2026-09-11 **against the published docs only** — no OpenAI key was available, so the fixtures are built from the documented response shape and the live tests are `#[ignore]` | | Implementation | `crates/puffinparse-core/src/providers/openai.rs` | This is not a document-parsing product: it is a general vision LLM asked, with a strict JSON schema, to transcribe a document page by page. That buys layout-aware markdown of figures, handwriting and messy scans, and costs you geometry — **there are no bounding boxes, no per-block types and no confidences**. Every page comes back as a single `text` block whose `content` is the page markdown. ## 2. Models exposed by PuffinParse | Model | Provider parameters PuffinParse sets | Estimated price (`pricing.json`) | |---|---|---| | `openai/gpt-5.6-luna` *(default)* | `model=gpt-5.6-luna`, `reasoning.effort=low` | ~$0.00114 / page | | `openai/gpt-5.6-terra` | `model=gpt-5.6-terra`, `reasoning.effort=low` | ~$0.0114 / page | | `openai/gpt-5.6-sol` | `model=gpt-5.6-sol`, `reasoning.effort=low` | ~$0.0200 / page | | `openai/gpt-6-astra` | `model=gpt-6-astra`, `reasoning.effort=low` | ~$0.0500 / page | **Pricing is per token, not per page**, so `pricing.json` holds an *estimate*: **1,500 input + 700 output tokens per page** (OpenAI bills a PDF page as extracted text *plus* a page image; 1.5k input tokens is the low end of the published 1,500–3,000 range, so a dense page costs more). Rates per 1M tokens (<https://developers.openai.com/api/docs/pricing>, 2026-09-11): luna $0.20/$1.20, terra $2/$12, sol $4/$20, astra $10/$50. The authoritative number for a call is `response.usage.provider_cost_usd`, which PuffinParse computes from the **actual** `usage.input_tokens`/`usage.output_tokens` and the per-token table embedded in `openai.rs` (`PRICES`, dated `2026-09-11`). It always wins over the per-page estimate, so `cost_usd` is exact even when the estimate is not. `metadata.openai_input_tokens` / `openai_output_tokens` carry the raw counts. Any other model can be reached without a registry change via `provider_options={"model": "gpt-4.1-mini"}`; the price table covers the gpt-5.x, gpt-4.1 and gpt-4o families as well, and cost falls back to `None` for a model it does not know. ## 3. Request flow PuffinParse uses One call. `POST {base}/v1/responses`, `Authorization: Bearer …`, JSON body: ```jsonc { "model": "gpt-5.6-luna", "input": [{ "role": "user", "content": [ // PDFs — the filename matters, the model sees it: {"type": "input_file", "filename": "invoice.pdf", "file_data": "data:application/pdf;base64,JVBER…"}, // …or, for images: {"type": "input_image", "image_url": "data:image/png;base64,iVBOR…", "detail": "high"} {"type": "input_text", "text": "Transcribe this document. …"} ] }], "text": {"format": {"type": "json_schema", "name": "puffinparse_pages", "schema": { … }, "strict": true}}, "max_output_tokens": 32000, "reasoning": {"effort": "low"}, "store": false } ``` * **Input handling.** Path and bytes inputs are read locally; a **URL input is downloaded by PuffinParse and inlined** as a `data:` URL (the API would only fetch public URLs, and the response is identical either way). MIME type comes from the URL's `Content-Type` when it is specific, else from the filename extension. `application/pdf` → `input_file`; `image/*` → `input_image`; anything else is an `input_error` before any network call. * **`parse` schema** (`name: puffinparse_pages`): `{"pages": [{"page_number": <int>, "markdown": <string>}]}`. * **`extract` schema** (`name: puffinparse_extraction`): the caller's JSON Schema, passed through [the strict sanitiser](#6-gotchas). * **`pages` and `language` are prompt-level**, not API parameters: the prompt says "transcribe only pages 1-3,7, keeping the original page numbers" and "the document is mainly in de". Page selection is therefore best-effort — the whole document is still uploaded and billed. * **`output="text"`** adds a rule telling the model to emit plain text instead of Markdown. * **`store: false`** by default so the document is not retained for the response store; set `provider_options={"store": true}` if you want to fetch the response later. * **`temperature: 0` is only sent to gpt-4.x models.** gpt-5.x and gpt-6 are reasoning models and return `400 Unsupported value: 'temperature' does not support 0 with this model` — PuffinParse sends `reasoning: {"effort": "low"}` to those instead (transcription is perception, not reasoning). * **`provider_options` is merged into the body verbatim** (deep merge, so `{"reasoning": {"effort": "medium"}}` replaces only the effort). The key `strict` is consumed by PuffinParse and never forwarded. ## 4. Response mapping | Responses API field | PuffinParse unified field | Notes | |---|---|---| | `id` (`resp_…`) | `ParseResponse.provider_job_id` | | | `output[].content[].text` where `type == "output_text"` | parsed as JSON → the structured answer | Parts are concatenated in order; `reasoning` items are skipped. A `refusal` part becomes a `provider` error. If nothing is found, the SDK-style `output_text` aggregate is used as a fallback. | | `pages[].page_number` | `Page.page_number` | Missing or `0` falls back to the array index + 1. | | `pages[].markdown` | `Page.markdown` (trimmed) | With `output="text"` this is already plain text. | | — | `Page.text` | `markdown_to_text(markdown)`, or the string itself in text mode. | | — | `Page.blocks` | Exactly one `text` block per page, `content` = page markdown, `bbox: None`, `confidence: None`. | | — | `Page.width` / `height` | Always `None` — the API reports no page geometry. | | number of returned pages | `Usage.pages` | The API never reports a page count; extract mode reports `0`. | | `usage.input_tokens` / `output_tokens` | `Usage.provider_cost_usd` (and `metadata.openai_*_tokens`) | Cost = tokens × the embedded per-token prices, so it is exact. | | — | `Usage.credits` | Always `None` — OpenAI has no credit unit. | | `status: "incomplete"` + `incomplete_details.reason` | `provider` error | Typically `max_output_tokens`. | | `status: "failed"` + `error.message` | `provider` error | | | everything else (`reasoning`, `annotations`, `output_tokens_details`, …) | — | Visible with `include_raw=True`. | `extract` mode returns the tool-free JSON object as `ExtractResponse.data`. `ExtractResponse.fields` is always empty: the Responses API has no per-field grounding, so `citations=True` is recorded as `metadata.openai_citations_unsupported = true` rather than silently pretending to support it. Trimmed response (`crates/puffinparse-core/tests/fixtures/openai_responses_parse.json`, built from the documented shape): ```json { "id": "resp_68f0a1b2c3d4e5f60123456789abcdef", "object": "response", "status": "completed", "model": "gpt-5.6-luna", "output": [ { "id": "rs_…", "type": "reasoning", "summary": [] }, { "id": "msg_…", "type": "message", "status": "completed", "role": "assistant", "content": [{ "type": "output_text", "annotations": [], "text": "{\"pages\":[{\"page_number\":1,\"markdown\":\"# Hello PuffinParse\\n\\nInvoice #1234…\"}]}" }] } ], "text": { "format": { "type": "json_schema", "name": "puffinparse_pages", "strict": true } }, "usage": { "input_tokens": 3120, "output_tokens": 412, "output_tokens_details": { "reasoning_tokens": 64 }, "total_tokens": 3532 }, "store": false } ``` ## 5. Errors, status codes, rate limits, timeouts The error envelope is `{"error": {"message", "type", "code", "param"}}`; `Error::from_http` picks up the nested `error.message` and classifies by status: | Status | Trigger | PuffinParse `ErrorKind` | |---|---|---| | 400 | invalid schema for `text.format`, unsupported parameter (`temperature` on a reasoning model), file too large / unreadable | `bad_request` | | 401 | missing or revoked key | `authentication` | | 403 | key/region or model not permitted for the org | `authentication` | | 404 | unknown model id | `bad_request` | | 429 | rate limit **or** insufficient quota (`code: "insufficient_quota"` — not actually retryable) | `rate_limit` (retried) | | 5xx | `server_error`, `502`, `503` | `provider` (retried) | **Retries.** `max_retries` (default 2), exponential backoff with full jitter, on rate-limit, network and 500/502/503/504 only. 4xx is never retried. **Rate limits.** Per-model RPM/TPM quotas by usage tier; responses carry `x-ratelimit-remaining-*` and `retry-after` headers, which PuffinParse does **not** read yet — backoff is blind. **Limits.** A file must stay under 50 MB, and all files in one request under 50 MB total; a PDF page is billed as extracted text *plus* a page image, so context, not page count, is the practical limit. `max_output_tokens` defaults to 32,000 (~20 dense pages); a longer document truncates with `status: "incomplete"`, `reason: "max_output_tokens"`. **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline and also caps the single HTTP request. Big documents on a reasoning model are slow: budget minutes, not seconds. ## 6. Gotchas * **Strict structured outputs restrict JSON Schema.** With `strict: true`, every object must set `"additionalProperties": false` and list **every** property in `required`; there are no optional fields — express "may be absent" as a nullable type (`"type": ["string", "null"]`). The root must be an object. PuffinParse applies `sanitize_strict_schema()` to the caller's `extract` schema recursively (through `properties`, `items`, `prefixItems`, `$defs`/`definitions`, `anyOf`/`oneOf`/ `allOf`, `if`/`then`/`else`, `not`, `contains`), rewriting `required` to the full property list and forcing `additionalProperties: false`. **This widens `required`**: fields you marked optional come back as `null` rather than missing. * **Some keywords are still rejected in strict mode** (`minimum`/`maximum`, `minLength`/`maxLength`, `pattern`, recursive `$ref` beyond `#`). If the API answers `400 Invalid schema`, either drop those keywords or turn the feature off with `provider_options={"strict": false}` — PuffinParse then sends your schema untouched and asks for JSON by prompt instead. * **`temperature` is not accepted by gpt-5.x / gpt-6.** See §3. If you force it through `provider_options`, expect a 400. * **Reasoning tokens are billed as output tokens** and are invisible in the text. `effort: "low"` keeps them small; raise it with `provider_options={"reasoning": {"effort": "medium"}}` if a dense or handwritten document transcribes badly. * **No geometry, ever.** `Block.bbox`, `Block.confidence` and `Page.width/height` are `None`, and every page holds exactly one `text` block. Anything that draws overlays needs a layout provider (Reducto, Extend, Azure, …). `ocr` mode is derived from `parse` by the default trait method, so its `words[]` carry no boxes either, and `metadata.puffinparse_derived_from = "parse"` says so. * **`pages` is a prompt instruction, not an API parameter.** The whole document is uploaded and billed even when you ask for one page, and the model may ignore the restriction on a bad day. Split the PDF client-side if this matters. * **Page numbers come from the model.** They are usually the document's own, but a model can repeat or skip one; PuffinParse falls back to the array position when the number is missing or `0`, and sorts pages by number when assembling the document markdown. * **Hallucination is the failure mode.** A layout parser drops text it cannot read; an LLM can invent plausible text instead. The prompt forbids it ("transcribe verbatim, never invent"), but for high-stakes extraction prefer a provider that returns citations. * **`store: false` is PuffinParse's default** — responses are not kept in OpenAI's response store, which also means `previous_response_id` chaining is unavailable unless you opt back in. * **`detail: "high"` is set on image inputs** for legibility of small print; `{"detail": "low"}` through `provider_options` on the input block is not possible (the block is built by PuffinParse) — use the `openai/gpt-5.6-luna` model and a downsampled image instead if you need to save tokens. ## 7. Useful `provider_options` passthrough ```python # 1. Longer documents: raise the output ceiling (the default 32k covers ~20 dense pages). puffinparse.parse("contract.pdf", model="openai/gpt-5.6-luna", provider_options={"max_output_tokens": 64000}) # 2. Harder documents: more reasoning (billed as output tokens). puffinparse.parse("handwritten.pdf", model="openai/gpt-5.6-terra", provider_options={"reasoning": {"effort": "medium"}}) # 3. A model that is not in the registry (pricing falls back to the embedded token table). puffinparse.parse("scan.png", model="openai/gpt-5.6-luna", provider_options={"model": "gpt-4.1-mini", "temperature": 0}) # 4. Extraction with a schema that uses keywords strict mode rejects. puffinparse.extract("invoice.pdf", schema=schema_with_patterns, model="openai/gpt-5.6-luna", provider_options={"strict": False}) # 5. Keep the response in OpenAI's response store, and tag it. puffinparse.parse("report.pdf", model="openai/gpt-5.6-sol", provider_options={"store": True, "metadata": {"run": "bench-2026-09"}}) ``` ## 8. Links * API reference: <https://developers.openai.com/api/docs/api-reference/responses> * PDF / file inputs: <https://developers.openai.com/api/docs/guides/pdf-files> * Images and vision: <https://developers.openai.com/api/docs/guides/images-vision> * Structured outputs: <https://developers.openai.com/api/docs/guides/structured-outputs> * Reasoning and `effort`: <https://developers.openai.com/api/docs/guides/reasoning> * Models: <https://developers.openai.com/api/docs/models> · Pricing: <https://developers.openai.com/api/docs/pricing> * Rate limits: <https://developers.openai.com/api/docs/guides/rate-limits> --- # Anthropic (Claude) <!-- source: docs/providers/anthropic.md | url: https://puffinparse.com/docs/providers/anthropic/ --> > **Status: docs-only.** Implemented from Anthropic's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `anthropic` | | Base URL | `https://api.anthropic.com` (override: `base_url` on the request, or `ANTHROPIC_BASE_URL`) | | API key | `ANTHROPIC_API_KEY` (or `api_key` on the request) — sent as `x-api-key: sk-ant-…` | | Docs | <https://platform.claude.com/docs> (the old `docs.anthropic.com/en/*` links 301 here) | | Endpoint | `POST /v1/messages` with `anthropic-version: 2023-06-01` — one synchronous call, no job id, no polling | | Modes | `parse`, `ocr` (derived from `parse`), `extract` | | Checked against | 2026-09-11 **against the published docs only** — no Anthropic key was available, so the fixtures are built from the documented response shape and the live tests are `#[ignore]` | | Implementation | `crates/puffinparse-core/src/providers/anthropic.rs` | Like `openai`, this is a general vision LLM rather than a document-parsing product: Claude is given the PDF (each page as text *plus* a page image) and asked, through a **forced tool call**, to return one markdown transcription per page. You get layout-aware markdown of tables, figures and handwriting, and you lose geometry — **no bounding boxes, no per-block types, no confidences**. ## 2. Models exposed by PuffinParse | Model | Provider parameters PuffinParse sets | Estimated price (`pricing.json`) | |---|---|---| | `anthropic/claude-sonnet-5` *(default)* | `model=claude-sonnet-5`, `output_config.effort=low` | ~$0.0100 / page | | `anthropic/claude-haiku-4-5` | `model=claude-haiku-4-5`, `temperature=0` | ~$0.0050 / page | | `anthropic/claude-opus-5` | `model=claude-opus-5`, `output_config.effort=low` | ~$0.0250 / page | **Pricing is per token, not per page**, so `pricing.json` holds an *estimate*: **1,500 input + 700 output tokens per page** — Anthropic's own guidance is 1,500–3,000 text tokens per page *plus* the page image's visual tokens, so a dense page costs more than the estimate. Rates per 1M tokens (<https://platform.claude.com/docs/en/about-claude/pricing>, 2026-09-11): Sonnet 5 $2/$10, Haiku 4.5 $1/$5, Opus 5 $5/$25. `response.usage.provider_cost_usd` is computed from the **actual** `usage.input_tokens` / `usage.output_tokens` and the per-token table embedded in `anthropic.rs` (`PRICES`, dated `2026-09-11`), so `cost_usd` is exact even where the per-page estimate is not. `metadata.anthropic_input_tokens` / `anthropic_output_tokens` carry the raw counts. Other model ids work without a registry change via `provider_options={"model": "claude-opus-4-8"}`; the embedded table covers the Opus 4.6–5, Sonnet 4.6/5, Haiku 4.5 and Fable 5/5.1 ids, and cost falls back to `None` for anything it does not know. **Claude Fable 5.1 / Mythos 5.1 do not work here**: they reject forced `tool_choice` with a 400 (see §6). ## 3. Request flow PuffinParse uses One call. `POST {base}/v1/messages` with `x-api-key`, `anthropic-version: 2023-06-01`, `content-type: application/json`: ```jsonc { "model": "claude-sonnet-5", "max_tokens": 32000, "system": "You are a precise document transcription and extraction engine. …", "messages": [{ "role": "user", "content": [ // PDFs (documents before text — Claude does better that way): {"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "JVBER…"}}, // …or, for images: {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "iVBOR…"}} {"type": "text", "text": "Transcribe this document. …"} ] }], "tools": [{"name": "emit_pages", "description": "…", "input_schema": { … }}], "tool_choice": {"type": "tool", "name": "emit_pages"}, "output_config": {"effort": "low"} } ``` * **Structured output is a forced tool call.** One tool is defined whose `input_schema` is the wanted shape, and `tool_choice` pins it, so the model must answer with a `tool_use` block whose `input` is the JSON. (`output_config.format` structured outputs would also work on these models; the tool route keeps `extract` and `parse` on one code path.) * `parse` → tool `emit_pages`, schema `{"pages": [{"page_number": <int>, "markdown": <string>}]}`. * `extract` → tool `record_extraction`, schema = the caller's JSON Schema. * **Input handling.** Path and bytes inputs are read locally; a **URL input is downloaded by PuffinParse and inlined** as base64 (Claude also accepts `source: {"type": "url"}`, but only for publicly reachable URLs). `application/pdf` → `document` block; `image/jpeg|png|gif|webp` → `image` block; anything else is an `input_error` before any network call. * **`pages` and `language` are prompt-level**, not API parameters: the prompt says "transcribe only pages 2-4, keeping the original page numbers" and "the document is mainly in fr". Page selection is best-effort — the whole document is still uploaded and billed. * **`output="text"`** adds a rule telling the model to emit plain text instead of Markdown. * **`temperature: 0` is only sent to Claude 4.5-era and older models** (`claude-haiku-4-5`, `claude-sonnet-4-5`, `claude-opus-4-5`, `claude-3*`). Claude 4.6 and later removed sampling parameters and return a 400 if you send them. * **`output_config: {"effort": "low"}` is only sent to models that support effort** (Opus 4.6–5, Sonnet 4.6/5). Haiku 4.5 rejects it. Effort controls how much (billed) thinking Claude does; transcription is perception, not reasoning, so `low` is the default. * **`thinking` is never sent.** Adaptive thinking (on by default on Sonnet 5 / Opus 5) is compatible with forced tool use; *manual* extended thinking (`thinking: {"type": "enabled"}`) is not and would break the forced call, so do not add it through `provider_options`. * **`provider_options` is merged into the body verbatim** (deep merge). The key `strict` is consumed by PuffinParse and never forwarded. ## 4. Response mapping | Messages API field | PuffinParse unified field | Notes | |---|---|---| | `id` (`msg_…`) | `ParseResponse.provider_job_id` | The `request-id` response header is not read. | | `content[]` block with `type == "tool_use"` and the expected `name` → `.input` | the structured answer | Any `thinking` / `text` blocks before it are skipped. | | `pages[].page_number` | `Page.page_number` | Missing or `0` falls back to the array index + 1. | | `pages[].markdown` | `Page.markdown` (trimmed) | With `output="text"` this is already plain text. | | — | `Page.text` | `markdown_to_text(markdown)`, or the string itself in text mode. | | — | `Page.blocks` | Exactly one `text` block per page, `content` = page markdown, `bbox: None`, `confidence: None`. | | — | `Page.width` / `height` | Always `None` — the API reports no page geometry. | | number of returned pages | `Usage.pages` | The API never reports a page count; extract mode reports `0`. | | `usage.input_tokens` / `output_tokens` | `Usage.provider_cost_usd` (and `metadata.anthropic_*_tokens`) | Cost = tokens × the embedded per-token prices, so it is exact. Cache and tier fields are not modelled. | | — | `Usage.credits` | Always `None` — Anthropic has no credit unit. | | `stop_reason: "max_tokens"` | `provider` error | "response truncated … raise provider_options.max_tokens". | | `stop_reason: "refusal"` (+ `stop_details.category`) | `provider` error | | | no `tool_use` block at all | `provider` error | The beginning of the model's prose answer is included in the message. | `extract` mode returns the tool input as `ExtractResponse.data`. `ExtractResponse.fields` is always empty: Claude's citations feature grounds *text* answers and cannot be combined with the forced tool call, so `citations=True` is recorded as `metadata.anthropic_citations_unsupported = true`. Trimmed response (`crates/puffinparse-core/tests/fixtures/anthropic_messages_parse.json`, built from the documented shape): ```json { "id": "msg_01XhT9Eq8bQ7vPz2kKcM4dLp", "type": "message", "role": "assistant", "model": "claude-sonnet-5", "content": [ { "type": "tool_use", "id": "toolu_01A09q90qw90lq917835lq9", "name": "emit_pages", "input": { "pages": [ { "page_number": 1, "markdown": "# Hello PuffinParse\n\nInvoice #1234…" } ] } } ], "stop_reason": "tool_use", "stop_sequence": null, "usage": { "input_tokens": 3210, "cache_creation_input_tokens": 0, "cache_read_input_tokens": 0, "output_tokens": 389, "service_tier": "standard" } } ``` ## 5. Errors, status codes, rate limits, timeouts The error envelope is `{"type": "error", "error": {"type", "message"}, "request_id"}`; `Error::from_http` picks up the nested `error.message` and classifies by status: | Status | Provider error type | PuffinParse `ErrorKind` | |---|---|---| | 400 | `invalid_request_error` — bad schema, `max_tokens` above the model's ceiling, `temperature` on a 4.6+ model, forced `tool_choice` on Fable 5.1 | `bad_request` | | 401 | `authentication_error` | `authentication` | | 402 | `billing_error` | `bad_request` | | 403 | `permission_error` | `authentication` | | 404 | `not_found_error` (unknown model id) | `bad_request` | | 413 | `request_too_large` (the request is over 32 MB) | `bad_request` | | 429 | `rate_limit_error` | `rate_limit` (retried) | | 500 / 504 | `api_error` / `timeout_error` | `provider` (retried) | | **529** | `overloaded_error` | `provider` (retried) — **reported as HTTP 503** | **529 is rewritten to 503.** PuffinParse's shared retry policy only treats 500/502/503/504 as transient, so `map_http_error()` reports an overloaded 529 as status `503` and prefixes the message with `overloaded (HTTP 529):`. `Error.status_code` is therefore `503` for this case — the real status is in the message. **Retries.** `max_retries` (default 2), exponential backoff with full jitter, on rate-limit, network, 500/502/503/504 and (via the rewrite) 529. The `retry-after` header is **not** read yet. **Limits.** 32 MB per request; 600 PDF pages per request (100 on models with a context window under 1M tokens); images up to 8000×8000 px and 10 MB each, JPEG/PNG/GIF/WebP only; no password-protected PDFs. `max_tokens` defaults to 32,000 (~20 dense pages) — a longer document stops with `stop_reason: "max_tokens"` and PuffinParse turns that into a provider error rather than returning half a document. **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline and also caps the single HTTP request. Anthropic recommends streaming or the Batch API beyond ~10 minutes; PuffinParse does neither, so keep documents small enough to finish inside the deadline. ## 6. Gotchas * **Forced tool use is not universal.** Claude Fable 5.1 and Mythos 5.1 reject `tool_choice: {"type": "tool"}` with `400 tool_choice: type "tool" and "any" are not supported for this model`, so they are deliberately absent from the registry. Manual extended thinking (`thinking: {"type": "enabled"}`) has the same restriction — do not add it via `provider_options`. * **Sampling parameters are gone on Claude 4.6+.** `temperature`, `top_p` and `top_k` return a 400 on Sonnet 5 / Opus 5 and the 4.6+ family. PuffinParse only sends `temperature: 0` to the older models that still accept it; determinism on the newer ones comes from the schema and the prompt, not sampling. * **`output_config.effort` is model-gated.** Haiku 4.5 and older reject it; PuffinParse sends it only to models in `EFFORT_MODELS`. Raise it (`{"output_config": {"effort": "medium"}}`) for hard scans. * **Thinking tokens are billed as output tokens.** On Sonnet 5 / Opus 5 adaptive thinking is on by default (display omitted, so you never see it); `effort: "low"` keeps the bill down. * **Strict tool use is opt-in here.** Unlike the OpenAI provider, PuffinParse does *not* set `strict: true` on the tool by default, because Claude accepts many JSON Schema keywords that strict mode rejects (`minimum`/`maximum`, `minLength`/`maxLength`, recursive schemas). Pass `provider_options={"strict": true}` to turn it on; PuffinParse then also sanitises your schema recursively — `additionalProperties: false` everywhere and **every** property moved into `required`, which means fields you marked optional come back as `null` instead of missing. * **No geometry, ever.** `Block.bbox`, `Block.confidence` and `Page.width/height` are `None`, and every page holds exactly one `text` block. `ocr` mode is derived from `parse` by the default trait method (`metadata.puffinparse_derived_from = "parse"`), so its `words[]` carry no boxes either. * **`pages` is a prompt instruction, not an API parameter.** The whole document is uploaded and billed even when you ask for one page. Split the PDF client-side if that matters. * **Page numbers come from the model**, so a model can repeat or skip one; PuffinParse falls back to the array position when the number is missing or `0` and sorts pages by number. * **Hallucination is the failure mode.** A layout parser drops what it cannot read; an LLM can invent plausible text. The prompt forbids it ("transcribe verbatim, never invent"), but do not use this provider for high-stakes extraction without review. * **Claude will not identify people in images** and refuses documents that violate the AUP; such a response arrives as HTTP 200 with `stop_reason: "refusal"`, which PuffinParse maps to a provider error. * **Prompt caching is not used.** Every call re-uploads the document; for repeated parses of the same file, add `cache_control` through `provider_options` yourself. ## 7. Useful `provider_options` passthrough ```python # 1. Longer documents: raise the output ceiling (the default 32k covers ~20 dense pages). puffinparse.parse("contract.pdf", model="anthropic/claude-sonnet-5", provider_options={"max_tokens": 64000}) # 2. Harder documents: more thinking (billed as output tokens). puffinparse.parse("handwritten.pdf", model="anthropic/claude-opus-5", provider_options={"output_config": {"effort": "medium"}}) # 3. A model that is not in the registry (pricing falls back to the embedded token table). puffinparse.parse("scan.png", model="anthropic/claude-sonnet-5", provider_options={"model": "claude-opus-4-8"}) # 4. Strict tool use for extraction (schema is sanitised: all fields become required). puffinparse.extract("invoice.pdf", schema=invoice_schema, model="anthropic/claude-sonnet-5", provider_options={"strict": True}) # 5. Cache the document across repeated calls (5-minute ephemeral cache). puffinparse.parse("handbook.pdf", model="anthropic/claude-haiku-4-5", provider_options={"cache_control": {"type": "ephemeral"}}) ``` ## 8. Links * Messages API: <https://platform.claude.com/docs/en/api/messages> * PDF support: <https://platform.claude.com/docs/en/build-with-claude/pdf-support> * Vision (images, limits, visual tokens): <https://platform.claude.com/docs/en/build-with-claude/vision> * Tool use / forcing a tool: <https://platform.claude.com/docs/en/agents-and-tools/tool-use/define-tools> * Structured outputs: <https://platform.claude.com/docs/en/build-with-claude/structured-outputs> * Errors and request-size limits: <https://platform.claude.com/docs/en/api/errors> * Models: <https://platform.claude.com/docs/en/models/overview> · Pricing: <https://platform.claude.com/docs/en/about-claude/pricing> * Rate limits: <https://platform.claude.com/docs/en/api/rate-limits> --- # Mathpix <!-- source: docs/providers/mathpix.md | url: https://puffinparse.com/docs/providers/mathpix/ --> > **Status: docs-only.** Everything below is taken from the official Mathpix documentation > (read 2026-09-11) and from the implementation in `crates/puffinparse-core/src/providers/mathpix.rs`. > No live call has been made — this repository has no Mathpix credentials. The fixtures under > `crates/puffinparse-core/tests/fixtures/mathpix_*.json` are hand-built from the documented response > shapes, not captured traffic. Mark this page **verified** only after the `#[ignore]`d live tests > in `providers/mathpix.rs` pass with a real key pair. > Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `mathpix` | | Base URL | `https://api.mathpix.com` (override: `base_url` on the request, or `MATHPIX_BASE_URL`) | | API key | **two** values: `MATHPIX_APP_ID` + `MATHPIX_APP_KEY`, sent as the `app_id` and `app_key` **headers** (no `Authorization:` header). `api_key` on the request overrides `MATHPIX_APP_KEY`; `provider_options={"app_id": …}` overrides `MATHPIX_APP_ID` | | Docs | <https://docs.mathpix.com> | | API version | Path-versioned (`/v3/...`). No version header; the model is reported per response as `version` (e.g. `SuperNet-200`) | | Checked against | **not live-verified** (see banner) — documentation read 2026-09-11 | | Implementation | `crates/puffinparse-core/src/providers/mathpix.rs` | Mathpix is an OCR engine rather than a layout parser: it is built for STEM content (printed *and* handwritten math, tables, chemistry diagrams) and its native output is **Mathpix Markdown (MMD)**, a markdown superset that can contain LaTeX (`$…$`, `\begin{tabular}`, `\section*{}`, `<smiles>…</smiles>`). PuffinParse passes MMD through unchanged. ## 2. Models exposed by PuffinParse | Model | Endpoint PuffinParse calls | Modes | List price (`pricing.json`) | |---|---|---|---| | `mathpix/pdf` *(default)* | `POST /v3/pdf` for documents; automatically `POST /v3/text` when the input is an image | `parse`, `ocr` | $0.005 / page | | `mathpix/text` | always `POST /v3/text` (one image = one request) | `parse`, `ocr` | $0.002 / image | Prices from <https://mathpix.com/pricing/api>: `v3/pdf` $5 per 1 000 pages (falling to $3.50 above 1M pages/month), `v3/text` $0.002 per image (falling to $0.0015 above 1M). A one-time **$19.99 setup fee** activates the first API key, and the asynchronous Files API (`files/v1/*`, $1.50 per 1 000 pages) is **not** used by PuffinParse. **Why two models.** `/v3/pdf` accepts documents and ebooks only (PDF, EPUB, DOCX, DOC, PPTX, AZW/ AZW3/KFX, MOBI, DJVU, WPD, ODT), while images (JPEG, PNG, BMP, JP2, WebP, PBM/PGM/PPM, PFM, Sun raster, TIFF, OpenEXR, HDR) are only accepted by `/v3/text`. `mathpix/pdf` therefore routes by the input's guessed MIME type — `image/*` goes to `/v3/text`, everything else to `/v3/pdf` — so one model string works for a mixed workload. `mathpix/text` is the explicit escape hatch when you want the cheaper per-image rate and snippet behaviour; sending it a PDF fails with `image_decode_error`. ## 3. Request flow PuffinParse uses ### 3.1 Documents (`mathpix/pdf` with a non-image input) 1. **Submit.** `POST {base}/v3/pdf` with `app_id` + `app_key`. * path / bytes input → `multipart/form-data` with a `file` part and **all options as one stringified JSON field, `options_json`** (this is how Mathpix takes options on multipart); * URL input → a JSON body `{"url": "…", …options}` (Mathpix downloads the file itself). Options PuffinParse always sends: `math_inline_delimiters: ["$","$"]` and `math_display_delimiters: ["$$","$$"]` (markdown-friendly instead of the default `\(…\)` / `\[…\]`). `pages` becomes `page_ranges` (see §4). `provider_options` are deep-merged last and win. Response: `{"pdf_id": "2026_01_15_abc123def456"}`. 2. **Poll.** `GET {base}/v3/pdf/{pdf_id}` every 2 s, backing off ×1.5 to 10 s, until `status == "completed"` (or `"error"`). Intermediate statuses are `received`, `loaded`, `split`. PuffinParse deliberately ignores `percent_done` / `num_pages_completed`: both reach 100 % while the output files are still being assembled, and a download at that moment 404s. 3. **Download line data.** `GET {base}/v3/pdf/{pdf_id}.lines.json` — per-page lines with polygons. A `404` (body = the status object) or `202` means "not ready yet" and PuffinParse keeps polling for it until the deadline; any other non-2xx is an error. 4. **Download markdown** (`parse` mode only). `GET {base}/v3/pdf/{pdf_id}.mmd` — the assembled Mathpix Markdown for the whole document. `ocr` mode skips this call. `.mmd` and `.lines.json` are generated automatically for every document ("Availability: Always"); they are *not* valid `conversion_formats` keys and cost nothing extra. ### 3.2 Images (`mathpix/text`, or `mathpix/pdf` with an image input) One synchronous call: `POST {base}/v3/text`, multipart `file` + `options_json`, or a JSON body with `{"src": "<url>"}` for URL inputs. Options PuffinParse sends: | Option | Value | Why | |---|---|---| | `formats` | `["text"]` | Mathpix Markdown output. Add `"data"`/`"html"` through `provider_options` if you want TSV/LaTeX/MathML alongside | | `include_line_data` | `true` | line polygons + per-line text (the source of `Block`s / `Line`s) | | `enable_document_layout` | `true` — **`parse` mode only** | full-page layout recognition (nested lists, pseudocode). Off by default on `/v3/text`, which otherwise assumes a snippet | | `include_word_data` | `true` — **`ocr` mode only** | word polygons for `TextPage.words` | | `math_inline_delimiters` / `math_display_delimiters` | `["$","$"]` / `["$$","$$"]` | markdown-friendly math | `enable_document_layout` and `include_word_data` **cannot be combined** (documented, rejected by the API), which is exactly why the two modes send different option sets. ## 4. Response mapping ### 4.1 Documents — `.lines.json` + `.mmd` | Mathpix field | PuffinParse unified field | Notes | |---|---|---| | `pdf_id` | `ParseResponse.provider_job_id` / `TextResponse.provider_job_id` | | | `pages[].page` | `Page.page_number` / `TextPage.page_number` | Already 1-based. | | `pages[].page_width` / `page_height` | `Page.width` / `Page.height` | Pixel coordinate space the polygons live in. | | `pages[].lines[]` | `Page.blocks[]` (parse) · `TextPage.lines[]` (ocr) | Selection rules below. | | `lines[].type` (+ `subtype`) | `Block.type` | See the table below. | | `lines[].text_display` | `Block.content` | The line's MMD, exactly as it appears in the assembled `.mmd`. Falls back to `text` when empty. With `output="text"`, `text` is used instead. | | `lines[].text` | `Block.text` / `Line.text` | Searchable plain text; falls back to `markdown_to_text(text_display)`. | | `lines[].cnt` | `Block.bbox` / `Line.bbox` | Polygon in page pixels, `[TL, TR, BR, BL]`; PuffinParse takes the enclosing axis-aligned box and divides by `page_width`/`page_height`. | | `lines[].confidence` | `Block.confidence` / `Line.confidence` | 0–1, the product of per-token OCR confidence. `confidence_rate` (geometric mean) is not mapped. | | `.mmd` body | `ParseResponse.markdown` | Mathpix's own rendering of the whole document wins over the pages joined together. Page-level `markdown` stays line-derived so text and boxes agree. | | status `num_pages` | `Usage.pages` | Falls back to the number of pages in `lines.json`. | | `region`, `is_printed`, `is_handwritten`, `links`, `children_ids`, `parent_id` | — | Not mapped; visible with `include_raw=True`. | **Which lines become blocks.** A line is kept when `conversion_output` is true (missing ⇒ true) and it has content; then any line whose `parent_id` points at another kept line is dropped, so a table's `table_cell` children are not emitted next to the table they already belong to. `ocr` mode does not apply that filter — every line with text is a `Line`, including ones excluded from the MMD, because their geometry is still real. **Words.** `/v3/pdf` has no word-level output, so `TextPage.words` is empty for documents. Only `/v3/text` reports `word_data`. ### 4.2 Images — `/v3/text` | Mathpix field | PuffinParse unified field | Notes | |---|---|---| | `request_id` | `provider_job_id` | | | `text` | `Page.markdown` (page 1) | The whole image's MMD. | | `image_width` / `image_height` | `Page.width` / `Page.height` | Pixel space for `cnt`. | | `line_data[]` | `Page.blocks[]` / `TextPage.lines[]` | Blocks keep only `conversion_output`/`included` lines; OCR lines keep all of them. | | `line_data[].text` | `Block.content` | Per-line MMD. | | `line_data[].cnt` | `Block.bbox` / `Line.bbox` | Normalised by `image_width`/`image_height`. | | `word_data[]` | `TextPage.words[]` | `text` + polygon + `confidence`. | | `confidence` | `metadata.mathpix_confidence` | Whole-image confidence. | | `latex_styled`, `data[]`, `html`, `detected_alphabets`, `auto_rotate_*` | — | Not mapped; visible with `include_raw=True`. | | — | `Usage.pages` | Always `1`: one image is one billed request. | ### 4.3 Block types Mathpix line types (the same vocabulary for `line_data` and PDF lines data): | Mathpix `type` | `BlockType` | |---|---| | `title` | `title` | | `section_header` | `section_header` | | `text`, `abstract`, `authors`, `quote`, `code`, `pseudocode`, `form_field`, `multiple_choice_block`, `multiple_choice_option`, `table_of_contents_row`, `table_of_contents_item`, `column` | `text` | | `math` | `formula` | | `table`, `table_cell` | `table` | | `diagram`, `chart` | `figure` | | `diagram_info`, `chart_info`, `figure_label` | `caption` | | `footnote` | `footnote` | | `page_info` (headers, footers, page numbers, stamps, QR codes) | `header` — except `subtype: "margin_note"` → `footnote` | | everything else (`equation_number`, `qed_symbol`, `rotated_container`, `table_of_contents_container`, …) | `other` | ### 4.4 `pages` selection `pages="1-3,7,10-"` becomes `page_ranges: "1-3,7,10--1"`. Mathpix's `page_ranges` is 1-based like PuffinParse's, and negative indices count from the end, so an open-ended range is closed with `-1` (the last page). For `/v3/text` the option does not exist — an image is a single page — and the selection is ignored. ## 5. Errors, status codes, rate limits, timeouts **The v3 API answers most errors with HTTP 200** and an error body: ```json {"error": "Image has no content", "error_info": {"id": "image_no_content", "message": "Image has no content"}} ``` Only `http_unauthorized` (401) and `http_max_requests` (429) use a real status code. PuffinParse therefore inspects **every** 200 payload (submit, status poll, `.lines.json`, `/v3/text`) and maps `error_info.id`: | `error_info.id` | PuffinParse `ErrorKind` | |---|---| | `http_unauthorized`, `account_disabled`, `expired_license`, `unauthorized_token_request` | `authentication` | | `http_max_requests` (monthly page/image quota **or** per-minute rate) | `rate_limit` (retried) | | `sys_exception`, `connection_closed` | `provider` | | `image_no_content`, `math_confidence`, `math_syntax`, `strokes_no_content` | `provider` (content the engine could not read — a router may fall back to another provider) | | `opts_*`, `json_syntax`, `pdf_missing`, `pdf_encrypted`, `pdf_unknown_id`, `pdf_page_limit_exceeded`, `image_*`, `file_missing`, `sys_request_too_large` | `bad_request` | | an `error` string with no `error_info.id` | `provider` | HTTP-level failures still go through `Error::from_http`, which picks up the `error` / `message` fields and classifies 401/403 → `authentication`, 429 → `rate_limit`, other 4xx → `bad_request`, 5xx → `provider`. **Size and time limits** (from the endpoint reference): `/v3/pdf` accepts files up to **1 GB**; `/v3/text` accepts a 5 MB JSON body, a 2 MB base64 image, a 10 MB image download from `src`, and gives that download 15 s. Per-account page limits per document exist (`pdf_page_limit_exceeded`) and `http_max_requests` carries `limit_name` / `limit_value` / `count` in `error_info`. The published per-minute request ceiling is on the *Limits & Quotas* page, which renders client-side and could not be read here — treat the exact number as unverified. **Retention.** Text outputs (MMD, JSON lines) are kept for up to **90 days**, uploaded source files and CDN image crops for **30 days**. `DELETE /v3/pdf/{pdf_id}` removes everything at once; PuffinParse never deletes on your behalf. **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline: submit + status polling + `.lines.json` + `.mmd`, and it caps each individual HTTP request. ## 6. Gotchas * **Errors hide behind HTTP 200** (see §5). A client that only checks the status code will treat `{"error_info": {"id": "pdf_encrypted"}}` as a successful parse with no pages. * **Two credentials, not one.** `app_id` *and* `app_key`, both as plain headers. `api_key` on the request only replaces the key; the id still comes from `MATHPIX_APP_ID` or `provider_options.app_id`. A missing id is an `authentication` error before any network call. * **`/v3/pdf` rejects images and `/v3/text` rejects PDFs.** `mathpix/pdf` handles that by routing on the input's MIME type; `mathpix/text` does not, by design. * **Poll `status`, not `percent_done`.** `percent_done` reaches 100 % when OCR finishes, which is before the outputs are assembled; downloading then returns `404` with the status object as the body. PuffinParse treats such a `404` (and a `202`, used while a conversion format is still running) as "not ready" and keeps polling. * **MMD is not plain markdown.** Expect `$…$` / `$$…$$` math (PuffinParse asks for those delimiters instead of the default `\(…\)`), `\begin{tabular}` or `\begin{array}` for complex tables, `\section*{}` headings on some documents, `<smiles>…</smiles>` for chemistry, and `\pagebreak` markers if you set `include_page_breaks`. Benchmark scoring against plain-markdown ground truth will punish this; it is the provider's format, not a bug. * **Table cells are children of the table line.** Both the `table` line and its `table_cell` children carry `conversion_output: true`, so PuffinParse drops any line whose `parent_id` is itself kept. Without that rule every table would appear twice. * **`include_page_info` defaults differ per endpoint**: `true` on `/v3/text`, `false` on `/v3/pdf`. Running heads and page numbers are therefore in image output but not in document output unless you ask for them (`provider_options={"include_page_info": true}`). * **`include_word_data` + `enable_document_layout` is rejected**, so parse and ocr modes send different options for images — an `ocr`-mode image call gets snippet-style layout. * **Images with more than 12 rows of text may be billed at the `v3/pdf` per-page rate**, so `mathpix/text` on a full page is not reliably $0.002. * **`conversion_output` supersedes `included`.** `/v3/text` still emits both; `/v3/pdf` lines only carry `conversion_output`. PuffinParse reads `conversion_output` first and defaults to keeping a line when neither is present. * **No `language` support.** Mathpix takes `alphabets_allowed` (which alphabets to *exclude*), not a language hint, so `language` on the request is ignored. Use `provider_options={"alphabets_allowed": {"ru": false}}` if you need it. * **Streaming exists but is unused.** `streaming: true` + `GET /v3/pdf/{id}/stream` (SSE) delivers pages as they finish; PuffinParse polls instead, because the unified response is whole-document. * **The Files API is a different product** (`files/v1/*`, $1.50 per 1 000 pages, results written to your own S3/GCS/Azure bucket) with a *different* error model — real HTTP status codes and a closed error-code set. PuffinParse does not use it. ## 7. Useful `provider_options` passthrough ```python # 1. Credentials in code instead of the environment (app_id is stripped from the request body). puffinparse.parse("paper.pdf", model="mathpix/pdf", api_key=MATHPIX_APP_KEY, provider_options={"app_id": MATHPIX_APP_ID}) # 2. Keep running heads, page numbers and QR codes, and mark page boundaries in the MMD. puffinparse.parse("book.pdf", model="mathpix/pdf", provider_options={"include_page_info": True, "include_page_breaks": True}) # 3. Idiomatic LaTeX for equation-heavy papers, with equation numbers preserved. puffinparse.parse("paper.pdf", model="mathpix/pdf", provider_options={"idiomatic_eqn_arrays": True, "include_equation_tags": True, "math_inline_delimiters": ["\\(", "\\)"]}) # 4. Plain markdown fences and flat lists instead of lstlisting / itemize environments. puffinparse.parse("manual.pdf", model="mathpix/pdf", provider_options={"disable_lstlisting": True, "disable_itemize": True}) # 5. Chemistry + table data on a single image, with the raw payload attached. puffinparse.parse("reaction.png", model="mathpix/text", include_raw=True, provider_options={"include_smiles": True, "formats": ["text", "data"], "data_options": {"include_table_html": True, "include_tsv": True}}) # 6. Ask for a DOCX conversion alongside the parse (downloaded separately from # GET /v3/pdf/{id}.docx once its conversion_status is completed — PuffinParse does not fetch it). puffinparse.parse("report.pdf", model="mathpix/pdf", provider_options={"conversion_formats": {"docx": True}}) ``` ## 8. Links * Docs home: <https://docs.mathpix.com> * Process Documents (`v3/pdf`, status, `.lines.json`, `.mmd`): <https://docs.mathpix.com/reference/post-v3-pdf> * Process Images (`v3/text`, `line_data`, `word_data`): <https://docs.mathpix.com/reference/post-v3-text> * Error handling (the HTTP-200 error model): <https://docs.mathpix.com/reference/error-handling> * Supported formats: <https://docs.mathpix.com/reference/supported-formats> * Limits & quotas: <https://docs.mathpix.com/reference/limits-quotas> * Pricing: <https://mathpix.com/pricing/api> * Mathpix Markdown spec: <https://mathpix.com/docs/mathpix-markdown/overview> * Console (keys, usage): <https://console.mathpix.com> --- # Datalab (Marker) <!-- source: docs/providers/datalab.md | url: https://puffinparse.com/docs/providers/datalab/ --> > **Status: docs-only.** Everything below comes from the official Datalab documentation and > OpenAPI spec (read 2026-09-11), from the open-source Marker renderer that produces the payloads, > and from `crates/puffinparse-core/src/providers/datalab.rs`. No live call has been made — this > repository has no Datalab key. `crates/puffinparse-core/tests/fixtures/datalab_convert.json` is > hand-built from the documented shapes. Mark this page **verified** only after the `#[ignore]`d > live test in `providers/datalab.rs` passes with a real key. > Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `datalab` | | Base URL | `https://www.datalab.to` (override: `base_url` on the request, or `DATALAB_BASE_URL`) | | API key | `DATALAB_API_KEY` (or `api_key` on the request) — sent as the `X-API-Key` header | | Docs | <https://documentation.datalab.to> | | API version | Path-versioned (`/api/v1/convert`). The response echoes the engine versions in `versions` (`marker`, `surya`) | | Checked against | **not live-verified** (see banner) — documentation read 2026-09-11 | | Implementation | `crates/puffinparse-core/src/providers/datalab.rs` | Datalab is the hosted version of **Marker** (plus Surya and Chandra), the open-source PDF → markdown pipeline. The response shapes below are Marker's own: `markdown`, an HTML-carrying block tree (`json`), pre-chunked blocks (`chunks`), and a `metadata` dictionary with `page_stats`. ## 2. Models exposed by PuffinParse | Model | Provider parameter | Modes | List price (`pricing.json`) | |---|---|---|---| | `datalab/fast` | `mode=fast` | `parse`, `ocr` (derived) | $0.004 / page | | `datalab/balanced` *(default)* | `mode=balanced` | `parse`, `ocr` (derived) | $0.004 / page | | `datalab/accurate` | `mode=accurate` | `parse`, `ocr` (derived) | $0.010 / page | From the rate card at <https://www.datalab.to/pricing>: *Convert — fast / balanced* is **$4 per 1 000 pages**, *Convert — accurate* is **$10 per 1 000 pages**. `balanced` is the default because the docs recommend it ("balance of speed and accuracy (recommended)"); `fast` suits clean digital PDFs at high throughput, `accurate` suits scans, dense layouts and complex tables. **Naming note.** These names track the documented `mode` parameter of the **current** `/api/v1/convert` endpoint. The older `/api/v1/marker` endpoint (with `use_llm`, `force_ocr`, `format_lines`) is marked deprecated in the API reference in favour of `/convert`, `/extract`, `/segment` and `/agent`, so there is no `datalab/marker` / `datalab/marker-llm` pair: `use_llm` no longer exists as a request field, and its role is taken by `mode=accurate`. Add-ons that change the bill are **not** enabled by PuffinParse and must be opted into through `provider_options`: `word_bboxes` (+$3/1k pages), `extras="table_cell_bboxes"` or `"list_item_bboxes"` (+$6/1k each, word prediction included once), `chart_understanding` (+$3/1k), `infographic` (+$4/1k), `merge_cross_page` (variable compute, ~$0.50/document), and EU `processing_location` (1.25× usage). ## 3. Request flow PuffinParse uses 1. **Submit.** `POST {base}/api/v1/convert`, `multipart/form-data`, header `X-API-Key`: | Field | Value | |---|---| | `mode` | the model name (`fast` \| `balanced` \| `accurate`) | | `output_format` | `json,markdown` — **both formats in one conversion**, one page charge | | `paginate` | `true` (page delimiters in the markdown) | | `page_range` | from `pages`, **converted 1-based → 0-based**: `"1-3,5"` → `"0-2,4"`; an open range `"10-"` becomes `"9-6999"` (7 000 pages is the per-request ceiling) | | `file` | the document bytes with filename and guessed MIME type — **path / bytes input only** | | `file_url` | the URL string — **URL input only**, in place of `file`; Datalab downloads it | Response: `{"success": true, "request_id": "…", "request_check_url": "https://www.datalab.to/api/v1/convert/…", "versions": {…}}`. 2. **Poll.** `GET` the check URL every 2 s, backing off ×1.5 to 10 s, with the same `X-API-Key`. Done when `status == "complete"`; a failure can also appear as `success == false` (with `status` still `"complete"`) or `status == "failed"`, so all three are terminal for PuffinParse. The returned `request_check_url` is re-hosted on the configured base URL (path only), so a `base_url` override or proxy keeps working. 3. **Download, when regional.** If the poll body carries `result_url`, the document content lives there instead of inline (EU processing, and any other region that requires it). PuffinParse fetches it **without** the API key — the signed URL authorises by itself — and merges: downloaded body first, then every non-null field from the poll response on top, because billing and score fields can be updated after the document was stored. 4. **Normalise.** `json` gives the block tree (types + polygons), `markdown` gives Marker's own page rendering. **Where `provider_options` are merged:** the request is a multipart form, so options are flattened into extra text fields exactly like the built-in ones. Strings pass through; booleans become `"true"`/`"false"`; numbers are stringified; objects/arrays are serialised as JSON text (which is what `additional_config` expects); `null` is skipped. A key PuffinParse already set is **replaced**, so `provider_options={"output_format": "markdown"}` really does turn the JSON block tree off (and with it, all `Block`s). ## 4. Response mapping | Datalab field | PuffinParse unified field | Notes | |---|---|---| | `request_id` | `ParseResponse.provider_job_id` | | | `page_count` | `Usage.pages` | Falls back to the number of reconstructed pages. | | `json.children[]` | one `Page` each | The top-level `json` object is Marker's `Document` block; its children are `Page` blocks. | | page block `id` (`"/page/10/Page/366"`) | `Page.page_number` | The `/page/<n>/` segment is the **0-based page index in the original document**; PuffinParse adds 1. Falls back to the position in `children`. | | page block `polygon` | `Page.width` / `Page.height` | Marker pages start at the origin, so the polygon's max x/y are the page size (PDF points for digital PDFs). | | page block `children[]` | `Page.blocks[]` | Group blocks (`TableGroup`, `FigureGroup`, `ListGroup`, `PictureGroup`) have HTML made only of `<content-ref src=…>` placeholders, so PuffinParse descends into their children instead of emitting the group. | | block `block_type` | `Block.type` | See the table below. | | block `html` | `Block.content` | Converted to markdown best-effort: `<h1>`–`<h6>` → `#`s, `<li>` → `- `, `<table>` → a markdown table when it is simple (rectangular, no `colspan`/`rowspan`, not nested) and the original HTML otherwise, `<math>` → `$$…$$`, everything else → tag-stripped text. With `output="text"` the tag-stripped text is used. | | block `html` (stripped) | `Block.text` | | | block `polygon` | `Block.bbox` | 4 points in page units; PuffinParse takes the enclosing box and divides by the page polygon's size. | | — | `Block.confidence` | Marker reports no per-block confidence. `parse_quality_score` is document-level. | | `markdown` (paginated) | `Page.markdown` | Split on Marker's page markers; preferred over the blocks joined together. `Page.text` is `markdown_to_text` of it. | | `parse_quality_score` | `metadata.datalab_parse_quality_score` | 0–5; < 3.0 is Datalab's own "retry with `accurate`" threshold. | | `cost_breakdown` | `metadata.datalab_cost_breakdown` | Documented as "cost in cents", shape unspecified, so it is passed through verbatim rather than mapped to `Usage.provider_cost_usd`. | | `checkpoint_id` | `metadata.datalab_checkpoint_id` | Only when `save_checkpoint=true` was requested. | | `metadata.failed_pages` | `metadata.datalab_failed_pages` | Only when non-empty. 0-based original page numbers. | | `images`, `metadata.table_of_contents`, `metadata.page_stats`, `versions`, `html`, `chunks`, `runtime` | — | Not mapped; visible with `include_raw=True`. | **Page markdown splitting.** With `paginate=true` Marker writes `\n\n{<page_id>}` + 48 dashes + `\n\n` **before** each page's content (`marker/renderers/markdown.py`), where `page_id` is the same 0-based index used in block ids. PuffinParse recognises any `{n}` + ≥ 8 dashes line, so a custom `page_separator` still works. **Block types** (Marker's vocabulary, `marker/schema/__init__.py`): | Marker `block_type` | `BlockType` | |---|---| | `SectionHeader` rendered as `<h1>` | `title` | | `SectionHeader` (`<h2>`…`<h6>`) | `section_header` | | `Text`, `TextInlineMath`, `Handwriting`, `Form`, `Code`, `Reference`, `Span`, `Line` | `text` | | `ListItem`, `ListGroup` | `list` | | `Table`, `TableGroup`, `TableCell` | `table` | | `Figure`, `FigureGroup`, `Picture`, `PictureGroup` | `figure` | | `Caption` | `caption` | | `Footnote` | `footnote` | | `PageHeader` | `header` | | `PageFooter` | `footer` | | `Equation` | `formula` | | `TableOfContents`, `ComplexRegion`, `Document`, `Page`, anything else | `other` | Marker has no `Title` type: the document title is a `SectionHeader` rendered as `<h1>`, which PuffinParse maps to `title` so the block vocabulary matches the other providers. **`ocr` mode** is derived from `parse` (`TextResponse::from_parse`): lines are block text split on newlines, carrying the block's box; `words` have no geometry. Datalab does have a real word-level product (`word_bboxes=true`, +$3 per 1 000 pages) but it annotates **HTML output** with `data-bbox` / `data-confidence` spans, which PuffinParse does not request or parse. ## 5. Errors, status codes, rate limits, timeouts Every HTTP error is `{"detail": "message"}` (or a FastAPI validation array), which `Error::from_http` picks up: | Status | Datalab type | PuffinParse `ErrorKind` | |---|---|---| | 400 | `invalid_request_error` (bad file type, file too large) | `bad_request` | | 401 | `authentication_error` (`"Invalid API key provided. Set the X-API-Key header…"`) | `authentication` | | 402 | `spend_cap_error` (30-day spend cap reached) | `bad_request` — **not** an auth error, watch for it | | 403 | `permission_error` (no active subscription, expired plan, failed payment) | `authentication` | | 404 | `not_found_error` (request id expired — results live **1 hour**) | `bad_request` | | 413 | `request_too_large` (> 200 MB) | `bad_request` | | 422 | validation error | `bad_request` | | 429 | `rate_limit_error` (requests/min or concurrency) | `rate_limit` (retried) | | 500 / 529 | `api_error` / `overloaded_error` | `provider` (500 retried; 529 is **not** in the retry set) | **Job-level failures** come back with HTTP 200: `{"success": false, "error": "…"}`. PuffinParse turns those into a `provider` error carrying the message and the `request_id` — except the **page concurrency limit**, which is enforced during processing rather than at submission (`"Page rate limit exceeded. Your team has … pages in flight …"`) and is classified as `rate_limit` so a router can back off or fall back. **Limits.** 200 MB per file, 7 000 pages per request, 1 hour result retention. Rate limits are per plan: free tier 25 requests/min and 25 concurrent (the API-limits page still says 10/5), Team 400/400. The page-concurrency ceiling is 5 000 pages in flight per team by default. **Timeouts.** `timeout_secs` (default 300) covers submit + polling + the `result_url` download and caps each request. ## 6. Gotchas * **`/api/v1/marker` is deprecated.** The reference points at `/convert`, `/extract`, `/segment` and `/agent`. `use_llm`, `force_ocr` and `format_lines` are gone; `mode` (`fast`/`balanced`/`accurate`) replaces them. * **`status: "complete"` does not mean success.** Check `success`; a failed document is reported as complete with `success: false` and an `error` string. * **EU results are not inline.** When `result_url` is present the content is only there, and the download must go out **without** the `X-API-Key` header. Keep the non-null poll fields on top of the downloaded body: billing and confidence numbers can change after the document was stored. (Datalab's own Python SDK 0.5.0 does not follow `result_url` for `convert()`.) * **Results are deleted one hour after processing.** There is no way to re-fetch afterwards. * **Page numbers are original, not sequential.** With `page_range="5-7"` the block ids stay `/page/5/…`, so PuffinParse's pages are 6, 7, 8 — deliberately, so boxes and page numbers still refer to the input document. `Usage.pages` is `page_count`, i.e. the pages actually converted. * **`page_range` is 0-based** while PuffinParse's `pages` is 1-based; for spreadsheets the same field selects **sheet** indices instead. * **Group blocks carry no content.** `TableGroup`, `FigureGroup`, `ListGroup` and `PictureGroup` have `html` consisting only of `<content-ref src='…'>` placeholders. PuffinParse replaces them with their children; a naive client that reads group HTML gets empty blocks. * **Blocks are HTML, pages are markdown.** The `json` output never contains markdown unless you set `include_markdown_in_chunks=true`. PuffinParse requests both formats so blocks keep their types and boxes while the page text stays Marker's own rendering. * **Caching is on by default.** A repeated conversion of the same file can be served from cache and `checkpoint_reused: true` means the conversion step is not re-billed — good for cost, fatal for latency benchmarks. Pass `provider_options={"skip_cache": True}` when measuring. * **Spreadsheets bill by cells, not pages** (2 500 cells per page capped at $0.60/sheet in simple mode, 500 cells per page in advanced), so `Usage.pages × per_page_usd` is wrong for `.xlsx` input. * **Images count as one page each**, and multi-page TIFF frames count individually. * **HTML prettifying changes spacing.** The default HTML is indented per tag, which browsers render as stray spaces around inline tags (`( 23 )`); `disable_html_prettify=true` fixes it. It affects `Block.content` for blocks kept as HTML. * **`merge_cross_page` bills a variable compute surcharge** (~$0.50/document) that the price table cannot model, and it applies on every run. ## 7. Useful `provider_options` passthrough ```python # 1. Benchmarking: defeat the result cache so latency and cost are real. puffinparse.parse("doc.pdf", model="datalab/balanced", provider_options={"skip_cache": True}) # 2. Per-cell and per-list-item boxes (HTML output, +$6 per 1k pages each). puffinparse.parse("statement.pdf", model="datalab/accurate", include_raw=True, provider_options={"extras": "table_cell_bboxes,list_item_bboxes", "word_bboxes": True, "output_format": "json,html", "disable_html_prettify": True}) # 3. Keep running headers and footers, and preserve spreadsheet formatting. puffinparse.parse("report.pdf", model="datalab/balanced", provider_options={"additional_config": {"keep_pageheader_in_output": True, "keep_pagefooter_in_output": True, "keep_spreadsheet_formatting": True}}) # 4. Stitch tables and paragraphs that continue across page breaks (beta, compute-billed). puffinparse.parse("annual-report.pdf", model="datalab/accurate", provider_options={"merge_cross_page": True}) # 5. Token-efficient markdown for an LLM pipeline, no images or synthetic captions. puffinparse.parse("doc.pdf", model="datalab/fast", provider_options={"token_efficient_markdown": True, "disable_image_extraction": True, "disable_image_captions": True}) # 6. EU data residency (requires file_url or a pre-uploaded datalab:// reference, 1.25x usage). puffinparse.parse("https://example.com/doc.pdf", model="datalab/balanced", provider_options={"processing_location": "eu"}) # 7. Save a checkpoint so a later /extract or /segment call skips re-parsing. puffinparse.parse("doc.pdf", model="datalab/balanced", provider_options={"save_checkpoint": True}) ``` ## 8. Links * Docs home: <https://documentation.datalab.to> * Convert API guide: <https://documentation.datalab.to/docs/recipes/conversion/conversion-api-overview> * `POST /api/v1/convert` reference: <https://documentation.datalab.to/api-reference/convert-document> * Result polling reference: <https://documentation.datalab.to/api-reference/convert-result-check> * OpenAPI spec: <https://www.datalab.to/openapi.json> * Error codes: <https://documentation.datalab.to/platform/errors> * Limits and rate limiting: <https://documentation.datalab.to/docs/common/limits> * Billing (what counts as a page): <https://documentation.datalab.to/platform/billing> * Pricing rate card: <https://www.datalab.to/pricing> * Marker (the open-source engine, JSON/chunks/metadata shapes): <https://github.com/datalab-to/marker> * Machine-readable docs index: <https://documentation.datalab.to/llms.txt> --- # Unstructured <!-- source: docs/providers/unstructured.md | url: https://puffinparse.com/docs/providers/unstructured/ --> > **Status: docs-only.** Everything below comes from the official Unstructured documentation and > the live OpenAPI spec at `https://api.unstructuredapp.io/general/openapi.json` (read 2026-09-11), > and from `crates/puffinparse-core/src/providers/unstructured.rs`. No live call has been made — this > repository has no Unstructured key. `crates/puffinparse-core/tests/fixtures/unstructured_elements.json` > is hand-built from the documented element shapes. Mark this page **verified** only after the > `#[ignore]`d live test in `providers/unstructured.rs` passes with a real key. > Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `unstructured` | | Base URL | `https://api.unstructuredapp.io` (override: `base_url` on the request, or `UNSTRUCTURED_BASE_URL`). **Business accounts get their own URL** at sign-up and must use it | | API key | `UNSTRUCTURED_API_KEY` (or `api_key` on the request) — sent as the `unstructured-api-key` header | | Docs | <https://docs.unstructured.io/api-reference/partition/overview> | | API version | Path-versioned: `POST /general/v0/general`. The deployed spec reports its own build (`1.5.99` when read) | | Checked against | **not live-verified** (see banner) — documentation read 2026-09-11 | | Implementation | `crates/puffinparse-core/src/providers/unstructured.rs` | The Partition Endpoint is the only synchronous, single-file API Unstructured offers, and it is the one PuffinParse uses. Unstructured labels it **legacy** and steers production users to the Pipelines / Workflow API (`/api/v1/jobs`, connectors, chunking, embeddings), which is a different product shape that does not fit a one-document-in / one-document-out call. ## 2. Models exposed by PuffinParse | Model | Provider parameter | Modes | List price (`pricing.json`) | |---|---|---|---| | `unstructured/hi_res` *(default)* | `strategy=hi_res` | `parse`, `ocr` (derived) | $0.015 / page | | `unstructured/fast` | `strategy=fast` | `parse`, `ocr` (derived) | $0.015 / page | | `unstructured/auto` | `strategy=auto` | `parse`, `ocr` (derived) | $0.015 / page | Pricing is **flat per page for every strategy**: <https://unstructured.io/pricing> lists Pay-As-You-Go at **$0.015 per page** with the first 10 000 pages free ("Let's Go"), all features included; Business is custom. There is no cheaper rate for `fast` — the saving is latency, not money. What the strategies do (<https://docs.unstructured.io/concepts/partitioning>): * `hi_res` — layout model + OCR; the only strategy that reliably produces `coordinates`, `text_as_html` for tables and `detection_class_prob`. Default here for that reason. * `fast` — rule-based text extraction from the file's own text layer. No OCR, so it cannot read scans, and it is **rejected for image files**. * `auto` — routes each page to fast / hi_res / VLM at runtime. `ocr_only`, `od_only` and `vlm` are also accepted by the endpoint but are not exposed as models; reach them with `provider_options={"strategy": "vlm", "vlm_model_provider": …, "vlm_model": …}` (which overrides the model's strategy — see §3). ## 3. Request flow PuffinParse uses One synchronous call; there is no job to poll. `POST {base}/general/v0/general`, `multipart/form-data`, headers `unstructured-api-key` and `accept: application/json`: | Field | Value | |---|---| | `files` | the document bytes with filename and guessed MIME type (note the plural — it is `files`, not `file`) | | `strategy` | the model name (`hi_res` \| `fast` \| `auto`) | | `output_format` | `application/json` | | `coordinates` | `true` — **off by default**, and without it no element has a bounding box | | `include_page_breaks` | `true` — emits `PageBreak` elements, which is how page numbers are recovered for file types with no page metadata | | `languages` | the request's `language`, when set (repeatable field of Tesseract codes) | Response: a JSON **array** of element objects (no envelope). **URL inputs are downloaded first.** The Partition Endpoint has no remote-URL parameter, so a `DocumentInput::Url` is fetched by PuffinParse (respecting the call deadline) and uploaded as bytes. A non-2xx download, or an empty body, is an `input` error. **`pages` is applied client-side.** The endpoint has no page-selection parameter (`starting_page_number` only renumbers pre-split PDFs), so PuffinParse filters elements by `metadata.page_number` after the fact and sets `metadata.unstructured_pages_filtered_client_side = true`. **You are still billed for the whole document**, so `Usage.pages` reports every page the API processed, not the filtered subset. **Where `provider_options` are merged:** the request is a multipart form, so options become extra text fields. Strings pass through; booleans become `"true"`/`"false"`; numbers are stringified; arrays become **repeated fields** (which is what `languages`, `extract_image_block_types` and `skip_infer_table_types` expect); `null` is skipped; objects are serialised as JSON text. A key PuffinParse already set is replaced, so `provider_options` can override `strategy`, `coordinates` or `output_format`. ## 4. Response mapping | Unstructured field | PuffinParse unified field | Notes | |---|---|---| | `type` (+ `metadata.category_depth`) | `Block.type` | See the table below. | | `text` | `Block.text` | Always the plain text of the element; for tables, the cell text run together. | | `text` / `metadata.text_as_html` | `Block.content` | Markdown rendering: `Title` → `# `, deeper headings → `##`… by `category_depth`, `ListItem` → `- `, `Table` → `text_as_html` converted to a markdown table when it is simple (rectangular, no `colspan`/`rowspan`, not nested) and left as HTML otherwise. With `output="text"` the raw `text` is used. | | `metadata.page_number` | `Block.page_number` → `Page.page_number` | Missing page numbers fall back to a counter that increments on every `PageBreak`. | | `metadata.coordinates.points` | `Block.bbox` | Four polygon corners **in pixels**, top-left origin, listed counter-clockwise from the top-left. PuffinParse takes the enclosing axis-aligned box. | | `metadata.coordinates.layout_width` / `layout_height` | `Page.width` / `Page.height`, and the bbox divisor | Also read from `coordinates.system` when a payload nests them there. No dimensions ⇒ `bbox = None`. | | `metadata.detection_class_prob` | `Block.confidence` | Only produced by `hi_res`. | | — | `Page.markdown` / `Page.text` | Assembled from the page's blocks in order (`pages_from_blocks`); Unstructured returns no page-level rendering of its own. | | — | `Usage.pages` | Number of distinct pages seen in the response (before any `pages` filter), at least 1. | | `metadata.filetype` | `metadata.unstructured_filetype` | | | `element_id`, `metadata.parent_id`, `languages`, `filename`, `last_modified`, `emphasized_text_*`, `image_base64`, `links`, `orig_elements` | — | Not mapped; visible with `include_raw=True`. | | `PageBreak` elements | — | Consumed as page delimiters, never emitted as blocks. | **Element types** (<https://docs.unstructured.io/concepts/document-elements>): | Unstructured `type` | `BlockType` | |---|---| | `Title` with `category_depth` 0 / absent | `title` | | `Title` with `category_depth` ≥ 1, `SectionHeader`, `Headline`, `Subtitle` | `section_header` | | `NarrativeText`, `UncategorizedText`, `Text`, `CompositeElement`, `Address`, `EmailAddress`, `CodeSnippet`, `FormKeysValues`, `Field-Name`, `Value`, `Abstract`, `Threading` | `text` | | `ListItem`, `List-item`, `BulletedText` | `list` | | `Table`, `TableChunk` | `table` | | `Image`, `Picture`, `Figure` | `figure` | | `Header`, `PageHeader` | `header` | | `Footer`, `PageFooter` | `footer` | | `Footnote` | `footnote` | | `FigureCaption`, `Caption` | `caption` | | `Formula` | `formula` | | `PageNumber` and anything unlisted | `other` | | `PageBreak` | *(dropped)* | **`ocr` mode** is derived from `parse` (`TextResponse::from_parse`): each block's text becomes one or more lines carrying the block's box; `words` have no geometry. Unstructured's own output is element-level — there is no word-level API on this endpoint. ## 5. Errors, status codes, rate limits, timeouts | Status | Body | PuffinParse `ErrorKind` | |---|---|---| | 401 / 403 | `{"detail": "…"}` (missing or invalid `unstructured-api-key`) | `authentication` | | 402 | payment required / quota exhausted | `bad_request` — **not** an auth error | | 422 | `{"detail": [ValidationError…]}` (bad enum, unparsable field) | `bad_request` | | 429 | rate limited | `rate_limit` (retried) | | 4xx other | `{"detail": "…"}` | `bad_request` | | 5xx | `{"detail": "An error occurred"}` (`ServerError`) | `provider` (500/502/503/504 retried) | `Error::from_http` reads `detail` (stringifying the validation-array form), so the provider message is never swallowed. **Rate limits and quotas** are not published as numbers in the docs; the troubleshooting page only describes `HTTP 402 Payment Required` and `HTTP 429 Too Many Requests` and tells you to slow down or upgrade. Assume blind backoff — PuffinParse's default (`max_retries=2`, exponential with full jitter). **Timeouts.** The call is synchronous and can take minutes for a large `hi_res` PDF; the quickstart says so explicitly. `timeout_secs` (default 300) covers the URL download plus the single request. Raise it for long scans rather than lowering `max_retries` — a retry re-runs (and re-bills) the whole document. **Billing pages** are counted as: one page per page/slide/image for `.pdf`, `.pptx`, `.tiff`; the page metadata for `.docx`; **file size ÷ 100 KB for everything else**. PuffinParse's `Usage.pages` is a page *count from the response*, so for HTML, email, text and similar inputs it will not match the billed number. ## 6. Gotchas * **This endpoint is officially "legacy".** Unstructured recommends the Pipelines / Workflow API for production (`/api/v1/jobs`, "latest and highest-performing models"). The Partition Endpoint is single-file, synchronous, and the only thing that fits PuffinParse's one-call contract today. * **`coordinates` is off by default.** Without `coordinates=true` every element comes back without geometry, silently. PuffinParse always sends it. * **The file field is `files`, not `file`.** A `file` part is ignored and the request fails validation. * **`fast` cannot read images.** Sending a PNG/JPG with `strategy=fast` is an error documented as its own support page; use `hi_res`, `auto` or `vlm` for images. * **Tables come back as HTML, not markdown.** `text` is the cell text run together (lossy, no row structure) and `metadata.text_as_html` is the structured form. Unstructured's own examples emit `<thead><th>…</th></thead>` **without a `<tr>`**, so a naive row splitter produces nothing; PuffinParse's converter treats a bare `<thead>` run as one row, and falls back to keeping the HTML whenever the table is ragged, spanning or nested. * **`Title` is used for both the document title and every heading.** `category_depth` is the only way to tell them apart, and it is not always present — expect the occasional section header mapped to `title`. * **No page selection.** `pages` is applied client-side and you are billed for the whole document; `starting_page_number` only renumbers a PDF you split yourself. * **No remote URL input**, so PuffinParse downloads and re-uploads, which doubles the bytes on the wire for URL inputs. * **Coordinates are pixels of the rendered page**, not PDF points, and the origin is top-left with `y` increasing downwards (the same convention as `BBox`), but the `points` are listed **counter-clockwise** from the top-left — only the enclosing box is stable, which is what PuffinParse stores. * **`detection_class_prob` only exists under `hi_res`**, so `Block.confidence` is `None` for `fast` and often for `auto`. * **Chunking changes the element vocabulary.** Passing `chunking_strategy` through `provider_options` replaces elements with `CompositeElement` / `TableChunk` chunks (mapped to `text` / `table`), and page numbers become chunk-level. Leave it off unless you want chunks. * **`ocr_languages` is deprecated** in favour of `languages`; both exist in the spec. * **Business accounts have their own base URL**, handed out at account creation. The documented `https://api.unstructuredapp.io` default is the serverless SaaS host; set `UNSTRUCTURED_BASE_URL` if yours differs. ## 7. Useful `provider_options` passthrough ```python # 1. VLM strategy (overrides the model's strategy), e.g. for handwriting-heavy scans. puffinparse.parse("scan.pdf", model="unstructured/hi_res", provider_options={"strategy": "vlm", "vlm_model_provider": "openai", "vlm_model": "gpt-4o"}) # 2. Pick the hi_res layout model and keep table inference on for every file type. puffinparse.parse("report.pdf", model="unstructured/hi_res", provider_options={"hi_res_model_name": "yolox", "pdf_infer_table_structure": True}) # 3. Multi-language OCR (array values become repeated form fields). puffinparse.parse("contract.pdf", model="unstructured/hi_res", provider_options={"languages": ["eng", "deu"]}) # 4. Base64 crops of images and tables in the raw payload. puffinparse.parse("paper.pdf", model="unstructured/hi_res", include_raw=True, provider_options={"extract_image_block_types": ["Image", "Table"]}) # 5. Chunk on the server for a RAG pipeline (changes the element vocabulary — see gotchas). puffinparse.parse("handbook.pdf", model="unstructured/auto", provider_options={"chunking_strategy": "by_title", "max_characters": 2000, "combine_under_n_chars": 500, "include_orig_elements": False}) # 6. Skip table inference for speed, and give every element a UUID. puffinparse.parse("minutes.docx", model="unstructured/fast", provider_options={"skip_infer_table_types": ["docx"], "unique_element_ids": True}) ``` ## 8. Links * Partition Endpoint overview: <https://docs.unstructured.io/api-reference/partition/overview> * Live OpenAPI spec: <https://api.unstructuredapp.io/general/openapi.json> · Swagger UI: <https://api.unstructuredapp.io/general/docs> * Document elements and metadata: <https://docs.unstructured.io/concepts/document-elements> * Partitioning strategies: <https://docs.unstructured.io/concepts/partitioning> * Quota / billing / rate limiting: <https://docs.unstructured.io/support/issues/quota-billing-rate-limiting> * Pricing: <https://unstructured.io/pricing> * Element type definitions (source of truth): <https://github.com/Unstructured-IO/unstructured/blob/main/unstructured/documents/elements.py> * Machine-readable docs index: <https://docs.unstructured.io/llms.txt> --- # Upstage Document Parse <!-- source: docs/providers/upstage.md | url: https://puffinparse.com/docs/providers/upstage/ --> > **Status: docs-only.** Implemented from Upstage's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `upstage` | | Base URL | `https://api.upstage.ai` (override: `base_url` on the request, or `UPSTAGE_BASE_URL`) | | API key | `UPSTAGE_API_KEY` (or `api_key` on the request) — sent as `Authorization: Bearer <key>` | | Docs | <https://console.upstage.ai/docs/capabilities/document-digitization/document-parsing> · agent-oriented dump: <https://console.upstage.ai/api/docs/for-agents/raw> | | Modes | `parse` (native), `ocr` (derived from `parse`) | | Checked against | 2026-09-11, **from documentation only** — no key was available, so the `#[ignore]`d live test has not been run | | Implementation | `crates/puffinparse-core/src/providers/upstage.rs` | Document Parse turns a document into layout **elements** (`paragraph`, `heading1`, `table`, `figure`, `chart`, …) with HTML, Markdown and plain text per element, plus four relative corner coordinates. PuffinParse uses the synchronous endpoint by default and the async batch endpoint on request. ## 2. Models exposed by PuffinParse | Model | Provider parameters PuffinParse sets | List price (`pricing.json`) | |---|---|---| | `upstage/document-parse` *(default)* | `model=document-parse` | $0.01 / page | | `upstage/document-parse-nightly` | `model=document-parse-nightly` | $0.01 / page (billed as standard) | Prices are the public per-page rates from <https://www.upstage.ai/pricing>: Document Parse **standard $0.01/page**, **enhanced $0.03/page**. PuffinParse does not set `mode`, so the account default (`standard`) applies; `provider_options={"mode": "enhanced"}` or `"auto"` changes the engine **and the price**, which `cost_usd` will then under-report. The response's `usage.enhanced` list is surfaced as `metadata.upstage_enhanced_pages` so an enhanced-mode run is visible after the fact. The separate Document **OCR** product (`model=ocr`, $0.0015/page, word boxes) is *not* wired up: PuffinParse's `ocr` mode for this provider is derived from the parse blocks. See §6. ## 3. Request flow PuffinParse uses Every request carries `Authorization: Bearer $UPSTAGE_API_KEY` and `accept: application/json`. 1. **Load the document.** Path and bytes inputs are uploaded as-is. Document Parse has **no URL input**, so a URL input is downloaded by PuffinParse first and then uploaded. 2. **Submit.** `POST {base}/v1/document-digitization`, `multipart/form-data`: | field | value PuffinParse sends | |---|---| | `document` | the file part (filename + guessed MIME) | | `model` | `document-parse` or `document-parse-nightly` | | `output_formats` | `["html","markdown","text"]` — all three, so blocks get Markdown *and* a provider plain text | | `ocr` | `auto` (digital-born PDFs keep their embedded text; images are OCR'd) | | `coordinates` | `true` | Anything in `provider_options` is appended as extra form fields and **overrides** a default with the same name. Non-string values are serialised: booleans/numbers as text, arrays/objects as JSON (`base64_encoding=["table"]`). The key `async` is consumed locally and never sent. 3. **Async (opt-in).** With `provider_options={"async": true}` the same multipart body goes to `POST {base}/v1/document-digitization/async` → `{"request_id": …}`. PuffinParse then polls `GET {base}/v1/document-digitization/requests/{request_id}` (2 s, backing off ×1.5 to 15 s) until `status` is `completed` or `failed`, and finally downloads every `batches[].download_url` (plain `GET`, **no auth header**, 15-minute expiry) in `batches[].id` order. Each batch payload has exactly the shape of the sync response, so they are concatenated into one result. 4. **Normalise.** See §4. `pages` is applied **client-side**: Document Parse has no page-range parameter, so PuffinParse parses the whole document and then drops the pages outside the selection, recording `metadata.upstage_pages_filtered_client_side`. `usage.pages` still reports the whole billed document. `language` is ignored (Document Parse auto-detects; there is no language parameter). ## 4. Response mapping ```json { "apiVersion": "1.1", "model": "document-parse-260630", "elements": [ { "id": 0, "category": "heading1", "page": 1, "content": { "html": "<h1 id='0'>Hello PuffinParse</h1>", "markdown": "# Hello PuffinParse", "text": "Hello PuffinParse" }, "coordinates": [ {"x":0.125,"y":0.0525}, {"x":0.425,"y":0.0525}, {"x":0.425,"y":0.1052}, {"x":0.125,"y":0.1052} ] } ], "content": { "html": "…", "markdown": "…", "text": "…" }, "usage": { "pages": 2, "standard": [1, 2] } } ``` (The full fixture is `crates/puffinparse-core/tests/fixtures/upstage_document_parse.json`.) | Upstage field | PuffinParse unified field | Notes | |---|---|---| | `elements[]` | `Page.blocks[]` | Grouped into pages by `element.page`, reading order preserved. | | `elements[].category` | `Block.type` | Mapping below. | | `elements[].content.markdown` | `Block.content` | `output="text"` uses `content.text` instead (falling back to `markdown_to_text`). Empty Markdown falls back to `text`, then `html`. | | `elements[].content.text` | `Block.text` | | | `elements[].coordinates[]` | `Block.bbox` | Four corner points, **already relative 0–1, top-left origin**; PuffinParse takes min/max x/y and clamps. | | `elements[].page` | `Block.page_number` | 1-based on the wire and in PuffinParse. | | joined element Markdown | `Page.markdown` / `Page.text` | Per page, blank-line separated (`pages_from_blocks`). | | `usage.pages` | `Usage.pages` | Falls back to the number of reconstructed pages. | | `model` | `metadata.upstage_model_version` | Resolved snapshot, e.g. `document-parse-260630`. | | `usage.enhanced` | `metadata.upstage_enhanced_pages` | Only when non-empty (`mode=auto`/`enhanced`). | | async `request_id` | `provider_job_id` | `None` for sync calls — the sync response carries no id. | | `content.{html,markdown,text}` | — | Only used as a fallback when `elements` is empty; otherwise page content is rebuilt from elements. | | `elements[].sub_category`, `base64_encoding`, `apiVersion` | — | Visible with `include_raw=True`. | | — | `Page.width` / `Page.height` | **Never set**: Document Parse reports no page dimensions (coordinates are already relative). | | — | `Block.confidence` | **Never set**: Document Parse reports no per-element confidence. | Block types: `paragraph` → `text`; `heading1` → `title`; `table` → `table`; `figure`, `chart` → `figure`; `caption` → `caption`; `list` → `list`; `header` → `header`; `footer` → `footer`; `footnote` → `footnote`; `equation` → `formula`; `code` → `text` (it is still text; the fenced block is kept in `content`); `index` and anything new → `other`. ## 5. Errors, status codes, rate limits, timeouts Error body: `{"error": {"message": …, "type": …, "code": …}}` — `Error::from_http` picks up `error.message`. | Status | Cause | PuffinParse `ErrorKind` | |---|---|---| | 400 | malformed request, unknown model, no document | `bad_request` | | 401 | invalid API key | `authentication` | | 403 | **insufficient credit** (also an expired async `download_url`) | `authentication` | | 404 | wrong path | `bad_request` | | 405 | `http://` instead of `https://` | `bad_request` | | 413 | file over 50 MB (async) | `bad_request` | | 415 | unsupported file format | `bad_request` | | 422 | corrupted / damaged document | `bad_request` | | 429 | rate limit | `rate_limit` (retried) | | 500/502/503/504 | server error | `provider` (retried) | Async failures come back as HTTP 200 with `status: "failed"` (plus `failure_message`), or with a per-batch `"status": "failed"`; PuffinParse raises `provider` errors carrying the `request_id` as `job_id`. **Limits.** 50 MB per file, 200 megapixels per page, PDF/JPEG/PNG/BMP/TIFF/HEIC/DOCX/PPTX/XLSX/HWP/HWPX. Sync: **100 pages** (pages beyond 100 are silently dropped). Async: **1 000 pages**, processed in 10-page batches; results are stored 30 days, each `download_url` expires after ~15 minutes. **Rate limits (tier 0).** Document Parse sync 1 RPS / 300 pages-per-minute; async 2 RPS / 1 200 PPM. Limits rise with the commitment tier. No `X-RateLimit-*` headers, so backoff is blind. There is no batch endpoint for multiple documents — send them one at a time. **Timeouts.** `timeout_secs` (default 300) covers download + upload + polling + batch fetches and caps each individual request. The async queue can hold a job for **up to 72 hours** at peak, so async runs need a `timeout` far beyond the default (or a webhook-style poll of your own). ## 6. Gotchas (documentation-derived; not yet live-verified) * **No page selection.** There is no `pages`/`page_range` parameter, so PuffinParse filters pages after the fact and you are billed for the whole document. * **No page dimensions and no confidences.** Coordinates are relative, which is what PuffinParse wants, but `Page.width`/`height` and `Block.confidence` stay `None` for this provider. * **Async page numbering is assumed global.** Batches cover 10-page ranges (`start_page`/`end_page`). PuffinParse trusts `element.page` as a document-level page number, but if a batch numbers its own pages from 1 while starting later in the document, the offset (`start_page - 1`) is added. Worth re-checking against a real 20+ page async run. * **`download_url` expires in ~15 minutes** and returns 403 afterwards; PuffinParse fetches it right after polling, so this only bites very slow clients. Re-fetching the request status mints a fresh URL and costs nothing. * **`ocr=auto` vs `force`.** Digital-born PDFs keep their embedded text layer with `auto`; a scanned PDF that still contains a bad text layer needs `provider_options={"ocr": "force"}`. * **`mode` changes the price** (standard $0.01 → enhanced $0.03) without changing the response shape. * **Charts degrade to figures.** With `chart_recognition` on (the default), a recognised chart comes back as `category: "chart"` with a Markdown table; when recognition fails it silently becomes a `figure` with OCR text only. Both map to `figure` in PuffinParse. * **Equations are LaTeX in `markdown`/`html` but raw (often wrong) OCR in `text`** — prefer `output="markdown"` for documents with formulas. * **`content.markdown` is not the same as joining the elements**: PuffinParse rebuilds page Markdown from elements so that pages, blocks and boxes stay consistent. Use `include_raw=True` if you want the provider's own whole-document HTML. * **The separate Document OCR product is cheaper** ($0.0015 vs $0.01 per page) and returns word-level `boundingBox.vertices` in pixels. Adding it as a native `ocr` model (e.g. `upstage/ocr`) is an obvious follow-up; today `mode="ocr"` on `upstage/document-parse` derives lines from parse blocks. * **File names matter**: ≤ 900 chars (≤ 300 for Korean), no path components, extension must match the real format, or documents can hang in `started`. ## 7. Useful `provider_options` passthrough ```python # 1. Scanned PDFs: force OCR instead of trusting an embedded text layer. puffinparse.parse("scan.pdf", model="upstage/document-parse", provider_options={"ocr": "force"}) # 2. Complex tables / charts / low-quality scans (enhanced mode is $0.03/page). puffinparse.parse("report.pdf", model="upstage/document-parse", provider_options={"mode": "enhanced"}) # 3. Long documents (up to 1 000 pages) through the async batch API. puffinparse.parse("book.pdf", model="upstage/document-parse", timeout=3600, provider_options={"async": True}) # 4. Tables merged across page breaks, plus cropped table images in the raw payload. puffinparse.parse("financials.pdf", model="upstage/document-parse", include_raw=True, provider_options={"merge_multipage_tables": True, "base64_encoding": ["table"]}) # 5. Word-level OCR boxes alongside the layout elements (raw only). puffinparse.parse("form.png", model="upstage/document-parse", include_raw=True, provider_options={"words": True}) ``` ## 8. Links * Document Parse: <https://console.upstage.ai/docs/capabilities/document-digitization/document-parsing> * Full API reference for agents (single markdown file): <https://console.upstage.ai/api/docs/for-agents/raw> * Pricing: <https://www.upstage.ai/pricing> · rate limits: <https://console.upstage.ai/docs/guides/rate-limits> * Document OCR (word boxes, $0.0015/page): `model=ocr` on the same endpoint. * Information Extraction (`POST /v1/information-extraction`) is a separate product and is **not** used by PuffinParse's `extract` mode for this provider. --- # Landing AI ADE <!-- source: docs/providers/landingai.md | url: https://puffinparse.com/docs/providers/landingai/ --> > **Status: docs-only.** Implemented from Landing AI's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `landingai` | | Base URL | `https://api.va.landing.ai` (override: `base_url`, or `LANDINGAI_BASE_URL`). EU: `https://api.va.eu-west-1.landing.ai` | | API key | `LANDINGAI_API_KEY`, falling back to `VISION_AGENT_API_KEY` (the name Landing AI's own libraries use), or `api_key` on the request — sent as `Authorization: Bearer <key>` | | Docs | <https://docs.landing.ai/ade/ade-overview> · <https://docs.landing.ai/ade/parse> · OpenAPI: <https://docs.landing.ai/ade/va_openapi_ade2.json> | | Modes | `parse` (native), `ocr` (derived from `parse`), `extract` (parse → extract, two calls) | | Checked against | 2026-09-11, **from documentation and the published OpenAPI spec only** — no key was available, so the `#[ignore]`d live test has not been run | | Implementation | `crates/puffinparse-core/src/providers/landingai.rs` | ADE parses a document into reading-order Markdown plus semantic **chunks**, each grounded to a page and a normalised bounding box, with separate grounding entries for tables and individual table cells. Field extraction is a *second* API that consumes the parse Markdown, so PuffinParse's `extract` mode is a parse followed by an extract, with chunk references resolved back into page + box citations. ## 2. Models exposed by PuffinParse | Model | Provider parameters PuffinParse sets | List price (`pricing.json`) | |---|---|---| | `landingai/dpt-2` *(default)* | `model=dpt-2-latest`, `split=page` | parse $0.03 / page · extract $0.04 / page | Credits cost **$0.01** each on the Explore and Team plans. Parsing costs **3 credits/page** (+1 credit/page with Zero Data Retention); spreadsheets are 1 credit/sheet plus 3 credits per embedded image. Extraction is charged on characters — `(input chars ÷ 5 000) + (output chars ÷ 1 000)`, rounded up to a tenth of a credit — which for a typical page of Markdown is around 1 credit, hence the $0.04/page estimate for the two-call `extract` mode (3 credits parse + ~1 credit extract). Treat the extract price as an order-of-magnitude estimate: it scales with document *length*, not page count. Pin a snapshot with `provider_options={"model": "dpt-2-20260410"}`; available values are `dpt-2`, `dpt-2-latest` and the dated snapshots (`dpt-2-20250919`, `-20251103`, `-20260302`, `-20260410`). `dpt-1` and `dpt-2-mini` are **deprecated and rejected by the API**. ## 3. Request flow PuffinParse uses 1. **Parse.** `POST {base}/v1/ade/parse`, `multipart/form-data`: | field | value PuffinParse sends | |---|---| | `document` | the file part — **or** `document_url` with the URL when the input is a URL (ADE downloads it itself) | | `model` | `dpt-2-latest` (or the `provider_options.model` override) | | `split` | `page` — one split per page, which gives authoritative per-page Markdown | `provider_options` are appended as extra form fields and override defaults of the same name (`password`, `custom_prompts`, `model`, `split`, …). Two keys are consumed locally and never sent: `auth_scheme` (see §6) and `extract_model`. 2. **Extract (mode `extract` only).** `POST {base}/v1/ade/extract`, `multipart/form-data` with `markdown` = the **raw** parse Markdown (anchors included — the extractor's references point at them), `schema` = the request's JSON Schema serialised to a string, and `model` = `extract-latest` (override with `provider_options.extract_model`). `ExtractRequest.instructions`, if set, is written into the schema's top-level `description` (ADE has no separate instruction field). 3. **Normalise.** See §4. `pages` is applied **client-side** — ADE Parse has no page-range parameter — and records `metadata.landingai_pages_filtered_client_side`; `usage.pages` stays at the billed page count for the whole document. `language` is not forwarded: ADE detects languages automatically and exposes no hint. PuffinParse always uses the **synchronous** endpoint. Parse Jobs (`POST /v1/ade/parse/jobs` + `GET /v1/ade/parse/jobs/{job_id}`) lift the limit from 100 pages to 6 000 pages / 1 GB and are the obvious next addition; the sync endpoint's own gateway timeout is 475 s. ## 4. Response mapping ```json { "markdown": "<a id='2831e56d-…'></a>\n\n# Hello PuffinParse\n\n…", "chunks": [ { "markdown": "<a id='2831e56d-…'></a>\n\n# Hello PuffinParse", "type": "text", "id": "2831e56d-94f5-4ec4-b001-6e16e188119b", "grounding": { "box": { "left": 0.017, "top": 0.038, "right": 0.463, "bottom": 0.212 }, "page": 0 } } ], "splits": [ { "class": "page", "identifier": "page_0", "pages": [0], "markdown": "…", "chunks": ["2831e56d-…"] } ], "grounding": { "2831e56d-…": { "box": {…}, "page": 0, "type": "chunkText", "confidence": 0.97 }, "1-1": { "box": {…}, "page": 1, "type": "table" }, "1-4": { "box": {…}, "page": 1, "type": "tableCell", "position": { "row": 1, "col": 0, … } } }, "metadata": { "filename": "…", "org_id": null, "page_count": 2, "duration_ms": 7861, "credit_usage": 6.0, "job_id": "job_…", "version": "dpt-2-20260410", "failed_pages": [] } } ``` (The full fixtures are `crates/puffinparse-core/tests/fixtures/landingai_parse.json` and `landingai_extract.json`.) ### parse / ocr | ADE field | PuffinParse unified field | Notes | |---|---|---| | `chunks[]` | `Page.blocks[]` | Reading order preserved. | | `chunks[].markdown` | `Block.content` | The `<a id='…'></a>` grounding anchors are stripped; table HTML (`<table id="0-1">…`) is kept verbatim. `output="text"` runs it through `markdown_to_text`. | | `chunks[].type` | `Block.type` | Mapping below. | | `chunks[].grounding.page` | `Block.page_number` | **Zero-indexed on the wire**, +1 in PuffinParse. | | `chunks[].grounding.box` | `Block.bbox` | `{left, top, right, bottom}`, already normalised 0–1 with a top-left origin → `{x0, y0, x1, y1}`. | | `grounding[<chunk id>].confidence` | `Block.confidence` | Only text-ish chunks carry one. | | `splits[]` with `class == "page"` | `Page.markdown` / `Page.text` | Preferred over joining the chunks; `pages[0] + 1` is the page number. | | `metadata.page_count` | `Usage.pages` | Falls back to the number of reconstructed pages. | | `metadata.credit_usage` | `Usage.credits` | Only when > 0. | | `metadata.job_id` | `provider_job_id` | | | `metadata.version` | `metadata.landingai_version` | Resolved snapshot, e.g. `dpt-2-20260410`. | | `metadata.failed_pages` | `metadata.landingai_failed_pages` | Converted to 1-based. Present on `206 Partial Content`. | | `grounding[<table/cell id>]` | — | Table and cell boxes are not modelled by PuffinParse's `Block`; use `include_raw=True`. | | — | `Page.width` / `Page.height` | **Never set**: ADE reports no page dimensions (boxes are already relative). | Chunk types: `text` → `text`, unless the chunk's Markdown starts with a heading — `# ` → `title`, `##`+ → `section_header`; `table` → `table`; `figure`, `logo`, `card`, `attestation` → `figure`; `marginalia` (headers, footers, page numbers) and `scan_code` (barcode/QR) → `other`. Legacy names (`title`, `caption`, `list`, `header`, `footer`, `footnote`, `equation`) are still mapped. ### extract | ADE field | PuffinParse unified field | Notes | |---|---|---| | `extraction` | `ExtractResponse.data` | Exactly the object ADE returns, shaped by your schema. | | `extraction_metadata.…{value, references}` | `ExtractResponse.fields["<json pointer>"]` | The metadata mirrors the schema; every leaf becomes one `FieldInfo` keyed by an RFC 6901 pointer into `data` (`/invoice/total`, `/items/0/description`). | | `references[]` (chunk / table-cell ids) | `FieldInfo.citations[]` | Resolved through the **parse** response's `grounding` map → `{page_number (1-based), bbox}`; for chunk ids the chunk's Markdown is attached as `Citation.text`. Unresolvable ids are dropped. | | parse `metadata.page_count` | `Usage.pages` | Extraction itself is not page-billed. | | parse + extract `credit_usage` | `Usage.credits` | Summed across both calls. | | `metadata.job_id` (extract) | `provider_job_id` | | | `metadata.schema_violation_error` | `metadata.landingai_schema_violation_error` | Non-null means the output does not fully conform to the schema (HTTP 206). | | `metadata.version` | `metadata.landingai_extract_version` | | With `include_raw=True` the extract response's `raw` is `{"parse": <parse payload>, "extract": <extract payload>}`. ## 5. Errors, status codes, rate limits, timeouts ADE is a FastAPI service: errors are `{"detail": "…"}` or `{"detail": [ValidationError, …]}`, which `Error::from_http` surfaces directly. | Status | Cause | PuffinParse `ErrorKind` | |---|---|---| | 200 | success | — | | 206 | **partial content** — some pages failed; `metadata.failed_pages` lists them (zero-indexed) | success, with `metadata.landingai_failed_pages` | | 400 | document download failed, unsupported/deprecated model version | `bad_request` | | 401 | missing or invalid API key | `authentication` | | 402 | **out of credits** | `bad_request` (not a fallback-eligible error for the router) | | 422 | input validation failed (bad schema, missing `document`/`document_url`) | `bad_request` | | 429 | pages-per-hour rate limit exceeded | `rate_limit` (retried) | | 500 | all pages failed to process | `provider` (retried) | | 504 | processing exceeded the 475 s gateway timeout | `provider` (retried) | **Limits.** ADE Parse (sync): **100 pages** per PDF. Parse Jobs: 6 000 pages or 1 GB. Accepted: PDF, JPEG/JPG/PNG/APNG and other images, Office documents and spreadsheets (XLSX/CSV are billed per sheet). Rate limits are **pages per hour per organization**, by plan, spread evenly across the hour; Extract Jobs have their own hourly budget where each job counts as one page-equivalent. **Timeouts.** `timeout_secs` (default 300) is the whole-call deadline — for `extract` that covers *both* HTTP calls — and also caps each individual request. Note the server-side 475 s ceiling on sync parses: a large document needs `timeout` ≥ 500 or the Parse Jobs API. ## 6. Gotchas (documentation-derived; not yet live-verified) * **Two generations of the API exist.** PuffinParse targets **ADE Gen1** (`api.va.landing.ai/v1/ade/*`, DPT-2), which is current and documented. A newer **Gen2** (`api.ade.landing.ai/v2/parse`, DPT-3, character-span grounding, a structure tree instead of chunks) is live with a different response shape; supporting it means a second model (`landingai/dpt-3`) and a second normaliser, not a base-URL switch. The *legacy* `v1/tools/agentic-document-analysis` endpoint and the `agentic-doc` library are deprecated and return errors. * **Auth header ambiguity.** Every code sample in the docs uses `Authorization: Bearer <key>` (what PuffinParse sends), but the OpenAPI security scheme is named "Basic Auth" with `bearerFormat: Basic`, and the troubleshooting page mentions an `apikey` header. If a key is rejected with 401, try `provider_options={"auth_scheme": "basic"}`, which sends `Authorization: Basic <key>` verbatim (any other string is used as the scheme as-is). * **Pages are zero-indexed everywhere** on the wire — `grounding.page`, `splits[].pages`, `failed_pages`, and the `page_0` identifiers. PuffinParse converts all of them to 1-based. * **Chunk grounding is a single object, not a list.** The Gen1 schema is `grounding: {box, page}`; the legacy endpoint used a list of groundings with `{l, t, r, b}` keys. PuffinParse's deserialiser accepts both shapes and both key spellings. * **Markdown carries anchors.** Every chunk's Markdown begins with `<a id='<chunk id>'></a>`, and table cells carry `id` attributes — that is how extraction references locations. PuffinParse strips the anchors from `content`/`markdown` but sends the *unstripped* Markdown to the extractor, which is what makes citations resolvable. * **Tables come back as HTML**, not Markdown pipes, and cell ids are `"{page}-{base62}"` (page zero-indexed). `markdown_to_text` strips the tags for `Block.text` and `output="text"`. * **There are no headings in the chunk vocabulary.** Titles and section headers are `text` chunks whose Markdown happens to start with `#`; PuffinParse maps those to `title` / `section_header`. * **`marginalia` is lossy.** It merges what other providers split into `header`, `footer`, `page_number` and `footnote`, so it maps to `other`. * **Extraction is length-priced, not page-priced**, and a very long document can cost far more to extract than to parse. `cost_usd` for `extract` is a per-page estimate and will drift. * **`split=page` is PuffinParse's default**, because without it ADE returns a single `full` split and per-page Markdown would have to be stitched from chunk groundings. `page` is the only documented value, and `provider_options` nulls are skipped, so the default cannot be unset — that is deliberate. (The `split` *parameter* is unrelated to the ADE **Split API**, which classifies sub-documents.) * **Password-protected files** are supported through `provider_options={"password": "…"}`. ## 7. Useful `provider_options` passthrough ```python # 1. Pin a model snapshot so results do not move under you. puffinparse.parse("doc.pdf", model="landingai/dpt-2", provider_options={"model": "dpt-2-20260410"}) # 2. Password-protected PDF. puffinparse.parse("locked.pdf", model="landingai/dpt-2", provider_options={"password": "s3cret"}) # 3. Tell the figure captioner what you care about. puffinparse.parse("chart.png", model="landingai/dpt-2", provider_options={"custom_prompts": '{"figure": "Describe axis labels in detail."}'}) # 4. Schema-driven extraction with citations (parse → extract under the hood). puffinparse.extract("invoice.pdf", model="landingai/dpt-2", citations=True, schema={"type": "object", "properties": { "invoice": {"type": "object", "properties": { "number": {"type": "string"}, "total": {"type": "number"}}}}}) # 5. If a key is rejected with 401, switch the authorization scheme. puffinparse.parse("doc.pdf", model="landingai/dpt-2", provider_options={"auth_scheme": "basic"}) # 6. EU data residency (key must come from the EU console). puffinparse.parse("doc.pdf", model="landingai/dpt-2", base_url="https://api.va.eu-west-1.landing.ai") ``` ## 8. Links * Overview: <https://docs.landing.ai/ade/ade-overview> · docs index: <https://docs.landing.ai/llms.txt> * Parse: <https://docs.landing.ai/ade/parse> · JSON response: <https://docs.landing.ai/ade/ade-json-response> · chunk types: <https://docs.landing.ai/ade/ade-chunk-types> * Extract: <https://docs.landing.ai/ade/ade-extract> · response: <https://docs.landing.ai/ade/ade-extract-response> * OpenAPI (Gen1): <https://docs.landing.ai/ade/va_openapi_ade2.json> · (Gen2 / DPT-3): <https://docs.landing.ai/dpt3/openapi-adev2.json> * Credits: <https://docs.landing.ai/ade/ade-credit-consumption> · plans: <https://docs.landing.ai/ade/ade-pricing> · rate limits: <https://docs.landing.ai/ade/ade-rate-limits> * Parse Jobs (async, 6 000 pages): <https://docs.landing.ai/ade/ade-parse-async> --- # Google Document AI <!-- source: docs/providers/google-documentai.md | url: https://puffinparse.com/docs/providers/google-documentai/ --> > **Status: docs-only.** Implemented from Google Cloud's published API documentation and tested > against fixture payloads built from it. It has not yet been run against the live API, so > expect wire-format differences. Help verify it: [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). ## 1. Summary | | | |---|---| | Provider name | `google_documentai` | | Base URL | `https://{location}-documentai.googleapis.com` — built from the configured location (override: `base_url`, or `GOOGLE_DOCUMENTAI_BASE_URL`) | | Credential | an **OAuth 2.0 access token** in `GOOGLE_DOCUMENTAI_ACCESS_TOKEN`, `provider_options.access_token`, or `api_key` — sent as `Authorization: Bearer <token>` | | Required config | `GOOGLE_DOCUMENTAI_PROJECT`, `GOOGLE_DOCUMENTAI_PROCESSOR_ID`, optional `GOOGLE_DOCUMENTAI_LOCATION` (`us` default, `eu`, …) — each overridable via `provider_options` | | Docs | <https://cloud.google.com/document-ai/docs/reference/rest/v1/projects.locations.processors/process> | | Modes | `ocr` (**native**), `parse`, `extract` (entities) | | Checked against | 2026-09-11, **from documentation only** — no credentials were available, so the `#[ignore]`d live test has not been run | | Implementation | `crates/puffinparse-core/src/providers/google_documentai.rs` | Document AI is a fleet of *processors* you create in your own Google Cloud project. PuffinParse makes one synchronous `:process` call against the processor named in your configuration; the PuffinParse model name only selects **how the response is read**. > **Service-account key exchange is out of scope.** PuffinParse does not sign JWTs or talk to > `oauth2.googleapis.com`. Mint a token yourself — `export GOOGLE_DOCUMENTAI_ACCESS_TOKEN=$(gcloud auth print-access-token)`, > a metadata-server token on GCE/Cloud Run, or your own service-account exchange — and remember that > tokens expire after ~1 hour. The required IAM permission is > `documentai.processors.processOnline` (scope `https://www.googleapis.com/auth/cloud-platform`). ## 2. Models exposed by PuffinParse | Model | Expected processor type | Modes | List price (`pricing.json`) | |---|---|---|---| | `google_documentai/ocr` *(default)* | Enterprise Document OCR (`OCR_PROCESSOR`) | `parse`, `ocr` | $0.0015 / page | | `google_documentai/layout` | Layout Parser (`LAYOUT_PARSER_PROCESSOR`) | `parse`, `ocr` | $0.01 / page | | `google_documentai/form` | Form Parser (`FORM_PARSER_PROCESSOR`) | `parse`, `ocr`, `extract` | $0.03 / page | | `google_documentai/prebuilt` | Invoice / W2 / Expense / Custom Extractor | `parse`, `ocr`, `extract` | $0.03 / page | Prices from <https://cloud.google.com/document-ai/pricing>: Enterprise Document OCR **$1.50 / 1 000 pages** (first 1 000 pages/month free, $0.60 above 5 M/month); Layout Parser **$10 / 1 000 pages**; Form Parser and Custom Extractor **$30 / 1 000 pages** ($20 above 1 M/month). Volume tiers and the free allowance are not modelled — `cost_usd` uses the first paid tier. The model **must match the processor you configured**: sending a Layout Parser id while asking for `google_documentai/ocr` yields a document with no `pages[]`, and the parse falls back to `documentLayout` if it is present. A bare `model="google_documentai"` resolves to `ocr` for `parse` and `ocr`, and to `form` for `extract` (the first extract-capable entry in the registry). ## 3. Request flow PuffinParse uses One call, no polling: ``` POST https://{location}-documentai.googleapis.com/v1/projects/{project}/locations/{location}/processors/{processorId}:process Authorization: Bearer <access token> Content-Type: application/json ``` With `provider_options.processor_version` the path becomes `…/processors/{processorId}/processorVersions/{version}:process` (e.g. `pretrained-ocr-v2.0-2023-06-02`). Body built by `build_body()`: ```json { "rawDocument": { "content": "<base64 of the file>", "mimeType": "application/pdf" }, "skipHumanReview": true, "processOptions": { "individualPageSelector": { "pages": [1, 2, 5] }, "ocrConfig": { "hints": { "languageHints": ["de"] } } } } ``` * **Input.** Path and bytes inputs are base64-encoded into `rawDocument`. Document AI accepts no http(s) URL, so a URL input is downloaded by PuffinParse and sent inline. (`gcsDocument` is reachable through `provider_options` if your file is already in Cloud Storage — pass `{"gcsDocument": {"gcsUri": "gs://…", "mimeType": "application/pdf"}}`; the `rawDocument` key stays in the body, so remove it via a full `provider_options` body override if Google rejects both.) * **`pages`** → `processOptions.individualPageSelector.pages`, 1-based, expanded, sorted and de-duplicated. **Open-ended ranges (`"3-"`) are rejected** with an `input` error before any network call: the selector needs explicit page numbers (`fromStart`/`fromEnd` are reachable through `provider_options`). * **`language`** → `processOptions.ocrConfig.hints.languageHints`, only for the `ocr` and `form` models: the Layout Parser returns an error if `ocrConfig` is set at all. * **`provider_options`** are deep-merged into the body verbatim, after the five configuration keys (`project`, `location`, `processor_id`, `processor_version`, `access_token`) are removed. So `{"processOptions": {"ocrConfig": {"enableNativePdfParsing": true}}}` and `{"imagelessMode": true}` both work, and an explicit value always wins over PuffinParse's default. ## 4. Response mapping The response is `{"document": {…}, "humanReviewStatus": {…}}`. `document.text` holds the whole document's text; everything else points into it with `textAnchor.textSegments[{startIndex, endIndex}]` (UTF-8 offsets, sent as **strings** because they are `int64`). PuffinParse slices `document.text` by byte offset, falling back to character offsets when the indices are not byte boundaries. Fixtures: `crates/puffinparse-core/tests/fixtures/google_documentai_{ocr,layout,form}.json`. ### parse — `ocr`, `form`, `prebuilt` | Document AI field | PuffinParse unified field | Notes | |---|---|---| | `pages[].paragraphs[]` | `Page.blocks[]` (type `text`) | Text via `layout.textAnchor`; empty paragraphs dropped. | | `pages[].tables[]` | `Page.blocks[]` (type `table`) | Rendered as a Markdown table from `headerRows`/`bodyRows` cell anchors; pipes and newlines inside cells are escaped. Emitted *before* the page's paragraphs. | | — | (paragraph suppression) | A paragraph whose box centre falls inside a table's box is skipped, so table text is not duplicated. | | `layout.boundingPoly.normalizedVertices[]` | `Block.bbox` | Min/max of the vertices, already 0–1 with a top-left origin. `vertices[]` (absolute pixels) are divided by `dimension`; without dimensions, `bbox` is `None`. | | `layout.confidence` | `Block.confidence` | 0–1. | | `pages[].pageNumber` | `Block.page_number` / `Page.page_number` | 1-based; falls back to the array index + 1. | | `pages[].dimension.{width,height}` | `Page.width` / `Page.height` | In `dimension.unit` (usually points). | | `pages.len()` | `Usage.pages` | | | `document.error` | error | A non-zero `error.code` becomes a `provider` error. | ### parse — `layout` (Layout Parser) | Document AI field | PuffinParse unified field | Notes | |---|---|---| | `documentLayout.blocks[]` | `Page.blocks[]` | Flattened depth-first, reading order preserved. | | `textBlock.type` | `Block.type` + Markdown prefix | `heading-1` → `title` (`# `), `heading-2`/`subtitle` → `section_header` (`## `), `heading-3/4/5` → `section_header` (`###`…), `header` → `header`, `footer` → `footer`, everything else → `text`. | | `textBlock.blocks[]` | nested blocks | Children are emitted after their parent. | | `tableBlock` | `Block` (type `table`) | Rendered as a Markdown table; `caption` is prepended when present. | | `listBlock` | `Block` (type `list`) | `- item` / `1. item` per `listEntries`, depending on `type`. | | `imageBlock` | `Block` (type `figure`) | Content is `imageText` (OCR/alt text). | | `pageSpan.pageStart` | `Block.page_number` | 1-based. A block spanning pages is filed under its first page. | | `boundingBox.normalizedVertices` | `Block.bbox` | | | max `pageSpan.pageEnd` | `Usage.pages` | The Layout Parser returns no `pages[]`, so page count comes from the spans. | | `chunkedDocument.chunks[]` | — | Not mapped; visible with `include_raw=True`. | | — | `Page.width` / `Page.height` | Not available in `documentLayout`. | ### ocr (native) `mode="ocr"` on `ocr`, `form` and `prebuilt` reads geometry straight off the page: `pages[].lines[]` → `TextPage.lines` (text, box, confidence), `pages[].tokens[]` → `TextPage.words` (trimmed, so trailing `detectedBreak` whitespace does not leak into the word), and the page's own `layout.textAnchor` → `TextPage.text` (falling back to the joined lines). For `layout` there are no lines or tokens, so `ocr` is derived from the parsed blocks (`puffinparse_derived_from=parse`). ### extract — `form`, `prebuilt` | Document AI field | PuffinParse unified field | Notes | |---|---|---| | `entities[].type` | key in `ExtractResponse.data` | Google's own names (`invoice_id`, `total_amount`, `line_item/description`). | | `entities[].properties[]` | nested object | Recursively, keyed by the child's `type`. | | repeated `type` | JSON array | Two `line_item` entities become `data["line_item"] == [ {...}, {...} ]`. | | `normalizedValue` | value | `booleanValue`/`integerValue`/`floatValue`/`signatureValue` are used as-is; `moneyValue`/`dateValue`/`datetimeValue`/`addressValue` are kept as objects with `normalizedValue.text` merged in; otherwise `normalizedValue.text`, else `mentionText`. | | `entities[].confidence` | `FieldInfo.confidence` | Keyed by JSON pointer (`/invoice_id`, `/line_item/1`). | | `entities[].pageAnchor.pageRefs[]` | `FieldInfo.citations[]` | `page` is a **0-based index into `document.pages`** (and is omitted when 0) → `page_number = page + 1`; `boundingPoly` → `bbox`; `mentionText` → `Citation.text`. | | `pages.len()` | `Usage.pages` | | | — | `metadata.google_documentai_schema_source = "processor"` | See the gotcha below. | **`ExtractRequest.schema` is not sent.** A Document AI processor extracts the schema it was trained on; there is no request-time JSON Schema. PuffinParse returns every entity the processor found and leaves the schema as documentation of intent — filter or rename on your side. (`processOptions.schemaOverride` exists but takes Google's `DocumentSchema` proto, not JSON Schema; it is reachable through `provider_options` if your processor version supports it.) ## 5. Errors, status codes, rate limits, timeouts Google's standard envelope — `{"error": {"code", "message", "status", "details": []}}` — is picked up by `Error::from_http` (it reports `error.message`). | Status | `status` | Typical cause | PuffinParse `ErrorKind` | |---|---|---|---| | 400 | `INVALID_ARGUMENT` / `FAILED_PRECONDITION` | page limit exceeded, unsupported MIME type, `ocrConfig` on a Layout Parser, bad page selector | `bad_request` | | 401 | `UNAUTHENTICATED` | missing / **expired** access token | `authentication` | | 403 | `PERMISSION_DENIED` | no `documentai.processors.processOnline`, API not enabled, wrong project | `authentication` | | 404 | `NOT_FOUND` | wrong processor id, or a processor in another location | `bad_request` | | 429 | `RESOURCE_EXHAUSTED` | per-project QPS / pages-per-minute quota | `rate_limit` (retried) | | 500/503 | `INTERNAL` / `UNAVAILABLE` | transient backend failure | `provider` (retried) | A 200 response can still carry `document.error` (a `google.rpc.Status`); a non-zero code becomes a `provider` error. **Limits.** Online `:process` accepts **40 MB** per request (batch: 1 GB) and, for almost every processor, **15 pages** — 30 with `imagelessMode: true`, and only when the pages are contiguous from page 1. Identity/driver-licence processors cap at 2 pages, Expense at 10. Images are capped at 40 megapixels. Bigger documents need `batchProcess` (async, Cloud Storage in and out), which PuffinParse does not implement: an obvious follow-up. **Timeouts.** `timeout_secs` (default 300) covers the whole call and caps the single HTTP request. Because there is no polling, a document that is too large fails fast with a 400 rather than hanging. ## 6. Gotchas (documentation-derived; not yet live-verified) * **Access tokens expire in about an hour.** A long-running process must refresh `GOOGLE_DOCUMENTAI_ACCESS_TOKEN` (or pass `provider_options.access_token` per call); a stale token is a plain 401. * **Location is part of the hostname *and* the resource path.** `us` and `eu` are separate endpoints; a processor created in `us` is `NOT_FOUND` on `eu-documentai.googleapis.com`. * **The model is a *reading strategy*, not a processor selector.** Both come from your configuration: point `GOOGLE_DOCUMENTAI_PROCESSOR_ID` at a Layout Parser and use `google_documentai/layout`; point it at an Invoice Parser and use `google_documentai/prebuilt`. Mismatches produce empty output rather than an error (PuffinParse logs a warning when `extract` finds no entities). * **15 pages online.** This is the single biggest practical limit; Document AI is the only provider here whose sync ceiling is that low. * **`int64` fields are JSON strings.** `startIndex`, `endIndex` and `pageRefs[].page` arrive as `"14"`, not `14` — and a value of `0` is **omitted entirely** (proto3 default), which is why `pageRefs[]` without a `page` means *page 1*. * **Text offsets are UTF-8 byte offsets** in the proto sense. PuffinParse slices bytes when the indices land on char boundaries and falls back to character slicing otherwise, so non-ASCII documents do not panic or truncate mid-codepoint. Worth re-checking against a real CJK document. * **The Document OCR processor returns no tables** — `pages[].tables` is a Form Parser (and specialised processor) feature, so `parse` with `google_documentai/ocr` yields paragraphs only. * **`normalizedVertices` vs `vertices`.** Most processors emit both; a few emit only absolute `vertices`, which are only convertible with `dimension`. The Layout Parser's `boundingBox` has no page dimensions at all, so absolute vertices there yield `bbox = None`. * **`skipHumanReview` is documented as deprecated** but is still the field on `ProcessRequest`; PuffinParse sends `true` so a human-review-enabled processor does not silently queue work. * **Prices differ by 20× across processors** ($1.50 vs $30 per 1 000 pages), so the model string is a cost decision, not just an output-shape decision. ## 7. Useful `provider_options` passthrough ```python # 1. Everything by configuration, nothing in the environment. puffinparse.parse("scan.pdf", model="google_documentai/ocr", provider_options={"project": "my-proj", "location": "eu", "processor_id": "1a2b3c4d5e6f7890", "access_token": token}) # 2. Pin a processor version for reproducible output. puffinparse.parse("scan.pdf", model="google_documentai/ocr", provider_options={"processor_version": "pretrained-ocr-v2.0-2023-06-02"}) # 3. Better text from digital-born PDFs, plus image-quality diagnostics. puffinparse.parse("report.pdf", model="google_documentai/ocr", provider_options={"processOptions": {"ocrConfig": { "enableNativePdfParsing": True, "enableImageQualityScores": True}}}) # 4. Push the online page limit from 15 to 30 (contiguous pages from page 1). puffinparse.parse("long.pdf", model="google_documentai/layout", provider_options={"imagelessMode": True}) # 5. Layout Parser chunking, for RAG pipelines (chunks land in `raw`). puffinparse.parse("handbook.pdf", model="google_documentai/layout", include_raw=True, provider_options={"processOptions": {"layoutConfig": {"chunkingConfig": { "chunkSize": 1000, "includeAncestorHeadings": True}}}}) # 6. Entities from a prebuilt Invoice Parser, with citations. puffinparse.extract("invoice.pdf", model="google_documentai/prebuilt", citations=True, schema={"type": "object", "properties": {"invoice_id": {"type": "string"}}}) # 7. A file already in Cloud Storage. puffinparse.parse("placeholder.pdf", model="google_documentai/ocr", provider_options={"gcsDocument": {"gcsUri": "gs://bucket/doc.pdf", "mimeType": "application/pdf"}}) ``` ## 8. Links * `processors.process` REST reference: <https://cloud.google.com/document-ai/docs/reference/rest/v1/projects.locations.processors/process> * `Document` (response shape): <https://cloud.google.com/document-ai/docs/reference/rest/v1/Document> · `ProcessOptions`: <https://cloud.google.com/document-ai/docs/reference/rest/v1/ProcessOptions> * Processor catalogue: <https://cloud.google.com/document-ai/docs/processors-list> · Layout Parser: <https://cloud.google.com/document-ai/docs/layout-parse-chunk> * Pricing: <https://cloud.google.com/document-ai/pricing> · limits: <https://cloud.google.com/document-ai/limits> · quotas: <https://cloud.google.com/document-ai/quotas> * Authentication: <https://cloud.google.com/docs/authentication/rest> — `gcloud auth print-access-token`. * Not implemented here: `batchProcess` (async, >15 pages, Cloud Storage in/out) and service-account token exchange. --- # Tesseract (local) <!-- source: docs/providers/tesseract.md | url: https://puffinparse.com/docs/providers/tesseract/ --> > **Status: verified locally** (a self-hosted engine, so there is no hosted API to verify against). > Tesseract 5.3.4 and poppler-utils 24.02 (Ubuntu 24.04 > packages) were installed in the development sandbox on 2026-09-24; the `#[ignore]`d live test > in `providers/tesseract.rs` passes and the fixture > `crates/puffinparse-core/tests/fixtures/tesseract_headings.tsv` is real `tesseract ... tsv` output for > `benchmark/datasets/synthetic-v1/docs/headings_001.png`. ## 1. Summary | | | |---|---| | Provider name | `tesseract` | | Runs | locally, by shelling out to the `tesseract` binary (no C bindings, no FFI) | | Binaries | `tesseract` (≥ 4; 5.x recommended), plus `pdftoppm` from poppler for PDFs | | Configuration | `TESSERACT_CMD` (default `tesseract`), `PDFTOPPM_CMD` (default `pdftoppm`) | | API key | none — `puffinparse providers` shows `local` in the Key column | | Price | $0 per page (`pricing.json` source `self-hosted`); the cost is your CPU time | | Docs | <https://tesseract-ocr.github.io/tessdoc/> | | Implementation | `crates/puffinparse-core/src/providers/tesseract.rs` (+ shared helpers in `providers/local.rs`) | Install: ```bash sudo apt-get install -y tesseract-ocr poppler-utils # Debian / Ubuntu brew install tesseract poppler # macOS # extra languages: apt-get install tesseract-ocr-deu tesseract-ocr-fra ... ``` Tesseract is the zero-key, zero-cost baseline: useful offline, in CI, and as the open reference point in the benchmark. It has **no layout model** — no headings, tables, figures or reading-order analysis beyond its own page segmentation — so expect good character accuracy on clean scans and a near-zero table score. ## 2. Models exposed by PuffinParse | Model | Modes | List price | |---|---|---| | `tesseract/default` *(default)* | `ocr` (native), `parse` (derived) | $0 | ## 3. Request flow PuffinParse uses 1. The input is loaded (a URL input is downloaded first) and sniffed by magic bytes, then by extension. Anything that is not an image or a PDF is an `input` error. 2. **Images** (PNG, JPEG, TIFF incl. multi-page, BMP, GIF, WebP, PNM, JP2) are passed to Tesseract directly; a path input is used in place, bytes/URL inputs go to a private scratch directory that is deleted afterwards. 3. **PDFs** are rasterised with `pdftoppm -r <dpi> -png [-f first -l last] input.pdf page` (default 300 dpi), one PNG per page, then each page is OCR'd in order. 4. Per image: `tesseract <image> stdout [-l <lang>] [--psm N] [--oem N] [--dpi N] [-c k=v ...] tsv`. The whole call honours `timeout_secs`; the child process is killed if the deadline passes. Page selection (`pages="1-3,7"`) limits rasterisation to the covering span and drops pages outside the selection; for multi-page TIFFs it filters Tesseract's `page_num`. ## 4. Response mapping Tesseract's TSV renderer emits one row per page (level 1), block (2), paragraph (3), line (4) and word (5), each with a pixel box `left top width height`, and a 0–100 confidence on word rows. | Unified field | Source | |---|---| | `TextPage.width/height`, `Page.width/height` | level-1 row (pixels of the image Tesseract read; for PDFs, the raster at `dpi`) | | `Word.text/bbox/confidence` | level-5 rows with non-empty text; `conf / 100`; `-1` → `None` | | `Line.text` | the line's words joined by a space | | `Line.bbox` / `Line.confidence` | level-4 box; mean of its word confidences | | `TextPage.text` | lines joined by `\n` | | `Block` (parse mode) | one `text` block per paragraph (level 3); `content` = its lines joined by `\n`; box from the paragraph row; confidence = mean word confidence | | `Page.markdown` | paragraphs joined by a blank line (no markdown syntax is invented) | | `Usage.pages` | pages OCR'd | | metadata | `tesseract_lang`, `tesseract_psm` (when set), `tesseract_pdf_dpi` (PDFs); parse responses also carry `puffinparse_derived_from: "ocr"` | | `raw` (`include_raw=True`) | `{"engine": "tesseract", "format": "tsv", "pages": [{"page_number", "tsv"}]}` | Boxes are normalised by the page's pixel size, origin top-left, clamped to 0..1. ## 5. Errors and limits | Situation | Error | |---|---| | `tesseract` / `pdftoppm` not found | `provider_error`: "tesseract binary 'tesseract' not found on PATH: install Tesseract (apt install tesseract-ocr / brew install tesseract) or set TESSERACT_CMD" (same shape for `pdftoppm`, naming poppler-utils and `PDFTOPPM_CMD`). `provider` kind so a router can fall back to another model | | Non-zero exit (unknown language, unreadable image, bad `-c` variable) | `provider_error` with Tesseract's / pdftoppm's stderr verbatim | | Deadline exceeded | `timeout_error`; the child process is killed | | Not an image or PDF (e.g. `.docx`) | `input_error` | | Bad `provider_options` (non-integer `psm`, non-object `config`) | `input_error` | There is no concurrency limit beyond your CPU. Tesseract uses OpenMP threads by default, which oversubscribe the CPU when several pages run in parallel (one small PNG took 70 s instead of 0.7 s on 4 cores), so PuffinParse starts `tesseract` with `OMP_THREAD_LIMIT=1` unless `OMP_THREAD_LIMIT` is already set in the environment. Set it yourself (e.g. `OMP_THREAD_LIMIT=4`) to give a single large document more threads. ## 6. Gotchas * **Language.** Default is Tesseract's `eng`. `language="de"` is mapped to `deu` (common ISO 639-1 codes are mapped; unknown ones pass through); `provider_options.lang` takes Tesseract's own syntax, e.g. `"eng+deu"`. The matching `tesseract-ocr-<lang>` package must be installed. * **Page segmentation.** The default `--psm 3` (automatic) handles multi-column pages well; `6` (single uniform block) can help receipts and forms; `11`/`12` for sparse text. * **Resolution.** Low-resolution scans OCR poorly; for PDFs raise `dpi` (e.g. 400) rather than upscaling images yourself. Images without DPI metadata make Tesseract guess — pass `dpi` if you know it. * **Skew** is not corrected; heavily rotated pages lose accuracy (the `skewed` category of `synthetic-v1` is its weakest). * **No tables.** Table cells come out as lines of text in reading order; the table score is ~0. ## 7. `provider_options` examples ```python import puffinparse puffinparse.ocr("scan.png", model="tesseract/default") # defaults puffinparse.ocr("scan.png", model="tesseract", provider_options={"lang": "eng+fra", "psm": 6}) puffinparse.parse("book.pdf", model="tesseract", provider_options={"dpi": 400, "oem": 1}) puffinparse.ocr("form.png", model="tesseract", provider_options={"config": {"preserve_interword_spaces": 1}}) # -c k=v puffinparse.ocr("scan.png", model="tesseract", provider_options={"cmd": "/opt/tesseract/bin/tesseract"}) ``` | Option | Meaning | |---|---| | `lang` | Tesseract language string (`eng`, `eng+deu`, `chi_sim`) | | `psm` | page segmentation mode (0–13) | | `oem` | OCR engine mode (0–3; 1 = LSTM only) | | `dpi` | PDF raster resolution (default 300); for images, passed to Tesseract as `--dpi` | | `config` | object of Tesseract variables, each passed as `-c key=value` | | `cmd` / `pdftoppm_cmd` | binary paths (override `TESSERACT_CMD` / `PDFTOPPM_CMD`) | ## 8. Links * Command-line usage: <https://tesseract-ocr.github.io/tessdoc/Command-Line-Usage.html> * Improving quality (psm, dpi, preprocessing): <https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html> * `pdftoppm(1)`: <https://manpages.debian.org/pdftoppm> --- # Docling (self-hosted) <!-- source: docs/providers/docling.md | url: https://puffinparse.com/docs/providers/docling/ --> > **Status: verified locally** (a self-hosted engine, so there is no hosted API to verify against). > docling-serve 1.35.0 (docling 2.130.0, docling-core 2.98.0) > was installed with `pip install docling-serve` (CPU torch) and run with `docling-serve run` in > the development sandbox on 2026-09-24. The `#[ignore]`d live test in `providers/docling.rs` > passes, and both fixtures (`docling_multipage.json` for `multipage_001.pdf`, > `docling_headings.json` for `headings_001.png`, from `benchmark/datasets/synthetic-v1`) are real > `GET /v1/result/{task_id}` responses from that server. ## 1. Summary | | | |---|---| | Provider name | `docling` | | Runs | on your own [docling-serve](https://github.com/docling-project/docling-serve) (IBM's open-source Docling as an HTTP service) | | Base URL | `http://localhost:5001` (override: `base_url` on the request, or `DOCLING_BASE_URL`) | | API key | none by default. If the server runs with `DOCLING_SERVE_API_KEY`, set `DOCLING_API_KEY` (or `api_key`); it is sent as `X-Api-Key` | | Price | $0 per page (`pricing.json` source `self-hosted`); the cost is your compute | | API version | docling-serve v1 (`/v1/...`); the older `/v1alpha` paths and the `file_sources` / `http_sources` body fields are gone — current servers want `sources: [{kind, ...}]` | | Implementation | `crates/puffinparse-core/src/providers/docling.rs` | Start a server: ```bash docker run -p 5001:5001 quay.io/docling-project/docling-serve # or docling-serve-cpu / -cu128 # without Docker (≈2.3 GB with CPU-only torch; models download on first use): pip install docling-serve --extra-index-url https://download.pytorch.org/whl/cpu docling-serve run --port 5001 ``` Docling runs a layout model, TableFormer for table structure and an OCR engine for bitmap text, all locally. On a 4-core CPU a one-page image took ~30 s and a 2-page digital PDF ~20 s (first-request model loading excluded); a GPU image is much faster. ## 2. Models exposed by PuffinParse | Model | Modes | List price | |---|---|---| | `docling/default` *(default)* | `parse`, `ocr` (derived) | $0 | `default` is docling's standard pipeline with the server's defaults (OCR on, table structure `accurate`). Other pipelines and presets (`pipeline: "vlm"`, `ocr_preset`, `table_mode: "fast"`) are reachable through `provider_options`. ## 3. Request flow PuffinParse uses Asynchronous, because the synchronous `/v1/convert/source` is capped by the server's `DOCLING_SERVE_MAX_SYNC_WAIT` (120 s by default) and long documents exceed it. 1. `POST {base}/v1/convert/source/async`, JSON: ```json { "options": { "to_formats": ["json"], "image_export_mode": "placeholder", "include_images": false, "page_range": [first, last], "ocr_lang": ["<language>"] }, "sources": [{"kind": "file", "base64_string": "<base64>", "filename": "<name>"}] } ``` URL inputs are sent as `{"kind": "http", "url": "..."}` and fetched by the server. `page_range` is only set when `pages` is given (docling takes one span; PuffinParse requests the span covering the selection and drops the other pages); `ocr_lang` only when `language` is set. `provider_options` are deep-merged into `options`. → `{"task_id", "task_status": "pending", "task_position", ...}` 2. `GET {base}/v1/status/poll/{task_id}` every 0.5 s growing to 5 s until `task_status` is `success` / `partial_success` / `failure` / `skipped`. 3. `GET {base}/v1/result/{task_id}` → `ConvertDocumentResponse`: `{document: {filename, md_content, json_content, ...}, status, errors[], processing_time, timings, confidence}`. ## 4. Response mapping `json_content` is a [DoclingDocument](https://docling-project.github.io/docling/concepts/docling_document/): `body.children` (and `furniture.children`) are the reading order as `{"$ref": "#/texts/3"}` pointers into `texts[]`, `tables[]`, `pictures[]` and `groups[]`. PuffinParse walks `body` depth-first through groups, so blocks come out in docling's reading order. | DoclingDocument | Unified block | |---|---| | `texts[]` label `title` | `title`, `# text` | | `section_header` (with `level`) | `section_header`, `#` × (level + 1) — matches docling's own markdown (`##` for level 1) | | `text`, `paragraph`, `reference`, `checkbox_*` | `text` | | `list_item` (inside a `list` group) | `list`, one block per item, `- text` (or its `marker` when `enumerated`) | | `caption` / `footnote` | `caption` / `footnote` | | `page_header` / `page_footer` (in `furniture`) | `header` / `footer`, placed first / last on their page | | `formula` | `formula`, `$$ … $$` | | `code` | `other`, fenced | | `tables[]` | `table`; markdown built from `data.grid` (spans expanded), else from `data.table_cells` offsets; `text` = one line per row. Captions follow the table | | `pictures[]` | `figure` with empty content (text docling found inside the picture is not emitted, as in docling's markdown); captions follow | | `groups[]` label `inline` | one `text` block (the formatted runs joined) | | other groups (`list`, `key_value_area`, `form_area`, ...) | walked through for their children | * **Boxes**: each item's `prov[].bbox` is `{l, t, r, b, coord_origin}`. With the usual `coord_origin: "BOTTOMLEFT"` (PDF convention, y up) PuffinParse converts `y0 = (H − t) / H`, `y1 = (H − b) / H`; `TOPLEFT` boxes are only divided. `H`/`W` come from `pages[n].size` (PDF points for PDFs, pixels for images). * An item with several `prov` entries (a paragraph continuing on the next page) is split into one block per page using each entry's `charspan`. * `Page.markdown` is the page's blocks joined; docling's own `md_content` is not used because it has no page boundaries. * `Usage.pages` = pages in `json_content.pages` (after page selection); pages without content are kept as empty pages. * Metadata: `docling_status`, `docling_processing_time_s`, `docling_confidence` (docling's `layout_score` / `ocr_score` / `mean_grade` report) and `docling_errors` on `partial_success`. * Block `confidence` is not set: docling reports quality per page/document, not per item. * `provider_job_id` = the docling-serve `task_id`; `raw` = the whole result JSON. ## 5. Errors and limits | Situation | Error | |---|---| | Server not running / wrong port | `network_error` naming docling-serve, the start command and `DOCLING_BASE_URL` | | 401/403 (server has an API key) | `authentication_error` | | 422 (bad `options`) | `bad_request_error` with FastAPI's `detail` | | `task_status: failure` | `provider_error` with the task's `error_message` / `failure` | | result `status: failure` / `skipped` | `provider_error` with `errors[].error_message` joined | | `partial_success` | success; the errors are in `metadata.docling_errors` | | deadline | `timeout_error` (`timeout_secs` covers submit + polling + result) | 5xx and network errors are retried with backoff like every HTTP provider. Results are **single-use** by default (`DOCLING_SERVE_SINGLE_USE_RESULTS=true`, removed `DOCLING_SERVE_RESULT_REMOVAL_DELAY` = 300 s after completion), so a result fetch that is retried after a dropped connection can 404; rerun the call. `DOCLING_SERVE_MAX_DOCUMENT_TIMEOUT` bounds processing time server-side. ## 6. Gotchas * The first request after start-up downloads and loads the models and can take minutes; the benchmark's latency numbers exclude that only if you warm the server first. * The docs at `docs/usage.md` in docling-serve still show `file_sources` / `http_sources` in some examples; servers ≥ 1.x reject those with 422. PuffinParse sends `sources` with `kind`. * `include_images: false` is sent so picture crops are not embedded in the JSON (they are unused and make results large). Override it in `provider_options` if you want them in `raw`. * Heading levels are flat (`level 1`) unless you pass `do_pdf_heading_hierarchy: true`. ## 7. `provider_options` examples ```python import puffinparse puffinparse.parse("report.pdf", model="docling") # standard pipeline puffinparse.parse("report.pdf", model="docling", provider_options={"table_mode": "fast"}) puffinparse.parse("scan.pdf", model="docling", provider_options={"force_ocr": True, "ocr_preset": "tesseract"}) puffinparse.parse("paper.pdf", model="docling", provider_options={"do_formula_enrichment": True, "do_pdf_heading_hierarchy": True}) puffinparse.parse("doc.pdf", model="docling", base_url="http://gpu-box:5001") ``` Any `ConvertDocumentsOptions` field from the docling-serve API reference is accepted. ## 8. Links * docling-serve: <https://github.com/docling-project/docling-serve> — usage: `docs/usage.md`, configuration: `docs/configuration.md`; live OpenAPI at `{base}/docs` * DoclingDocument format: <https://docling-project.github.io/docling/concepts/docling_document/> --- # PaddleOCR (self-hosted) <!-- source: docs/providers/paddleocr.md | url: https://puffinparse.com/docs/providers/paddleocr/ --> > **Status: docs-only.** Implemented from the PaddleOCR 3.x serving API reference (the > "Service-Based Deployment" sections of `docs/version3.x/pipeline_usage/OCR.en.md` and > `PP-StructureV3.en.md` in PaddlePaddle/PaddleOCR) and the PaddleX serving schemas > (`paddlex/inference/serving/infra/models.py`: `DataInfo`, `ImageInfo`, `PDFInfo`), read > 2026-09-24. No PaddleOCR server was run: the sandbox is CPU-only with limited disk and the > PaddlePaddle + PaddleX serving stack and models were not installed. The fixtures > `paddleocr_ocr.json` and `paddleocr_layout_parsing.json` are hand-built from those documented > shapes. Mark this page **verified** after the `#[ignore]`d live test in `providers/paddleocr.rs` > passes against a real server. > Tracked in [issue #17](https://github.com/ajinkyashejul/puffinparse/issues/17). ## 1. Summary | | | |---|---| | Provider name | `paddleocr` (aliases `paddle`, `paddle_ocr`, `paddlex`) | | Runs | on your own PaddleOCR / PaddleX "basic serving" endpoints | | Base URL | `http://localhost:8080` (override: `base_url` on the request, or `PADDLEOCR_BASE_URL`) | | Parse base URL | `PADDLEOCR_PARSE_BASE_URL`, falling back to `PADDLEOCR_BASE_URL` (a `base_url` on the request wins for both modes) | | API key | none | | Price | $0 per page (`pricing.json` source `self-hosted`) | | Implementation | `crates/puffinparse-core/src/providers/paddleocr.rs` | Each served pipeline is its own HTTP service, so ocr and parse usually run on two ports: ```bash pip install "paddleocr[all]" # or paddlex; plus paddlepaddle (CPU) or paddlepaddle-gpu paddlex --install serving paddlex --serve --pipeline OCR --port 8080 # → POST /ocr paddlex --serve --pipeline PP-StructureV3 --port 8081 # → POST /layout-parsing export PADDLEOCR_BASE_URL=http://localhost:8080 PADDLEOCR_PARSE_BASE_URL=http://localhost:8081 ``` ## 2. Models exposed by PuffinParse | Model | Modes | Endpoint | List price | |---|---|---|---| | `paddleocr/default` *(default)* | `ocr` (native) | `POST /ocr` — general OCR pipeline (PP-OCRv5 by default) | $0 | | | `parse` | `POST /layout-parsing` — PP-StructureV3 (layout, tables, formulas, reading order) | $0 | ## 3. Request flow PuffinParse uses One synchronous JSON call per document: ```json {"file": "<base64 of the file, or a URL the server can fetch>", "fileType": 0, "visualize": false} ``` `fileType` is `0` for PDF and `1` for images (from magic bytes, or from the URL's extension; omitted when unknown, and the server infers it). `visualize: false` stops the server from returning base64 visualisation images. `provider_options` are deep-merged into the body, so any documented field (`useDocOrientationClassify`, `useDocUnwarping`, `useTextlineOrientation`, `textDetLimitSideLen`, `textRecScoreThresh`, `useTableRecognition`, `returnMarkdownImages`, ...) passes through. Response envelope (both endpoints): ```json {"logId": "<uuid>", "errorCode": 0, "errorMsg": "Success", "result": {"ocrResults" | "layoutParsingResults": [ ...one per page... ], "dataInfo": {...}}} ``` `dataInfo` is `{"width", "height", "type": "image"}` or `{"numPages", "pages": [{"width", "height"}], "type": "pdf" | "tiff"}` — sizes of the images the pipeline actually ran on (PDF pages are rendered), i.e. the same pixel space as every box. ## 4. Response mapping **ocr** — `result.ocrResults[i].prunedResult` (page `i + 1`): | Unified | Source | |---|---| | `Line.text` / `confidence` | `rec_texts[j]` / `rec_scores[j]`; empty texts dropped | | `Line.bbox` | `rec_boxes[j]` (`[x_min, y_min, x_max, y_max]`), else the enclosing box of `rec_polys[j]`, normalised by the page size from `dataInfo` | | `Word` | each line split on whitespace; no per-word geometry (PaddleOCR recognises lines), confidence = the line's | | `TextPage.text` | lines joined by `\n`, in PaddleOCR's order | **parse** — `result.layoutParsingResults[i].prunedResult.parsing_res_list[]`, already in reading order: | `block_label` | Block type / markdown | |---|---| | `doc_title` | `title`, `# …` | | `paragraph_title` | `section_header`, `## …` | | `text`, `content`, `abstract`, `reference`, `reference_content`, `aside_text`, unknown | `text` | | `table` | `table`; `block_content` is HTML → converted to a markdown table when simple (no spans), else kept as HTML; `text` is one line per row | | `image`, `chart`, `seal`, `header_image`, `footer_image` | `figure` | | `figure_title`, `table_title`, `chart_title` | `caption` | | `formula` | `formula`, wrapped in `$$ … $$` unless already delimited | | `header` / `footer`, `number` | `header` / `footer` | | `footnote`, `vision_footnote` | `footnote` | | `algorithm` | `other`, fenced | | `formula_number` | `other` | `block_bbox` (`[x_min, y_min, x_max, y_max]` pixels) is normalised by the page size from `dataInfo` (falling back to `prunedResult.width/height`). Page markdown is built from the blocks; PaddleOCR's own `markdown.text` is not used because it embeds tables and images as HTML. Both modes: `Usage.pages` = pages returned (after `pages` selection, applied client-side); metadata `paddleocr_log_id`, and `paddleocr_pages_truncated: {returned, document_pages}` when the server returned fewer pages than `dataInfo.numPages` (see §6); `raw` = the whole envelope. ## 5. Errors and limits Failures come back as `{"logId", "errorCode": <HTTP status>, "errorMsg": "..."}`. PuffinParse maps the HTTP status (or `errorCode` if a 200 carries a non-zero code) through the usual table — 401/403 → `authentication_error`, 422/4xx → `bad_request_error`, 5xx → `provider_error` — with `errorMsg` as the message, verbatim. A refused connection is a `network_error` that names the serving command and `PADDLEOCR_BASE_URL` / `PADDLEOCR_PARSE_BASE_URL`. Non-image, non-PDF input is an `input_error` before any request. ## 6. Gotchas * **10-page limit.** By default the serving layer processes only the first 10 pages of a PDF or multi-page TIFF. Set `Serving: extra: max_num_input_imgs: null` in the pipeline config to lift it; PuffinParse flags truncation in `metadata.paddleocr_pages_truncated`. * **Two servers.** `ocr` and `parse` hit different pipelines. If only the OCR pipeline is running, `parse` gets a 404; set `PADDLEOCR_PARSE_BASE_URL`. * **Image payloads.** Without `visualize: false` (PuffinParse sends it) the server returns several base64 JPEGs per page. PP-StructureV3 also returns markdown images unless `returnMarkdownImages: false` — pass it in `provider_options` to shrink responses. * `pages` is applied after the call (the serving API has no page-range field), so every page up to the server's limit is processed. ## 7. `provider_options` examples ```python import puffinparse puffinparse.ocr("scan.png", model="paddleocr") puffinparse.ocr("photo.jpg", model="paddleocr", provider_options={"useDocOrientationClassify": True, "useTextlineOrientation": True}) puffinparse.parse("report.pdf", model="paddleocr", provider_options={"returnMarkdownImages": False, "useChartRecognition": False}) puffinparse.parse("report.pdf", model="paddleocr", base_url="http://gpu-box:8081") ``` ## 8. Links * OCR pipeline, serving API: <https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/OCR.html> * PP-StructureV3, serving API: <https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PP-StructureV3.html> * Serving deployment guide: <https://www.paddleocr.ai/latest/en/version3.x/deployment/serving.html> --- # Benchmark <!-- source: benchmark/README.md | url: https://puffinparse.com/docs/benchmark/ --> A reproducible benchmark that ranks OCR / document-parsing providers on **accuracy**, **latency** and **cost**, using the same unified client as the SDK. Results are committed to `results/` and rendered into [`LEADERBOARD.md`](/docs/benchmark/leaderboard/index.md). ## Principles 1. **Exact ground truth.** Documents in `synthetic-v1` are rendered from the same source the truth markdown is written from, so there is no annotation noise. The generator is seeded and byte-reproducible (`python benchmark/generate_synthetic.py`). 2. **Deterministic metrics.** No LLM judge is needed. Everything is computed in Rust (`puffinparse_core::bench`) from the prediction and the truth after normalisation. 3. **Tied to a dataset revision.** Every result file records the SHA-256 of the manifest plus every input and truth file, the PuffinParse version, the models and the normalisation options. 4. **Three axes.** Accuracy, latency (p50 / p95 / ms per page as observed from the client, which includes upload and polling), and cost per 1,000 pages from the public list price of each model. ## Metrics Normalisation: NFKC, markdown syntax and HTML tags stripped (headings, emphasis, list bullets, table pipes and separator rows), HTML entities decoded, curly quotes and dashes straightened, whitespace collapsed, lowercased (unless `--case-sensitive`). | Metric | Definition | |---|---| | **Overall** | `100 × summary.headline`: the mean over documents of each document's headline metric (`char_similarity`, or `table_score` for `table-only`, or `rule_pass_rate` for `kind: rules`); a failed call scores 0 | | `char_similarity` | `1 − levenshtein(pred, truth) / max(|pred|, |truth|)`; the summary value is the literal mean, never the headline | | `cer` | `levenshtein(pred, truth) / |truth|` | | `wer` | word-level Levenshtein over whitespace tokens `/ |truth words|` | | `word_recall`, `word_precision`, `word_f1` | bag-of-words overlap | | `order_score` | Kendall-τ-style fraction of concordant pairs among lines present in both texts (reading order) | | `table_score` | `char_similarity` restricted to table rows (only when the truth has tables). Markdown pipe tables and HTML `<table>`s (with `colspan`/`rowspan` repeated into every slot, like the ParseBench truth) are both read into rows of cells | | `teds_grid` | TEDS (tree-edit-distance similarity, Zhong et al. 2020) on the `table > row > cell` grid: `1 − TED / max(nodes)`, cell renames cost their normalised Levenshtein distance. Structure-aware where `table_score` is not: a merged or split row or column costs here even when the text is all there. It is TEDS without `thead`/`tbody` and span attributes, since the truth is markdown; each truth table is compared with its best-matching predicted table | | `rule_pass_rate` | `passed / total` over a rule-scored document's assertions (only for `kind: rules`) | ## Document kinds A dataset document declares how it is scored (`kind` in the manifest, default `transcript`; see [`docs/benchmarks/adapters.md`](/docs/benchmark/adapters/index.md)): | `kind` | Truth | Scored by | |---|---|---| | `transcript` | `truth`, a markdown file | the metrics above, headlined by `char_similarity` | | `rules` | `rules`, a JSON list of machine-checkable assertions (`present`, `absent`, `order`, `table_cell`, `bag_of_sentences`) | `rule_pass_rate = passed / total`, which takes the place of `char_similarity` so the document aggregates with the rest | Two per-document adjustments follow from that: - A **`rules` document has no markdown truth.** The runner reads its rule file instead, and a document whose rules cannot be read or parsed fails with `rules unreadable: …` — without spending a provider call. - A **`table-only` document** (ParseBench's table split: the truth is the page's table, the prediction is the whole page) is headlined by `table_score` instead of `char_similarity`, so `Overall` means the same thing for it as for every other document. `char_similarity`, `cer` and `wer` are still recorded, and are still systematically bad on those documents by construction. ## Running ```bash cargo build --release -p puffinparse-cli ./target/release/puffinparse bench run \ --dataset benchmark/datasets/synthetic-v1 \ --models reducto/standard reducto/r-1 extend/parse_performance extend/parse_light \ llamaparse/fast llamaparse/cost_effective llamaparse/agentic \ --concurrency 4 --save-outputs benchmark/runs/outputs ./target/release/puffinparse bench report benchmark/results/*.json > benchmark/LEADERBOARD.md ``` `--save-outputs` writes each model's markdown per document so mistakes can be inspected. Committed runs keep them under `results/outputs/<run_id>/<model>/<doc_id>.md`; the static site under `site/` renders them next to the input and the truth with a word-level diff. `--filter <substring>` and `--limit N` select a subset of documents. Ids of a combined dataset carry their source (`synthetic/plain_001`), so the saved output path keeps that directory level and `--filter synthetic` runs one source. Before spending money, check the plan and put a ceiling on it: ```bash ./target/release/puffinparse bench run --dataset benchmark/datasets/combined-v1 \ --models reducto/standard llamaparse/agentic --out benchmark/results/combined.json --dry-run # | Model | Calls | Skipped (resumed) | Est. pages | $/page | Est. cost | … no provider is called ./target/release/puffinparse bench run … --max-cost 5 # aborts before the first call if the estimate is higher ``` The estimate is manifest `pages` × list price (`pricing.json`), so it is only as good as the manifest's page counts; `--max-cost` refuses to run a model that has no list price. Runs survive interruptions. Every finished (model, document) call is appended to `<out>.partial.jsonl` and flushed immediately; the final JSON is assembled from it at the end and the log is removed. After a crash, Ctrl-C or a batch of provider failures, rerun the same command with `--resume` (and the same `--out`: the default path contains today's date). Pairs that already succeeded — in the partial log or in an existing result JSON — are not called again; missing and failed ones are. The resumed run keeps the original `run_id`, so `--save-outputs` files of the first attempt stay valid, and it refuses to mix in records from another dataset revision or normalisation. `--retries N` re-issues a document after a retryable error (rate limit, 5xx, timeout, network) with backoff; it is off by default because a retried provider job can be billed twice. Each document record is auditable: `provider_job_id` (look the job up in the provider's dashboard), `cache_hit` (`false` when caches were disabled, the default), `attempts`, `started_at`, and `error_kind` + `error` for failures. The run ends with a summary line: calls made, resumed, failed, total cost and wall time. `bench report` renders the leaderboard table (**TEDS** is the mean `teds_grid` over documents whose truth has a table, a **Rules** column shows the mean rule pass rate, `–` for datasets with no rule documents), then the per-category breakdown, and — for datasets whose ids carry a `<source>/` prefix — a per-source breakdown of documents and `Overall`. ### Result JSON for consumers Each model's `summary` carries `headline` (0–1; rank on this, `overall = 100 × headline`), `char_similarity` (literal), `table_score`, `teds_grid`, `rule_pass_rate`, and each document its own `headline`. The run records `scorer_version` (`puffinparse_core::bench::SCORER_VERSION`, currently `2`). Files without it are scorer v1: read `headline` as `overall / 100`; their `summary.char_similarity` held the headline, and they have no `teds_grid`. Re-score them offline: ```bash puffinparse bench rescore benchmark/results/2026-09-11-combined-v1.json \ --outputs benchmark/results/outputs/run-20260911T111039Z # [--dataset <dir>] [--out <path>] ``` `rescore` reads the saved per-document outputs, scores them with the current scorer, keeps latency, cost, pages and errors as measured, and sets `scorer_version` and `rescored_at`. It makes no network calls and refuses to guess: a missing output for a successful document is an error, unless `--keep-missing` is given — then that document keeps its recorded score and the count is recorded as `rescore_kept_docs`. That is how runs over research-only sources (OmniDocBench), whose outputs are not committed, are re-scored. Scoring a single pair without any network access: ```bash puffinparse bench score prediction.md truth.md ``` or from Python: `puffinparse.score(prediction, truth)`. ## Datasets | Dataset | Docs | Categories | Source | |---|---|---|---| | [`synthetic-v1`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/synthetic-v1/README.md) | 39 | plain, invoice, table, two_column, headings, noisy_scan, low_res, multipage, skewed, dense, faded, receipt, complex_table | generated, CC0 | | [`parsebench`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/parsebench/README.md) | 40 committed (1,009 indexed) | tables (transcript, `table-only`), text pages as rule assertions (`kind: rules`) | LlamaIndex ParseBench, Apache-2.0, pinned upstream commit | | [`olmocr`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/olmocr/README.md) | 40 committed (824 indexed) | headers_footers, long_tiny_text, multi_column, old_scans, table_tests — all `kind: rules` (205 assertions) | AI2 olmOCR-bench, ODC-BY-1.0, pinned upstream commit | | [`omnidocbench`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/omnidocbench/README.md) | 40 **indexed, fetched at run time** | 10 document types (book, newspaper, exam paper, slides, notes, …), English + Chinese, transcript | OpenDataLab OmniDocBench, research-only / non-commercial — not redistributed; `python -m benchmark.adapters omnidocbench` materialises it | | [`dpbench`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/dpbench/README.md) | 40 committed (200 convertible) | table, text, chart, figure, equation, list, index (dominant layout feature); reading-order transcript with headers/footers kept, tables as pipe tables | Upstage DP-Bench, MIT, pinned upstream commit | | [`combined-v1`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/combined-v1/README.md) | 79 | synthetic-v1 + parsebench, source-prefixed ids (frozen: has committed results) | per source | | [`combined-v2`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/combined-v2/README.md) | 159 | synthetic-v1 + parsebench + olmocr + omnidocbench (frozen once it has committed results) | per source | | [`combined-v3`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/combined-v3/README.md) | 199 | combined-v2 + dpbench | per source | Adding a dataset: create `benchmark/datasets/<name>/manifest.json` with `{name, version, description, license, documents:[{id, file, truth, pages, category, tags}]}`, put inputs under `docs/` and truth markdown under `truth/`. Public benchmarks are converted by adapters (`python -m benchmark.adapters <name>`; see [`docs/benchmarks/adapters.md`](/docs/benchmark/adapters/index.md)), which also introduce `kind: rules` documents scored by machine-checkable assertions instead of a transcript. The [academic benchmark survey](/docs/benchmark/academic-benchmarks/index.md) covers olmOCR-bench, OmniDocBench, DP-Bench, READoc and others; READoc (a long-document track) is the next candidate. ## Caveats - Synthetic documents are cleaner than most real-world scans. Treat `synthetic-v1` as a floor for basic fidelity, reading order and table structure, not as the last word on hard documents. - Latency is measured from the client through the public API and includes upload, queueing and polling. Run from a different region or under load and numbers will move. - Prices are list prices. Volume discounts, batch queues and cache hits change real cost. - **A rule pass rate is a floor, not an accuracy.** ParseBench's assertions are generated from its own reference extraction, so a rule's text can carry that extraction's artifacts. Scorer v2 neutralises the commonest one — punctuation spaced as separate tokens (`this " agreement "`) — by ignoring spaces next to punctuation on both sides, but others remain (two lines of the page fused into one "sentence", a stray footnote marker), and `present` / `order` still match exact substrings, so such a rule fails on correct output. The effect is the same for every model, so it moves the absolute number far more than the ranking. The one `bag_of_sentences` rule per ParseBench document matches each sentence fuzzily (≥ 0.8 similar) and passes at 80 % of the page's sentences: at the adapter's original 1.0 a single fused "sentence" failed it for every model (see [`docs/benchmarks/findings.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/docs/benchmarks/findings.md)). - **An empty parse is scored, not failed.** A provider that returns HTTP 200 with no text (Reducto and Extend on ParseBench `text_multicolumns_2col`, whose page is one Form XObject with a degenerate `/BBox`) scores 0 on that document but is not counted in **Failed**; it is flagged `empty_output: true` and shown as `(+N empty)` next to the failure count. - **olmOCR `max_diffs` is honoured** since scorer v2 (fuzzy `present` / `absent` / `order` / `table_cell`, as upstream); its skipped tests (math, positional absences, vertical table neighbours, baseline) are counted in `datasets/olmocr/conversion-stats.json`. Its `absent-only` documents pass for an empty parse. - **OmniDocBench must be fetched** before a run (`python -m benchmark.adapters omnidocbench`); otherwise its documents fail as file-not-found and the dataset `sha256` does not cover them. - **DP-Bench truth keeps page headers and footers**, because DP-Bench's own NID scores them; OmniDocBench truth drops them, because OmniDocBench does not. A parser that strips page furniture loses a little on `dpbench` and nothing on `omnidocbench`. - **A combined score mixes datasets, licences and document kinds.** Read `combined-v1` / `combined-v2` / `combined-v3` next to the per-source table under the leaderboard, not instead of it. --- # Leaderboard <!-- source: benchmark/LEADERBOARD.md | url: https://puffinparse.com/docs/benchmark/leaderboard/ --> Generated from the result files in `benchmark/results/` with `puffinparse bench report` (one section per dataset; each section is that command's output for one result file). Higher **Overall** is better (100 = character-exact after normalisation, or every rule passing). Latency is measured from the client through the public API, including upload and polling, with provider result caches disabled. Prices are public pay-as-you-go list prices. Methodology and caveats: [`benchmark/README.md`](/docs/benchmark/index.md). Every document, output, diff and rule check is browsable at [puffinparse.com/benchmark-results](https://puffinparse.com/benchmark-results/). **Headline: `combined-v3`** (run 2026-09-25, scorer v2) — 199 documents from five sources, each scored by its own ground truth: `synthetic-v1` (exact transcripts), a ParseBench subset (rules and table truth), an olmOCR-bench subset (its unit-test style rules, `max_diffs` honoured), an OmniDocBench subset (reading-order transcripts; English and Chinese) and a DP-Bench subset (Upstage's document-parsing benchmark: reading-order transcripts and table truth, MIT). Compare models within a source column rather than across sources. 1,194 API calls, 0 failures, $14.65 at list price. The free `tesseract/default` baseline (Tesseract 5.5.1, `eng`, one OpenMP thread per process, 10 documents at a time on a 10-core Mac) was added to the same run on 2026-10-08 with `bench run --resume`, scored by the same scorer v2: 199 documents, 0 failures, 2 min 42 s wall. Its latency is local CPU time on a shared machine, not comparable to the API rows. OmniDocBench is research-only, so its per-page outputs are not committed (scores are). Older runs are kept for comparison: `combined-v2` (159 documents, 2026-09-24), `combined-v1` (79 documents, 2026-09-11) and `synthetic-v1` (39 documents, 2026-09-11); the 2026-09-11 runs were re-scored offline with scorer v2 on 2026-09-24 (`puffinparse bench rescore`; latency and cost as originally measured). See [`docs/benchmarks/findings.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/docs/benchmarks/findings.md) for what scorer v2 changed. ## combined-v3 (headline) | Rank | Model | Overall | Char sim | CER | WER | Word F1 | Order | Table | TEDS | Rules | p50 latency | p95 latency | ms/page | $/1k pages | Failed | Dataset | |---:|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---| | 1 | `llamaparse/cost_effective` | **84.35** | 0.802 | 0.492 | 0.584 | 0.787 | 0.911 | 0.857 | 0.865 | 70.2% | 9452 ms | 19494 ms | 11154 | $3.75 | 0/199 | combined-v3 v3.0.0 | | 2 | `llamaparse/agentic` | **83.49** | 0.790 | 0.523 | 0.651 | 0.791 | 0.920 | 0.876 | 0.868 | 70.9% | 14127 ms | 29269 ms | 15368 | $12.50 | 0/199 | combined-v3 v3.0.0 | | 3 | `reducto/r-1` | **82.40** | 0.780 | 0.548 | 0.680 | 0.779 | 0.921 | 0.897 | 0.878 | 70.7% | 3459 ms | 9305 ms | 4329 | $10.00 | 0/199 | combined-v3 v3.0.0 | | 4 | `reducto/standard` | **79.78** | 0.756 | 0.567 | 0.703 | 0.755 | 0.913 | 0.872 | 0.859 | 66.0% | 3019 ms | 8443 ms | 3797 | $15.00 | 0/199 (+1 empty) | combined-v3 v3.0.0 | | 5 | `extend/parse_performance` | **77.43** | 0.734 | 0.678 | 0.875 | 0.748 | 0.920 | 0.885 | 0.856 | 67.4% | 21965 ms | 33103 ms | 25309 | $25.00 | 0/199 | combined-v3 v3.0.0 | | 6 | `extend/parse_light` | **76.46** | 0.725 | 0.690 | 0.890 | 0.735 | 0.914 | 0.870 | 0.867 | 65.4% | 32038 ms | 52619 ms | 33067 | $6.25 | 0/199 | combined-v3 v3.0.0 | | 7 | `tesseract/default` | **56.36** | 0.620 | 0.685 | 1.116 | 0.639 | 0.841 | 0.007 | 0.008 | 45.4% | 5381 ms | 26893 ms | 7749 | $0.00 | 0/199 (+4 empty) | combined-v3 v3.0.0 | ### Overall score by category | Model | academic_literature | book | chart | colorful_textbook | complex_table | dense | equation | exam_paper | faded | figure | headers_footers | headings | historical_document | index | invoice | list | long_tiny_text | low_res | magazine | multi_column | multipage | newspaper | noisy_scan | note | old_scans | plain | ppt2pdf | receipt | research_report | skewed | table | table_tests | text | two_column | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | `llamaparse/cost_effective` | 77.6 | 83.8 | 70.7 | 85.9 | 100.0 | 100.0 | 94.4 | 77.0 | 100.0 | 88.2 | 25.0 | 100.0 | 70.9 | 92.9 | 100.0 | 98.8 | 94.3 | 100.0 | 72.3 | 71.2 | 100.0 | 86.9 | 100.0 | 93.1 | 66.3 | 100.0 | 89.3 | 100.0 | 84.6 | 100.0 | 87.7 | 72.9 | 87.2 | 100.0 | | `llamaparse/agentic` | 83.4 | 82.8 | 71.2 | 73.9 | 100.0 | 100.0 | 94.7 | 54.9 | 100.0 | 84.9 | 19.8 | 100.0 | 57.3 | 80.8 | 100.0 | 98.8 | 89.4 | 100.0 | 80.4 | 82.5 | 100.0 | 85.7 | 100.0 | 92.3 | 62.5 | 100.0 | 79.6 | 100.0 | 83.8 | 100.0 | 88.5 | 72.9 | 89.7 | 100.0 | | `reducto/r-1` | 75.0 | 81.2 | 58.4 | 70.4 | 100.0 | 100.0 | 95.5 | 46.7 | 100.0 | 79.8 | 14.6 | 100.0 | 76.4 | 82.7 | 100.0 | 99.6 | 94.7 | 100.0 | 58.2 | 83.1 | 100.0 | 64.3 | 100.0 | 93.4 | 66.3 | 100.0 | 88.7 | 100.0 | 83.4 | 100.0 | 89.1 | 85.4 | 83.4 | 100.0 | | `reducto/standard` | 78.9 | 76.6 | 58.3 | 63.1 | 100.0 | 99.9 | 93.1 | 54.3 | 99.9 | 80.3 | 11.5 | 100.0 | 16.8 | 81.5 | 99.9 | 99.5 | 84.1 | 100.0 | 67.3 | 83.1 | 99.9 | 64.4 | 99.8 | 90.7 | 62.6 | 100.0 | 81.0 | 100.0 | 86.8 | 100.0 | 88.7 | 79.5 | 80.1 | 100.0 | | `extend/parse_performance` | 78.2 | 76.9 | 38.3 | 57.4 | 99.6 | 99.9 | 93.2 | 47.4 | 100.0 | 60.1 | 5.2 | 100.0 | 41.5 | 80.9 | 100.0 | 97.4 | 90.9 | 100.0 | 53.2 | 83.1 | 100.0 | 61.5 | 99.7 | 91.8 | 58.8 | 100.0 | 77.6 | 99.9 | 73.7 | 83.9 | 84.4 | 92.7 | 79.1 | 100.0 | | `extend/parse_light` | 74.5 | 77.0 | 39.1 | 56.4 | 100.0 | 99.9 | 92.8 | 46.6 | 100.0 | 60.9 | 8.3 | 100.0 | 45.1 | 81.8 | 100.0 | 97.1 | 77.7 | 100.0 | 52.7 | 73.8 | 100.0 | 59.4 | 99.4 | 91.4 | 61.3 | 100.0 | 74.3 | 99.9 | 74.0 | 81.0 | 83.2 | 84.4 | 83.1 | 100.0 | | `tesseract/default` | 14.9 | 50.6 | 66.3 | 44.3 | 99.9 | 99.6 | 93.8 | 32.9 | 99.2 | 85.4 | 28.1 | 100.0 | 5.8 | 68.8 | 100.0 | 95.2 | 80.2 | 99.9 | 59.9 | 61.3 | 99.5 | 50.6 | 97.9 | 2.1 | 38.8 | 100.0 | 50.6 | 97.7 | 45.2 | 80.1 | 31.4 | 0.0 | 69.3 | 100.0 | ### `combined-v3` by source | Source | Docs | Model | Overall | |---|---:|---|---:| | dpbench | 40 | `llamaparse/cost_effective` | 90.29 | | dpbench | 40 | `llamaparse/agentic` | 88.77 | | dpbench | 40 | `reducto/standard` | 87.65 | | dpbench | 40 | `reducto/r-1` | 86.90 | | dpbench | 40 | `tesseract/default` | 86.77 | | dpbench | 40 | `extend/parse_light` | 79.63 | | dpbench | 40 | `extend/parse_performance` | 79.45 | | olmocr | 40 | `reducto/r-1` | 68.82 | | olmocr | 40 | `extend/parse_performance` | 66.15 | | olmocr | 40 | `llamaparse/cost_effective` | 65.96 | | olmocr | 40 | `llamaparse/agentic` | 65.43 | | olmocr | 40 | `reducto/standard` | 64.15 | | olmocr | 40 | `extend/parse_light` | 61.09 | | olmocr | 40 | `tesseract/default` | 41.67 | | omnidocbench | 40 | `llamaparse/cost_effective` | 82.14 | | omnidocbench | 40 | `llamaparse/agentic` | 77.39 | | omnidocbench | 40 | `reducto/r-1` | 73.77 | | omnidocbench | 40 | `reducto/standard` | 68.00 | | omnidocbench | 40 | `extend/parse_performance` | 65.93 | | omnidocbench | 40 | `extend/parse_light` | 65.15 | | omnidocbench | 40 | `tesseract/default` | 35.70 | | parsebench | 40 | `llamaparse/agentic` | 86.27 | | parsebench | 40 | `llamaparse/cost_effective` | 83.74 | | parsebench | 40 | `reducto/r-1` | 82.96 | | parsebench | 40 | `reducto/standard` | 79.66 | | parsebench | 40 | `extend/parse_light` | 78.52 | | parsebench | 40 | `extend/parse_performance` | 77.44 | | parsebench | 40 | `tesseract/default` | 20.73 | | synthetic | 39 | `reducto/r-1` | 100.00 | | synthetic | 39 | `llamaparse/agentic` | 100.00 | | synthetic | 39 | `llamaparse/cost_effective` | 100.00 | | synthetic | 39 | `reducto/standard` | 99.96 | | synthetic | 39 | `extend/parse_performance` | 98.70 | | synthetic | 39 | `extend/parse_light` | 98.48 | | synthetic | 39 | `tesseract/default` | 97.98 | ## combined-v2 | Rank | Model | Overall | Char sim | CER | WER | Word F1 | Order | Table | TEDS | Rules | p50 latency | p95 latency | ms/page | $/1k pages | Failed | Dataset | |---:|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---| | 1 | `llamaparse/cost_effective` | **83.79** | 0.786 | 0.578 | 0.683 | 0.757 | 0.903 | 0.876 | 0.880 | 73.0% | 9535 ms | 19846 ms | 11864 | $3.75 | 0/159 | combined-v2 v2.0.0 | | 2 | `llamaparse/agentic` | **83.05** | 0.774 | 0.609 | 0.768 | 0.759 | 0.905 | 0.892 | 0.879 | 71.7% | 14269 ms | 26848 ms | 15701 | $12.50 | 0/159 | combined-v2 v2.0.0 | | 3 | `reducto/r-1` | **81.32** | 0.758 | 0.633 | 0.784 | 0.748 | 0.895 | 0.899 | 0.867 | 70.6% | 3654 ms | 8437 ms | 4841 | $10.00 | 0/159 | combined-v2 v2.0.0 | | 4 | `reducto/standard` | **77.87** | 0.727 | 0.660 | 0.817 | 0.715 | 0.887 | 0.877 | 0.860 | 65.7% | 3570 ms | 8253 ms | 4364 | $15.00 | 0/159 (+1 empty) | combined-v2 v2.0.0 | | 5 | `extend/parse_performance` | **77.35** | 0.720 | 0.729 | 0.966 | 0.720 | 0.895 | 0.893 | 0.853 | 67.4% | 9089 ms | 22019 ms | 10592 | $25.00 | 0/159 | combined-v2 v2.0.0 | | 6 | `extend/parse_light` | **75.92** | 0.710 | 0.744 | 0.979 | 0.707 | 0.888 | 0.866 | 0.846 | 65.4% | 5709 ms | 14298 ms | 7478 | $6.25 | 0/159 | combined-v2 v2.0.0 | | 7 | `tesseract/default` | **48.68** | 0.556 | 0.815 | 1.344 | 0.572 | 0.786 | 0.007 | 0.007 | 45.4% | 2136 ms | 19769 ms | 3841 | $0.00 | 0/159 (+4 empty) | combined-v2 v2.0.0 | ### Overall score by category | Model | academic_literature | book | colorful_textbook | complex_table | dense | exam_paper | faded | headers_footers | headings | historical_document | invoice | long_tiny_text | low_res | magazine | multi_column | multipage | newspaper | noisy_scan | note | old_scans | plain | ppt2pdf | receipt | research_report | skewed | table | table_tests | text | two_column | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | `llamaparse/cost_effective` | 79.2 | 82.9 | 77.9 | 100.0 | 100.0 | 73.0 | 100.0 | 28.1 | 100.0 | 73.8 | 100.0 | 94.3 | 100.0 | 84.9 | 82.5 | 100.0 | 77.4 | 100.0 | 93.2 | 69.1 | 100.0 | 91.0 | 100.0 | 85.2 | 100.0 | 86.9 | 72.9 | 82.7 | 100.0 | | `llamaparse/agentic` | 83.5 | 82.6 | 73.4 | 100.0 | 100.0 | 55.0 | 100.0 | 19.8 | 100.0 | 73.2 | 100.0 | 91.2 | 100.0 | 80.3 | 82.5 | 100.0 | 86.8 | 99.9 | 91.9 | 66.4 | 100.0 | 79.7 | 100.0 | 83.8 | 100.0 | 89.3 | 72.9 | 85.3 | 100.0 | | `reducto/r-1` | 75.0 | 82.6 | 70.5 | 100.0 | 100.0 | 46.7 | 100.0 | 14.6 | 100.0 | 76.3 | 100.0 | 93.6 | 100.0 | 59.3 | 83.1 | 100.0 | 65.9 | 100.0 | 93.4 | 66.3 | 100.0 | 88.7 | 100.0 | 83.4 | 100.0 | 88.6 | 85.4 | 75.8 | 100.0 | | `reducto/standard` | 78.9 | 76.4 | 69.1 | 100.0 | 99.9 | 54.4 | 99.9 | 11.5 | 100.0 | 17.3 | 99.9 | 84.1 | 100.0 | 67.2 | 83.1 | 99.9 | 64.9 | 99.8 | 90.6 | 60.1 | 100.0 | 81.0 | 100.0 | 86.8 | 100.0 | 86.5 | 79.5 | 71.2 | 100.0 | | `extend/parse_performance` | 78.4 | 77.0 | 57.4 | 99.6 | 99.9 | 47.6 | 100.0 | 5.2 | 100.0 | 44.3 | 100.0 | 90.9 | 100.0 | 53.2 | 83.1 | 100.0 | 61.5 | 99.7 | 91.7 | 58.8 | 100.0 | 77.6 | 99.9 | 73.8 | 83.9 | 85.5 | 92.7 | 70.7 | 100.0 | | `extend/parse_light` | 74.7 | 77.3 | 57.9 | 100.0 | 99.9 | 47.6 | 100.0 | 8.3 | 100.0 | 48.5 | 100.0 | 77.7 | 100.0 | 53.2 | 73.8 | 100.0 | 60.3 | 99.4 | 91.4 | 61.3 | 100.0 | 77.4 | 99.9 | 73.4 | 81.0 | 81.7 | 84.4 | 76.7 | 100.0 | | `tesseract/default` | 15.0 | 50.5 | 44.1 | 99.8 | 99.6 | 33.2 | 99.2 | 26.0 | 100.0 | 6.2 | 100.0 | 75.2 | 99.9 | 59.9 | 67.5 | 99.4 | 50.5 | 97.8 | 2.5 | 39.9 | 100.0 | 50.5 | 97.7 | 43.2 | 80.0 | 10.7 | 0.0 | 55.3 | 100.0 | ### `combined-v2` by source | Source | Docs | Model | Overall | |---|---:|---|---:| | olmocr | 40 | `llamaparse/cost_effective` | 69.39 | | olmocr | 40 | `reducto/r-1` | 68.61 | | olmocr | 40 | `llamaparse/agentic` | 66.57 | | olmocr | 40 | `extend/parse_performance` | 66.15 | | olmocr | 40 | `reducto/standard` | 63.65 | | olmocr | 40 | `extend/parse_light` | 61.09 | | olmocr | 40 | `tesseract/default` | 41.72 | | omnidocbench | 40 | `llamaparse/cost_effective` | 81.85 | | omnidocbench | 40 | `llamaparse/agentic` | 79.02 | | omnidocbench | 40 | `reducto/r-1` | 74.18 | | omnidocbench | 40 | `reducto/standard` | 68.66 | | omnidocbench | 40 | `extend/parse_performance` | 66.25 | | omnidocbench | 40 | `extend/parse_light` | 66.18 | | omnidocbench | 40 | `tesseract/default` | 35.56 | | parsebench | 40 | `llamaparse/agentic` | 87.04 | | parsebench | 40 | `llamaparse/cost_effective` | 84.33 | | parsebench | 40 | `reducto/r-1` | 82.95 | | parsebench | 40 | `reducto/standard` | 79.75 | | parsebench | 40 | `extend/parse_performance` | 78.82 | | parsebench | 40 | `extend/parse_light` | 78.49 | | parsebench | 40 | `tesseract/default` | 20.73 | | synthetic | 39 | `reducto/r-1` | 100.00 | | synthetic | 39 | `llamaparse/cost_effective` | 100.00 | | synthetic | 39 | `llamaparse/agentic` | 99.99 | | synthetic | 39 | `reducto/standard` | 99.96 | | synthetic | 39 | `extend/parse_performance` | 98.70 | | synthetic | 39 | `extend/parse_light` | 98.48 | | synthetic | 39 | `tesseract/default` | 97.94 | ## combined-v1 | Rank | Model | Overall | Char sim | CER | WER | Word F1 | Order | Table | TEDS | Rules | p50 latency | p95 latency | ms/page | $/1k pages | Failed | Dataset | |---:|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---| | 1 | `llamaparse/agentic` | **93.08** | 0.821 | 0.916 | 0.872 | 0.865 | 0.922 | 0.922 | 0.859 | 85.4% | 14912 ms | 29307 ms | 16179 | $12.50 | 0/79 | combined-v1 v1.0.0 | | 2 | `reducto/r-1` | **91.37** | 0.802 | 0.921 | 0.872 | 0.842 | 0.894 | 0.924 | 0.872 | 75.9% | 3532 ms | 6565 ms | 3767 | $10.00 | 0/79 | combined-v1 v1.0.0 | | 3 | `llamaparse/cost_effective` | **90.38** | 0.803 | 0.904 | 0.859 | 0.852 | 0.893 | 0.886 | 0.850 | 81.3% | 14005 ms | 29394 ms | 13057 | $3.75 | 0/79 | combined-v1 v1.0.0 | | 4 | `reducto/standard` | **89.91** | 0.794 | 0.944 | 0.900 | 0.829 | 0.885 | 0.913 | 0.868 | 71.2% | 2825 ms | 4670 ms | 2931 | $15.00 | 0/79 | combined-v1 v1.0.0 | | 5 | `extend/parse_light` | **88.02** | 0.775 | 1.020 | 0.975 | 0.822 | 0.905 | 0.892 | 0.863 | 71.2% | 5913 ms | 21784 ms | 8190 | $6.25 | 0/79 | combined-v1 v1.0.0 | | 6 | `extend/parse_performance` | **87.80** | 0.777 | 1.006 | 0.962 | 0.821 | 0.905 | 0.889 | 0.853 | 70.7% | 8920 ms | 21986 ms | 10012 | $25.00 | 0/79 | combined-v1 v1.0.0 | ### Overall score by category | Model | complex_table | dense | faded | headings | invoice | low_res | multipage | noisy_scan | plain | receipt | skewed | table | text | two_column | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | `llamaparse/agentic` | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.9 | 100.0 | 100.0 | 100.0 | 88.3 | 85.4 | 100.0 | | `reducto/r-1` | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 88.6 | 75.9 | 100.0 | | `llamaparse/cost_effective` | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 82.9 | 81.3 | 100.0 | | `reducto/standard` | 100.0 | 99.9 | 99.9 | 100.0 | 99.9 | 100.0 | 99.9 | 99.8 | 100.0 | 100.0 | 100.0 | 87.0 | 71.2 | 100.0 | | `extend/parse_light` | 100.0 | 99.9 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.4 | 100.0 | 99.9 | 81.0 | 83.7 | 71.2 | 100.0 | | `extend/parse_performance` | 99.6 | 99.9 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.7 | 100.0 | 99.9 | 80.6 | 83.5 | 70.7 | 100.0 | ### `combined-v1` by source | Source | Docs | Model | Overall | |---|---:|---|---:| | parsebench | 40 | `llamaparse/agentic` | 86.34 | | parsebench | 40 | `reducto/r-1` | 82.95 | | parsebench | 40 | `llamaparse/cost_effective` | 81.00 | | parsebench | 40 | `reducto/standard` | 80.11 | | parsebench | 40 | `extend/parse_light` | 77.82 | | parsebench | 40 | `extend/parse_performance` | 77.42 | | synthetic | 39 | `reducto/r-1` | 100.00 | | synthetic | 39 | `llamaparse/cost_effective` | 100.00 | | synthetic | 39 | `llamaparse/agentic` | 99.99 | | synthetic | 39 | `reducto/standard` | 99.96 | | synthetic | 39 | `extend/parse_light` | 98.48 | | synthetic | 39 | `extend/parse_performance` | 98.44 | ## synthetic-v1 | Rank | Model | Overall | Char sim | CER | WER | Word F1 | Order | Table | TEDS | Rules | p50 latency | p95 latency | ms/page | $/1k pages | Failed | Dataset | |---:|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---| | 1 | `reducto/r-1` | **100.00** | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | – | 3262 ms | 8778 ms | 4231 | $10.00 | 0/39 | synthetic-v1 v1.1.0 | | 2 | `llamaparse/cost_effective` | **100.00** | 1.000 | 0.000 | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | – | 9157 ms | 14872 ms | 8910 | $3.75 | 0/39 | synthetic-v1 v1.1.0 | | 3 | `reducto/standard` | **99.96** | 1.000 | 0.000 | 0.001 | 0.999 | 1.000 | 0.998 | 0.986 | – | 2848 ms | 4347 ms | 2669 | $15.00 | 0/39 | synthetic-v1 v1.1.0 | | 4 | `llamaparse/agentic` | **99.75** | 0.997 | 0.003 | 0.003 | 0.998 | 1.000 | 1.000 | 1.000 | – | 9666 ms | 15199 ms | 10658 | $12.50 | 0/39 | synthetic-v1 v1.1.0 | | 5 | `extend/parse_performance` | **98.70** | 0.987 | 0.013 | 0.018 | 0.998 | 1.000 | 0.999 | 0.949 | – | 5680 ms | 9077 ms | 6184 | $25.00 | 0/39 | synthetic-v1 v1.1.0 | | 6 | `extend/parse_light` | **98.48** | 0.985 | 0.015 | 0.022 | 0.998 | 1.000 | 1.000 | 1.000 | – | 5517 ms | 8932 ms | 5364 | $6.25 | 0/39 | synthetic-v1 v1.1.0 | | 7 | `llamaparse/fast` | **87.58** | 0.876 | 0.124 | 0.144 | 0.926 | 1.000 | 0.161 | 0.170 | – | 5834 ms | 9286 ms | 5995 | $1.25 | 0/39 | synthetic-v1 v1.1.0 | ### Overall score by category | Model | complex_table | dense | faded | headings | invoice | low_res | multipage | noisy_scan | plain | receipt | skewed | table | two_column | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | `reducto/r-1` | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | | `llamaparse/cost_effective` | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | | `reducto/standard` | 100.0 | 99.9 | 99.9 | 100.0 | 99.9 | 100.0 | 99.9 | 99.8 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | | `llamaparse/agentic` | 100.0 | 100.0 | 100.0 | 96.8 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | | `extend/parse_performance` | 99.6 | 99.9 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.7 | 100.0 | 99.9 | 83.9 | 100.0 | 100.0 | | `extend/parse_light` | 100.0 | 99.9 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.4 | 100.0 | 99.9 | 81.0 | 100.0 | 100.0 | | `llamaparse/fast` | 99.2 | 99.9 | 99.9 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 | 100.0 | 99.4 | 11.1 | 100.0 | 29.3 | --- # Vendor benchmarks <!-- source: docs/benchmarks/vendor-benchmarks.md | url: https://puffinparse.com/docs/benchmark/vendor-benchmarks/ --> Researched 2026-09-11. All three benchmarks were **verified by actually downloading a sample** into a scratch directory. Nothing in this repository was modified. PuffinParse manifest target format (from `benchmark/README.md`): `benchmark/datasets/<name>/manifest.json` = `{name, version, description, license, documents:[{id, file, truth, pages, category, tags}]}` with inputs in `docs/` and **ground-truth markdown per document** in `truth/`. Scoring is deterministic char/word Levenshtein + reading order + table sub-score against that truth markdown. --- ## Summary table | Name | Public data? | Location | License | Size | GT format | Convertible to PuffinParse manifest? | |---|---|---|---|---|---|---| | **ParseBench** (LlamaIndex) | Yes, fully | HF `llamaindex/ParseBench`; code `github.com/run-llama/ParseBench` | Apache-2.0 (data card + code LICENSE) — redistribution permitted | 592 MB total (517 MB docs, 2,079 files; 71 MB rule JSONL) | 169,011 **rule** assertions in 5 JSONL files; **no reference markdown** except 503 HTML tables | **Partially.** Table split → direct (503 docs, HTML→md truth). Text splits → only a *reconstructable approximation* from `bag_of_sentence` + pairwise `order` rules. Chart/layout/formatting splits → not convertible. | | **RealDoc-Bench** (Extend) — QA track | Yes | HF `Extend-AI/RealDoc-Bench`; code `github.com/extend-hq/realdoc-bench` | Annotations **CC-BY-4.0**; **source PDFs explicitly excluded from that license**, rights vary (471/581 "not_established") | 540 MB (538 MB = 581 PDFs; `qa_bank.json` 1.1 MB) | 1,356 question → typed `gold_dict` JSON pairs (+137 capability tags) | **No** for char-level markdown scoring — there is no page text truth at all. Only usable as a *separate QA-over-parse* track. PDFs are **download-on-demand only**, do not redistribute. | | **RealDoc-Bench-Layout** (Extend) | Yes | HF `Extend-AI/RealDoc-Bench-Layout` | Annotations CC-BY-4.0; page images keep original per-source licenses | 374 MB (1,500 PNG/JPG = 371 MB; annotations 2.3 MB) | COCO bbox + 9 block classes per page. **No text in annotations** | **No.** Bboxes only, zero transcription. | | **LongExtractBench-50** (micro1, commissioned by Reducto) | Yes (50-doc subset of a 225-doc corpus) | HF `micro1-inc/longextract-bench-50`; code `github.com/micro1-research/longextract-bench` | Labels **CC-BY-4.0** (micro1); **`document.pdf` retains original source rights** — "verify before redistributing" | 375 MB / 50 folders × 3 files (largest folder 84 MB; median GT ~640 KB) | `schema.json` (JSON Schema) + `ground_truth.json` (nested JSON extraction) | **No** for markdown scoring — GT is a structured extraction, not page text. Excellent as a *long-document extraction* track; PDFs are 1–200+ pages so they are also a good latency/robustness stress corpus. | **Bottom line for a "combined open benchmark dataset" in PuffinParse's current manifest shape (input file + expected markdown per document): only ParseBench's `table` split drops in cleanly (503 documents, Apache-2.0, redistributable).** Everything else measures either field extraction or layout, and would need a second manifest/scorer kind. --- ## 1. ParseBench — LlamaIndex / LlamaParse - Website `parsebench.ai` · Paper arXiv:2604.08538 (Zhang, Acosta, Carlson, Bron, Doulcet, Ospina, Suo, 2026) · Code `github.com/run-llama/ParseBench` (Apache-2.0) · Data `huggingface.co/datasets/llamaindex/ParseBench`. ### What it measures Parsing/OCR fidelity (not extraction), split into **five capability dimensions**: | Dimension | Metric | Pages | Docs | Rules | |---|---|---:|---:|---:| | Tables | GTRM = mean(GriTS, TableRecordMatch) | 503 | 284 | — (continuous) | | Charts | ChartDataPointMatch | 568 | 99 | 4,864 | | Content Faithfulness | Content Faithfulness Score | 506 | 506 | 141,322 | | Semantic Formatting | Semantic Formatting Score | 476 | 476 | 5,997 | | Layout / Visual Grounding | Element Pass Rate (IoA + class + attribution) | 500 | 321 | 16,325 | | **Total unique** | | **2,078** | **1,211** | **169,011** | Content Faithfulness and Semantic Formatting share the same 507–508 text pages with different rule sets. ### Documents & languages Publicly-sourced enterprise documents: insurance (SERFF filings), finance (10-K, proxy statements, annual reports), government publications (UN E-Government Survey, OECD), plus newspapers, timetables, contracts. **One page per document file** — every `docs/*` entry is a single-page PDF/JPG/PNG cut from a larger source. Card language field is `en`, but the text split has a `multilang` tag covering **20+ languages / all major scripts (47 docs)**; also `handwritting` (13), `ocr` scans (119), `multicolumns` (97), `dense` (14), `sparse` (14), `simple` (170), `misc` (33). ### Data availability — VERIFIED Public, no auth, no gating. Downloaded: ``` parsebench/ 7.3 MB downloaded ├── README.md, eval.yaml ├── chart.jsonl 1,591,287 B ├── table.jsonl 3,087,333 B ├── text_formatting.jsonl 1,785,910 B ├── tc_head.jsonl 300,001 B (first 300 KB of text_content.jsonl via HTTP Range) └── docs/chart/(Web_version)_E-Government_Survey_2024_1392024_p101.pdf 183,569 B (1 page) ``` Full repo: **592.2 MB**, 2,113 files — `docs/` 2,079 files / 517 MB (chart 568, text 508, table 503, layout 500; 2,037 `.pdf`, 23 `.jpg`, 19 `.png`), `text_content.jsonl` 55.4 MB, `layout.jsonl` 9.6 MB, `table.jsonl` 3.1 MB, `text_formatting.jsonl` 1.8 MB, `chart.jsonl` 1.6 MB, `thumbnails/` 3.5 MB. Command that worked (no TLS games needed, just point the CA bundle at the proxy): ```bash export REQUESTS_CA_BUNDLE=/root/.ccr/ca-bundle.crt SSL_CERT_FILE=/root/.ccr/ca-bundle.crt python -c "from huggingface_hub import snapshot_download; snapshot_download( 'llamaindex/ParseBench', repo_type='dataset', local_dir='parsebench', allow_patterns=['README.md','eval.yaml','*.jsonl','docs/chart/…p101.pdf'])" ``` ### License / redistribution `license: apache-2.0` in the dataset card front-matter, and the card's Copyright Statement: *"All documents are sourced from public online channels. The dataset is released under the Apache 2.0 License. If there are any copyright concerns, please contact us via the GitHub repository."* SPDX: **Apache-2.0**. Redistribution in another repo is therefore **permitted by the publisher's terms** — this is the only one of the three where that is true. Caveat worth recording in our own README: the underlying pages are third-party corporate and government documents that LlamaIndex re-licensed unilaterally; for a low-risk posture we could still ship only the manifest + SHA-256 and fetch on demand. ### Ground-truth format — VERIFIED One JSONL line per **test rule**, identical schema across all five files: ```json {"pdf":"docs/chart/report_p41.pdf","category":"chart","id":"b17e5e98d6fc2763", "type":"chart_data_point","rule":"{...json-encoded payload...}","page":null, "expected_markdown":null,"tags":["need_estimate"]} ``` Real rows pulled from the download: - **table.jsonl** (the only one with reference content): `id: "0000027_page1_expected_markdown"`, `type: "expected_markdown"`, `rule: {}`, `tags: ["easy"]`, and `expected_markdown: "<table>\n<tr><th>Business activity</th><th>Registered office</th>…<tr><td colspan=\"6\"><strong>JOINT VENTURES CONSOLIDATED USING THE EQUITY METHOD</strong></td></tr>…"` — i.e. **ground-truth HTML tables** (with `colspan`/`rowspan`/`<br/>`/`<strong>`), 503 of them. - **chart.jsonl**: `rule: {"labels":["IF","193 UN Member States"],"max_diffs":0, "normalize_numbers":true,"value":"0.8079"}` — a spot-check data point, no page text. - **text_formatting.jsonl**: `rule: {"text":"BAOTOU 包头","level":1}`, `type: "is_title"`, `tags: ["dense","hard"]`. - **text_content.jsonl** (11 rule types seen in the first 300 KB — counts in that window: `missing_specific_word` 520, `missing_specific_sentence` 96, `order` 71, plus one each of the aggregate types). The aggregate rules carry the actual reference content as **bags**: - `{"bag_of_sentence": {"BAOTOU 包头":1, "Baotou is the largest city within the Inner Mongolia Autonomous Region in China":1, …}}` - `{"bag_of_word": {"10":1,"12":2,…}}` - `{"bag_of_digit": {"0":60,"1":46,…}}` - `{"before":"Baotou is the largest city…","after":"It's population of more than 1.6 million…","max_diffs":0}` Metadata: `category` (chart/table/text/layout), `tags` = difficulty (`easy`/`hard`) + document type (`dense`,`sparse`,`simple`,`multicolumns`,`ocr`,`multilang`,`misc`,`handwritting`) + chart flags (`need_estimate`, `3d_chart`). `page` is 1-indexed, used by layout rules only. Images/PDFs are plain files under `docs/<category>/<name>.{pdf,jpg,png}`, referenced by the relative `pdf` field. ### Scoring / official eval code `github.com/run-llama/ParseBench` (Apache-2.0), a full harness with 180+ pipeline configurations (OpenAI, Anthropic, Google, LlamaParse, specialised parsers), parallel runs, per-dimension scoring and cross-pipeline comparison; `eval.yaml` declares six tasks (`mean`, plus one per split). **Deterministic and rule-based — no LLM judge by default.** Metrics: GTRM (GriTS + TableRecordMatch, bag-of-records, order-insensitive) for tables; ChartDataPointMatch (orientation-insensitive, numeric-tolerant) for charts; rule pass-rate scores for content faithfulness and formatting; Element Pass Rate (IoA localisation + classification + attribution) for layout. Leaderboard headline: LlamaParse Agentic Plus 90.20, LlamaParse Agentic 87.01 (vendor-run). ### Conversion into the PuffinParse manifest - **`table` split → clean fit.** 503 single-page PDFs, each with one ground-truth HTML table. Convert HTML → markdown pipe table for `truth/<id>.md`, `pages: 1`, `category: "table"`, `tags: ["easy"|"hard","parsebench"]`. Our `table_score` metric applies directly; our `char_similarity` becomes a table-fidelity proxy. **Lost:** GriTS structural scoring, merged-cell/colspan semantics (markdown pipe tables cannot express `colspan`/`rowspan`, so hierarchical headers get flattened) — this is a real fidelity loss on the `hard` tables. - **`text` split (508 pages) → approximate fit only.** There is no ordered reference markdown. You could reconstruct a pseudo-truth by taking `bag_of_sentence` and topologically sorting it with the pairwise `order` rules, but ordering is only partially constrained, sentence boundaries are the annotator's, and all formatting/whitespace is gone — so CER/WER against it would be systematically wrong. Not recommended as ground truth; better to keep ParseBench's own rule scorer for that split. - **`chart`, `layout`, `text_formatting` → not convertible.** Data-point assertions, bboxes and style flags have no text-similarity analogue. - Net: **~503 of 2,078 pages (24%) usable in the current manifest shape.** --- ## 2. RealDoc-Bench — Extend (extend.ai) - Blog: `extend.ai/resources/realdocbench` and `/resources/parse-2-and-realdocbench-launch` · Paper arXiv:2606.07401 (CC BY 4.0 on arXiv) · Code `github.com/extend-hq/realdoc-bench` (Apache-2.0, `pip install realdoc-bench`) · Data `huggingface.co/datasets/Extend-AI/RealDoc-Bench` and `huggingface.co/datasets/Extend-AI/RealDoc-Bench-Layout`. ### What it measures Two tracks, neither of which is OCR-similarity: 1. **QA track** — *field-level extraction accuracy through a parser*. The parser produces markdown; an LLM reader (Gemini 3 Flash in the official harness) answers each question using **only** that markdown; the answer is scored against a typed gold dict. Reported as **per-field accuracy** and **strict per-question accuracy**, plus cost and latency. 1,356 questions / 3,742 fields / 581 documents. 2. **Layout track** — bounding-box + block-type detection over 1,500 page images. Hungarian matcher with adjacency-aware split/merge recovery; strict **F1**, **adjusted F1** (allows merging adjacent same-type fragments), **mAP**, with per-class breakdowns. Headline (vendor-run): Extend Parse 2.0 96.0% per-field / 90.9% per-question, LlamaParse (Agentic) 92.2% / 84.5%; layout 0.781 strict F1 / 0.847 adjusted F1. ### Documents & languages Four domains — counted from the downloaded `qa_bank.json`: **mortgage 478, finance 378, supply_chain 319, medical_healthcare 181** questions. Document types: hospital intake forms, EOBs, tax and ACORD insurance forms, mortgage packets, bills of lading, "systems-of-record" documents; dense forms, checkboxes, handwriting, stamps, barcodes, messy scans. Language: `en` only per the card. The two PDFs I fetched were **1 page each** (461 KB and 27 KB), so the QA corpus is mostly short form-like documents (581 docs / 538 MB ≈ 0.93 MB each). Layout track domains include `government`, `billing`, etc. (from `manifest.csv`). ### Data availability — VERIFIED Public, no auth. Downloaded: ``` realdocbench/ 2.2 MB downloaded ├── README.md ├── qa_bank.json 1.09 MB (1,356 items) ├── manifest.json 0.44 MB (581 document provenance records) └── docs/finance_1.pdf 460,998 B (1 page) ; docs/mortgage_1.pdf 26,930 B (1 page) realdocbench_layout/ 1.4 MB downloaded ├── README.md, manifest.csv (0.43 MB) ├── annotations/0008ff17-….json └── images/0008ff17-….png ``` Full repos: QA **539.6 MB** / 585 files (`docs/` 581 PDFs = 538 MB); Layout **373.9 MB** / 3,003 files (`images/` 1,500 = 371 MB, `annotations/` 1,500 = 2.3 MB). ### License / redistribution — **the blocker** Card: *"The QA bank and gold answers are licensed under CC BY 4.0. **Source documents in `docs/` are excluded from this annotation license.** Their applicable rights and reuse terms vary… Some source rights remain unverified. A source URL, public availability, or AI modification does not by itself establish permission to redistribute or relicense a document."* `manifest.json` (`schema_version 1.0`, `metadata_as_of 2026-09-09`, `annotation_license CC-BY-4.0`) makes this concrete — aggregated over all 581 documents: | `source_rights.status` | count | |---|---:| | `not_established` | 471 | | `public_domain_us_federal_work` | 58 | | `copyright_notice_no_reuse_license` | 31 | | `explicit_license` | 8 | | `agency_reuse_policy_with_conditions` | 8 | | `distribution_restrictions_no_open_license` | 5 | `document_license` is **`null` for all 581**. `ai_status`: `ai_generated_edit` 291, `collected_real_document` 196, `unknown` 94 — i.e. **half the corpus is synthetically edited**, which matters if we claim "real documents". Each record carries `sha256`, `source.url`, `source.type` (`original_document` 290 / `original_template` 288), `source.match`, `source.availability`. Monthly takedown process with `takedowns/removed_ids.jsonl`. **Verdict: annotations SPDX CC-BY-4.0 and redistributable with attribution; PDFs are download-on-demand only.** We can ship a manifest + sha256 + HF path, never the bytes. The layout images are the same story (annotations CC-BY-4.0, images keep per-source licenses). ### Ground-truth format — VERIFIED `qa_bank.json` = `{name, domains:["finance","medical_healthcare","mortgage","supply_chain"], items:[…1,356…]}`. A real item: ```json { "question_id": "finance_q1", "source_file": "finance_1", "domain": "finance", "question": "In the PRIOR CARRIER INFORMATION (continued) table, for the year 201 entry with an expiration date of 12/31/2024, what is the premium for the automobile category?", "response_format": "Return exactly: automobile_premium=<number>", "gold_answer": "automobile_premium=12800", "gold_dict": {"automobile_premium": 12800}, "capabilities": ["field_value_pairing","multi_column_grid","repeated_labels","row_binding","table_structure"] } ``` (the HF card mentions a `template` field; in the shipped file it is absent/`null` for all 1,356 items — the typing lives in `response_format` + `gold_dict`.) **137 distinct capability tags** across 8 buckets; most common: `field_value_pairing` 502, `checkbox_state` 385, `column_alignment` 298, `row_binding` 288, `table_structure` 272, `form_region` 222, `parallel_columns` 203, `line_binding` 180, `scanned_form` 160, `multi_column_grid` 150, `handdrawn_check` 127, `blank_field` 122. These are excellent difficulty/category metadata and map well onto our `tags` field. Layout annotation (`annotations/<pageId>.json`, COCO-style): ```json {"image":{"id":20,"file_name":"human/0008ff17-….png","width":640,"height":1102,"domain":"government"}, "annotations":[{"id":198,"image_id":20,"category_id":0,"bbox":[23,24,58,13]}, …], "categories":[{"id":0,"name":"text"},{"id":1,"name":"heading"},{"id":2,"name":"section_heading"}, {"id":3,"name":"header"},{"id":4,"name":"footer"},{"id":5,"name":"page_number"}, {"id":6,"name":"figure"},{"id":7,"name":"table"},{"id":8,"name":"key_value"}], "page_info":{…}} ``` Note: **no `content`/text field on the annotations** — pure geometry + class. `manifest.csv` is the canonical row list: `image_id,file_name,domain,pageId,match_status,originalImageUrl,sourceUrl`. ### Official eval code `pip install realdoc-bench`; pipeline `download → parse → score → report`, per-parser scoping with cached intermediates: ```bash realdoc-bench evaluate download --run-dir runs/v1 --dataset Extend-AI/Realdoc-Bench realdoc-bench evaluate run --run-dir runs/v1 -p extend_performance_v2_0_0_advanced realdoc-bench evaluate run --run-dir runs/smoke -p pymupdf --limit 20 ``` Layout normaliser at `realdoc_bench/layout/normalizers/coco.py`. Apache-2.0. Scoring is **exact-match over typed gold dicts**, but the *answering* step uses an LLM reader, so runs are not bit-reproducible and cost money — that conflicts with PuffinParse's "no LLM judge, deterministic metrics" principle #2. ### Conversion into the PuffinParse manifest **Not convertible as-is.** There is no page transcription anywhere in either repo, so nothing can populate `truth/<id>.md`. Options: - Add a **second manifest kind** (`qa`) — `{id, file, questions:[{question, response_format, gold_dict, capabilities}], pages, category, tags}` — plus a scorer that runs the parse, feeds the markdown to a reader model, and exact-matches `gold_dict`. That is a genuinely different (and non-deterministic, paid) axis from `bench run`. - Or use the layout track as a third kind for bbox F1. - Either way `docs/` stays **download-on-demand** (manifest records `sha256` + HF repo path; a `puffinparse bench fetch` step pulls 540 MB / 374 MB on first use). - **Lost if forced into the markdown shape:** everything — you'd be inventing truth. --- ## 3. LongExtractBench — micro1 (commissioned by Reducto) - Site `micro1.ai/benchmark/long-extraction` · Code `github.com/micro1-research/longextract-bench` (**MIT**) · Data `huggingface.co/datasets/micro1-inc/longextract-bench-50` (CC-BY-4.0) · PR: "Reducto Deep Extract Ranks First Overall in LongExtractBench" (PRNewswire). ### What it measures **Schema-driven structured extraction, not OCR.** Given `document.pdf` + `schema.json`, a system must emit JSON validating against the schema; it is compared cell-by-cell to `ground_truth.json`. Three capabilities: extraction fidelity, schema conformance, and long-document handling. Seven providers evaluated: Reducto, Extend, LlamaExtract, OpenAI, Anthropic Claude, Google Gemini, Datalab. Reducto's reported result: 99.6% recall, 99.6% precision, 99.3% leaf accuracy, 0 failures, only provider at 100% coverage. ### Documents, pages, languages Public HF release is a **50-document curated subset** of a **225-document** corpus (benchmark run dated 2026-06-26) — the other 175 are not released. Stratified by page count (short → multi-hundred pages), schema complexity (flat → deeply nested multi-array), and domain: government/public-sector statistics, financial filings (10-K / 10-Q / DEF 14A proxy), healthcare & clinical reporting, regulatory/compliance, energy, education, census & demographics. **Predominantly English, with a few German and Dutch documents** (e.g. the folder `b9489a19__Statistisch Jaarboek 2025`). Verified page count on the sample I pulled: `06_19_Bankruptcy_Filings_Statistics/document.pdf` = **202 pages** in 823 KB — these are long, dense, table-heavy, mostly born-digital PDFs, not scans. ### Data availability — VERIFIED Public, no auth, no gating. Downloaded: ``` longextractbench/ 1.6 MB downloaded ├── README.md └── 06_19_Bankruptcy_Filings_Statistics/ ├── document.pdf 822,849 B (202 pages) ├── ground_truth.json 652,812 B └── schema.json 3,505 B ``` Full repo: **374.7 MB**, 152 files = 50 folders × 3 files + `.gitattributes`. Largest folders: `UK_asylum-applications-datasets-mar-2023` 84.3 MB, `Capital_improvement_plan___CIP_project_budget_report` 36.3 MB, `std__Annual_report_10-K_-_wfc-20201231_d2` 19.9 MB, `06_19_Government_zoning_and_land_use_geospatial_datasets` 17.8 MB. Folder slugs carry an informal difficulty prefix (`std__`, `hard__`, `m1__`, date prefixes) that could feed our `tags`. ### License / redistribution Card front-matter `license: cc-by-4.0`, but the License section is precise: *"The label files (`schema.json`, `ground_truth.json`) are released by micro1. The underlying `document.pdf` files originate from public sources and retain their original rights — verify the terms of an individual document before redistributing it."* SPDX: **CC-BY-4.0 for labels; unspecified/per-source for PDFs.** Code is **MIT**. Same posture as RealDoc-Bench: **labels redistributable, PDFs download-on-demand.** Note the disclosure: ground truth is **model-assisted** (drafted by a frontier model, reconciled by humans) — "high quality but not guaranteed error-free, and may share blind spots with the LLMs being evaluated" — and the benchmark was **commissioned by Reducto**, who then won it. Both facts belong in any README we write. ### Ground-truth format — VERIFIED `schema.json` is a standard JSON Schema: `type: "object"`, `additionalProperties: false`, a `required` list, and a **natural-language `description` on every field** that pins down formatting and null-handling. From the downloaded sample: ```json {"title":"Bankruptcy Filing Statistics Data Export Schema","type":"object", "additionalProperties":false, "properties":{ "covered_year_end":{"type":"integer","description":"The largest year value appearing in the printed 'year' column … Emit as a four-digit integer with no quotes, commas, or decimal point…"}, "covered_year_start":{"type":"integer","description":"…"}, "filing_count_records":{"type":"array","description":"One item for each non-header data row … Preserve the document's row order from the first data row through the final data row, continuing across page breaks…", "items":{"type":"object","additionalProperties":false, "required":["year","chapter","district","case_count"], "properties":{"year":{"type":"integer","description":"…"}, "chapter":{"type":"integer","description":"…"}, "district":{"type":"string","description":"…preserving lowercase letters and any printed alphabetic suffix…"}, "case_count":{"type":"integer","description":"…remove thousands separators…"}}}}}, "required":["covered_year_start","covered_year_end","filing_count_records"]} ``` `ground_truth.json` for that document (653 KB) begins: ```json {"covered_year_end": 2025, "covered_year_start": 2008, "filing_count_records": [ {"case_count": 245, "chapter": 11, "district": "akbk", "year": 2008}, {"case_count": 17, "chapter": 11, "district": "almbk", "year": 2008}, {"case_count": 84, "chapter": 11, "district": "alnbk", "year": 2008}, … ]} ``` So: document-level scalars + one or more arrays of row objects mirroring tables, median GT ~640 KB. No per-page metadata, no categories/difficulty field (only the folder-slug prefixes). ### Scoring / official eval code `github.com/micro1-research/longextract-bench` (MIT), runnable CLI, dataset downloads on first run via `src/longextract_bench/dataset.py`. Metrics: - **Precision / Recall over array rows**, matched to GT rows by a **key the grader infers per array** (not by position) — precision penalises hallucinated/duplicated/extra rows, recall penalises misses. - **Leaf accuracy** = fraction of scalar leaf values exactly matching. - **Completion is a first-class result**: accuracy is computed only over completed documents and always reported next to a completion count, with failure rate + reasons and latency tracked separately — explicitly to stop systems from looking good by silently dropping hard docs. Deterministic, no LLM judge. ### Conversion into the PuffinParse manifest **Not convertible to `truth/*.md`.** The ground truth is a nested JSON extraction of selected fields, not a transcription — most of the 202-page PDF's text is deliberately *not* in the GT. Realistic uses: - **A third manifest kind (`extract`)**: `{id, file, schema, truth_json, pages, category, tags}` with a scorer implementing row-keyed precision/recall + leaf accuracy. That is a well-specified, deterministic, LLM-judge-free metric — a good fit for PuffinParse's principle #2, and it exercises the `extract`-style endpoints our providers expose (Reducto Deep Extract, Extend, LlamaExtract) rather than the parse endpoints. - **A latency/robustness corpus for the existing parse benchmark**: 50 PDFs of 1–200+ pages is exactly the "ms per page, p95, failure rate" stress we currently lack (synthetic-v1 tops out at `multipage`). We'd run parse on them and report latency/cost/failure only — **accuracy would have to be omitted**, since there is no text truth. - **Lost:** if you tried to synthesise markdown truth from `ground_truth.json` you'd get a tiny table fragment vs. a 202-page document; CER would be meaningless. - PDFs stay **download-on-demand** (374.7 MB); labels (schema + GT, a few MB) are CC-BY-4.0 and could be vendored. --- ## Recommendations 1. **ParseBench `table` split is the one drop-in win** — 503 single-page PDFs + HTML table truth, Apache-2.0, redistributable. Write an `adapters/parsebench_table.py` that pulls `table.jsonl` + the 503 `docs/table/*.pdf` (~130 MB of the 517 MB) and emits `manifest.json` + `truth/*.md`. Flag in the dataset README that colspan/rowspan is flattened. 2. **Do not vendor any PDFs from RealDoc-Bench or LongExtractBench.** Both explicitly carve the source documents out of their CC-BY-4.0 annotation licence. A `puffinparse bench fetch` that snapshot_downloads from HF and verifies `sha256` (RealDoc-Bench ships them; LongExtractBench does not, so we'd record our own) keeps us clean and matches the manifest's existing "downloaded by the user and converted" plan. 3. **Two new manifest kinds would unlock the rest**: `extract` (JSON Schema + expected JSON, deterministic row-keyed P/R + leaf accuracy — LongExtractBench, and LlamaIndex's companion ExtractBench) and `qa` (question + typed gold dict, needs a reader model — RealDoc-Bench). Only the `extract` kind preserves PuffinParse's "no LLM judge" principle. 4. **Provenance caveats to carry through** into any combined dataset README: RealDoc-Bench is ~50% `ai_generated_edit` and 471/581 documents have `not_established` rights; LongExtractBench was commissioned by the vendor that won it and its labels are model-drafted; ParseBench annotations are frontier-VLM auto-labelled with targeted human correction. All three are vendor-published and each vendor leads its own leaderboard. ## Reproduce the downloads ```bash cd "$SCRATCH"/benchres export REQUESTS_CA_BUNDLE=/root/.ccr/ca-bundle.crt SSL_CERT_FILE=/root/.ccr/ca-bundle.crt pip install huggingface_hub python - <<'PY' from huggingface_hub import snapshot_download as d d("llamaindex/ParseBench", repo_type="dataset", local_dir="parsebench", allow_patterns=["README.md","eval.yaml","chart.jsonl","table.jsonl","text_formatting.jsonl", "docs/chart/(Web_version)_E-Government_Survey_2024_1392024_p101.pdf"]) d("Extend-AI/RealDoc-Bench", repo_type="dataset", local_dir="realdocbench", allow_patterns=["README.md","qa_bank.json","manifest.json","docs/finance_1.pdf","docs/mortgage_1.pdf"]) d("Extend-AI/RealDoc-Bench-Layout", repo_type="dataset", local_dir="realdocbench_layout", allow_patterns=["README.md","manifest.csv","annotations/0008ff17-5fec-43a0-9889-2a9e664700f7.json", "images/0008ff17-5fec-43a0-9889-2a9e664700f7.*"]) d("micro1-inc/longextract-bench-50", repo_type="dataset", local_dir="longextractbench", allow_patterns=["README.md","06_19_Bankruptcy_Filings_Statistics/*"]) PY # 55 MB text_content.jsonl sampled without a full fetch: curl -sSL --cacert /root/.ccr/ca-bundle.crt -H "Range: bytes=0-300000" \ https://huggingface.co/datasets/llamaindex/ParseBench/resolve/main/text_content.jsonl \ -o parsebench/tc_head.jsonl ``` Total downloaded here: **12.5 MB** across the four repos (full corpora would be 592 + 540 + 374 + 375 MB ≈ **1.88 GB**). ## Sources - [llamaindex/ParseBench (HF)](https://huggingface.co/datasets/llamaindex/ParseBench) - [run-llama/ParseBench (GitHub)](https://github.com/run-llama/ParseBench) - [ParseBench paper, arXiv:2604.08538](https://arxiv.org/abs/2604.08538) - [Extend-AI/RealDoc-Bench (HF)](https://huggingface.co/datasets/Extend-AI/RealDoc-Bench) - [Extend-AI/RealDoc-Bench-Layout (HF)](https://huggingface.co/datasets/Extend-AI/RealDoc-Bench-Layout) - [extend-hq/realdoc-bench (GitHub)](https://github.com/extend-hq/realdoc-bench) - [RealDocBench paper, arXiv:2606.07401](https://arxiv.org/html/2606.07401v1) - [Parse 2.0 and RealDoc-Bench launch (Extend blog)](https://www.extend.ai/resources/parse-2-and-realdocbench-launch) - [micro1-inc/longextract-bench-50 (HF)](https://huggingface.co/datasets/micro1-inc/longextract-bench-50) - [micro1-research/longextract-bench (GitHub)](https://github.com/micro1-research/longextract-bench) - [LongExtractionBench (micro1)](https://www.micro1.ai/benchmark/long-extraction) - [Reducto tops LongExtractBench (PRNewswire)](https://www.prnewswire.com/news-releases/reducto-deep-extract-ranks-first-overall-in-longextractbench-an-independent-benchmark-for-complex-document-extraction-302815264.html) --- # Academic benchmarks <!-- source: docs/benchmarks/academic-benchmarks.md | url: https://puffinparse.com/docs/benchmark/academic-benchmarks/ --> Researched 2026-09-24. Companion to the [vendor benchmark review](/docs/benchmark/vendor-benchmarks/index.md), which covers ParseBench, RealDoc-Bench and LongExtractBench. Every licence below was read from the dataset card on Hugging Face (`https://huggingface.co/api/datasets/<id>` plus the README at the pinned revision) or from the repository's `LICENSE` file. Revisions are the commits that were current when this was written. **olmOCR-bench** and **OmniDocBench** now have adapters (see [adapters.md](/docs/benchmark/adapters/index.md)). For the others this page recommends what to do next. PuffinParse scores two document kinds (SPEC §10): `transcript` (reference markdown, scored with char/word edit distance, reading order and `table_score`) and `rules` (machine-checkable `present` / `absent` / `order` / `table_cell` / `bag_of_sentences` assertions). "Maps to" below means into one of those two kinds. ## Summary | Benchmark | Measures | Size | Licence (verified) | Redistributable in PuffinParse (MIT)? | GT format | Maps to | Recommendation | |---|---|---|---|---|---|---|---| | **olmOCR-bench** (AI2) | PDF → markdown unit tests | 1,403 PDFs, 7,010 tests | ODC-BY-1.0 (card) | **Yes**, with attribution | JSONL unit tests | `rules`: 3,040 of 7,019 tests (43%) | **Adapted**: 40-doc subset committed | | **OmniDocBench** (OpenDataLab) | end-to-end page parsing: text, tables, formulas, reading order | 1,651 pages, 10 doc types, EN + ZH | none on the card; "research purposes only and not for commercial use" | **No**, index only | one JSON: blocks + order + text/LaTeX/HTML | `transcript` (1,443 of 1,651 pages) | **Adapted**: 40-page index, fetched at run time | | **DP-Bench** (Upstage) | element serialisation (NID), table structure (TEDS / TEDS-S) | 200 single-page PDFs, 36 MB | MIT (card) | Yes (see the caveat below) | `reference.json`: elements with category, coordinates, text/html/markdown | `transcript` | **Adapted** (`benchmark/datasets/dpbench`, in `combined-v3`) | | **READoc** (ISCAS) | realistic PDF → markdown on whole documents | 2,233 docs (arXiv + GitHub) | MIT (card) | Ground truth yes; PDFs doubtful | one markdown file per document | `transcript`, multi-page | Adapter for a long-document track; fetch PDFs at run time | | **Nanonets IDP leaderboard** | aggregator: olmOCR-bench, OmniDocBench, "IDP Core" (KIE, VQA, OCR, tables, classification) | 6,406 IDP Core samples plus the two above | harness MIT; datasets mixed | Per dataset | per-dataset | — | No adapter of its own; its two public page benchmarks are covered above | | **Fox** (UCAS) | fine-grained, region/line/colour-focused page OCR, EN + ZH | 1 zip (`focus_benchmark_test.zip`) | CC-BY-NC-SA-4.0 (card) | **No** (non-commercial, share-alike) | prompt → text pairs | partly `transcript` | Skip: focus-prompted task, NC licence | | **CC-OCR** (Alibaba / Qwen) | LMM literacy: scene text, multilingual, doc parsing, KIE | 39 subsets, 7,058 images | card says MIT; README says only "the source code is licensed under MIT" | Unclear for the images | TSV with base64 images | `doc_parsing` track → `transcript` | Maybe later, fetched at run time | | **OCRBench v2** | LMM OCR QA across 31 scenarios | 10,000 QA pairs, 1.1 GB | MIT (card); images drawn from many source datasets | Unclear for the images | QA pairs + eval type | — (QA, not page transcription) | Skip | Pinned revisions: | Benchmark | Data | Revision | Scorer / code | Revision | |---|---|---|---|---| | olmOCR-bench | HF `allenai/olmOCR-bench` | `54a96a6fb6a2bd3b297e59869491db4d3625b711` | `github.com/allenai/olmocr` (Apache-2.0), `olmocr/bench/tests.py` | `f7cfe4c22098b154c76b6ec950d1c0a464eecf8d` | | OmniDocBench | HF `opendatalab/OmniDocBench` | `aa1ee96d106dbe53d0ae59474d75c6e6d9b53fec` | `github.com/opendatalab/OmniDocBench` (Apache-2.0) | `f133a71e9e91c3621c7ce8994200a7b394a06eb3` | | DP-Bench | HF `upstage/dp-bench` (data and `evaluate.py` together) | `24702c61a2fb13325534be664653bc6e60250d13` | same repo | same | | READoc | HF `lazyc/READoc` | `782bf954e4d9a31ae2bcfe14ff59f1f4b2467592` | `github.com/icip-cas/READoc` | not pinned (not adapted) | | Fox | HF `ucaslcl/Fox_benchmark_data` | `d6c6f5202d61bf28c8cf9b7b777754dc9bbdc0e0` | `github.com/ucaslcl/Fox` | not pinned | | CC-OCR | HF `wulipc/CC-OCR` | `c64517e92179991d509776064174776700cdd5a2` | `github.com/AlibabaResearch/AdvancedLiterateMachinery` (`Benchmarks/CC-OCR`) | not pinned | | OCRBench v2 | HF `ling99/OCRBench_v2` | `c7e7cdf23bdb6774661e9b0caf0d9935a42feb8b` | `github.com/Yuliang-Liu/MultimodalOCR` | not pinned | | IDP leaderboard | — | — | `github.com/NanoNets/idp-leaderboard-benchmarks` (MIT) | not pinned | --- ## 1. olmOCR-bench (Allen Institute for AI): adapted - Paper: arXiv:2502.18443 ("olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models"). Code: `allenai/olmocr` (Apache-2.0). Data: `allenai/olmOCR-bench`. **What it measures.** Whether a PDF → markdown converter gets specific, checkable facts right. It does not compare the output against a reference transcript. Each test is a unit test: this sentence is present, this running header is absent, this paragraph precedes that one, this table cell sits under that heading, this equation renders the same as the reference. **Size.** 1,403 single-page PDFs and 7,010 tests in seven JSONL files, one per source. The runner also adds one implicit `baseline` test per PDF. There are 9 explicit baseline tests, so the files hold 7,019 rows. | Split | PDFs | Tests | Test types | |---|---:|---:|---| | `arxiv_math` | 522 | 2,927 | math | | `old_scans_math` | 36 | 458 | math | | `table_tests` | 188 | 1,020 (+2 baseline) | table | | `old_scans` | 98 | 526 | present 279, absent 70, order 177 | | `headers_footers` | 266 | 753 (+7 baseline) | absent | | `multi_column` | 231 | 884 | order | | `long_tiny_text` | 62 | 442 | present | **Licence.** The card front-matter says `license: odc-by`, and the README says the dataset is "licensed under ODC-BY-1.0 … intended for research and educational use in accordance with AI2's Responsible Use Guidelines". ODC-BY allows redistribution with attribution. The underlying pages are third-party (arXiv, Library of Congress, Internet Archive, crawled PDFs), and every test carries the page's origin `url`. The adapter keeps that URL as `source_url` on each document. **Ground truth and scorer.** One JSON object per test: `{pdf, page, id, type, max_diffs, checked, url, …}` plus type-specific keys: `text`, `case_sensitive`, `first_n`, `last_n` (present/absent); `before`, `after` (order); `cell`, `up`, `down`, `left`, `right`, `top_heading`, `left_heading` (table); `math` (LaTeX). Semantics, read from `olmocr/bench/tests.py`: - Both sides go through `normalize_text`: whitespace collapsed, `**`/`__`/`*`/`_` emphasis stripped, NFC, and typographic quotes and dashes folded to ASCII. - `present` / `absent` use `rapidfuzz.partial_ratio` with threshold `1 - max_diffs/len(text)`. `first_n` / `last_n` restrict the search to the start or end of the output. `present` is case-sensitive by default. - `order` uses `fuzzysearch.find_near_matches` with `max_l_dist = max_diffs` and passes if any `before` match starts before any `after` match. It is case-sensitive. - `table` parses markdown and HTML tables, finds a cell by `fuzz.ratio`, then checks the immediate `up`/`down`/`left`/`right` neighbour or the column/row heading. - `math` renders both equations with KaTeX and compares the rendered symbol layout. - `baseline` fails on empty output, long repeated n-grams, or CJK and emoji characters. The leaderboard averages pass rates per split, not per document. **Mapping to PuffinParse.** Every test is already a rule, so documents are `kind: rules`: | Upstream | → PuffinParse | Converted | Skipped | Fidelity | |---|---|---:|---:|---| | `present` | `present` | 721 | 0 | exact substring, stricter than fuzzy when `max_diffs > 0` | | `absent` (no `first_n`/`last_n`) | `absent` | 622 | — | exact, looser than fuzzy when `max_diffs > 0` | | `absent` with `first_n`/`last_n` | — | 0 | 201 | a whole-page absence would fail a page number that also appears in the body | | `order` | `order` (case-sensitive) | 1,061 | 0 | exact; 94% of the `multi_column` tests have `max_diffs > 0`, so this is stricter than upstream | | `table` with `top_heading` / `left_heading` / no relation | `table_cell` with `col_header` / `row_header` | 309 | 0 | exact cell equality (upstream: `fuzz.ratio ≥ max(0.5, …)`) | | `table` with `left` / `right` | `table_cell` with `row_header` = neighbour | 327 | 0 | **relaxed**: "immediately left/right of" becomes "in the same row as" | | `table` with `up` / `down` | — | 0 | 384 | the schema has no column-adjacency relation, and "value exists" would inflate scores | | `math` | — | 0 | 3,385 | KaTeX-rendered equivalence has no text-rule analogue | | `baseline` | — | 0 | 9 (+1,394 implicit) | repetition and charset heuristic | | **Total** | | **3,040** | **3,979** | 824 of 1,403 PDFs keep at least one rule | `max_diffs` is kept on every converted rule, so a fuzzy scorer can honour it later. **Recommendation.** Done. See [`benchmark/datasets/olmocr/README.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/olmocr/README.md). Two scorer extensions would recover most of what is skipped or approximated: fuzzy matching driven by `max_diffs`, and `up`/`down`/`left`/`right` neighbour fields on `table_cell`. Math is deliberately out of scope until PuffinParse has a formula metric. ## 2. OmniDocBench (OpenDataLab / Shanghai AI Laboratory): adapted as an index - Paper: arXiv:2412.07626. Code: `opendatalab/OmniDocBench` (Apache-2.0). Data: `opendatalab/OmniDocBench`, not gated. **What it measures.** End-to-end parsing of diverse real pages: text (normalised edit distance), tables (TEDS), display formulas (CDM and edit distance) and reading order (edit distance over block order). It also has layout-detection and single-module tracks. **Size.** 1,651 page images (PNG/JPG, over 1 GB; the first 1,000 alone are 1,019 MB) plus `OmniDocBench.json` (42 MB), which includes a 296-page hard subset added 2026-04-09. Page attributes: - `data_source` (10): book 276, PPT2PDF 253, academic_literature 215, exam_paper 193, colorful_textbook 159, newspaper 151, magazine 149, research_report 132, note 118, historical_document 5. - `language`: simplified_chinese 765, english 755, en_ch_mixed 116, traditional_chinese 13, other 2. - `layout`: single / double / three column, mixed, other. - `special_issue`: watermark, fuzzy_scan, colorful_background, table flags, and so on. - `subset`: v1.5, table_hard, layout_hard, equation_hard. **Licence.** The card has **no licence field** (`cardData` is null in the API). The only terms are the Copyright Statement: *"The PDFs are collected from public online channels and community user contributions. Content that is not allowed for distribution has been removed. The dataset is for research purposes only and not for commercial use."* That grants no redistribution right and forbids commercial use, so PuffinParse, an MIT repository, **does not vendor any of it**: not the images, and not the truth derived from the annotations. The evaluation code is Apache-2.0. **Ground truth.** Per page: `page_info` (image path, size, attributes) and `layout_dets`. Each block has a `category_type` (28 block classes), a polygon, a reading `order`, and `text` / `latex` / `html`. `extra.relation` holds `truncated` links, which join paragraphs split across columns, and `parent_son` links, which attach captions. Upstream's `tools/json2md.py` shows how to turn this into markdown. **Mapping to PuffinParse.** One `transcript` document per page. Blocks are emitted in `order`, with truncated chains merged. Titles become `#` headings, tables become pipe tables converted from the HTML (merged cells flattened, tagged `merged-cells`), display formulas stay as `$$…$$`, and headers, footers, page numbers, page footnotes, `abandon` regions and figures are dropped. Those are the page furniture OmniDocBench itself does not score. Pages with `*_mask` regions (126) are skipped because their content is deliberately unannotated. So are pages whose truth is under 120 characters (82). That leaves **1,443 of 1,651 pages** convertible. `data_source` is the `category`. Language, layout, subset, special issues and `has-table` / `has-formula` are tags. What is lost: TEDS table structure (PuffinParse's `table_score` compares cell text row by row), CDM formula matching (LaTeX is compared as text), and OmniDocBench's block-level matching, which forgives reading-order differences. PuffinParse's `char_similarity` over the whole page does not. **Recommendation.** Done as a fetch-at-run-time index. See [`benchmark/datasets/omnidocbench/README.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/omnidocbench/README.md). If OpenDataLab ever publishes an explicit licence that permits redistribution, the same adapter can commit the files by deleting its `.gitignore`. ## 3. DP-Bench (Upstage): adapted - Data, inference scripts and `evaluate.py` all live in HF `upstage/dp-bench`. There is no paper. **What it measures.** Two things: - **NID** (normalised indel distance: edit distance without substitutions) over the text elements serialised in reading order. Tables, figures and charts are excluded. - **TEDS / TEDS-S** over the 55 tables. **Size.** 200 single-page PDFs (36 MB, largest 4.3 MB): 90 from the Library of Congress, 90 from Open Educational Resources, 20 from Upstage internal documents. The documents contain 1,822 layout elements in 12 classes: Paragraph 804, Heading1 194, Footer 168, Caption 154, Header 101, List 91, Chart 67, Footnote 63, Equation 58, Figure 57, Table 55, Index 10. **Licence.** The card front-matter says `license: mit`. The Library of Congress pages are largely public domain and OER pages are openly licensed. The 20 Upstage internal pages are covered only by the MIT declaration, so record that in the dataset README. **Ground truth.** `dataset/reference.json` (1.5 MB) is `{<pdf>: {elements: [{category, coordinates, id, page, content: {text, html, markdown}}]}}`. Tables are HTML, equations LaTeX, and `id` is the reading order. **Mapping.** A `transcript` per page: elements in `id` order, tables as pipe tables via `html_table_to_markdown`, and headers and footers either dropped (NID excludes nothing but tables, figures and charts, so they could stay) or tagged. This is structurally the same as the OmniDocBench adapter, so it is roughly a day's work, and the whole dataset (36 MB) is small enough that a 40–60 page subset commits in about 5 MB. **Status.** Adapted by `benchmark/adapters/dpbench.py`: headers and footers are kept (NID scores them), figures and charts dropped, 40 pages committed (4.5 MB), folded into `combined-v3`. See [`adapters.md`](/docs/benchmark/adapters/index.md#dp-bench-mapping). ## 4. READoc (Institute of Software, CAS) - Paper: arXiv:2409.05137. Code: `icip-cas/READoc`. Data: `lazyc/READoc`. **What it measures.** Realistic document structured extraction: a *whole* PDF (arXiv papers, GitHub READMEs rendered to PDF) converted to one markdown file. It scores text, headings, tables, formulas and reading order after a standardisation and segmentation step (the "DSE Evaluation Suite"). **Size.** 2,233 documents: `arxiv_ground_truth/` has 1,009 markdown files and `github_ground_truth/` has 1,224. The PDFs come as `arxiv.zip` (1.19 GB), `github.zip` (0.59 GB) and `zenodo.zip` (1.45 GB). **Licence.** The card says `license: mit`. The arXiv PDFs keep their authors' licences, most often arXiv's non-exclusive distribution licence, which does not let third parties redistribute. Treat the markdown truth as MIT and the PDFs as fetch-only. **Ground truth.** One markdown file per document, e.g. `arxiv_ground_truth/0705.4297.md` (91 KB). **Mapping.** `transcript`, multi-page, exactly PuffinParse's existing kind, but with documents of 10 to 40 pages. That makes it a latency and cost stress test as much as an accuracy one. **Recommendation.** Worth an adapter as a separate long-document track (`readoc-arxiv`, about 30 documents). Ship the index plus truth and fetch the PDFs from the zips at run time. Selective extraction needs HTTP range reads of the zip central directory; downloading 1.2 GB for 30 documents is wasteful. Keep it out of `combined-*` because its per-document cost dwarfs the single-page sets. ## 5. Nanonets IDP leaderboard - Site `idp-leaderboard.org`. Harness `NanoNets/idp-leaderboard-benchmarks` (MIT). Toolkit `NanoNets/docext`. **What it measures.** It is an aggregator rather than a dataset. It runs **olmOCR-bench** (from `allenai/olmOCR-bench`), **OmniDocBench** (from `opendatalab/OmniDocBench`) and its own "IDP Core": key-information extraction, VQA, OCR, document classification, long-document processing, table extraction and confidence calibration. IDP Core has 6,406 samples drawn from existing datasets (Nanonets KIE, DocILE, handwritten forms, ChartQA, DocVQA, OCR handwriting and diacritics sets, table-extraction sets), loaded through `docext` from Hugging Face. **Licence.** The harness is MIT. The IDP Core datasets keep their own terms, and several need an agreement or are research-only (DocILE, DocVQA). **Not verified dataset by dataset here.** **Recommendation.** No adapter of its own. Its two page-parsing benchmarks are now PuffinParse adapters, so PuffinParse can report comparable per-benchmark numbers directly. The KIE and table tracks belong with a future `extract`-mode benchmark and would need a per-dataset licence audit. ## 6. Fox (UCAS / MEGVII) - Data `ucaslcl/Fox_benchmark_data`, a single `focus_benchmark_test.zip`. **What it measures.** Fine-grained, focus-prompted document understanding in English and Chinese: OCR of a region given a box, a line, or a colour highlight, plus page-level OCR and multi-page variants. Scoring uses edit distance, F1, BLEU and METEOR. **Licence.** CC-BY-NC-SA-4.0 (card and tags). The non-commercial and share-alike terms are incompatible with vendoring into an MIT repository. **Recommendation.** Skip. The task is prompt-conditioned, and PuffinParse's providers expose no "focus this box" input. Only the plain page-OCR slice would map to `transcript`, and the licence keeps it fetch-only. ## 7. CC-OCR (Alibaba / Qwen team) - Paper: arXiv:2412.02210. Data `wulipc/CC-OCR`, the TSV version used by VLMEvalKit. **What it measures.** Literacy of large multimodal models across four tracks: multi-scene text reading, multilingual text reading, **document parsing** (documents, tables, formulas; scored with normalised edit distance and TEDS), and key-information extraction. It has 39 subsets and 7,058 images, of which 41% come from real applications. **Licence.** The card front-matter says `mit`, but the README's licence section says only that "the source code is licensed under the MIT License". The images' own terms are not stated. **Mapping.** The `doc_parsing` subsets give an image and a reference markdown, HTML or LaTeX string, which maps to `transcript`, with tables via HTML → pipe table. **Recommendation.** Possible later, as a fetch-at-run-time index like OmniDocBench, until the data licence is clarified. Low priority: the parsing track is small and mostly overlaps OmniDocBench. ## 8. OCRBench v2 - Paper: arXiv:2501.00321. Data `ling99/OCRBench_v2` (parquet, 1.1 GB, 10,000 QA pairs). The first OCRBench (`echo840/OCRBench`) carries no licence on its card. **What it measures.** LMM OCR ability as question answering over 31 scenarios and 23 tasks: text recognition, referring, spotting, key-information extraction, parsing, reasoning. Scoring is task-specific (exact match, ANLS, TEDS, IoU, and so on). **Licence.** The card says `license: mit`. The images are collected from many public datasets that keep their own terms. **Recommendation.** Skip. It is QA over images, not page transcription, and it conflicts with the "score `parse` output" design exactly as RealDoc-Bench's QA track does. --- ## What changes in PuffinParse because of this survey 1. `benchmark/adapters/olmocr.py` and `benchmark/adapters/omnidocbench.py` exist, and `combined-v2` includes both. `combined-v1` is unchanged because it has committed results. 2. Scorer work that would raise fidelity, in priority order: - `max_diffs` fuzzy matching for `present`, `absent`, `order` and `table_cell`. The field is already in the rule files. - Neighbour relations (`up`, `down`, `left`, `right`) on `table_cell`. - HTML tables in predictions (in progress elsewhere). - A structural table metric (TEDS) next to `table_score`. 3. Next adapter: READoc as a long-document track (DP-Bench is done: `combined-v3`). --- # Dataset adapters <!-- source: docs/benchmarks/adapters.md | url: https://puffinparse.com/docs/benchmark/adapters/ --> How public OCR benchmarks become PuffinParse datasets. [ADR-10](/docs/project/decisions/index.md) says PuffinParse does not author a competing benchmark: it runs every public benchmark through one harness. An **adapter** is the piece that makes that true — it fetches one upstream benchmark at a pinned revision and rewrites it into `benchmark/datasets/<name>/manifest.json`. Datasets whose license permits redistribution are vendored into the repository; the rest stay an index plus a `sha256`, fetched on demand. Everything lives in [`benchmark/adapters/`](https://github.com/ajinkyashejul/puffinparse/tree/main/benchmark/adapters): | File | What it is | |---|---| | `base.py` | the framework: `Adapter`, the manifest model, the rule schema, shared helpers | | `parsebench.py` | LlamaIndex ParseBench → `benchmark/datasets/parsebench/` | | `olmocr.py` | AI2 olmOCR-bench → `benchmark/datasets/olmocr/` (`kind: rules`) | | `omnidocbench.py` | OpenDataLab OmniDocBench → `benchmark/datasets/omnidocbench/` (`kind: transcript`, index only, fetched at run time) | | `dpbench.py` | Upstage DP-Bench → `benchmark/datasets/dpbench/` (`kind: transcript`, 40 PDFs committed) | | `combined.py` | union of the datasets → `benchmark/datasets/combined-v1/` (adapter `combined`, frozen), `benchmark/datasets/combined-v2/` (adapter `combined-v2`) and `benchmark/datasets/combined-v3/` (adapter `combined-v3`: v2 + dpbench) | | `__main__.py` | the `python -m benchmark.adapters` CLI | ## CLI ```bash python -m benchmark.adapters --list python -m benchmark.adapters <name> [--limit N] [--out DIR] [--cache DIR] [--seed N] [--no-download] ``` | Flag | Meaning | |---|---| | `--limit N` | cap the number of documents built (selection stays deterministic) | | `--out DIR` | output dataset directory (default: the adapter's `default_out`) | | `--cache DIR` | download cache (default: `$PUFFINPARSE_BENCH_CACHE`, else `$XDG_CACHE_HOME/puffinparse/benchmarks`, else `~/.cache/puffinparse/benchmarks`) — never inside the repo | | `--seed N` | selection seed, default `1234` | | `--no-download` | build from an existing cache, never hit the network | Downloads go over HTTPS through whatever proxy the environment configures. Behind the agent proxy, point `huggingface_hub` at the CA bundle (the adapter sets these if they are unset) — never disable TLS verification: ```bash export REQUESTS_CA_BUNDLE=/root/.ccr/ca-bundle.crt SSL_CERT_FILE=/root/.ccr/ca-bundle.crt ``` The CLI prints a build summary: documents by kind, committed size, and the adapter's own counters (rules converted, rules skipped by upstream type, documents skipped and why). ## Manifest extension The manifest stays exactly what [`benchmark/README.md`](/docs/benchmark/index.md) and [`docs/SPEC.md`](/docs/project/spec/index.md) §10.2 describe: ```json {"name", "version", "description", "license", "documents": [{"id", "file", "truth", "pages", "category", "tags"}]} ``` Adapters add optional keys. **Every added key is additive and ignored by the current Rust `Manifest` / `ManifestDoc`** (serde ignores unknown fields by default), so old manifests keep working and new manifests load in an unmodified CLI. ### Document-level additions | Key | Type | Meaning | |---|---|---| | `kind` | `"transcript"` \| `"rules"` | how the document is scored. **Default `"transcript"`** — absent means transcript, which is what every pre-existing manifest is | | `rules` | path | for `kind: "rules"`, the assertion file, relative to the dataset directory | | `source_id` | string | the upstream id, kept verbatim when `id` had to be slugified | | `upstream_path` | string | the document's path inside the upstream repo, for on-demand fetch | | `sha256` | hex | the document file's hash | | `license` | SPDX | per-document license, when a dataset mixes them | | `attribution` | string | the citation this document must carry | | `source_url` | string | where the page originally came from (olmOCR-bench records one URL per test) | | `truth_sha256` | hex | hash of a truth file that is generated locally and not committed (OmniDocBench) | `truth` is **always emitted**, even for `kind: "rules"` documents, where it is the empty string. That is not cosmetic: the Rust `ManifestDoc` declares `pub truth: String` without `#[serde(default)]`, so a missing key fails to deserialise the *whole* manifest. ### Manifest-level additions `generator`, `upstream` (`{repo_id, revision, url, repo_type}`), `attribution`, `notes`, and — for combined datasets — `sources`: ```json "sources": [{"name", "version", "license", "path", "documents", "manifest_sha256", "attribution"}] ``` ### `kind: "transcript"` The existing behaviour: `truth` is markdown, scored by `puffinparse_core::bench` with `char_similarity` / `cer` / `wer` / `word_f1` / `order_score` / `table_score`. ### `kind: "rules"` `truth` is empty and `rules` points at a JSON **list** of machine-checkable assertions. One schema covers every rule-based upstream benchmark we care about (ParseBench today, olmOCR-bench next): ```json { "id": "text_dense_baoutou_order_623", "type": "present" | "absent" | "order" | "table_cell" | "bag_of_sentences", "text": "…", // present / absent "before": "…", "after": "…", // order "cell": {"row_header": "…", "col_header": "…", "value": "…"}, // table_cell "sentences": ["…"], // bag_of_sentences "threshold": 0.8, // bag_of_sentences: required pass fraction "case_sensitive": false, "max_diffs": 0, // optional: upstream fuzzy allowance (olmOCR-bench) "source": "text_dense__baoutou_order_623" } ``` | Type | Passes when the parsed markdown… | |---|---| | `present` | contains `text` (after the run's normalisation) | | `absent` | does **not** contain `text` | | `order` | contains both `before` and `after`, and the first occurrence of `before` precedes the first occurrence of `after` (with `max_diffs > 0`: some fuzzy match of `before` starts before some fuzzy match of `after`, as olmOCR does) | | `table_cell` | has a table (markdown pipe table or HTML `<table>`, spans repeated into every slot) with a row matching `cell.row_header` and a column matching `cell.col_header` whose cell equals `cell.value` | | `bag_of_sentences` | contains at least `threshold` (fraction, default `0.8`) of `sentences`; a sentence counts when some window of the output is ≥ 0.8 similar to it (`1 − edit distance / length`) | Matching normalisation (scorer v2): the run's [`normalize`](/docs/benchmark/index.md#metrics), then every space adjacent to punctuation is dropped on **both** sides, so a tokenised rule such as `(this " agreement ")` matches the printed `(this "Agreement")` and vice versa. Words still need their spaces. `case_sensitive` is always present and always explicit. `source` is the upstream rule id, so any score can be pushed back to the publisher's own harness for cross-checking. `max_diffs` is optional and records how many Levenshtein edits the upstream scorer tolerates. Since scorer v2 the Rust scorer honours it: `present` / `absent` use a fuzzy substring search (Sellers' algorithm, like upstream's `fuzzysearch.find_near_matches(max_l_dist=max_diffs)`), `order` uses the fuzzy rule above, and `table_cell` tolerates `max_diffs` edits in the header matches and the value. `0` or absent means exact matching after normalisation. A document's rule score is `passed / total`; a dataset's is the mean over its rule documents. That is deliberately the same shape as an accuracy in `[0, 1]`, so it slots next to `char_similarity` in the existing `Summary`. ## ParseBench mapping Upstream: [`llamaindex/ParseBench`](https://huggingface.co/datasets/llamaindex/ParseBench), pinned to commit `2805a1d940f95a203e0ae4b88be9934f7765b3fc`, Apache-2.0, 2,078 single-page documents and 169,011 assertions in five JSONL files. ### Rule conversion, whole upstream dataset | Upstream file | Upstream type | Rules | → PuffinParse | Converted | Skipped | |---|---|---:|---|---:|---:| | `table.jsonl` | `expected_markdown` | 503 | `kind: transcript`, HTML table → markdown | 503 | 0 | | `text_content.jsonl` | `missing_specific_word` | 105,369 | `present` (`rule.word`) | 105,369 | 0 | | | `missing_specific_sentence` | 18,768 | `present` (`rule.sentence`) | 18,768 | 0 | | | `order` | 13,087 | `order` (`rule.before` / `rule.after`) | 13,087 | 0 | | | `missing_sentence_percent` | 503 | `bag_of_sentences` (`rule.bag_of_sentence` keys, `threshold: 0.8`) | 503 | 0 | | | `unexpected_sentence_percent` | 503 | — | 0 | 503 | | | `too_many_sentence_occurence_percent` | 503 | — | 0 | 503 | | | `missing_word_percent` | 506 | — | 0 | 506 | | | `unexpected_word_percent` | 506 | — | 0 | 506 | | | `too_many_word_occurence_percent` | 506 | — | 0 | 506 | | | `bag_of_digit_percent` | 486 | — | 0 | 486 | | | `is_header` | 278 | — | 0 | 278 | | | `is_footer` | 307 | — | 0 | 307 | | `text_formatting.jsonl` | `is_title`, `is_bold`, `is_italic`, `is_sup`, `is_sub`, `is_underline`, `is_strikeout`, `is_mark`, `is_latex`, `is_code_block`, `title_hierarchy_percent` | 5,997 | — | 0 | 5,997 | | `chart.jsonl` | `chart_data_point` | 4,864 | — | 0 | 4,864 | | `layout.jsonl` | element bounding boxes | 16,325 | — | 0 | 16,325 | | **Total** | | **169,011** | | **138,230** | **30,781** | Why the skips: - **`unexpected_*` / `too_many_*` / `*_percent` counters** are precision-direction assertions: "the output must not contain sentences the page does not have". Checking one needs the *complete* reference text, and ParseBench never publishes it for the text split — only bags. A `bag_of_*` compared against an incomplete reference would punish correct output. - **`bag_of_digit_percent`** is a digit histogram, which no `present`/`absent` assertion expresses. - **`is_header` / `is_footer`** assert a string's *role* on the page, not its presence. - **`text_formatting.jsonl`** asserts style (bold, italic, superscript, LaTeX, heading level). Markdown carries some of this, but our normalisation strips it before scoring, so a rule about emphasis would be unscoreable. A future `formatting` kind could pick these up. - **`chart.jsonl`** asserts a value read off a chart; **`layout.jsonl`** asserts bounding boxes with IoA matching. Neither has a text-similarity or substring analogue. Net: **138,230 of 169,011 upstream assertions (81.8%)** are expressible, covering the `table` and `text_content` dimensions — 1,009 of 2,078 pages (48.6%). ### Committed subset The full conversion would vendor 517 MB of PDFs, so only a curated subset is committed (`benchmark/datasets/parsebench/`, **5.7 MB** total, 4.0 MB of PDFs): | | Docs | Rules | Notes | |---|---:|---:|---| | `kind: transcript` (table split) | 25 | — | 18 tagged `merged-cells`, 8 tagged `hard`, all tagged `table-only` | | `kind: rules` (text split) | 15 | 3,703 | `present` 3,303 · `order` 385 · `bag_of_sentences` 15 | Selection is deterministic for a given `--seed`: table pages are taken at most one per source document with roughly a third `hard`; text pages round-robin across all eight ParseBench document-type buckets (`simple`, `ocr`, `multicolumns`, `multilang`, `misc`, `dense`, `sparse`, `handwritting`). Documents over 400 KB, over the 11 MB budget, or with more than one page are skipped and counted. `manifest.full.json` in the same directory indexes all 1,009 convertible documents with their `upstream_path` and is *not* runnable as-is; `subset-committed.txt` lists the 40 committed ids. ### The `table-only` caveat ParseBench's table truth is the page's **table**, not the page. A parser returns the **whole page**. Measured on `apple_10_k_page1` with `reducto/standard`: | Metric | Value | Reading | |---|---:|---| | `table_score` | **0.990** | the real signal — markdown table rows only | | `char_similarity` | 0.311 | meaningless here: the prediction also has the rest of the page | | `cer` / `wer` | 2.212 / 1.938 | same | | `overall` | 31.13 | derived from `char_similarity`, so also meaningless | So: **score `table-only` documents with `table_score`.** Every such document carries the `table-only` tag precisely so a scorer can pick the right primary metric. `merged-cells` marks the 18 documents where `colspan`/`rowspan` had to be flattened by repeating cells — cell content survives, table structure (ParseBench's GriTS) does not. ## Framework reference (`base.py`) ```python class Adapter(ABC): name: str # CLI argument and dataset name license: str # SPDX for the redistributable part upstream: Upstream # repo_id + pinned revision + url default_out: str # repo-relative output directory def download(self, cache_dir: Path) -> Path: ... def build(self, out_dir: Path, limit: int | None = None, seed: int = 1234) -> Manifest: ... ``` Shared helpers: | Helper | What it does | |---|---| | `sha256_file`, `sha256_bytes` | streamed hashing for manifest provenance | | `write_manifest`, `write_rules`, `write_json` | stable 2-space JSON with a trailing newline; returns the SHA-256 written | | `slugify` | ASCII, lowercase, `_`-separated ids safe for filenames and `--filter` | | `html_table_to_markdown` | HTML `<table>` → GitHub pipe table; flattens `colspan`/`rowspan` by repeating the cell into every position it covers, and reports `merged_cells` so the caller can add the `merged-cells` tag. Inline markup is dropped, `<br>` becomes a space, `\|` is escaped (backslashes are kept verbatim so LaTeX in cells survives), short rows are padded | | `pdf_page_count` | `pypdf` when importable, otherwise a byte scan: the `/Count` of the root `/Type /Pages` node (also inside inflated object streams), falling back to *distinct* `/Type /Page` object numbers so an incrementally-updated PDF is not double counted | | `default_cache_dir` | `$PUFFINPARSE_BENCH_CACHE` → `$XDG_CACHE_HOME/puffinparse/benchmarks` → `~/.cache/puffinparse/benchmarks` | `pypdf` is optional on purpose. It pulls in `cryptography`, whose Rust extension can raise a `pyo3` `PanicException` (a `BaseException`, not an `Exception`) on a mismatched wheel, so the import is guarded broadly and cached; a broken optional dependency degrades to the byte scan instead of aborting a build. ## Adding an adapter 1. Create `benchmark/adapters/<name>.py`. 2. Subclass `Adapter`, set `name`, `license`, `default_out`, `description` and `upstream` with a **pinned commit** — not a branch. 3. Implement `download(cache_dir)` (fetch only what conversion needs; keep the big media trees for a per-document fetch) and `build(out_dir, limit, seed)`. 4. Decorate the class with `@register` and import the module from `benchmark/adapters/__init__.py`. 5. Use `Doc` / `Manifest` / `Rule` and `write_manifest` / `write_rules` so output stays byte-stable, and record every skip with `self.bump(...)` so the summary is honest. 6. If the dataset belongs in the combined benchmark, add a new `SOURCES_V<n>` and `CombinedV<n>Adapter` in `combined.py` (as `combined-v3` did for DP-Bench). Never change the sources of an existing combined version: a result is only reproducible against the exact manifest it scored. Write a `README.md` in the dataset directory with the license, and list the dataset in `benchmark/README.md`. 7. Add the dataset to `crates/puffinparse-core/tests/benchmark_datasets.rs`. Every rule must pass against a witness built from its own document, and every transcript must score 1.0 against itself. Not adaptable into either kind, per the [vendor benchmark survey](/docs/benchmark/vendor-benchmarks/index.md): - RealDoc-Bench: QA over a parse. It needs an LLM reader, which conflicts with the deterministic-metrics principle, and its source PDFs are not redistributable. - RealDoc-Bench-Layout: bounding boxes, no text. - LongExtractBench: schema-driven extraction, a different mode. The [academic benchmark survey](/docs/benchmark/academic-benchmarks/index.md) covers olmOCR-bench, OmniDocBench, DP-Bench, READoc, the Nanonets IDP leaderboard, Fox, CC-OCR and OCRBench v2. DP-Bench is now adapted (below); READoc, as a long-document track, is the next candidate. ## olmOCR-bench mapping Upstream: [`allenai/olmOCR-bench`](https://huggingface.co/datasets/allenai/olmOCR-bench) at `54a96a6fb6a2bd3b297e59869491db4d3625b711`, ODC-BY-1.0. It has 1,403 single-page PDFs and 7,010 tests, plus 9 explicit baseline tests. The semantics were read from `olmocr/bench/tests.py` at `f7cfe4c22098b154c76b6ec950d1c0a464eecf8d`. Every document is `kind: rules`. | Upstream test | → PuffinParse | Converted | Skipped | Note | |---|---|---:|---:|---| | `present` | `present` (upstream `case_sensitive`, default true) | 721 | 0 | | | `absent` | `absent` | 622 | 201 | positional (`first_n` / `last_n`) absences are skipped: a whole-page absence would be the wrong assertion | | `order` | `order`, `case_sensitive: true` | 1,061 | 0 | | | `table` | `table_cell` | 636 | 384 | `top_heading` → `col_header`, `left_heading` → `row_header`; `left`/`right` → `row_header` (327, relaxed to "same row"); `up`/`down` skipped | | `math` | — | 0 | 3,385 | KaTeX-rendered equivalence | | `baseline` | — | 0 | 9 | plus 1,394 implicit per-PDF baseline tests the upstream runner adds | | **Total** | | **3,040** | **3,979** | 824 of 1,403 PDFs keep ≥ 1 rule | All rule text has whitespace runs collapsed. Upstream's `normalize_text` does the same on both sides, so nothing is lost, and it matters for table headers annotated across two lines. What is committed: - 40 documents, 8 from each split that has convertible tests: 6.0 MB of PDFs and 205 rules. - `manifest.full.json`, an index of all 824 convertible documents. - `conversion-stats.json`, which records every count. Documents whose rules are all `absent` carry the `absent-only` tag, because an empty parse passes them. Details and caveats: [`benchmark/datasets/olmocr/README.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/olmocr/README.md). ## OmniDocBench mapping Upstream: [`opendatalab/OmniDocBench`](https://huggingface.co/datasets/opendatalab/OmniDocBench) at `aa1ee96d106dbe53d0ae59474d75c6e6d9b53fec`. It has 1,651 page images and one annotation JSON with blocks, reading order and text / LaTeX / HTML. Each page becomes a `kind: transcript` document. The truth is built the way upstream's `tools/json2md.py` builds it: - blocks are emitted in reading order, with `truncated` chains merged; - titles become `#` headings, tables are converted from HTML to pipe tables, and display formulas stay as `$$…$$`; - headers, footers, page numbers, page footnotes, `abandon` regions and figures are dropped, because OmniDocBench does not score them either. Pages with `*_mask` regions (126) or less than 120 characters of truth (82) are skipped, which leaves 1,443 convertible pages. `category` is the upstream `data_source`. Language, layout, subset, special issues, `has-table`, `has-formula` and `merged-cells` are tags. **Licence.** The card has no licence and says "research purposes only and not for commercial use". So only `manifest.json`, with the image `sha256` and `truth_sha256`, is committed. `docs/` and `truth/` are excluded from version control and produced by `python -m benchmark.adapters omnidocbench`. Every document carries the `fetch-required` tag. Details: [`benchmark/datasets/omnidocbench/README.md`](https://github.com/ajinkyashejul/puffinparse/blob/main/benchmark/datasets/omnidocbench/README.md). ## DP-Bench mapping Upstream: [`upstage/dp-bench`](https://huggingface.co/datasets/upstage/dp-bench) at `24702c61a2fb13325534be664653bc6e60250d13` — data and `evaluate.py` pinned together. 200 single-page PDFs and `dataset/reference.json`, which lists every layout element of a page (12 categories) in reading order (`id`), with `content.text` and, for tables, `content.html`. **Licence.** The card front-matter says `license: mit`; the repository has no `LICENSE` file at that revision. The adapter reads the card at the pinned revision and refuses to build if it no longer says MIT. The 20 Upstage-internal pages are covered only by that declaration. **Truth.** Each page becomes one `kind: transcript` document, following what upstream scores: | Upstream category | → truth | Why | |---|---|---| | `Heading1` | `# heading` | | | `Table` | pipe table from `content.html` (merged cells repeated, tag `merged-cells`) | upstream scores tables with TEDS; `table_score` / `teds_grid` do the same job | | `Equation` | LaTeX inside `$$ … $$` (tags `has-equation`, `has-formula`) | | | `Figure`, `Chart` | dropped | upstream's NID ignores `figure`, `table`, `chart` (`--ignore-classes-for-layout`) | | `Paragraph`, `Caption`, `List`, `Footnote`, `Index`, **`Header`, `Footer`** | `content.text` verbatim | NID scores every one of them, so headers and footers stay — unlike OmniDocBench, whose own matching drops them | No `rules` are emitted: a document has one kind, and the table truth inside the transcript is already scored by `table_score` and `teds_grid`. **Selection.** `category` is the page's dominant layout feature, in priority order `table` > `equation` > `chart` > `figure` > `index` > `list` > `text`. Upstream has 42 / 18 / 43 / 38 / 10 / 9 / 40 pages in those buckets; the committed subset takes 10 / 5 / 6 / 5 / 3 / 4 / 7 (40 pages, 4.5 MB of PDFs, 4 with merged-cell tables). Pages over 600 KB or past the 8 MB budget are skipped and counted. All 200 upstream pages convert (none skipped). Every element category present is a tag (`has-table`, `has-header`, `has-chart` …). The reference has no per-page source (Library of Congress / OER / Upstage) or domain, so neither can be tagged. Counts: `benchmark/datasets/dpbench/conversion-stats.json`. **Not comparable to published DP-Bench numbers**: upstream NID concatenates element text with newlines *removed* (not replaced by spaces) and compares with `rapidfuzz.fuzz.ratio`, on text only; PuffinParse scores a markdown transcript with its own normalisation, tables included. Upstream also discards predicted text that lies inside a ground-truth figure or chart region (`--filter-by-gt-area`). A transcript has no regions, so chart labels a parser transcribes count as extra text here: `tesseract/default` scores 65 on `chart` pages and 99 on `text` pages. ## Self-checks `crates/puffinparse-core/tests/benchmark_datasets.rs` keeps the adapters honest: - `olmocr_rules_are_satisfiable` builds a *witness* prediction from each document's own assertions: every `present` text, every `order` pair, and each `table_cell` as a one-row table. Every rule, including the `absent` ones, must pass, so an adapter cannot emit a rule the scorer can never satisfy. It also requires an empty parse to fail every document that is not `absent-only`. - `omnidocbench_truth_scores_itself_perfectly` scores each locally built truth against itself and checks that the `has-table` tag agrees with `table_score`. - `dpbench_truth_scores_itself_perfectly` scores each committed truth against itself (1.0, `table_score` and `teds_grid` 1.0 on `has-table` pages, tag agreement) and requires an empty parse to score 0. - `combined_v2_references_resolve` / `combined_v3_references_resolve` check every committed `../` path of `combined-v2` / `combined-v3`. - Two `#[ignore]`d helpers: - `PUFFINPARSE_SELFCHECK_RULES_DIR` runs the witness check over any directory of rule files. On the full 824-document olmOCR conversion: 3,040 rules, 0 unsatisfiable. - `PUFFINPARSE_SELFCHECK_PRED_DIR` scores real extractions, such as `pdftotext` output. ## Licensing | Dataset | Redistributable? | What is committed | |---|---|---| | `synthetic-v1` | yes, CC0-1.0 | everything | | `parsebench` | yes, Apache-2.0 (publisher's terms) | 40 documents + truth + rules | | `olmocr` | yes, ODC-BY-1.0 (attribution; AI2 Responsible Use Guidelines) | 40 PDFs + rules | | `omnidocbench` | **no**: no licence; "research only, not for commercial use" | manifest index only (`sha256`, `truth_sha256`); fetched at run time | | `dpbench` | yes, MIT (dataset card) | 40 PDFs + truth | | RealDoc-Bench / LongExtractBench | annotations only | would be manifest + `sha256` only | Rules for any future adapter: - Record the SPDX id in the manifest `license`, and per document when a dataset mixes them. - Carry the publisher's citation in `attribution` on the manifest **and** on each document, so a combined manifest never loses it. - Never vendor bytes whose license does not clearly allow it — index them with `upstream_path` + `sha256` and fetch at run time. - Keep `upstream.revision` a commit hash. A result JSON is only reproducible if the input is. ## Known gaps in the Rust side Both gaps are closed: `ManifestDoc` carries `kind` / `rules`, rule files are hashed into the dataset SHA, `kind: rules` documents are scored with `puffinparse_core::bench::score_rules`, and `table-only` documents are headlined by `table_score` (`summarize_with`). --- # Specification <!-- source: docs/SPEC.md | url: https://puffinparse.com/docs/project/spec/ --> > **One API for every OCR / document-parsing provider.** Rust core, Python SDK, CLI, > and an open benchmark that ranks providers on accuracy, latency and cost. Status: `v0.1` — providers: **Reducto**, **Extend**, **LlamaParse**. --- ## 1. Goals and non-goals ### Goals 1. **Single call, any provider.** `puffinparse.parse("invoice.pdf", model="reducto/standard")` returns the same `ParseResponse` shape whether the backend is Reducto, Extend, LlamaParse, or anything added later. Switching provider is a one-string change, within a mode (§3.1). 2. **Fast and lite.** The core is a Rust library (`puffinparse-core`) with a small dependency set. Python only wraps it (PyO3). No provider SDKs are vendored; every provider is talked to over plain HTTPS with `reqwest`. 3. **Cost tracking.** Every response carries `usage` (pages, provider credits) and a computed `cost_usd` from an embedded, overridable pricing table. 4. **Reliability primitives.** Retries with backoff, per-call timeouts, and a `Router` with ordered fallbacks and simple load balancing. 5. **Open benchmark.** A reproducible harness + datasets + metrics that ranks providers, with results committed to the repo and published as a leaderboard. 6. **Open-source standards.** MIT license, CI (fmt, clippy, tests, Python lint + tests), semver, CHANGELOG, contributor docs, typed Python API, docstrings. ### Non-goals (v0.1) - A hosted service. (A self-hosted gateway now exists: `puffinparse serve`, §14.) - Bundling or running OCR models in-process. Local and self-hosted engines are supported only as providers that call *out* to them — the `tesseract` binary via `tokio::process`, a docling-serve or PaddleOCR serving endpoint over HTTP (§8.4) — never through C bindings or an embedded runtime. - A parsing-only scope. v0.1 specifies and wires three modes end to end (§3.1); which models serve `extract` is a registry fact reported by `list_models("extract")`, and the three original providers serve `parse` and `ocr` only. Classification and other vendor products remain out of scope. --- ## 2. Architecture ``` ┌────────────────────────────────────────────────────────────────────┐ │ Python SDK (python/puffinparse) CLI (crates/puffinparse-cli) │ │ parse / ocr / extract (+ a*) puffinparse parse | ocr | extract │ │ Router(models, mode=...) puffinparse providers | bench │ └───────────────┬────────────────────────────────┬───────────────────┘ │ PyO3 (crates/puffinparse-python) │ ┌───────────────▼────────────────────────────────▼───────────────────┐ │ puffinparse-core (Rust) │ │ ├─ modes parse → Parse / ocr → Text / extract → Extract │ │ ├─ types DocumentRequest / ParseResponse / TextResponse / │ │ │ ExtractResponse / Page / Block / Line / Word │ │ ├─ providers trait Provider { reducto, extend, llamaparse } │ │ ├─ router fallbacks, retries, strategy (ordered/round-robin) │ │ ├─ pricing embedded per-mode price table → cost_usd │ │ ├─ input path | bytes | url → DocumentInput │ │ └─ bench text normalisation + metrics (CER, WER, similarity)│ └────────────────────────────────────────────────────────────────────┘ ``` Crates: | Crate | Purpose | |---|---| | `crates/puffinparse-core` | Library. All provider logic, types, router, pricing, benchmark metrics. `#![forbid(unsafe_code)]`. | | `crates/puffinparse-cli` | `puffinparse` binary: `parse`, `ocr`, `extract`, `providers`, `bench run`, `bench report`. | | `crates/puffinparse-python` | PyO3 extension module `puffinparse._core`, built with maturin. | | `crates/puffinparse-server` | HTTP gateway behind `puffinparse serve` (axum): aliases, virtual keys, budgets, metrics. §14, `docs/SERVER.md`. | | `python/puffinparse` | Pure-Python public API, dataclasses, callbacks, typing. | | `benchmark/` | Datasets, manifests, ground truth, results, leaderboard generator. | --- ## 3. Modes and model naming ### 3.1 Modes Document-AI vendors sell three different products, and they are not interchangeable. PuffinParse makes that explicit: every call names a **mode**, and the mode decides both which providers can serve it and what comes back. | Mode | Entry point | Response | What it is for | |---|---|---|---| | `parse` | `puffinparse.parse` / `puffinparse_core::parse` | `ParseResponse` (§5.1) | layout-aware markdown + typed blocks: RAG chunks, tables, structure | | `ocr` | `puffinparse.ocr` / `puffinparse_core::ocr` | `TextResponse` (§5.2) | plain text with line/word boxes: search, redaction, overlays | | `extract` | `puffinparse.extract` / `puffinparse_core::extract` | `ExtractResponse` (§5.3) | a JSON object shaped by a schema, with per-field citations | **Providers are swappable only within a mode.** A one-string provider switch is only honest between models that do the same job: a markdown parse and a schema extraction are not substitutes for each other, and a router that silently fell back from one to the other would change the shape of the answer. So each model in the registry declares `modes: &[Mode]`, and `ModelRef::parse_for(model, mode)` rejects a model that does not serve the requested mode — before any network call, with a message naming the mode and the models that do serve it. `Mode` is a Rust enum (`Mode::{Parse, Ocr, Extract}`) with `FromStr` (`"ocr"`/`"text"`, `"extract"`/`"extraction"`, case-insensitive), `as_str`, and `Mode::ALL`. In Python it is the string literal type `puffinparse.Mode = Literal["parse", "ocr", "extract"]`. `list_models(mode)` (`list_models_for` in Rust) lists the models for one mode; with no mode it lists all of them. A provider that has no native OCR endpoint serves `ocr` from its own parse output (the default `Provider::ocr` implementation): lines come from block text, words from lines, and the response is tagged `metadata["puffinparse_derived_from"] = "parse"` so callers can tell native OCR geometry from derived geometry. Nothing is derived across any other mode pair. ### 3.2 Model naming Like LiteLLM, the `model` string selects provider and model: `"<provider>/<model>"`. | Provider | Models (v0.1) | Modes | Maps to | |---|---|---|---| | `reducto` | `reducto/standard` (default), `reducto/r-1`, `reducto/agentic` | parse, ocr | default settings / `settings.model="r-1"` / `enhance.agentic=[{scope:text},{scope:table}]` | | `extend` | `extend/parse_performance` (default), `extend/parse_light`, `extend/parse_auto` | parse, ocr | `config.engine` | | `llamaparse` | `llamaparse/fast`, `llamaparse/cost_effective` (default), `llamaparse/agentic`, `llamaparse/agentic_plus` | parse, ocr | `tier` form field (+ `version=latest`) | The three providers above serve `parse` and `ocr` only. `list_models("extract")` reports which models (if any) serve extraction; calling `extract` with a parse-only model raises `UnsupportedModelError` before any network call. Aliases: `llama`, `llama_parse`, `llamacloud` → `llamaparse`. Matching is case-insensitive. `model="reducto"` (no slash) selects the provider's default model **for the mode being called**. Unknown providers or models raise `UnsupportedModelError` before any network call, as does a known model asked for a mode it does not declare. Provider-specific knobs that do not fit the common request are passed through `provider_options` (a JSON object) and merged into the provider request body verbatim. This is the escape hatch; it never changes the response shape. --- ## 4. Unified request ### 4.1 The common document request Every mode takes the same document request; `parse` and `ocr` take nothing else. ```python puffinparse.parse( input, # str path | pathlib.Path | bytes | "https://..." URL model: str = "reducto", # "<provider>/<model>", must support this mode *, filename: str | None = None, # required when input is bytes pages: str | None = None, # "1-3,7" 1-based page selection (best effort per provider) language: str | None = None, # BCP-47 hint, forwarded if provider supports it output: Literal["markdown", "text"] = "markdown", # preferred `content` of blocks (parse only) output_format: str | None = None, # "reducto" | "extend" | "llamaparse": return that vendor's # native JSON shape instead of the unified response # (docs/COMPAT.md, ADR-13); None = unified provider_options: dict | None = None, include_raw: bool = False, # attach the provider's raw JSON to response.raw timeout: float = 300.0, # seconds, whole call including polling max_retries: int = 2, # on 429 / 5xx / network errors, exponential backoff api_key: str | None = None, # overrides env var base_url: str | None = None, # overrides provider base URL metadata: dict | None = None # echoed back, useful for callbacks/logging ) -> ParseResponse ``` `aparse(...)` is the `async def` equivalent. In Rust this is `DocumentRequest`, a builder over the same fields. `DocumentRequest` also carries `webhook_url: Option<String>`, used only when the request is *submitted* as a job (§15; Python `submit(..., webhook_url=...)`); `parse` ignores it. ### 4.2 `ocr` ```python puffinparse.ocr(input, model="reducto", *, ...same keywords, minus `output`...) -> TextResponse ``` `ocr` always returns plain text, so `output` does not apply; `aocr(...)` is the async form. ### 4.3 `extract` ```python puffinparse.extract( input, schema: dict, # JSON Schema (draft 2020-12 subset) object for the result *, model: str = "reducto", # must support the `extract` mode instructions: str | None = None, # extra natural-language guidance, forwarded if supported citations: bool = False, # ask for per-field page/box/source-text citations ...same common keywords as §4.1... ) -> ExtractResponse ``` `aextract(...)` is the async form. In Rust this is `ExtractRequest { document: DocumentRequest (flattened when serialised), schema, instructions, citations }`. A `schema` that is not a JSON object is rejected before any network call (`TypeError` in Python, `InputError` in the core). ### 4.4 Input handling Input handling (`DocumentInput`) is identical in all three modes: - **Path** → read bytes, sniff MIME from extension (`mime_guess`), upload. - **Bytes** → require `filename` (used for MIME + provider upload). - **URL** (`http(s)://`) → passed to the provider as a remote URL when the provider supports it (all three do); otherwise downloaded and uploaded. Supported document types are whatever the provider accepts; PuffinParse does not pre-validate beyond a non-empty body. --- ## 5. Unified responses One response type per mode. All three carry the same envelope — `id`, `provider`, `model`, `provider_job_id`, `usage`, `cost_usd`, `latency_ms`, `created_at`, `metadata`, `raw` — and differ only in the payload. ### 5.1 `parse` → `ParseResponse` ```python @dataclass class ParseResponse: id: str # puffinparse-generated uuid provider: str # "reducto" model: str # "reducto/standard" provider_job_id: str | None pages: list[Page] markdown: str # whole document, pages joined by "\n\n" text: str # plain text usage: Usage cost_usd: float | None # None if pricing unknown latency_ms: int # wall-clock for the whole call, incl. polling created_at: str # RFC 3339 metadata: dict raw: Any | None # provider payload if include_raw @dataclass class Page: page_number: int # 1-based width: float | None # points/pixels if provider reports it height: float | None markdown: str text: str blocks: list[Block] @dataclass class Block: type: BlockType # text | title | section_header | list | table | figure | # header | footer | footnote | caption | formula | other content: str # markdown (tables as markdown/HTML per provider) text: str | None # plain text if provider gives a separate one bbox: BBox | None # normalised 0..1 {x0, y0, x1, y1}, origin top-left confidence: float | None # 0..1 if provider reports one page_number: int @dataclass class Usage: pages: int # pages billed/processed credits: float | None # provider-native credit units, if any provider_cost_usd: float | None # if the provider reports $ directly ``` Rules: - `markdown`/`text` at document level are derived from pages in order. - A provider that only returns per-chunk (not per-page) content gets pages reconstructed from block page numbers; if that is impossible, a single page `page_number=1` is emitted and `Usage.pages` still reflects the billed count. - Block types are mapped from each provider's vocabulary (see §8). Unknown → `other`. - `bbox` is normalised so consumers can draw overlays without knowing the page size. Providers reporting absolute coordinates are divided by page dims. ### 5.2 `ocr` → `TextResponse` ```python @dataclass class TextResponse: id: str provider: str model: str provider_job_id: str | None pages: list[TextPage] text: str # whole document, pages joined by "\n\n" usage: Usage cost_usd: float | None latency_ms: int created_at: str metadata: dict # "puffinparse_derived_from": "parse" when derived (§3.1) raw: Any | None @dataclass class TextPage: page_number: int # 1-based width: float | None height: float | None text: str # plain text in reading order, lines separated by "\n" lines: list[Line] words: list[Word] @dataclass class Line: # and Word, identical shape text: str bbox: BBox | None # normalised 0..1, origin top-left confidence: float | None ``` No markdown, no block types: `ocr` is recognition, not layout analysis. When the result is derived from a parse (§3.1), lines carry their block's box and words carry none. ### 5.3 `extract` → `ExtractResponse` ```python @dataclass class ExtractResponse: id: str provider: str model: str provider_job_id: str | None data: Any # the extracted object, shaped by the request schema fields: dict[str, FieldInfo] # keyed by JSON pointer into `data`, e.g. "/invoice/total" usage: Usage cost_usd: float | None latency_ms: int created_at: str metadata: dict raw: Any | None @dataclass class FieldInfo: confidence: float | None citations: list[Citation] @dataclass class Citation: page_number: int # 1-based bbox: BBox | None # normalised 0..1, origin top-left text: str | None # source text the value was read from, if reported ``` `fields` is empty when the provider reports no per-field metadata; `citations` is only populated when the request asked for them and the provider supports them. `ExtractResponse.field_info(p)` and `.citations(p)` are Python conveniences over the pointer map. --- ## 6. Errors All errors derive from `puffinparse.PuffinParseError`: | Error | When | |---|---| | `AuthenticationError` | 401/403 from provider or missing API key | | `RateLimitError` | 429 (retried first; raised after `max_retries`) | | `BadRequestError` | 4xx other than auth/rate limit | | `ProviderError` | 5xx or provider-reported job failure | | `TimeoutError` | overall timeout (upload + poll) exceeded | | `UnsupportedModelError` | bad `model` string, or a model that does not serve the requested mode | | `InputError` | unreadable file, bytes without filename, empty body, bad mode name, router asked for another mode | Every error carries `provider`, `status_code` (if any), `message`, and `request_id`/`job_id` when available. Mode errors are raised before any network call. `UnsupportedModelError` from a mode mismatch names the offending model, the modes it does serve, and the models that serve the mode you asked for. --- ## 7. Router ```python router = puffinparse.Router( models=["reducto/standard", "llamaparse/agentic", "extend/parse_light"], mode="parse", # "parse" (default) | "ocr" | "extract" strategy="ordered", # "ordered" (fallback order) | "round_robin" fallback_on=("ProviderError", "RateLimitError", "TimeoutError", "NetworkError"), ) resp = router.parse("doc.pdf") # tries each in turn / rotates ``` A router is bound to one mode. Every model is validated against it at construction (`UnsupportedModelError` otherwise), and calling a method for a different mode — say `Router([...], mode="parse").ocr(...)` — raises `InputError` rather than answering with a different shape. `router.mode` reports it; `parse`/`ocr`/`extract` (and `aparse`/`aocr`/ `aextract`) are the per-mode calls, each taking the same arguments as the module-level function minus `model`. Semantics: `ordered` → try `models[0]`, on a fallback-eligible error move on. `round_robin` → rotate the starting index per call, then fallback in order. The router records per-model success/failure counts and average latency, exposed as `router.stats()`. Auth/BadRequest/Input errors never trigger fallback. When a fallback served the call, the response metadata carries `puffinparse_fallback_index` and `puffinparse_fallback_from_error`. --- ## 8. Provider mapping All three mappings below were verified against live API responses on 2026-09-11; the captured payloads live in `crates/puffinparse-core/tests/fixtures/` and drive unit tests. ### 8.1 Reducto - Base: `https://platform.reducto.ai`, header `Authorization: Bearer <key>` (`REDUCTO_API_KEY`, `REDUCTO_BASE_URL` override). - Flow: `POST /upload` (multipart `file`) → `{file_id: "reducto://…"}`; URLs are passed directly. Then `POST /parse` (sync, 900 s ceiling) with the v3 body `{input, retrieval:{chunking:{chunk_mode:"page"}}, formatting:{table_output_format:"md"}, settings:{…}}`. With `provider_options.async = true`: `POST /parse_async` → `{job_id}` → poll `GET /job/{id}` (`Pending` → `Completed` | `Failed`); the parse payload is `job.result`. - `pages` → `settings.page_range = [{start,end}]` (1-based). - Response: `result.type` is `"full"` (`chunks[]`) or `"url"` (fetch `result.url`; the body is the same `FullResult` object). `chunks[].blocks[]` carry `type` (values contain spaces, e.g. `"Section Header"`), `bbox{left,top,width,height,page,original_page}` already normalised to 0..1, `content`, `confidence` (`"high"|"low"`), `granular_confidence.parse_confidence`. `usage.num_pages`, `usage.credits` (null on per-product pricing accounts). - Block types: `Title→title`, `Section Header→section_header`, `Text`/`Key Value`/`Comment→text`, `List Item→list`, `Table→table`, `Figure→figure`, `Header→header`, `Footer→footer`, `Footnote→footnote`, `Caption→caption`, `Formula→formula`, else `other`. - Pages: with `chunk_mode=page` each chunk is one page; page number is taken from the chunk's blocks' `bbox.page` (chunks themselves have no page field). Chunk `content` is kept as the page markdown. - Errors: `{"error":{"code","name","message"},"detail"}`, `422` Pydantic arrays, a bare nginx HTML `403` when the header is missing, and a non-standard `442` for password-protected files. ### 8.2 Extend - Base: `https://api.extend.ai` (`EXTEND_API_KEY`, `EXTEND_BASE_URL`), headers `Authorization: Bearer <key>` and the **mandatory** `x-extend-api-version: 2026-02-09`. `provider_options.workspace_id` sets `x-extend-workspace-id` for org-scoped keys. - Flow: `POST /files/upload` (multipart) → `{id: "file_…"}`; URLs are passed as `file:{url,name}`. Then `POST /parse_runs` (async) → poll `GET /parse_runs/{id}` until `status ∈ {PROCESSED, FAILED}`. (The sync `POST /parse` has a 5-minute hard limit, so PuffinParse always uses runs.) - Body: `{file, config:{target:"markdown", chunkingStrategy:{type:"page"}, engine, blockOptions:{tables:{targetFormat:"markdown"}}, advancedOptions:{pageRanges}}}`. `provider_options` keys `target`, `chunkingStrategy`, `engine`, `engineVersion`, `blockOptions`, `advancedOptions` are merged into `config`; others (`metadata`, `dataRetention`) at top level. - Response: the run object itself: `output.chunks[]` (`type:"page"`, `content`, `metadata.pageRange`) with `blocks[]` (`type`, `content`, `metadata.page{number,width,height}`, `metadata.avgOcrConfidence`, `boundingBox{left,top,right,bottom}` in page pixels); `metrics.pageCount`, `usage.credits`. `responseType=url` results are fetched from `outputUrl`. - Block types: `heading→title`, `section_heading→section_header`, `text`/`key_value→text`, `table`/`table_head`/`table_cell→table`, `figure→figure`, `formula→formula`, `header`, `footer`, else `other` (`page_number`, `barcode`). - Errors: `{code, message, requestId, retryable}`; failed runs carry `failureReason`. ### 8.3 LlamaParse - Base: `https://api.cloud.llamaindex.ai` (`LLAMA_API_KEY`, `LLAMA_BASE_URL`; EU: `https://api.cloud.eu.llamaindex.ai`), header `Authorization: Bearer llx-…`. - Flow: `POST /api/v1/parsing/upload` (multipart `file` or `input_url`; form fields `tier`, `version=latest`, `language`, `target_pages` (0-based, converted from `pages`), plus any `provider_options` as extra form fields) → `{id, status}`; poll `GET /api/v1/parsing/job/{id}` until `SUCCESS | PARTIAL_SUCCESS | ERROR | CANCELLED`; then `GET /api/v1/parsing/job/{id}/result/json`. - Response: `pages[].{page (1-based), text, md, items[], width, height}`; items have `type` (`heading` with `lvl`, `text`, `table`), `md`, `value`, `bBox{x,y,w,h,confidence}` in page units. `job_metadata.job_pages`; `job_credits_usage` is `0` until billing settles and is therefore only reported when positive. - Block types: `heading` lvl 1 → `title`, other headings → `section_header`, `text→text`, `table→table`, else `other`. - Errors: FastAPI `{"detail": "…"}` / `{"detail": [ValidationError]}`. ### 8.4 Self-hosted engines (Tesseract, Docling, PaddleOCR) Listed in `model::SELF_HOSTED`; `ProviderInfo::self_hosted()` is `true`, no API key is required (`env_var` is empty, or names an optional key), prices are `0.0` with `source: "self-hosted"`, and `puffinparse providers` shows `local` in the Key column (`self_hosted` / `key_required` in `--json` and in Python `providers()`). - `tesseract/default`: `tesseract <image> stdout ... tsv`, PDFs rasterised with `pdftoppm`; `ocr` is native (TSV words/lines, confidences), `parse` is one `text` block per Tesseract paragraph. - `docling/default`: docling-serve `POST /v1/convert/source/async` → poll → `GET /v1/result`; DoclingDocument items mapped to blocks, bottom-left boxes flipped to top-left. - `paddleocr/default`: PaddleX serving `POST /ocr` (ocr) and `POST /layout-parsing` (PP-StructureV3, parse). Details, errors and limits: `docs/providers/{tesseract,docling,paddleocr}.md`. --- ## 9. Pricing `crates/puffinparse-core/pricing.json` (embedded via `include_str!`) maps `"<provider>/<model>"` → `{ "parse": float, "ocr": float, "extract": float, "source": url, "updated": date }` — a per-page price **per mode**, each optional, since vendors price parsing, OCR and extraction differently. `cost_usd = usage.pages * price_per_page(model, mode)` unless the provider reports credits with a known credit price, in which case `credits * per_credit_usd` is used. A mode with no price yields `cost_usd = None`. Users can override one mode at a time with `puffinparse.set_pricing({"reducto/standard": 0.01}, "parse")`, and ask for an estimate with `puffinparse.estimate_cost("reducto/standard", 1000, "ocr")`. Prices are best-effort public list prices; the benchmark reports them as such. --- ## 10. Benchmark ### 10.1 Principles 1. **Reproducible**: every run records provider, model, dataset hash, options, timestamp, and the raw provider outputs (optionally, gitignored). 2. **Machine-checkable ground truth**: text-based metrics that don't need an LLM. An optional LLM-judge is a plug-in, never required for the leaderboard. 3. **Three axes**: accuracy, latency (p50/p95 per page), cost per 1k pages. 4. **Open datasets only**: synthetic documents generated by the repo (so the ground truth is exact) plus adapters for public sets that users download themselves. Implemented adapters: ParseBench (LlamaIndex), olmOCR-bench (AI2), OmniDocBench (OpenDataLab, index only: research-only licence), DP-Bench (Upstage, MIT; folded into `combined-v3`). Planned: RealDocBench (Extend) and LongExtractBench (Reducto / micro_1) once public. Licences and mappings: `docs/benchmarks/academic-benchmarks.md`, `docs/benchmarks/adapters.md`. Vendor-published benchmarks are each won by their publisher; running all of them through one harness with one scoring pipeline is the point of the meta-benchmark. Each adapter is `benchmark/adapters/<name>.py`, emits a `manifest.json`, and records the upstream version/commit so results stay tied to a dataset revision. The combined leaderboard reports per-dataset scores and an unweighted mean across datasets. ### 10.2 Dataset format ``` benchmark/datasets/<name>/ manifest.json # {name, version, description, license, documents:[…]} docs/<id>.<ext> # input file truth/<id>.md # expected markdown (or .txt for text-only docs) ``` `manifest.documents[]`: `{id, file, truth, pages, tags:[...], category}`. Categories in the built-in `synthetic-v1` set: `plain`, `invoice`, `table`, `two_column`, `headings`, `noisy_scan`, `low_res`, `multipage`, `skewed`, `dense`, `faded`, `receipt`, `complex_table`. The generator (`benchmark/generate_synthetic.py`) is seeded and byte-reproducible; the truth is produced from the same source the pixels are rendered from. ### 10.3 Metrics (computed in Rust, `puffinparse_core::bench`) Given predicted `P` and truth `T` after **normalisation** (NFKC, strip markdown syntax and HTML tags, decode HTML entities, straighten quotes/dashes, collapse whitespace, lowercase for the `case_insensitive` variant): - `char_similarity = 1 - levenshtein(P, T) / max(|P|, |T|)` (the primary metric of a plain transcript document) - `cer = levenshtein(P, T) / |T|` - `wer = word_levenshtein(P_words, T_words) / |T_words|` - `word_recall` = fraction of truth word tokens present in prediction (bag-of-words) - `word_precision`, `word_f1` - `table_score` (when the truth has a table): char_similarity restricted to table rows. Tables are markdown pipe tables **or HTML `<table>`s** (thead/tbody, th/td, `colspan`/`rowspan` repeated into every slot they cover, entities decoded), each reduced to rows of normalised cells; a row is its cells joined by a space. `puffinparse_core::bench::tables`. - `teds_grid` (when the truth has a table): TEDS (Zhong et al. 2020) computed on the grid tree `table > row > cell` — `1 - TED / max(|Tp|, |Tt|)`, Zhang–Shasha tree edit distance, unit insert/delete, cell rename = normalised Levenshtein of the contents. It is TEDS without `thead`/`tbody` nodes and span attributes, because the ground truth is markdown. Each truth table is matched to its best predicted table; extra predicted tables are not penalised. - `order_score`: Kendall-τ–like agreement of the order of lines shared by both texts Per document all metrics are recorded, plus the document's **headline**: `table_score` for `table-only` documents, the rule pass rate for `kind: rules`, `char_similarity` otherwise. Aggregate = mean over docs (failures count 0), plus per category. `Summary.headline` = mean of the document headlines (0–1), **Overall** = `100 * headline`; `Summary.char_similarity` is the plain mean of the documents' `char_similarity` (a rule document has no transcript, so its `char_similarity` field carries the pass rate). Rule matching (`score_rules`) normalises both sides with the run's options, then drops every space adjacent to punctuation, so tokenised rule text (`(this " agreement ")`) matches the printed page. `bag_of_sentences` counts a sentence as present when a window of the prediction is ≥ 0.8 similar (`BAG_SENTENCE_MIN_SIMILARITY`); the rule passes when at least `threshold` of them are (default `BAG_DEFAULT_THRESHOLD = 0.8`). A rule's optional `max_diffs` (olmOCR-bench) allows that many Levenshtein edits in `present` / `absent` / `order` / `table_cell` matching. The scorer is versioned: `puffinparse_core::bench::SCORER_VERSION` (currently `2`) is written to every result as `scorer_version` and bumped whenever a change would move a committed score. ### 10.4 Runner and outputs ``` puffinparse bench run --dataset benchmark/datasets/synthetic-v1 \ --models reducto/standard extend/parse_performance llamaparse/agentic \ --out benchmark/results/<date>-synthetic-v1.json puffinparse bench report benchmark/results/*.json --format markdown > benchmark/LEADERBOARD.md ``` Result JSON: `{run_id, created_at, puffinparse_version, scorer_version, rescored_at?, dataset:{name, version, documents, sha256}, normalize, models:[{model, docs:[{id, category, kind, table_only, pages, metrics, headline, latency_ms, cost_usd, error}], summary:{documents, failed, headline, char_similarity, cer, wer, word_f1, order_score, table_score, teds_grid, rule_pass_rate, overall, latency_p50_ms, latency_p95_ms, latency_per_page_ms, total_pages, total_cost_usd, cost_per_1k_pages_usd, by_category}}]}`. The `sha256` covers the manifest plus every input, truth and rule file, so a result is tied to an exact dataset revision. Consumers (the viewer in `benchmark/site/`, `bench report`) rank on `summary.headline` (0–1) when present, else `overall / 100`, and show a document's `headline` when present. Compatibility with older files: a file without `scorer_version` is scorer v1; its `summary.headline` is read as `overall / 100`, `teds_grid` is absent, and — unlike v2 — its `summary.char_similarity` held the headline (table-only documents contributed `table_score`). `puffinparse bench rescore <result.json> --outputs <dir> [--dataset <dir>] [--out <path>]` re-scores a run from its saved per-document outputs with the current scorer, without network access: metrics, headlines and summaries are recomputed, `kind`/`table_only`/`category` are refreshed from the manifest, latency, cost, pages and errors are kept as measured, `scorer_version` and `rescored_at` are set, and the dataset `sha256` is updated (with a warning) if the dataset changed since the run. A missing output for a successful document is an error unless `--keep-missing` is given, which keeps that document's recorded scores and records how many in `rescore_kept_docs` (used for sources whose outputs are not committed, e.g. research-only datasets). A successful call that returned only whitespace is flagged `empty_output: true` and counted in `summary.empty_outputs`; it is scored, not failed. Each document record also carries audit fields (all optional, so older result files still load): `provider_job_id` (the provider's id for the call — Reducto job id, Extend parse run id, LlamaParse job id — taken from `ParseResponse.provider_job_id`, or from `Error.job_id` for a failure), `cache_hit` (`true` if the provider reported a result-cache hit via a `<provider>_cache_hit` metadata key, `false` if caches were disabled for the run, `null` unknown; always written), `attempts` (calls the runner issued, >1 only with `--retries`; retries inside the HTTP client are not counted), `started_at` (RFC 3339, first attempt) and, for failures, `error_kind` (the `ErrorKind` serde name, `"input"` for an unreadable truth/rule file) next to the `error` message. Resilience: while running, every finished (model, document) record is appended and flushed to `<out>.partial.jsonl` (first line `{"type":"header", run_id, created_at, dataset_sha256, normalize}`, then `{"type":"doc", model, doc}` per call). The final JSON is written from those records and the log is then deleted. `bench run --resume` reads the result JSON and/or the partial log, refuses them if the dataset `sha256` or normalisation differs, keeps the original `run_id`, and calls only the pairs without a successful record. A leftover log without `--resume` is an error, never silently overwritten. `--dry-run` prints the plan (calls, manifest pages, list-price estimate per model) without network access; `--max-cost <usd>` aborts before the first call when the estimate exceeds it (or when a planned model has no list price). The run ends with one line: calls made, resumed, failed, total cost, this invocation's cost and wall time. `LEADERBOARD.md` is regenerated from committed results and links to each run. Document kinds: a manifest document is `kind: "transcript"` (default; `truth` markdown, scored by the text metrics) or `kind: "rules"` (a `rules` file of machine-checkable assertions — `present`, `absent`, `order`, `table_cell`, `bag_of_sentences` — scored by `puffinparse_core::bench::score_rules`, reported as `rule_pass_rate` / `rules_passed` / `rules_total` in `Metrics` and `rule_pass_rate` in `Summary`). Documents tagged `table-only` are headlined by `table_score`. Result JSON documents carry `kind` and `table_only`; rule files are included in the dataset `sha256`. --- ## 11. Python SDK details - `python/puffinparse/__init__.py` exports the three modes — `parse`/`aparse`, `ocr`/`aocr`, `extract`/`aextract` — plus `Router`, the response dataclasses (`ParseResponse`, `Page`, `Block`, `BBox`, `Usage`, `TextResponse`, `TextPage`, `Line`, `Word`, `ExtractResponse`, `FieldInfo`, `Citation`), the `Mode` literal (`"parse" | "ocr" | "extract"`), errors, `set_pricing`, `estimate_cost`, `list_models`, `resolve_model`, `providers`, `modes` and the callback lists. - Mode arguments are plain strings everywhere (`list_models("ocr")`, `Router([...], mode="ocr")`, `set_pricing({...}, "ocr")`, `estimate_cost(m, 1000, "ocr")`), typed as `puffinparse.Mode`. - Callbacks: `puffinparse.success_callback: list[Callable[[Response], None]]` where `Response` is the union of the three response types, and `puffinparse.failure_callback` — both fire for every mode, sync and async; awaitables returned by a callback are awaited (on the caller's loop for `a*` calls, on a private loop otherwise). - Bytes never cross the FFI boundary as base64: the Python layer passes the document as a separate `bytes` argument and the extension builds `DocumentInput::Bytes`. `extract` follows the same convention — the request dict is `ExtractRequest` with the document fields flattened into it. - Errors from a mode mismatch are enriched in Python: an `UnsupportedModelError` from `extract`/`ocr` names the mode *and* the models that serve it (`list_models(mode)`), so a user who passes a parse-only model to `extract` is told what to use instead. Passing a non-dict `schema` raises `TypeError` before any FFI call. - Jobs (§15): `submit`/`asubmit` → `Job` (dataclass mirroring `JobHandle`), `retrieve`/`aretrieve` and `handle_webhook`/`ahandle_webhook` → `Job | ParseResponse` (`puffinparse.JobResult`), implemented in `python/puffinparse/jobs.py` over `_core.submit`, `_core.retrieve` and `_core.parse_webhook`. - `puffinparse.score`, `normalize_text`, `markdown_to_text` expose the benchmark metrics. - Logging: `PUFFINPARSE_LOG=debug` enables tracing in the core; Python uses `logging.getLogger("puffinparse")`. - Typing: fully typed, `py.typed` shipped; dataclasses mirror the Rust structs 1:1. - Build: maturin, `abi3-py39` wheels, `pip install puffinparse`. --- ## 12. CLI ``` puffinparse parse <file|url> [--model reducto/standard] [--format markdown|text|json] [--raw] puffinparse ocr <file|url> [--model reducto/standard] [--format text|json] puffinparse extract <file|url> --schema <file.json|inline JSON> [--instructions TEXT] [--citations] puffinparse providers [--mode parse|ocr|extract] [--json] # models, modes, per-mode pricing, key status puffinparse bench run|report|score|rescore # see §10 (parse mode; rescore is offline) puffinparse serve [--config puffinparse.toml] [--host H] [--port P] # HTTP gateway, see §14 ``` One subcommand per mode; `--model` must name a model that serves that subcommand's mode, and a bare provider name resolves to its default model *for that mode*. All three share the common options (`--pages`, `--language`, `--options`, `--timeout`, `--max-retries`, `--api-key`, `--base-url`, `--raw`). - `parse` prints markdown (default), plain text, or the whole `ParseResponse` as JSON. - `ocr` prints the plain text (default) or the whole `TextResponse` as JSON, which includes per-page `lines[]` and `words[]` with boxes. - `extract` always prints the `ExtractResponse` as JSON. `--schema` is either a path to a JSON file or inline JSON starting with `{`. - `providers` lists every model with the modes it serves and its price in each mode; `--mode` filters the table to one mode and shows a single price column. Non-JSON output prints a one-line summary to stderr (`[model] N page(s) in T ms, est. $X`). Exit code 0 on success, 1 on provider error, 2 on usage/config error (including a model that does not serve the requested mode). --- ## 13. Quality bar - Rust: `cargo fmt --check`, `cargo clippy -D warnings`, unit tests with fixtures for each provider's parser, no network in tests (live tests are `#[ignore]` and run only when keys are present). - Python: `ruff`, `mypy --strict` on the package, `pytest` with a fake core for unit tests; live tests skipped without keys. - CI: GitHub Actions on push/PR (Linux; wheels build matrix on tags). - Versioning: semver, single workspace version, `CHANGELOG.md` (Keep a Changelog). --- ## 14. Gateway server `puffinparse serve` (crate `puffinparse-server`) exposes the three modes over HTTP for clients that should not hold provider keys. Operator reference: [`SERVER.md`](/docs/gateway/index.md). Contract: - **Endpoints.** `POST /v1/parse | /v1/ocr | /v1/extract` take the §4 fields (`model`, `pages`, `language`, `output`, `output_format`, `provider_options`, `include_raw`, `timeout`, `max_retries`, `metadata`; `schema` / `instructions` / `citations` for extract) plus `fallbacks: [str]`, as JSON (`document_url`, or base64 `document` + `filename`) or multipart (`file` part + the same fields). `api_key`, `base_url` and local paths are rejected. They return the §5 response JSON unchanged, or the vendor shape for `output_format` (parse, extract). Also `GET /v1/models`, `GET /v1/usage`, `GET /health`, `GET /metrics` (Prometheus text). - **Jobs (§15).** `POST /v1/jobs` takes the `/v1/parse` body plus `webhook_url` and returns `202 {id, object: "job", status: "pending", job: JobHandle}` (without `base_url`). Auth, allow-lists, aliases, budget pre-check and `rpm` apply as for `/v1/parse`, but an alias submits to its first target and `fallbacks` is rejected (no fallback for jobs). `GET /v1/jobs/{id}` checks the provider once: `{id, object, status: pending|succeeded|failed, model, provider, provider_job_id, submitted_at, result?, error?}`, `result` in the submit-time `output_format` (or `?output_format=`), `error` the error object below. `id` is opaque and bound to the submitting key (other keys get `404 not_found`; the master key sees all). Cost is charged to that key once, when the job is first observed succeeded; polls count toward `rpm`, not the budget. Handles (no secrets) live in memory and the `state_file` for `job_retention_hours`. Optional `POST /v1/webhooks/{provider}` (`[webhooks] enabled`, shared `secret` as `?token=` or `x-puffinparse-webhook-secret`, else 404) resolves a provider webhook body with core `parse_webhook` (+ one retrieve when needed) against a stored job and settles it the same way. - **Config.** One TOML file: `[server]`, `master_key`, `[providers.<name>]` (`api_key`, `base_url`), `[[models]]` aliases (`name`, `targets`, `strategy`, `fallback_on`, with the §7 semantics and per-target credential overrides), `[[keys]]` virtual keys (`id`, `key`, `models` allow-list with `provider/*` wildcards, `monthly_budget_usd`, `rpm`). Secrets may be `env:VAR`. No master key and no keys means auth is off. - **Accounting.** Spend = response `cost_usd`, per key per UTC calendar month, checked before each call (`402` once spent ≥ budget); `rpm` is a sliding 60 s window (`429` + `Retry-After`). State is in memory, optionally persisted to a JSON `state_file`. - **Errors.** One body shape, `{"error": {type, message, provider, provider_status, job_id, request_id}}`. `ErrorKind` → HTTP: input / bad_request / unsupported_model → 400, rate_limit → 429, timeout → 504, provider / network / authentication → 502 (provider credentials are the operator's). Gateway-own types: `unauthorized` 401, `budget_exceeded` 402, `model_not_allowed` 403, `not_found` 404, `payload_too_large` 413, `key_rate_limited` 429. The provider's message is passed through. - **Logs.** One JSON line per request: `ts, request_id, key_id, method, path, mode, model, served_model, provider, fallback_index, pages, cost_usd, latency_ms, status, error_type, provider_status`, plus `job_id, job_status` on jobs API lines. Never document content or URLs, provider error text, or any secret. Metrics label jobs traffic `mode="job_submit" | "job_retrieve" | "webhook"` and count `puffinparse_jobs_total{event}`. ## 15. Asynchronous jobs and webhooks `parse` blocks until the provider is done (polling job-queue providers internally). The jobs API splits that call in two so the caller owns the waiting — for long documents, large batches, or pipelines driven by provider webhooks. An SDK cannot *receive* a webhook, so PuffinParse exposes the primitives and a parser for webhook bodies; your web handler does the receiving. `parse` mode only. ```rust puffinparse_core::submit_parse(DocumentRequest) -> Result<JobHandle> puffinparse_core::retrieve_parse(&JobHandle) -> Result<JobStatus> // env credentials puffinparse_core::retrieve_parse_with(&JobHandle, &RetrieveOptions) -> Result<JobStatus> puffinparse_core::parse_webhook(model, &serde_json::Value) -> Result<WebhookEvent> // no network puffinparse_core::resolve_webhook(model, &Value, &RetrieveOptions) -> Result<JobStatus> ``` ```python job = puffinparse.submit("big.pdf", model="reducto/standard", webhook_url="https://…/hook") # -> Job puffinparse.retrieve(job) # -> Job (still running) | ParseResponse; raises on failure puffinparse.handle_webhook(body, model="reducto") # -> Job | ParseResponse; raises on failure # asubmit / aretrieve / ahandle_webhook are the async equivalents ``` ```ts const job = await submit('big.pdf', { model: 'reducto/standard', webhookUrl }) // -> Job (camelCase JobHandle) await retrieve(job, { apiKey?, outputFormat? }) // -> the same Job | ParseResponse; rejects on failure await handleWebhook(body, { model: 'reducto' }) // -> Job | ParseResponse; rejects on failure ``` Over HTTP the gateway exposes the same pair as `POST /v1/jobs` / `GET /v1/jobs/{id}` (§14). **`JobHandle`** (`Job` in Python): `provider`, `model` (qualified), `job_id` (the provider's id), `submitted_at` (RFC 3339), `output`, `include_raw`, `base_url`, `provider_state`, `metadata`. It never contains a secret: API keys are resolved again at retrieve time (env var or `RetrieveOptions.api_key` / `retrieve(api_key=…)`), and `provider_state` holds only the non-secret options a later call needs (Extend `workspace_id`, `responseType`). Serialise it with serde / `Job.to_dict()` to store it or hand it to another process. **`JobStatus`**: `Pending` | `Succeeded(ParseResponse)` | `Failed(Error)`. A succeeded job is normalised exactly like `parse` (qualified model, cost, request metadata echoed, `raw` only with `include_raw`); `latency_ms` is the time since submission. A provider-side failure is `Ok(Failed(e))` with `e.job_id` set — `Err` means the status check itself failed (auth, network). Python raises the typed exception for `Failed`. **Webhooks.** `webhook_url` maps to each provider's own per-job setting; `parse_webhook` reads the body the provider POSTs: | Provider | Submit / retrieve | `webhook_url` → | Webhook body → `WebhookStatus` | |---|---|---|---| | `reducto` | `POST /parse_async` / `GET /job/{id}` | `async.webhook = {"mode": "direct", "url": …}` | `{"status", "job_id"}`: `Completed`/`Failed` → `Finished` (retrieve for result or reason), else `Pending` | | `extend` | `POST /parse_runs` / `GET /parse_runs/{id}` | **rejected** (`input` error): Extend only has workspace webhook endpoints | `{"eventType": "parse_run.*", "payload": parse_run_status}`: `PROCESSED` → `Finished` (or `Succeeded` if a full run with output), `FAILED` → `Failed` (reason + message), else `Pending` | | `llamaparse` | `POST /api/v1/parsing/upload` / `GET /api/v1/parsing/job/{id}` (+ `result/json`) | multipart field `webhook_url` | the `webhook_url` result push `{"txt","md","json":[pages]}` → `Succeeded`; a LlamaCloud event `{"event_type": "parse.*", "data": {"job_id"}}` → `Finished` / `Pending` | `WebhookStatus::Finished` means "terminal, but the body carries neither the result nor the error detail"; `resolve_webhook` (Python `handle_webhook`) then makes one retrieve. Verifying webhook authenticity (Extend and LlamaCloud HMAC signatures, a secret in Reducto `async.metadata`) is the caller's job and must happen before the body is trusted. Other providers return `unsupported_model` from `submit_parse`. --- # Native-format compatibility <!-- source: docs/COMPAT.md | url: https://puffinparse.com/docs/project/compat/ --> PuffinParse normalises every provider to one response shape. That is the right default, but it is a migration cost for anyone already integrated with a vendor: code that reads `result.chunks[].blocks[].bbox.left` has to be rewritten before a single request can be re-pointed at another provider. The compatibility layer removes that step. Ask for a vendor's shape and PuffinParse renders the unified response into that vendor's own JSON, **whatever provider actually produced it**: ```text any provider → ParseResponse (unified) → render_parse(Format::Reducto) → Reducto's parse JSON ``` Switch the model string, keep your parser. --- ## 1. Using it Rust: ```rust use puffinparse_core::{parse, DocumentRequest}; let resp = parse(DocumentRequest::from_path("invoice.pdf").model("extend/parse_performance")).await?; let reducto_shaped = resp.to_format("reducto")?; // serde_json::Value ``` or explicitly: ```rust use puffinparse_core::compat::{render_parse, Format}; let value = render_parse(&resp, Format::Reducto); ``` A request can carry the choice so the SDK/CLI can apply it for you: ```rust let req = DocumentRequest::from_path("invoice.pdf").model("llamaparse/agentic").output_format("reducto"); req.validate_output_format()?; // fails fast on a typo, before any provider call ``` Accepted names (case-insensitive, `-` and `_` interchangeable): | Value | Shape | |---|---| | `puffinparse` (also `unified`), or unset | PuffinParse's own `ParseResponse` JSON | | `reducto` | Reducto `POST /parse` response (`response_type: "parse"`) | | `extend` | Extend `parse_run` object (`GET /parse_runs/{id}`) | | `llamaparse` (also `llama`, `llama_parse`) | LlamaParse `…/result/json` payload | `output_format` is **independent of `output`**: `output` picks markdown vs plain text *inside* block content; `output_format` picks the JSON envelope around it. The core entry points (`parse`, `ocr`, `extract`) always return the unified structs — rendering is a pure function on the result, so nothing about routing, retries, pricing or fallbacks changes. --- ## 2. What is guaranteed **Structural fidelity, not semantic identity.** Guaranteed: - the **key set and nesting** of the vendor's payload, at the envelope, chunk/page and block/item levels; - **one chunk/page per unified page**, and one block/item per unified block, in reading order; - **content strings** (`content` / `md`) byte-identical to what the unified response carries; - **block types** drawn only from the vendor's own vocabulary (tables in §4); - **coordinates in the vendor's units and convention** (§5); - the **billed page count** (`usage.num_pages` / `metrics.pageCount` / `job_metadata.job_pages`). Not guaranteed: - byte equality with what the vendor would have returned for the same document — a different engine produced the text; - fields PuffinParse does not model. They are rendered as `null`, `[]`, `{}` or a stable constant (§3), never invented; - vendor-specific enrichments (chart data, figure crops, OCR word layers, layout add-ons, studio links, presigned URLs). These are always `null`/empty; - confidence semantics. Every vendor scores differently; the number is whatever the *source* provider reported, re-expressed in the target vendor's field. If your integration depends on a field listed as always-null below, the compatibility layer will not carry you — use the unified shape, or `include_raw=True` to get the source provider's own payload alongside. --- ## 3. Per-format field map ### Reducto (`output_format="reducto"`) Rendered envelope: every top-level key of a real `POST /parse` response. | Field | Value | |---|---| | `response_type` | `"parse"` | | `job_id` | source provider's job id, else PuffinParse's response `id` | | `duration` | `latency_ms / 1000` (whole PuffinParse call, including polling) | | `usage.num_pages` / `usage.credits` | unified `usage.pages` / `usage.credits` (`credits` is `null` when the source reports none) | | `result.type` | `"full"` (the URL variant is never emitted) | | `result.chunks[]` | one per page, `chunk_mode="page"` semantics; `content` and `embed` are both the page markdown | | `…blocks[].bbox` | `{left, top, width, height, page, original_page}`, already normalised 0..1 | | `…blocks[].confidence` | `"high"` when confidence ≥ 0.8, `"low"` below, `null` when unknown | | `…blocks[].granular_confidence.parse_confidence` | the numeric confidence, or `null` | | **Always `null`** | `pdf_url`, `studio_link`, `parse_mode`, `document_properties`, `usage.credit_breakdown`, `usage.page_billing_breakdown`, `usage.non_empty_cell_count`, `result.ocr`, `result.custom`, chunk `enriched`, block `image_url`, `chart_data`, `extra`, `granular_confidence.extract_confidence` | | **Always `false`** | chunk `enrichment_success` | `original_page` is set equal to `page`: PuffinParse does not track pre-split page numbers. ### Extend (`output_format="extend"`) | Field | Value | |---|---| | `object` | `"parse_run"` | | `id` | source job id, else PuffinParse's response `id` | | `status` | `"PROCESSED"` (a failure never reaches this code path — it is raised as an error) | | `output.chunks[]` | one per page, `type: "page"`, ids `chunk_<page>` | | `…blocks[]` | ids `block_<page>_<n>`, `details: {}` | | `…blocks[].metadata.page` | `{number, width, height}` in page pixels (§5) | | `…blocks[].boundingBox` | `{left, top, right, bottom}` in page pixels; `null` when the block has no box | | `…blocks[].polygon` | the four corners of that box, clockwise from top-left; `[]` when there is no box | | `…chunks[].metadata` | `pageRange {start, end}` (both the page number), plus `minOcrConfidence` / `avgOcrConfidence` over the page's blocks | | `output.metadata.pages[]` | `{number, rotationApplied: 0, originalPageWidth, originalPageHeight, dpi: null}` | | `metrics` | `{processingTimeMs: latency_ms, pageCount: usage.pages}` | | `config` | `{target: "markdown", chunkingStrategy: {type: "page"}, engine: null}` | | `usage` | `{credits, totalCredits: credits, breakdown: []}` | | `metadata` | `null`, **or** `{"puffinparse_synthetic_page_dims": true}` when page sizes had to be synthesised (§5). The run-level `metadata` map is free-form in Extend's API, so this is a legal place to say so. | | **Always `null`** | `file`, `failureReason`, `failureMessage`, `dataRetention`, `outputUrl`, `batchId`, `config.engine` | | **Absent** | `output.ocr` (the word layer), `config.blockOptions`, `config.advancedOptions`, `block.details` contents | ### LlamaParse (`output_format="llamaparse"`) | Field | Value | |---|---| | `pages[].page` / `text` / `md` | unified page number, text, markdown | | `pages[].width` / `height` | page units (§5) | | `pages[].items[]` | one per block: `{type, md, value, lvl, bBox, layoutAwareBbox: []}` | | `…items[].type` | `heading` \| `text` \| `table` only | | `…items[].lvl` | `1` for a title, `2` for a section header, `null` otherwise | | `…items[].value` | the block's plain-text variant when the source provided one, else `null` | | `…items[].bBox` | `{x, y, w, h, confidence}` in page units; `null` when the block has no box | | `pages[].confidence` | mean of the page's block confidences, or `null` | | `job_metadata` | `{credits_used, job_credits_usage, job_pages, job_auto_mode_triggered_pages: 0, job_is_cache_hit}` | | **Always empty** | `images`, `charts`, `links`, `layoutAwareBbox` | | **Always constant** | `status: "OK"`, `triggeredAutoMode: false`, `noStructuredContent: false`, `noTextContent: false`, `pageHeaderMarkdown`/`pageFooterMarkdown`/`printedPageNumber`: `""`, `parsingMode: null`, `structuredData: null` | | **Absent** | page: `originalOrientationAngle`, `layout`, `costOptimized`, `slideSpeakerNotes`, `slideSectionName` (add-ons PuffinParse does not model); table items: `csv`, `html`, `rows`, `isPerfectTable` (restatements of the markdown table already in `md`) | Figures have no LlamaParse item type (LlamaParse puts them in `images[]`/`charts[]`, which we cannot synthesise), so a figure block is emitted as a `text` item carrying its markdown rather than being dropped. --- ## 4. Block-type mapping Unified → native. Cells marked **lossy** have no exact counterpart in that vendor's vocabulary. | Unified | Reducto | Extend | LlamaParse | |---|---|---|---| | `title` | `Title` | `heading` | `heading` (`lvl: 1`) | | `section_header` | `Section Header` | `section_heading` | `heading` (`lvl: 2`) | | `text` | `Text` | `text` | `text` | | `list` | `List Item` | `text` **lossy** | `text` **lossy** | | `table` | `Table` | `table` | `table` | | `figure` | `Figure` | `figure` | `text` **lossy** | | `header` | `Header` | `header` | `text` **lossy** | | `footer` | `Footer` | `footer` | `text` **lossy** | | `footnote` | `Text` **lossy** | `text` **lossy** | `text` **lossy** | | `caption` | `Text` **lossy** | `text` **lossy** | `text` **lossy** | | `formula` | `Text` **lossy** | `formula` | `text` **lossy** | | `other` | `Text` **lossy** | `text` **lossy** | `text` **lossy** | Two consequences worth stating plainly: - **The mapping is not injective, so it is not invertible.** Extend's `key_value` normalises to `text` and comes back as `text`, not `key_value`. Round-trip tests therefore compare types *up to the provider's own forward mapping* (§7). - **Reducto's render deliberately uses a reduced vocabulary.** It emits only `Title`, `Section Header`, `Text`, `List Item`, `Table`, `Figure`, `Header`, `Footer`. Reducto itself also emits `Footnote`, `Caption`, `Formula`, `Page Number` and others (`docs/providers/reducto.md` §4); a Reducto → Reducto round trip over a document containing those blocks will see them arrive as `Text`. Widening the reverse map is a one-line change in `crates/puffinparse-core/src/compat/reducto.rs::block_type` if that fidelity is wanted. --- ## 5. Coordinates The unified `BBox` is `{x0, y0, x1, y1}` normalised to 0..1 with a top-left origin (ADR-3). Each renderer converts to the vendor's convention: | Format | Units | Conversion | |---|---|---| | Reducto | normalised 0..1, `left/top/width/height` | `left = x0`, `top = y0`, `width = x1 - x0`, `height = y1 - y0` — no page size needed | | Extend | page pixels, `left/top/right/bottom` | multiply by the page's width/height | | LlamaParse | page units, `x/y/w/h` | multiply by the page's width/height | **The synthetic-page rule.** Extend and LlamaParse express boxes in page units, so a page size is required. When the unified response carries `Page.width`/`height` (Extend and LlamaParse report them; Reducto and the vision-LLM providers do not) those are used verbatim. When it does not, the renderer assumes a **1000 × 1000** page. That keeps the numbers readable and makes the original normalised coordinates exactly recoverable by dividing by 1000 — but they are not real page dimensions, and the Extend render says so via `metadata.puffinparse_synthetic_page_dims = true`. The constant is `puffinparse_core::compat::SYNTHETIC_PAGE_DIM`. A block with no box at all renders as `{left: 0, top: 0, width: 0, height: 0, page, original_page}` for Reducto (whose `bbox` is not nullable in practice) and as `null` for Extend (`boundingBox`) and LlamaParse (`bBox`), matching each vendor's own optionality. --- ## 6. Migration examples ### Extend → Reducto You read Reducto's shape today and want Extend's engine. ```python # before resp = reducto_client.parse.run(document_url=url) for chunk in resp.result.chunks: for block in chunk.blocks: draw(block.bbox.left, block.bbox.top, block.type) # after — same parsing code, Extend doing the work doc = puffinparse.parse(url, model="extend/parse_performance", output_format="reducto") for chunk in doc["result"]["chunks"]: for block in chunk["blocks"]: draw(block["bbox"]["left"], block["bbox"]["top"], block["type"]) ``` What changes: `pdf_url`, `studio_link` and the billing breakdowns are `null`; Extend's `key_value` blocks arrive as `Text`; boxes are Extend's pixel boxes divided by the page size, so they line up with the same page image. ### Reducto → Extend ```python doc = puffinparse.parse(path, model="reducto/standard", output_format="extend") assert doc["object"] == "parse_run" and doc["status"] == "PROCESSED" for chunk in doc["output"]["chunks"]: text = chunk["content"] page = chunk["metadata"]["pageRange"]["start"] ``` What changes: Reducto reports no page dimensions, so the render uses a 1000 × 1000 page and sets `metadata.puffinparse_synthetic_page_dims = true`. `boundingBox` values are therefore *relative* coordinates × 1000, not PDF points. If you overlay boxes on a rendered page image, scale by `your_image_size / 1000` — or scale by `boundingBox / page.width`, which is correct in both cases and is the recommended form. ### LlamaParse → Reducto ```python doc = puffinparse.parse(path, model="llamaparse/agentic", output_format="reducto") pages = {b["bbox"]["page"] for c in doc["result"]["chunks"] for b in c["blocks"]} ``` What changes: LlamaParse items become Reducto blocks (`heading` + `lvl:1` → `Title`, `lvl:2` → `Section Header`, `table` → `Table`); boxes are divided by LlamaParse's real page size, so they land in Reducto's normalised 0..1 space; `job_metadata.credits_used` becomes `usage.credits` (`null` when LlamaParse has not settled billing yet). --- ## 7. Extract mode (best effort) `ExtractResponse::to_format(...)` / `compat::render_extract(...)` render the extract envelope of each vendor. This is explicitly **weaker** than the parse renderers: extract surfaces differ far more between vendors (schemas, citation objects, per-field metadata placement), and there are no captured extract fixtures for all three providers yet, so only the documented envelope and the citation/confidence placement are reproduced. | Format | Envelope | |---|---| | `reducto` | `{response_type: "v3_extract", job_id, usage: {num_pages, num_fields, credits}, result, studio_link: null}` — `result` rebuilds Reducto's `{value, citations}` leaf wrappers from the unified pointer-keyed `fields`, the exact inverse of the unwrapping done when normalising | | `extend` | `{object: "extract_run", id, status: "PROCESSED", output: {value, metadata: {<field>: {confidence, citations}}}, usage}` | | `llamaparse` | `{data, extraction_metadata: {field_metadata: {<field>: {confidence, citations}}, job_id}}` (LlamaExtract) | Caveat: an extract response carries no page dimensions, so citation boxes stay in the **unified normalised 0..1 space** for the Extend and LlamaExtract renders instead of being scaled to page units. Field keys come from the unified JSON pointers (`/invoice/total`) with the leading slash stripped; LlamaExtract uses dots (`invoice.total`). --- ## 8. How this is validated `crates/puffinparse-core/src/compat/roundtrip.rs` runs, for each of the three providers: ```text real fixture → provider::normalize → ParseResponse → render_parse(same format) → compare ``` The comparison is a `skeleton_diff` that reports the first mismatch with a path (`unit[1].block[0].content differs: …`) and checks: 1. top-level key set, exactly; 2. chunk/page and block/item key sets (no key the fixture has may be missing, except the documented absences in §3); 3. billed page count; 4. chunk/page count, then block/item count per unit; 5. `content` / `md` strings, trimmed; 6. block types, compared after mapping both sides through the provider's own forward map (so `key_value` → `text` counts as faithful); 7. boxes, converted back to the unified 0..1 space and compared within `1e-6`. Relaxations, and why: - **Trimmed string comparison.** `providers::llamaparse::normalize` trims page markdown, so the fixture's trailing newlines do not survive. Content is compared trimmed. - **Block types up to the forward map.** The mapping is not injective (§4). - **`originalPageWidth` is not compared.** Extend reports both a PDF-point page size in `output.metadata.pages[]` (596 × 842) and a rasterised pixel size in `blocks[].metadata.page` (1241 × 1754). The unified `Page` keeps the one the boxes are in (pixels), so the render emits that in both places. - **Cross-format tests do not compare block types**, since vendors share no vocabulary. Alongside those, the suite asserts that every render **deserialises with the target provider's own wire types** (`WireParseResponse`, `ParseRun`, `JsonResult`), that re-normalising a render and rendering again is a fixed point, that every `Format::ALL` renders a response whose pages have no dimensions and whose blocks have no boxes without panicking, and that `skeleton_diff` itself detects each class of mismatch it claims to. --- # Decisions <!-- source: docs/DECISIONS.md | url: https://puffinparse.com/docs/project/decisions/ --> Architecture decision records, newest last. One entry per decision that someone might otherwise re-litigate. Format: context → decision → consequences. Add an entry when you change direction; do not edit old entries, supersede them. ## ADR-1: Rust core with a thin Python wrapper **Context.** The brief asks for "fastest and litest" and a Python SDK. Provider calls are HTTP + JSON; the benchmark needs Levenshtein over long documents. **Decision.** All provider logic, normalisation, routing, pricing and metrics live in the `liteocr-core` crate. Python (`python/liteocr`) is a typed wrapper over a PyO3 module (`liteocr._core`) and holds no provider logic. The CLI is a separate crate on the same core. **Consequences.** One implementation to test per provider; a future gateway server and a TypeScript SDK reuse the core. Cost: Python contributors need a Rust toolchain (`maturin develop`), and wheels must be built per platform (handled by `release.yml`). ## ADR-2: `"<provider>/<model>"` model strings, LiteLLM style **Decision.** The `model` string selects provider and quality tier (`reducto/r-1`, `llamaparse/agentic`). A bare provider name selects its default. Provider-specific knobs go in `provider_options` and are merged verbatim into the provider request. **Consequences.** Switching provider is a one-string change; the response shape never depends on options. The registry in `crates/liteocr-core/src/model.rs` is the only place models are declared; pricing is keyed by the same string. ## ADR-3: Normalised, top-left-origin bounding boxes **Decision.** Every block bbox is `{x0, y0, x1, y1}` in `0..=1` of the page. Reducto already reports normalised boxes; Extend and LlamaParse boxes are divided by the page dimensions they report. Boxes are `None` when dimensions are unknown. **Consequences.** Consumers can draw overlays without knowing page size or DPI. Precision of Extend's DPI-scaled coordinates is preserved because the division uses the same units. ## ADR-4: Extend always via `/parse_runs`; Reducto sync by default **Decision.** Extend's sync `/parse` has a 5-minute hard limit and its docs call it "for onboarding"; LiteOCR always creates a run and polls. Reducto's sync `/parse` allows 15 minutes and returns inline, so it is the default; `provider_options.async = true` switches to `/parse_async` + `/job/{id}`. **Consequences.** One code path per provider is exercised by default; both are covered by the same unified deadline (`timeout`), so polling never exceeds the caller's budget. ## ADR-5: Page reconstruction rules **Decision.** Pages are the unit of the unified response. Reducto is asked for `chunk_mode=page` and pages are derived from `blocks[].bbox.page`; Extend uses page chunks and `metadata.page.number`; LlamaParse's JSON result is already per page. Chunk-level markdown from the provider is preferred over re-joining blocks, so the page markdown is what the provider itself renders. **Consequences.** Benchmarks compare provider-rendered markdown, not a LiteOCR re-rendering. ## ADR-6: Cost is computed from a list-price table, not from provider credits **Decision.** `cost_usd = pages × per_page_usd` from `pricing.json`, overridable at runtime. Provider credits are passed through in `usage.credits` but not converted: Reducto credits are null on new pricing plans, Extend credits depend on the plan's $/credit, and LlamaParse reports 0 credits until billing settles. **Consequences.** Costs are comparable across providers and stable, but are list prices; the benchmark labels them as such. ## ADR-7: Benchmark ground truth is exact by construction; metrics are deterministic **Decision.** `synthetic-v1` renders documents from the same source as the truth markdown, so there is no annotation noise. Metrics are text-only (char similarity, CER, WER, word F1, order, table) computed in Rust. No LLM judge is required for the leaderboard. **Consequences.** Fully reproducible and cheap to run, but synthetic documents are cleaner than real scans; the combined open dataset (see TASKS) is the answer to that, not a looser scorer. ## ADR-8: Benchmark runs disable provider result caches **Context.** LlamaParse caches results for 48 hours; a re-run returned in ~450 ms instead of ~9 s and would have misreported latency. **Decision.** `bench run` passes cache-busting options per provider (`do_not_cache` + `invalidate_cache` for LlamaParse) unless `--allow-cache` is given. **Consequences.** Latency in the leaderboard reflects real processing. Repeated runs cost real credits. ## ADR-9: (withdrawn) Repository housekeeping that no longer applies. ## ADR-10: Combined open benchmark instead of a new vendor benchmark **Context.** Each vendor publishes a benchmark it wins (ParseBench, RealDocBench, LongExtractBench). **Decision.** LiteOCR will not author a competing "vendor-neutral" opinion benchmark. It will run every public benchmark through one harness with one scoring pipeline and publish per-dataset and combined scores, with every input, truth and output inspectable online. Datasets whose license permits redistribution are vendored in the repo; others are downloaded by an adapter at run time with the upstream revision pinned in the manifest. **Consequences.** The site is the verification surface; the CLI is the reproduction surface. Some benchmarks (olmOCR-bench) need their own scorer implemented rather than transcript similarity. ## ADR-11: Modes — providers are only interchangeable within a mode **Context.** "OCR provider" covers different products: plain text recognition (Textract DetectDocumentText, Azure Read, Google Document OCR), layout-aware parsing to markdown and typed blocks (Reducto, Extend, LlamaParse, Mistral OCR, Datalab, Unstructured, Upstage, Landing AI), and schema-driven structured extraction (Reducto Extract, Extend Extract, LlamaExtract, Azure prebuilt models, Textract Queries/Forms, vision LLMs with structured output). Swapping a parse model for an extract model is not a like-for-like switch. **Decision.** The core defines `Mode::{Parse, Ocr, Extract}` with its own request/response type per mode (`DocumentRequest → ParseResponse`, `DocumentRequest → TextResponse`, `ExtractRequest → ExtractResponse`) and its own entry point (`parse`, `ocr`, `extract`). Every registry model declares the modes it supports; resolution (`ModelRef::parse_for`) and the `Router` reject models outside the requested mode. Pricing is per model *and* mode. A layout provider serves `ocr` by deriving lines/words from its parse output (marked with `liteocr_derived_from = "parse"`), so plain-text callers can still use it, but native OCR endpoints override that when they exist. **Consequences.** The Python API becomes `liteocr.parse` / `liteocr.ocr` / `liteocr.extract` (the pre-release `liteocr.ocr` that meant parse is renamed; nothing was published). Vision LLMs (Gemini, OpenAI, Anthropic) fit as parse/extract models without geometry, which the response makes explicit by returning blocks without boxes. Adding a fourth mode (classify, split) is additive: a new enum variant, types, trait method and entry point. ## ADR-12: Provider fan-out rules **Decision.** Each provider lives in one file (`crates/liteocr-core/src/providers/<name>.rs`), is declared up front in `providers/mod.rs`, and is implemented independently against the `Provider` trait. Shared registration points (`model.rs` registry, `pricing.json`, `build()`, README table, `.env.example`) are edited only by the integrator after a provider lands, from snippets in the provider's hand-off. Providers without a key in the environment are implemented from official docs with fixtures built from documented responses and `#[ignore]` live tests; the task board records which ones have been verified live. **Consequences.** Many providers can be built in parallel without merge conflicts; the cost is a short integration step per provider and, for unverified providers, a "docs-only" label until someone with a key runs the live tests. ## ADR-13: Native-format compatibility is a renderer over the unified response **Context.** The unified response is the product, but it is also the migration cost: a team already parsing Reducto's `result.chunks[].blocks[].bbox.left` cannot try another provider without rewriting the code that reads the result. The one thing that would make switching free is getting answers back in the shape they already parse. **Decision.** Add `crates/liteocr-core/src/compat/`: a pure, infallible renderer `render_parse(&ParseResponse, Format) -> serde_json::Value` (plus a best-effort `render_extract`) with one module per vendor shape — `Format::{Liteocr, Reducto, Extend, LlamaParse}`. It is applied *after* a call, not inside it: `parse`/`ocr`/`extract` keep returning the unified structs, `DocumentRequest.output_format` only records and validates the caller's choice (`validate_output_format`), and the SDK/CLI call `ParseResponse::to_format(&str)` on the result. So routing, retries, pricing, fallbacks and every existing test are untouched. What is promised is **structural fidelity, not semantic identity**: the vendor's key set, nesting, chunk/page and block counts, content strings, block-type vocabulary and coordinate units. Fields LiteOCR does not model are rendered `null`/empty and enumerated in `docs/COMPAT.md`, never invented. Extend and LlamaParse need a page size for their unit boxes; when the source provider reports none (Reducto, vision LLMs) the renderer assumes a 1000×1000 page and, for Extend, records `metadata.liteocr_synthetic_page_dims = true` in the run's free-form metadata map. The claim is enforced rather than asserted: for each provider, `compat/roundtrip.rs` runs `fixture → provider::normalize → render_parse(same format)` and compares against the original fixture with a `skeleton_diff` (key sets, counts, contents, types up to the provider's own forward mapping, boxes within 1e-6, billed pages), plus cross-format and no-geometry cases. The tests live in the crate because the `normalize` functions are `pub(crate)`. **Consequences.** Adding a fourth shape is one module plus one enum variant. Lossy edges are real and documented: type mappings are not injective (Extend's `key_value` returns as `text`), the Reducto render uses a reduced vocabulary (`footnote`/`caption`/`formula`/`other` → `Text`), and `render_extract` is explicitly weaker than the parse path until extract fixtures exist for all three providers. Because the renderer only reads the unified types, any provider added later gets all three native shapes for free — and any unified field a new provider cannot fill shows up as a `null` in someone's vendor-shaped payload, which is the honest outcome. ## ADR-14: The Node.js SDK is a napi-rs addon over the same core **Context.** A TypeScript SDK was on the roadmap once the Python surface stabilised. The options were a WASM build, a wrapper around the CLI, or a native N-API addon. **Decision.** `crates/liteocr-node` exposes the core through napi 3; the `js/` package is a thin layer. The addon only converts values: requests and responses cross as the core's serde JSON, errors as a prefixed JSON payload (`LITEOCR_CORE_ERROR:<json>`) that the JS layer rebuilds into typed `LiteOCRError` subclasses. The JS layer converts to camelCase with explicit per-type converters; `data`, `metadata` and `raw` are never renamed, and `index.d.ts` is hand-written against SPEC §5 (the napi-generated `native.d.ts` is internal). The crate is a normal workspace member: napi's `dyn-symbols` keeps `cargo test --workspace` free of Node. It uses `deny(unsafe_code)` because napi's macro expansion is incompatible with `forbid`, and declares `rust-version = "1.88"` (napi 3) while the rest of the workspace stays at 1.80. `timeout` is in seconds, as in the other SDKs, and `LiteOCRError.kind` uses the ErrorKind serde values. **Consequences.** One implementation behind three surfaces (Python, Node, CLI). No provider logic in JS. Prebuilt binaries per platform are required for `npm install` without a Rust toolchain; the release matrix builds them but publishing is not wired yet. ## ADR-15: The gateway is a thin axum layer over the core, TOML-configured, with no database **Context.** LiteLLM's proxy is what teams actually deploy: one endpoint, central keys, budgets and logs. LiteOCR needed the same without growing a second implementation of providers. **Decision.** `crates/liteocr-server` (axum + tower-http, which are HTTP frameworks, not provider SDKs) calls `liteocr_core::{parse, ocr, extract}` and runs its own fallback loop, because the core `Router` cannot give each target its own credentials or base URL. Configuration is TOML (already idiomatic in Rust, lighter than YAML). Usage and budget state live in memory with an optional JSON state file; a database is out of scope. Clients may not send `api_key` or `base_url`, so they cannot redirect the gateway's credentials, and local file paths are refused. Provider credential failures map to 502, not 401, because the caller's own key was valid. **Consequences.** Single binary, no infrastructure to run. Budgets can be overshot by requests in flight, and key changes need a restart. A database, HTTP key management, async job endpoints and metrics auth are follow-ups, not blockers. ## ADR-16: Vendor only what the licence permits; index the rest **Context.** The combined benchmark (ADR-10) pulls in public datasets whose licences differ: olmOCR-bench is ODC-BY-1.0, OmniDocBench has no licence and is marked research-only / non-commercial. **Decision.** A dataset is vendored into the repo (with attribution) only when its licence permits redistribution. Otherwise the repo holds a manifest with upstream paths, a pinned revision and image/truth hashes, and the adapter materialises the data locally. Combined datasets are versioned and never rewritten once results exist (`combined-v2` supersedes `combined-v1` for new runs). Upstream tests that cannot be expressed faithfully in the shared rule schema are skipped and counted, never weakened silently; the one relaxed mapping (olmOCR "left/right of" → same row) is counted as relaxed. **Consequences.** Anyone can reproduce every score, but index-only sources need a fetch step before a run (their documents carry a `fetch-required` tag). Stats files make the coverage of each conversion auditable. ## ADR-17: The benchmark results viewer stays vanilla JS **Context.** TASKS listed an open choice for `benchmark/site/` between a React app on Extend UI (PDF viewer and layout overlays out of the box) and the zero-build vanilla viewer. The viewer has to be where every benchmark claim can be checked: page rendering, side-by-side outputs, diffs, rule checklists, bbox overlays, charts, deep links. **Decision.** Keep static HTML/CSS/ES2018 in `benchmark/site/src/`: no framework, no bundler, no npm install; the build stays stdlib Python (Pillow optional). The one third-party runtime dependency is pdf.js, loaded lazily from cdnjs at a pinned version with SRI, only when a PDF is opened, with the build-time PNG as fallback. Overlays and charts are inline SVG. The per-document rule checklist uses a JS port of `liteocr-core`'s `score_rules`, and every rules page compares its count with the recorded Rust score and flags any disagreement. **Consequences.** Vercel and Pages build the site with one `uv run` command and nothing to audit. The data contract (`data/index.json`, `data/runs/`, `data/outputs/…/<doc>.{md,json}`) is independent of the front-end, so this can be revisited without touching the build. The JS port must follow scorer changes in `bench.rs`; the mismatch badge makes drift visible. ## ADR-18: Webhooks are exposed as primitives, not received **Context.** TASKS asked for "webhooks instead of polling". An SDK cannot host an HTTP endpoint, and providers differ: Reducto and LlamaParse accept a per-job webhook URL, Extend only has workspace-level webhook endpoints. **Decision.** LiteOCR offers `submit_parse` / `retrieve_parse` plus a pure `parse_webhook` / `resolve_webhook` that normalises a provider's webhook body into `JobStatus` (doing one retrieve when the body only names the job). `JobHandle` is serialisable and secret-free; credentials are resolved again at retrieve time. `DocumentRequest.webhook_url` maps to per-job provider webhooks where they exist and is rejected with an input error where they don't. Only `parse` mode has jobs for now. **Consequences.** Users wire LiteOCR into their own web handler; the gateway (ADR-15) can later add job endpoints on top of the same primitives. Webhook body shapes come from vendor docs until real deliveries are captured. ## ADR-19: Local and self-hosted engines are out-of-process providers **Context.** SPEC §1 listed local models as a v0.1 non-goal, but LiteOCR was unusable without a paid key and the benchmark had no open baseline. **Decision.** Support local and self-hosted engines only as providers that call out of process: a CLI binary via `tokio::process` (Tesseract, with `pdftoppm` for PDFs) or an HTTP server the user runs (docling-serve, PaddleOCR/PaddleX serving). No C bindings, FFI, embedded runtimes or model weights ship in LiteOCR. They are listed in `model::SELF_HOSTED`, need no key, and are priced at 0.0 with source "self-hosted". **Consequences.** `#![forbid(unsafe_code)]` and the no-SDK rule still hold. Benchmark latency for these engines depends on the user's hardware, so leaderboard rows need a hardware note. Live tests need a binary or a server rather than a key. ## ADR-20: The scorer is versioned; committed runs are re-scored offline **Context.** The first combined run exposed scorer defects (HTML tables scored 0, tokenised punctuation in rule text, a dead `bag_of_sentences` threshold) that moved the leaderboard by up to four points. Re-calling providers to fix a scoring bug would cost money and change latency and provider versions at the same time. **Decision.** Any scorer change that moves scores bumps `SCORER_VERSION` in `liteocr-core` and re-scores committed runs from their saved outputs with `liteocr bench rescore`, never by re-calling providers; measured latency and cost are kept. Result JSON records `scorer_version` (absent = 1) and `rescored_at`. The viewer's JS rule checker follows every scorer version and flags any document where it disagrees with the recorded Rust score. The `bag_of_sentences` threshold of 0.8 is justified by sentences the reference extraction fused together. The table structure metric is TEDS on the row/cell grid and is named `teds_grid`, not TEDS, because the truth has no header/body or span structure. **Consequences.** Leaderboards stay comparable across scorer fixes at no cost, and a reader can see which scorer produced a number. Saved outputs are part of every committed run (already required). ## ADR-21: Gateway jobs are owned, not routed **Context.** ADR-18 left gateway job endpoints for later; ADR-15 gave the gateway fallbacks and per-key budgets. **Decision.** `/v1/jobs` submits to exactly one deployment (an alias's first target, or the next in a round-robin rotation) with no fallback: failures only surface at retrieve time, and retrying would mean resubmitting. Jobs are stored under opaque gateway ids bound to the submitting key (the master key can read all, as with `/v1/usage`), without secrets; provider credentials are resolved from config on every poll. Cost is charged once, on the first observed success, under the usage lock. Provider webhooks are received only when the operator opts in with a shared secret. **Consequences.** Clients poll the gateway, never the provider. Providers can reuse job ids (LlamaParse returns the cached job for an identical upload), so one webhook may settle several gateway jobs. Vendor HMAC verification and automatic webhook registration remain open. ## ADR-22: One design language; the viewer renders PDF pages at build time **Context.** The docs, landing page and viewer each had their own palette (blue in the docs, vermilion on the landing page), and the viewer showed raw ids, eleven metric tiles per document and full-table heatmaps; PDF pages depended on pdf.js from a CDN and often did not render. **Decision.** `website/assets/tokens.css` is the single source of colour, type, space and shape, described in `docs/DESIGN.md`; every surface links it first and styles only through its custom properties. The viewer follows the language's principles (one number per view, evidence one click away, human names with raw ids on request) and renders PDF pages to WebP at build time with pypdfium2 (a pip wheel that works on Vercel); pdf.js remains only as a lazy fallback. Layout-box overlays are off by default. **Consequences.** A palette or type change is one edit. The site build needs pypdfium2 (added to `vercel.json` and `pages.yml`); without it the build falls back to Pillow page-1 extraction. ## ADR-23: The product is renamed PuffinParse **Context.** `liteocr` on PyPI belongs to an unrelated OCR engine, several GitHub projects already use the name, and "OCR" undersells a tool whose modes are parse, OCR and extract. The owner's preference, LiteParse, is a LlamaIndex product. About 110 names were checked against domains, PyPI/npm/crates.io and web collisions. **Decision.** The product, crates (`puffinparse-{core,cli,python,node,server}`), Python package (`puffinparse`, native module `puffinparse._core`), Node package, CLI binary, environment variables (`PUFFINPARSE_*`), metadata keys (`puffinparse_*`), error base class (`PuffinParseError`), the native output format value (`"puffinparse"`) and the site (`puffinparse.vercel.app`) all use the new name. PuffinParse had every checked domain (.com, .dev, .ai, .io) and every registry name free and no product collision; the puffin's black, white and orange match the existing palette. **Consequences.** Committed benchmark results, recorded provider fixtures (which contain "LiteOCR" in document text), the CHANGELOG history and earlier ADRs keep the old name. Result files written before the rename carry `liteocr_version`, which the CLI and site builder read as an alias. `liteocr.vercel.app` keeps serving the same project. ## ADR-24: A puffin mark and mascot; brand imagery only in brand moments **Context.** After the rename (ADR-23) the product had no logo: each surface improvised a different mark (a scan line, an orange square, plain text), and the favicon was still the pre-ADR-22 blue. The puffin's black, white and orange already matched the palette. **Decision.** Two brand assets. The **mark** (`website/assets/mark.svg`): a puffin head in a rounded-square tile, the beak carrying two stripes read as parsed lines; it is the favicon and sits beside the wordmark on every surface. The **puffin** (`website/assets/puffin.svg`): a mascot holding three document fish, used only in brand moments (landing hero, 404, empty states, social card, README), never beside data in the tools. Both are flat SVG in existing tokens plus three brand tokens (`--mark-tile`, `--mascot-body`, `--mascot-wing`) that lift in dark mode. A committed social card (`og.png`) is the `og:image` for every page. Alongside, two table rules tightened: headers are sentence case, and bold marks a best value only when it is unique as displayed. **Consequences.** "No decorative imagery" now reads "none in the tools". Inline copies of the mark live in `website/build.py` (`MARK`) and the viewer's `index.html`; the social card must be re-rendered (`website/og/render.py`, Playwright) when the card, mark or mascot changes. The launch videos still show the text wordmark until they are re-rendered. *Amended 2026-09-25:* the mark is C2 of the explored set (kept after comparing six alternatives), and the mascot became the waving puffin with three pages in its beak (instead of three document fish), which says "documents" more directly. `docs/DESIGN.md` gained a Brand motion section. --- # Contributing <!-- source: CONTRIBUTING.md | url: https://puffinparse.com/docs/project/contributing/ --> Thanks for helping build PuffinParse — one API for every OCR / document-parsing provider. This document covers local setup, the checks CI runs, and the contributions we get asked about most: **verifying or adding a provider** and **adding a benchmark dataset**. If you use a coding agent, point it at [AGENTS.md](https://github.com/ajinkyashejul/puffinparse/blob/main/AGENTS.md), which summarises the same rules. ## Good first issues Issues labelled [`good first issue`](https://github.com/ajinkyashejul/puffinparse/labels/good%20first%20issue) are scoped to one area and list a "done when" condition. Comment on the issue to claim it so two people don't do the same work. Another useful first contribution if you have a key for one of the **docs-only** providers (Mistral, Azure, Textract, Gemini, OpenAI, Anthropic, Mathpix, Datalab, Unstructured, Upstage, Landing AI, Google Document AI, PaddleOCR) is running its live tests and reporting what differs; see [issue #10](https://github.com/ajinkyashejul/puffinparse/issues/10). Questions and ideas that are not yet issues go to [Discussions](https://github.com/ajinkyashejul/puffinparse/discussions). By participating you agree to the [Code of Conduct](https://github.com/ajinkyashejul/puffinparse/blob/main/CODE_OF_CONDUCT.md). PuffinParse is MIT licensed; contributions are accepted under the same license. --- ## 1. Setup Prerequisites: - **Rust stable** (≥ 1.80 — the workspace `rust-version`), via [rustup](https://rustup.rs). - **Python 3.9+** (3.9 is the minimum we build abi3 wheels for). - A C toolchain (whatever `cc` your platform ships) for the native deps. - **Node.js 18+** only if you work on the TypeScript SDK in `js/`. ```bash git clone https://github.com/ajinkyashejul/puffinparse cd puffinparse # Rust side cargo build --workspace # Python side — use a virtualenv; maturin installs into the active one. python -m venv .venv && source .venv/bin/activate pip install maturin ruff mypy pytest maturin develop # builds crates/puffinparse-python, installs `puffinparse` ``` `maturin develop` compiles the PyO3 extension (`puffinparse._core`) and links it against the pure-Python package in `python/puffinparse`, so edits to the Python sources take effect immediately; re-run it after changing any Rust code. Use `maturin develop --release` when you care about speed (benchmarks, large documents) — debug builds of the core are slow. Provider keys go in a `.env` (copy `.env.example`) or in your shell: ```bash export REDUCTO_API_KEY=... export EXTEND_API_KEY=... export LLAMA_API_KEY=... # LlamaCloud / LlamaParse, starts with llx- ``` **You do not need any keys to contribute.** Every test that talks to a live provider skips itself when the relevant key is absent, and CI sets no secrets. Live tests are opt-in: `PUFFINPARSE_LIVE_TESTS=1 pytest python/tests -q` for Python and `cargo test -p puffinparse-core -- --ignored` for Rust. Never commit a `.env` file or paste a key into an issue, log or fixture. ### Debugging Set `PUFFINPARSE_LOG` to turn on the core's `tracing` output — request URLs, retry and backoff decisions, poll loops, router fallbacks: ```bash PUFFINPARSE_LOG=debug puffinparse parse invoice.pdf --model reducto/standard PUFFINPARSE_LOG=puffinparse_core::providers=trace pytest python/tests -q ``` On the Python side the same events also surface through `logging.getLogger("puffinparse")`. --- ## 2. Tests, lint, formatting Run everything the CI runs with `make lint test` (and `make test-node` for the TypeScript SDK), or piecemeal: ```bash # Tests cargo test --workspace pytest python/tests -q cd js && npm ci && npm run build:debug && npm test # Node SDK # Lint / format cargo fmt --all # or `cargo fmt --all --check` to only verify cargo clippy --workspace --all-targets -- -D warnings ruff check python/ benchmark/ examples/ ruff format --check python/ benchmark/ examples/ mypy python/puffinparse ``` Rules of the road: - `cargo fmt` output is authoritative; do not hand-format around it. - Clippy warnings are errors in CI. Prefer fixing over `#[allow]`; if an `#[allow]` is genuinely right, put a one-line comment saying why. - **No network in unit tests.** Provider parsing is tested against recorded JSON in `crates/puffinparse-core/tests/fixtures/`. Live tests are `#[ignore]`d and/or key-gated. - `puffinparse-core` is `#![forbid(unsafe_code)]`. Keep it that way. - The Python package is fully typed and `mypy --strict`-clean; new public API needs annotations and a docstring. --- ## 3. Adding a provider This is the highest-value contribution. A provider is a single file implementing one trait. Check for an existing [`new provider` issue](https://github.com/ajinkyashejul/puffinparse/issues) first, or open one from the template so we can agree on model naming before you write code. 1. **Implement the trait.** Create `crates/puffinparse-core/src/providers/<name>.rs` and implement `OcrProvider` (see `crates/puffinparse-core/src/provider.rs`). Use the shared HTTP helpers in `src/http.rs` so you inherit retries, backoff, deadlines and error classification — do not build your own `reqwest::Client`, and do not vendor a provider SDK. Map the provider's response onto the unified `OcrResponse` / `Page` / `Block` / `Usage` types in `src/types.rs`: - normalise `bbox` to 0..1 with a top-left origin; - map the provider's block vocabulary onto `BlockType`, unknown → `other`; - fill `Usage.pages` with the *billed* page count; - map provider errors onto the `Error` variants in `src/error.rs` (401/403 → auth, 429 → rate limit, 5xx / failed job → provider error); only `ProviderError`, `RateLimitError` and `TimeoutError` are fallback-eligible in the router, so classify carefully. 2. **Register it.** Add the module and a `match` arm to `build()` in `crates/puffinparse-core/src/providers/mod.rs`. 3. **Declare its models.** Add a `ProviderInfo` entry to `PROVIDERS` in `crates/puffinparse-core/src/model.rs`: `name`, `display_name`, `env_var`, `base_url`, `docs`, and one `ModelInfo` per mode with exactly one `default: true`. Model strings are `"<provider>/<model>"`; keep them short, lowercase and stable — they are public API. 4. **Add pricing.** Add `"<provider>/<model>"` entries to `crates/puffinparse-core/src/pricing.json` with `per_page_usd`, a `source` URL pointing at the public pricing page, and the `updated` date. Public list prices only. 5. **Add a fixture + normalisation test.** Save one real (redacted) response as `crates/puffinparse-core/tests/fixtures/<name>_<endpoint>.json` and add a unit test that parses it and asserts the normalised output: page count, block types, a bbox inside 0..1, `usage.pages`, and that the document-level `markdown` is the pages joined in order. Scrub keys, job ids, customer names and anything else non-public from the fixture. 6. **Document it.** Add `docs/providers/<name>.md` (same sections as the existing pages) with a status banner, a row in `docs/providers/README.md`, the API-key env var (and any `*_BASE_URL` override) in `.env.example`, the model table in `README.md`, and a line in the `## [Unreleased]` section of `CHANGELOG.md`. Label the provider **live-verified** only if its live tests passed against the real API; otherwise it is **docs-only**. 7. **Benchmark it.** If you have keys, run the benchmark with `--save-outputs`, commit the result JSON under `benchmark/results/` and the per-document outputs under `benchmark/results/outputs/<run_id>/`, and add a section to `benchmark/LEADERBOARD.md` (it is stitched per dataset by hand until [#13](https://github.com/ajinkyashejul/puffinparse/issues/13) lands). ### Verifying a docs-only provider Run its `#[ignore]`d live tests with your key (`cargo test -p puffinparse-core <provider> -- --ignored --nocapture`), fix any wire-format differences, replace the hand-built fixture with a redacted real response, and change the label to live-verified in `docs/providers/README.md`, the provider page's banner and `README.md`. One PR per provider. The Python SDK needs no changes: it forwards whatever model string the core accepts. ### Local and self-hosted engines Engines that run on the user's machine or their own server (`tesseract`, `docling`, `paddleocr`) follow the same steps, with these differences: - **No API key.** Add the provider name to `SELF_HOSTED` in `model.rs`; set `env_var` to `""` (or to an *optional* key, as Docling does) and `base_url` to the local default. `puffinparse providers` then shows `local` in the Key column instead of a missing-key cross. Read the base URL from `<NAME>_BASE_URL` via `provider::resolve_base_url`. - **Price 0.** `pricing.json` gets `0.0` for each mode with `"source": "self-hosted (...)"`. - **Local binaries** are run with `tokio::process` (never C bindings or FFI), with the call's deadline and `kill_on_drop`. A missing binary must produce an error that names the binary and how to install it or point at it (`TESSERACT_CMD`). Shared helpers (download a URL input, base64, file-type sniffing, a self-cleaning scratch directory) are in `crates/puffinparse-core/src/providers/local.rs`. - **Fixtures.** Capture a real output from a local install where you can (the Tesseract TSV and docling-serve fixtures are real); otherwise shape it from the server's documented schema and mark the doc page *docs-only*. The `#[ignore]`d live test needs the binary or server rather than a key. - **Document** how to install or start the engine in `docs/providers/<name>.md`. --- ## 4. Adding a benchmark dataset Datasets live in `benchmark/datasets/<name>/` and follow SPEC §10.2: ``` benchmark/datasets/<name>/ manifest.json # {name, version, description, license, documents:[…]} docs/<id>.<ext> # the input file truth/<id>.md # expected markdown (.txt for text-only documents) ``` Each entry in `manifest.documents[]` is: ```json { "id": "invoice_001", "file": "docs/invoice_001.png", "truth": "truth/invoice_001.md", "pages": 1, "category": "invoice", "tags": ["clean", "table"] } ``` - `id` is unique within the dataset and is what shows up in results JSON. - `category` groups documents for per-category scores. The built-in `synthetic-v1` categories are `plain`, `invoice`, `table`, `two_column`, `noisy_scan`, `handwriting_like`, `low_res`, `rotated`, `headings`, `multipage`. - `tags` are free-form; documents tagged `table` additionally get a `table_score`. - `license` in the manifest is required and must permit redistribution. **Only open datasets are committed.** For a public set that cannot be redistributed (olmOCR-bench, OmniDocBench), contribute a downloader/adapter that produces this layout locally instead of the files themselves. The same goes for outputs: per-document outputs of research-only datasets (OmniDocBench) are not committed, only their scores. Ground truth must be *exact* — that is why the built-in set is generated: `make dataset` (`python benchmark/generate_synthetic.py`) renders documents deterministically from the same source text it writes to `truth/`. Prefer extending the generator over hand-writing truth files. Verify with a cheap model before proposing the dataset: ```bash cargo run -p puffinparse-cli --release -- bench run \ --dataset benchmark/datasets/<name> \ --models llamaparse/cost_effective \ --out benchmark/results/$(date +%F)-<name>.json ``` --- ## 5. Commit messages and pull requests Write commit subjects in the **imperative mood**, ≤ 72 characters, no trailing period — `Add Mistral OCR provider`, not `Added…` / `Adds…`. [Conventional Commits](https://www.conventionalcommits.org) prefixes (`feat:`, `fix:`, `docs:`, `perf:`, `refactor:`, `test:`, `chore:`) are welcome but optional. Explain the *why* in the body; wrap it at 72 columns. Keep one logical change per commit and rebase rather than merge `main`. Before opening a PR: - [ ] `cargo fmt --all --check` is clean - [ ] `cargo clippy --workspace --all-targets -- -D warnings` is clean - [ ] `cargo test --workspace` passes - [ ] `ruff check` / `ruff format --check` / `mypy python/puffinparse` are clean - [ ] `pytest python/tests -q` passes - [ ] New behaviour has a test (fixture-based, no network) - [ ] Docs updated — `README.md`, `docs/SPEC.md` if the contract changed, `.env.example` for a new key - [ ] `CHANGELOG.md` `## [Unreleased]` has an entry - [ ] Public API changes are semver-appropriate and noted in the PR description Fill in the PR template (summary, test plan, checklist). Small, focused PRs get reviewed fastest. If you are planning something large — a new crate, a change to the unified response shape, a new metric — open an issue first so we can agree on the design. --- ## 6. Releasing (maintainers) 1. Update `CHANGELOG.md`: move `## [Unreleased]` items under a new `## [x.y.z] - YYYY-MM-DD` heading. 2. Bump `workspace.package.version` in the root `Cargo.toml`, run `cargo check --workspace` to refresh `Cargo.lock`, and commit. 3. Tag `vx.y.z` and push the tag. `.github/workflows/release.yml` builds wheels + an sdist + CLI archives, publishes to PyPI via trusted publishing, and creates the GitHub Release. --- # Changelog <!-- source: CHANGELOG.md | url: https://puffinparse.com/docs/project/changelog/ --> All notable changes to this project are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). The workspace crates (`puffinparse-core`, `puffinparse-cli`, `puffinparse-python`, …) and the Python package `puffinparse` share a single version. Entries before the rename say LiteOCR. ## [Unreleased] ### Added - Agent-friendly website: docs URLs return their Markdown for `Accept: text/markdown` (with `Vary: Accept`, and a Markdown 404 for unknown docs paths); one schema.org JSON-LD block per page (`SoftwareApplication`, `WebSite` with a docs-search `SearchAction` backed by `/docs/?q=`, `TechArticle`, and licensed `Dataset` entries on the benchmark viewer); `robots.txt` names the major AI crawlers explicitly; `llms.txt` gains when-to-use guidance and the benchmark JSON endpoints; and a new `/docs/agents/` page, also served as `/agents.md`, with a copyable onboarding prompt. ## [0.1.1] - 2026-10-08 ### Fixed - The source distribution now includes the LICENSE file its metadata names; PyPI rejected the 0.1.0 sdist for this, so 0.1.0 shipped wheels only. ## [0.1.0] - 2026-10-08 First public release (PyPI wheels, CLI binaries, gateway image). ### Changed - Provider verification labels are consistent across the README, the provider reference, the landing page and the docs site: **live-verified** (Reducto, Extend, LlamaParse), **verified locally** (Tesseract, Docling) and **docs-only** (the rest, implemented from documentation and tested against fixtures, not yet run live; tracked in issue #10). Tesseract and Docling were previously counted as live-verified on the landing page. - README roadmap now lists the open GitHub issues; added issue templates for benchmark and dataset suggestions, a shorter pull-request template, `AGENTS.md`, and private vulnerability reporting in `SECURITY.md`. - Benchmark and docs tables use sentence-case headers, and a column's best value is bold only when it is unique as displayed (ties are no longer bolded). - **Renamed to PuffinParse** (ADR-23). Crates `puffinparse-*`, Python package `puffinparse` (`import puffinparse`), Node package and CLI binary `puffinparse`, environment variables `PUFFINPARSE_*` (was `LITEOCR_*`), metadata keys `puffinparse_*`, `PuffinParseError`, and `output_format="puffinparse"` for the native shape. The site moves to puffinparse.vercel.app (liteocr.vercel.app still works). Old result files with `liteocr_version` still load. ### Added - **Release pipeline** (`docs/RELEASING.md`). A `v*` tag publishes abi3 wheels and an sdist to PyPI, `puffinparse` plus five prebuilt platform packages to npm, `puffinparse-core`/`-server`/`-cli` to crates.io, the gateway image to `ghcr.io/ajinkyashejul/puffinparse`, and CLI archives with `SHA256SUMS` to GitHub Releases, all through trusted publishing (no stored registry tokens). npm and crates.io are switched on per repository variable once their one-time setup is done. - **Tesseract baseline in `combined-v3`.** `tesseract/default` (Tesseract 5.5.1, one OpenMP thread per process) was added to the 2026-09-25 run with `bench run --resume`: 199 documents, 0 failures, Overall 56.36 (dpbench 86.77, olmocr 41.67, omnidocbench 35.70, parsebench 20.73, synthetic 97.98). The six API rows are unchanged. Licence audit notes in the DP-Bench README (5 of the 40 vendored pages are Upstage's own documents, kept under its MIT declaration) and the olmOCR README (redistribution with attribution complies with ODC-BY and AI2's guidelines). - **Puffin brand** (ADR-24, `docs/DESIGN.md` Brand): a puffin mark (favicon and header logo on the landing page, docs and benchmark viewer) and a waving puffin mascot with three pages in its beak, shown on the landing hero, the 404 page and the README. Every page now has an Open Graph / Twitter social card (`website/assets/og.png`). The landing demo shows a response on arrival instead of starting blank. - **First `combined-v3` results** (6 API models × 199 documents, adds a 40-page DP-Bench subset; 0 failures, $14.65): llamaparse/cost_effective leads (84.35) ahead of llamaparse/agentic (83.49) and reducto/r-1 (82.40). It is the new headline in `benchmark/LEADERBOARD.md` and the viewer; `combined-v2` keeps the Tesseract baseline row. - **Design language** (`docs/DESIGN.md`, `website/assets/tokens.css`): paper-and-ink neutrals, one scan accent, verdict and proof-mark colours, layout-box hues, a system-font type scale, the bounding-box and scan-line signatures; every text token passes WCAG AA in both themes. The docs, landing page and benchmark viewer all style through it. - **Benchmark viewer redesign.** Human document titles ("Headers & footers 3") with raw ids in Details; one primary number per view with secondary metrics, methodology and reproduce commands in disclosures; a compact leaderboard with per-source columns and an All metrics toggle; a calmer score-vs-cost chart with non-overlapping labels; plain-English check rows; a documents list with mean score and best model. PDF pages are rendered at build time with pypdfium2, so olmOCR-bench and ParseBench pages show without pdf.js. - **First `combined-v2` results** (7 models × 159 documents, 0 failures, $11.74): llamaparse/cost_effective leads (83.79) ahead of llamaparse/agentic (83.05) and reducto/r-1 (81.32); `tesseract/default` is the free baseline (48.68, 97.94 on synthetic). `benchmark/LEADERBOARD.md` now has one section per dataset with `combined-v2` as the headline. The viewer shows research-only documents (OmniDocBench) as scores only, with the command to fetch the data, instead of broken links. - **Jobs everywhere.** Gateway: `POST /v1/jobs` (parse body + `webhook_url`, 202 with a job id) and `GET /v1/jobs/{id}` (pending / succeeded / failed, `output_format` on retrieval); job ids are bound to the submitting key, cost is charged once when the job first succeeds, handles persist in `state_file` (`server.job_retention_hours`), and an opt-in `POST /v1/webhooks/{provider}` receiver (`[webhooks] enabled`, shared secret) settles jobs from provider webhook bodies. New `liteocr_jobs_total` metric and `job_id`/`job_status` log fields. Node SDK: `submit`, `retrieve`, `handleWebhook` with a typed `Job`. - **DP-Bench** (Upstage, MIT) adapter: 40 committed single-page PDFs with reading-order transcript truth (`python -m benchmark.adapters dpbench`, pinned at `24702c61`), and `combined-v3` = combined-v2 + DP-Bench (199 documents). Headers and footers stay in the DP-Bench truth because DP-Bench's own NID scores them; each source keeps its publisher's conventions. - `tesseract/default` runs `tesseract` with `OMP_THREAD_LIMIT=1` unless set, so parallel pages no longer oversubscribe the CPU (70 s → 0.7 s for a small page on 4 cores). - Benchmark results flag a successful call that returned no text as `empty_output: true` (counted in `summary.empty_outputs`, shown as `(+N empty)` in the leaderboard), and `bench rescore --keep-missing` keeps the recorded scores of documents whose outputs are not committed. - **Scorer v2** (`scorer_version: 2` in result JSON). HTML `<table>` output is scored like markdown tables (it used to score 0, which ranked Reducto r-1 last); a new `teds_grid` metric (TEDS on the row/cell grid); rule matching ignores spaces next to punctuation (ParseBench rule text is tokenised); `bag_of_sentences` matches sentences fuzzily (0.8) and passes at 0.8 (1.0 was dead signal); olmOCR `max_diffs` tolerances are honoured; HTML entities are decoded; and results carry an explicit `summary.headline` and per-document `headline`, so `char_similarity` means literal character similarity again. Python and Node `Metrics` gain `teds_grid` / `tedsGrid`, and the viewer's rule checker follows the same rules. - **`liteocr bench rescore`** re-scores a run from its saved outputs with the current scorer, no network calls, keeping measured latency and cost. Both committed runs are re-scored and `benchmark/LEADERBOARD.md` regenerated: on combined-v1, reducto/r-1 moves from last (87.19) to second (91.37); llamaparse/agentic stays first (93.08). - `docs/benchmarks/findings.md`: why four models returned nothing for `parsebench/text_multicolumns_2col` (a Form XObject with a ±2^1023 `/BBox` that 32-bit renderers clip to empty; fixing the file fixes both Reducto and Extend, no option does). - **Self-hosted engines, no key, $0/page.** `tesseract/default` (local `tesseract` binary, PDFs via `pdftoppm`; native OCR with word/line boxes and confidences), `docling/default` (docling-serve v1 async API; layout, tables, OCR) and `paddleocr/default` (PaddleOCR/PaddleX serving: `/ocr` and PP-StructureV3 `/layout-parsing`; docs-only). Tesseract and Docling are live-verified locally. `liteocr providers` shows `local` for them; `--json` and Python `providers()` include `self_hosted`. - README, docs site and landing page cover the TypeScript SDK, the gateway, jobs/webhooks, self-hosted engines and the new benchmark sources; the docs navigation gains the gateway, the self-hosted provider pages, the academic-benchmark survey and the adapter notes. - **Async jobs API and webhooks.** Split a parse into submit and retrieve so the caller owns the waiting (long documents, batches, webhook-driven pipelines). Rust: `submit_parse`, `retrieve_parse` / `retrieve_parse_with`, `parse_webhook`, `resolve_webhook`, `JobHandle` (serialisable, never holds a key), `JobStatus`, and `DocumentRequest.webhook_url`. Python: `liteocr.submit` / `asubmit`, `retrieve` / `aretrieve`, `handle_webhook` / `ahandle_webhook` and `liteocr.Job`. Covers Reducto, Extend and LlamaParse (live-verified); `webhook_url` maps to Reducto `async.webhook` and the LlamaParse `webhook_url` field, and is rejected for Extend, which only has workspace-level webhooks. SPEC §15. - **Benchmark viewer: verify every claim.** A per-document inspector at `#/<run>/<model>/<doc>` shows the page itself (pdf.js for PDFs, pinned with SRI, PNG fallback), every model's output side by side, a word diff against the truth, and a pass/fail rule checklist for rules documents that is checked against the recorded score. Layout-box overlays by block type appear when a run saved unified responses. The leaderboard ranks by `summary.headline` when present and adds p90 latency, a score-vs-cost/latency scatter with the Pareto frontier, and per-source and source × category breakdowns. Every view is a shareable link, with keyboard navigation (`j`/`k`, `m`, `1`–`4`, `d`, `o`, `?`), the product site's design and theme switch, and a mobile layout. - `liteocr bench run --save-outputs` also writes the unified response as `<doc>.json` next to `<doc>.md`, which the viewer uses for its layout overlay. - **More public benchmarks in the combined dataset.** `olmocr` (40 AI2 olmOCR-bench PDFs, 205 rules, ODC-BY-1.0) and `omnidocbench` (40 pages across 10 document types, English and Chinese; index only — images and truth are fetched at a pinned revision by `python -m benchmark.adapters omnidocbench` because the data is research-only), joined with synthetic-v1 and ParseBench as `combined-v2` (159 documents; `combined-v1` is unchanged). Upstream tests that cannot be expressed faithfully (math, baseline, positional absences, vertical table neighbours) are skipped and counted in `conversion-stats.json`, never dropped silently. A dataset self-check test proves every converted rule is satisfiable. Survey of academic benchmarks with licences at pinned revisions: `docs/benchmarks/academic-benchmarks.md`. - **Gateway server: `liteocr serve` (`crates/liteocr-server`).** An HTTP gateway in front of every provider, the LiteOCR equivalent of the LiteLLM proxy: `POST /v1/parse|ocr|extract` (JSON or multipart, `output_format`, `fallbacks`), `GET /v1/models`, `/v1/usage`, `/health` and Prometheus `/metrics`. A `liteocr.toml` config defines model aliases with ordered or round-robin fallback, provider keys as `env:` references, and virtual keys with model allow-lists, monthly USD budgets and per-minute rate limits. Every request is logged as one JSON line (never document content, provider error text or secrets), and all errors share one JSON body. Clients cannot override `api_key`/`base_url` or read local files. Ships a multi-stage distroless `Dockerfile` and `examples/server/liteocr.toml`; see [`docs/SERVER.md`](/docs/gateway/index.md). - **Node.js / TypeScript SDK.** The `liteocr` npm package in `js/` runs on a napi-rs addon over the same Rust core (`crates/liteocr-node`): async `parse` / `ocr` / `extract` (with `fallbacks`), `Router`, camelCase typed responses (`index.d.ts`), `LiteOCRError` subclasses mapped from the core `ErrorKind`, and the pricing, model and scoring helpers. CI builds the addon and runs the typecheck and `node:test` suite; the release workflow builds prebuilt `.node` binaries as artifacts (npm publishing not wired yet, so build from source for now). Docs at `/docs/typescript/`. - **`liteocr bench run --resume`**: every finished call is appended to `<out>.partial.jsonl`, so an interrupted or partly failed run continues where it stopped and only re-runs missing or failed (model, document) pairs. `--dry-run` prints the plan (calls, pages, list-price estimate) without calling any provider; `--max-cost <usd>` aborts before the first call when the estimate is higher; `--retries N` re-issues documents after retryable errors (off by default, since a retried job may be billed twice). Result documents record `provider_job_id`, `cache_hit`, `attempts`, `started_at` and `error_kind`, and the run ends with a summary line (calls, failures, cost, wall time). - **Native-format compatibility (`output_format`).** A response can be rendered in a provider's own JSON shape instead of the unified one, so an integration already written against Reducto, Extend or LlamaParse can switch the underlying provider without rewriting its parsing code: `ParseResponse::to_format("reducto")` in Rust (`liteocr_core::compat::{Format, render_parse, render_extract}`), with `DocumentRequest.output_format` carrying and validating the choice. What is guaranteed is structural fidelity — key set, chunk/page and block counts, content strings, block-type vocabulary, coordinate units and billed pages — not byte equality; the always-null fields and the lossy type mappings are enumerated in [`docs/COMPAT.md`](/docs/project/compat/index.md), and each provider's fixture is round-tripped through its own renderer in the test suite. Extract-mode rendering is best effort. See ADR-13. - **`output_format` in the Python SDK and the CLI.** `liteocr.parse(..., output_format="reducto")` (and `aparse`, `extract`, `aextract`, plus the matching `Router` methods) returns the vendor's own JSON as a `dict` instead of the dataclass; `output_format=None` (the default) or `"liteocr"` keeps the unified shape. The value is validated in the core before any network call — an unknown name raises `BadRequestError` listing `liteocr | reducto | extend | llamaparse` — the rendering happens in Rust (`liteocr._core.render_parse` / `render_extract`), and success callbacks still receive the dataclass. Overloads type the return (`None` → dataclass, `str` → `dict`), `liteocr.output_formats()` lists the accepted values, and `examples/switch_provider_keep_format.py` shows a provider swap with the parsing code untouched. The CLI gains `--output-format <vendor>` on `parse` (with `--format json`; ignored with a warning on stderr otherwise) and on `extract`, and `liteocr providers --json` now emits `{"providers": [...], "output_formats": [...]}` instead of a bare array. ### Changed - Internal: the vision-LLM providers (Gemini, OpenAI, Anthropic) share `providers/vlm.rs`; behaviour and request bodies are unchanged. - Adapter HTML→markdown table conversion no longer doubles backslashes, so LaTeX in table cells reaches the scorer as a parser would print it. - **Modes.** Every call now names a mode — `parse` (markdown + typed blocks), `ocr` (plain text with line/word boxes) or `extract` (a JSON object from a schema, with per-field confidence and citations) — and providers can only be swapped within a mode. The Python SDK exposes `parse`/`aparse`, `ocr`/`aocr` (now plain text, not markdown) and `extract`/`aextract`, the new `TextResponse` / `ExtractResponse` dataclasses, a mode-bound `Router(models, mode=...)`, and mode arguments on `list_models`, `resolve_model`, `set_pricing` and `estimate_cost`; pricing is now per page *per mode*. The CLI gains `liteocr ocr` and `liteocr extract`, and `liteocr providers` shows modes, per-mode prices and a `--mode` filter. The old `liteocr.ocr` (which returned markdown) is now `liteocr.parse`, and `OcrResponse` is now `ParseResponse`. Which models serve which mode is reported by `liteocr.list_models(mode)` and `liteocr providers --mode <mode>`. - **README model table.** The "Model names" section is regenerated from the registry and `pricing.json`: every provider group carries its env var and verification status (live-verified vs docs-only), every model its modes, the default per mode and the list price per page for `parse · ocr · extract`. ### Fixed - Extend `provider_options={"responseType": "url"}` is now sent as the `GET /parse_runs/{id}` query parameter; it used to go in the request body, so the presigned-output path never triggered. Reducto and Extend url-typed results are now covered by live-captured fixtures and loopback tests. ## Initial development - 2026-09-11 Initial release. ### Added - **Rust core (`liteocr-core`)** — a unified request/response contract for document parsing: `OcrRequest` accepting a path, bytes or URL, and an `OcrResponse` with `pages`, `blocks` (normalised block types and 0..1 bounding boxes), document-level `markdown`/`text`, `usage`, `cost_usd` and `latency_ms`, identical across every provider. `#![forbid(unsafe_code)]`. - **Providers** — Reducto, Extend and LlamaParse, each talked to over plain HTTPS with no vendored SDK, addressed by `"<provider>/<model>"` model strings (`reducto/standard`, `extend/parse_performance`, `llamaparse/cost_effective`, …) with a per-provider default and a `provider_options` pass-through escape hatch. - **Router** — ordered fallbacks and round-robin load balancing across models, with per-model success/failure and latency stats; only provider, rate-limit and timeout errors trigger fallback. - **Pricing and cost tracking** — an embedded, overridable `pricing.json` price table mapping model → per-page (or per-credit) USD, used to compute `cost_usd` on every response. - **Reliability primitives** — retries with exponential backoff on 429/5xx and network errors, per-call timeouts, and whole-call deadlines that cover upload plus polling. - **Benchmark metrics** — character similarity (the primary score), CER, WER, word-level F1/recall, reading-order agreement and a table-restricted score, all computed in Rust over NFKC-normalised text. - **CLI (`liteocr`)** — `parse` (markdown/text/json output, optional raw provider payload), `providers` (models, pricing and API-key status), and `bench run` / `bench report` / `bench score` for running the benchmark, scoring results and generating the leaderboard. - **Python SDK (`liteocr`)** — `ocr()`, the async `aocr()`, `Router`, typed dataclasses (`OcrResponse`, `Page`, `Block`, `BBox`, `Usage`), the full error hierarchy, `set_pricing()`, `list_models()`, and success/failure callbacks. Fully typed, ships `py.typed`, distributed as abi3 wheels for CPython 3.9+. - **`synthetic-v1` benchmark dataset** (v1.1.0) — 39 deterministically generated documents across 13 categories (plain, headings, invoice, table, two-column, noisy scan, low resolution, multi-page PDF, skewed, dense, faded, receipt, complex table) with exact markdown ground truth, plus the generator that produces them byte-for-byte. - **First leaderboard** (`benchmark/LEADERBOARD.md`) from a run over seven models across the three providers; `bench run` disables provider result caches by default (`--allow-cache` to opt out) so latency reflects real work. [Unreleased]: https://github.com/ajinkyashejul/puffinparse/compare/v0.1.0...HEAD [0.1.0]: https://github.com/ajinkyashejul/puffinparse/releases/tag/v0.1.0 --- # Security <!-- source: SECURITY.md | url: https://puffinparse.com/docs/project/security/ --> ## Supported versions PuffinParse is pre-1.0. Security fixes land on `main` and ship in the next release; only the latest released version is supported. | Version | Supported | |---|---| | 0.1.x | ✅ | | < 0.1 | ❌ | ## Reporting a vulnerability **Please do not open a public issue for a security problem.** Report it privately with GitHub's private vulnerability reporting, which is enabled for this repository: 1. Go to https://github.com/ajinkyashejul/puffinparse/security/advisories/new (repository → **Security** → **Report a vulnerability**). 2. Describe the issue, the affected version or commit, and — if you can — a minimal reproduction and the impact you believe it has. You will get an acknowledgement within 5 business days and an assessment with a fix timeline within 10. We will keep you updated while we work on a fix, credit you in the advisory and `CHANGELOG.md` unless you prefer otherwise, and publish the advisory once a fixed release is out. Please give us a reasonable window to ship a fix before disclosing publicly. If GitHub advisories are unavailable to you, open a public issue that says only that you have a security report and asks a maintainer to make contact — no details — and we will take it from there. ## What is in scope - The Rust crates (`puffinparse-core`, `puffinparse-cli`, `puffinparse-python`, `puffinparse-node`, `puffinparse-server`), the Python package `puffinparse` and the Node package `puffinparse`. - The self-hosted gateway (`puffinparse serve`): authentication with virtual keys, budget and rate-limit enforcement, and anything that could leak a provider key or document content through responses, logs or metrics. - The release and CI workflows in `.github/workflows/`, and the published artifacts (PyPI wheels/sdist, GitHub Release binaries). Out of scope: vulnerabilities in the third-party providers themselves (report those to the provider directly), and issues that require an already-compromised machine or a malicious local Rust/Python dependency you introduced. ## How PuffinParse handles credentials - **Provider API keys are only read from the environment** (each provider's own variable, such as `REDUCTO_API_KEY`; `.env.example` lists them all), passed explicitly as the `api_key` argument, or, for the gateway, referenced as `env:` values in its config file. PuffinParse never reads them from anywhere else, never writes them to disk, and never sends them anywhere but the provider's own base URL. - **Keys are never logged.** `PUFFINPARSE_LOG=debug` traces requests, retries and polling, but `Authorization` headers and key values are redacted; errors carry provider, status code, message and request/job id only. If you ever see a key in log output, in an error message, or in a serialized `OcrResponse`, that is a vulnerability — please report it. - Document bytes are sent only to the selected provider. PuffinParse has no telemetry and makes no network calls other than to the provider you choose. - Recorded test fixtures under `crates/puffinparse-core/tests/fixtures/` must be redacted; never commit a fixture containing a real key, token or job id tied to a live account. - Releases are published to PyPI with trusted publishing (OIDC), so no long-lived PyPI token exists in this repository's secrets.