PuffinParse docs
/ llms.txt GitHub

PuffinParse

One API for every document parser. Rust core, Python and TypeScript SDKs, CLI, gateway, and an open benchmark that ranks providers on accuracy, latency and cost.

import puffinparse
doc = puffinparse.parse("invoice.pdf", model="reducto/standard")
print(doc.markdown, doc.cost_usd)

Switch providers by changing one string. Same request, same response shape, same errors.

Get started pip install puffinparse Leaderboard GitHub

Open benchmark#

Accuracy, latency and cost, measured through the same client you would use. Exact ground truth, deterministic metrics, no LLM judge.

#ModelOverallp50 latency$/1k pages
1llamaparse/cost_effective84.359,452 ms$3.75
2llamaparse/agentic83.4914,127 ms$12.50
3reducto/r-182.403,459 ms$10.00
4reducto/standard79.783,019 ms$15.00
5extend/parse_performance77.4321,965 ms$25.00

combined-v3 v3.0.0 · 199 documents — full leaderboard · methodology and caveats · results viewer

For agents#

Every page on this site is also served as plain markdown at <page>/index.md, and is linked from each HTML page with <link rel="alternate" type="text/markdown">. Three entry points are meant for you:

  • /llms.txt — the site map, with a one-line description of every page.
  • /llms-full.txt — every page concatenated, in nav order.
  • /project/contributing/ — how to contribute: required checks and where each kind of change goes.

One API for every document parser: parse, OCR and extract. Rust core, Python and TypeScript SDKs, a CLI, a self-hosted gateway, and an open benchmark that ranks providers on accuracy, latency and cost.

import puffinparse

doc = puffinparse.parse("invoice.pdf", model="reducto/standard")     # or "extend/parse_performance", "llamaparse/agentic", ...
print(doc.markdown)                                              # unified markdown, every provider
print(doc.pages[0].blocks[0].bbox, doc.usage.pages, doc.cost_usd)

text = puffinparse.ocr("scan.png", model="llamaparse/fast")          # plain text + line/word boxes
print(text.text, text.pages[0].lines[0].bbox)

Switch providers by changing one string. Same request, same response shape, same errors.

Why#

Every document-parsing vendor has its own upload flow, polling loop, JSON layout, block vocabulary, coordinate system and billing unit. PuffinParse hides all of that behind one call, tracks cost per call, retries and falls back across providers, and ships a reproducible benchmark so you can pick a provider on evidence instead of marketing.

Two things make switching real rather than aspirational. Modes: every call names parse, ocr or extract, models declare the modes they serve, and a model that cannot serve the one you asked for fails before any network call — so a provider swap can never quietly change the shape of your answer. Native-format compatibility: if you are already integrated with Reducto, Extend or LlamaParse, output_format="reducto" (and friends) renders any provider's result into that vendor's own JSON, so you can re-point a request without touching your parsing code — see below.

Providers (v0.1) 18 providers · 60 models · 3 modes. Live-verified against the real APIs: Reducto, Extend, LlamaParse. Verified locally: the self-hosted Tesseract and Docling. Docs-only (implemented from the provider's API documentation and tested against fixture payloads, not yet run live): Mistral, Azure, Textract, Gemini, OpenAI, Anthropic, Mathpix, Datalab, Unstructured, Upstage, Landing AI, Google Document AI and the self-hosted PaddleOCR (help verify them). Full table
Modes parse (markdown + blocks), ocr (plain text + boxes), extract (JSON from a schema)
Core Rust (puffinparse-core): reqwest + tokio, no vendor SDKs, #![forbid(unsafe_code)]
SDKs Python 3.9+ (sync + async, fully typed) and Node.js / TypeScript (js/), both on the same Rust core
Gateway puffinparse serve: one HTTP endpoint with aliases, fallbacks, virtual keys, budgets, rate limits, JSON logs and Prometheus metrics (docs/SERVER.md)
Long documents submit / retrieve jobs and provider webhooks instead of a blocking call (below)
CLI puffinparse parse, puffinparse ocr, puffinparse extract, puffinparse providers, puffinparse bench
Reliability Retries with jittered backoff, whole-call deadlines, Router with ordered fallbacks / round-robin
Compatibility output_format renders any provider's result in Reducto's, Extend's or LlamaParse's own JSON, so an existing integration keeps its parser (docs/COMPAT.md)
Cost Embedded, overridable price table → cost_usd on every response
Benchmark One harness over synthetic data and public benchmarks (ParseBench, olmOCR-bench, OmniDocBench, DP-Bench); deterministic metrics and rule checks, latency, $/1k pages; every output inspectable at puffinparse.com/benchmark-results

Install#

pip install puffinparse              # Python SDK (abi3 wheels: Linux, macOS, Windows)
npm install puffinparse              # Node.js SDK (prebuilt for Linux x64/arm64 glibc, macOS, Windows x64)
cargo install puffinparse-cli        # CLI + gateway; or grab an archive from GitHub Releases
docker pull ghcr.io/ajinkyashejul/puffinparse   # gateway image

Set the keys for the providers you use:

export REDUCTO_API_KEY=...
export EXTEND_API_KEY=...
export LLAMA_API_KEY=llx-...      # LlamaCloud / LlamaParse

No key yet? The self-hosted engines work out of the box once installed:

sudo apt-get install tesseract-ocr poppler-utils     # or: brew install tesseract poppler
puffinparse ocr scan.png -m tesseract                    # free, local, word boxes + confidences

Every provider reads its own variable: .env.example lists all of them, the model tables say which belongs to which provider, and puffinparse providers shows which ones are set in your shell.

From source (Rust stable + Python 3.9+):

git clone https://github.com/ajinkyashejul/puffinparse && cd puffinparse
python -m venv .venv && . .venv/bin/activate
pip install maturin && maturin develop --release     # builds the extension into the venv
cargo build --release -p puffinparse-cli                 # ./target/release/puffinparse

Modes#

Document AI vendors sell three different products, and they are not interchangeable: layout parsing, plain-text OCR, and schema-driven extraction. A call picks one mode, and the mode decides the response type:

Mode Call Returns Use it for
parse puffinparse.parse(...) ParseResponse — markdown, pages[].blocks[] with types and boxes RAG chunks, tables, document structure
ocr puffinparse.ocr(...) TextResponse — text, pages[].lines[] / words[] with boxes search indexes, redaction, overlays
extract puffinparse.extract(..., schema) ExtractResponse — data shaped by your JSON Schema, plus per-field confidence and citations invoices, forms, anything with fields

Providers are swappable only within a mode. Every model declares the modes it serves, so a model that cannot do what you asked raises UnsupportedModelError before any network call instead of silently returning the wrong shape. puffinparse.list_models("ocr") lists the candidates for a mode; puffinparse providers --mode ocr does the same on the command line. Providers without a native OCR endpoint serve ocr from their parse output, flagged as resp.metadata["puffinparse_derived_from"] == "parse".

Which models serve which mode is a registry fact, not a guess — ask puffinparse.list_models("extract") or puffinparse providers --mode extract. Most providers serve parse and ocr; extract needs a model built for it (reducto/extract, extend/extraction_light, azure/invoice, the vision-LLM models, ...), and calling it with a parse-only model raises UnsupportedModelError naming the mode.

Usage#

One call, any provider#

import puffinparse

# path, URL, or bytes (+ filename)
doc = puffinparse.parse("contract.pdf", model="extend/parse_performance")
doc = puffinparse.parse("https://cdn.reducto.ai/samples/fidelity-example.pdf", model="reducto/r-1", pages="1-2")
doc = puffinparse.parse(open("scan.png", "rb").read(), filename="scan.png", model="llamaparse/agentic")

doc.markdown            # whole document
doc.text                # plain text
doc.pages[0].markdown   # per page
for block in doc.pages[0].blocks:
    block.type           # text | title | section_header | list | table | figure | header | footer | ...
    block.content        # markdown
    block.bbox           # BBox(x0, y0, x1, y1) normalised 0..1, origin top-left (or None)
    block.confidence     # 0..1 when the provider reports one
doc.usage.pages, doc.usage.credits, doc.cost_usd, doc.latency_ms

Plain text and boxes (ocr)#

text = puffinparse.ocr("scan.png", model="reducto/standard")

text.text                     # whole document, pages joined by a blank line
page = text.pages[0]
page.text                     # plain text in reading order
for line in page.lines:       # Line(text, bbox, confidence)
    x0, y0, x1, y1 = line.bbox.to_pixels(page.width, page.height)
for word in page.words:       # Word(text, bbox, confidence)
    ...

Structured extraction (extract)#

schema = {
    "type": "object",
    "properties": {"invoice_number": {"type": "string"}, "total": {"type": "number"}},
    "required": ["invoice_number", "total"],
}
result = puffinparse.extract("invoice.pdf", schema, model="...", citations=True)

result.data                          # {"invoice_number": "INV-42", "total": 1280.5}
result.fields["/total"].confidence   # per-field confidence, keyed by JSON pointer
result.citations("/total")           # [Citation(page_number=2, bbox=..., text="Total due 1,280.50")]

Async#

doc = await puffinparse.aparse("contract.pdf", model="llamaparse/cost_effective")
text = await puffinparse.aocr("scan.png", model="llamaparse/fast")

Runs on the Rust runtime; the event loop is never blocked.

Long documents: jobs and webhooks#

parse waits for the provider. For long documents, batches or webhook-driven pipelines, split it:

job = puffinparse.submit("annual-report.pdf", model="reducto/standard",
                     webhook_url="https://example.com/hooks/puffinparse")   # optional
store(job)                                  # a Job is plain data: serialisable, holds no key

result = puffinparse.retrieve(job)              # Job (still pending) or ParseResponse
# ...or, in your web handler, turn the provider's webhook body into a result:
result = puffinparse.handle_webhook(request.json(), model="reducto")  # verify the signature first

Reducto, Extend and LlamaParse support jobs; webhook_url maps to each provider's per-job webhook where one exists (Extend only has workspace-level webhooks, so it is rejected there). See docs/SPEC.md §15.

TypeScript / Node.js#

The same core as a napi-rs addon, with camelCase typed responses:

import { parse, Router } from "puffinparse";

const doc = await parse("invoice.pdf", { model: "reducto/standard", fallbacks: ["llamaparse/agentic"] });
console.log(doc.markdown, doc.usage.pages, doc.costUsd);

npm install puffinparse ships prebuilt binaries for Linux x64/arm64 (glibc), macOS and Windows x64; other platforms build from source (cd js && npm ci && npm run build), see js/README.md.

Model names#

"<provider>/<model>", like LiteLLM. A bare provider name picks that provider's default model for the mode you called — marked * below. puffinparse providers and puffinparse.list_models(mode) print the live list; the tables here are the built-in registry (crates/puffinparse-core/src/model.rs and pricing.json).

Prices are public pay-as-you-go list prices, per page, per mode, shown as parse · ocr · extract with — where a model does not serve that mode. Override one mode at a time with puffinparse.set_pricing({"reducto/standard": 0.012}, "parse"), and estimate with puffinparse.estimate_cost("reducto/standard", pages=1000, mode="ocr"). The vision-LLM providers (Gemini, OpenAI, Anthropic) bill tokens rather than pages, so their per-page numbers are estimates — see the source lines in pricing.json.

Each provider carries one of three verification labels:

  • live-verified: the live tests pass against the real API with a key, and the fixtures include redacted live responses (Reducto, Extend, LlamaParse).
  • verified locally: a self-hosted engine whose live tests pass against a local install (Tesseract, Docling).
  • docs-only: implemented from the provider's API documentation and tested against fixture payloads built from it, but not yet run against the live API. Expect wire-format differences until it is verified; issue #10 tracks this, and a run with your own key is a welcome contribution.

docs/providers/README.md tracks the state and links one reference page per provider.

Reducto · REDUCTO_API_KEY · live-verified — Layout parsing plus a schema extractor with citations; the default parse target.

Model Modes List price / page (parse · ocr · extract) Notes
reducto/standard parse *, ocr * $0.015 · $0.015 · — Reducto Parse with account-default model (legacy standard)
reducto/r-1 parse, ocr $0.01 · $0.01 · — Reducto Parse with settings.model=r-1 (newest model, cheaper)
reducto/agentic parse, ocr $0.03 · $0.03 · — Reducto Parse with agentic text+table enhancement (highest accuracy, 2x cost)
reducto/extract extract * — · — · $0.035 Reducto Extract: POST /extract with a JSON schema, citations on request
reducto/deep_extract extract — · — · $0.055 Reducto Deep Extract (settings.deep_extract = true) for long or complex documents

Extend · EXTEND_API_KEY · live-verified — Parse engines that trade accuracy for cost per page, and two extraction processors.

Model Modes List price / page (parse · ocr · extract) Notes
extend/parse_performance parse *, ocr * $0.025 · $0.025 · — Extend engine=parse_performance (highest accuracy)
extend/parse_light parse, ocr $0.00625 · $0.00625 · — Extend engine=parse_light (fast, cheap, digital-native docs)
extend/parse_auto parse, ocr $0.025 · $0.025 · — Extend engine=parse_auto (picks light or performance per page)
extend/extraction_performance extract * — · — · $0.0625 Extend Extract, baseProcessor=extraction_performance (runs parse_performance)
extend/extraction_light extract — · — · $0.015 Extend Extract, baseProcessor=extraction_light (runs parse_light)

LlamaParse (LlamaCloud) · LLAMA_API_KEY · live-verified — Credit-priced tiers from plain text extraction to agentic parsing; extract from cost_effective up.

Model Modes List price / page (parse · ocr · extract) Notes
llamaparse/fast parse, ocr $0.00125 · $0.00125 · — LlamaParse tier=fast (text extraction, no OCR of images)
llamaparse/cost_effective parse *, ocr *, extract * $0.00375 · $0.00375 · $0.01 LlamaParse tier=cost_effective
llamaparse/agentic parse, ocr, extract $0.0125 · $0.0125 · $0.03125 LlamaParse tier=agentic
llamaparse/agentic_plus parse, ocr, extract $0.05625 · $0.05625 · $0.11875 LlamaParse tier=agentic_plus (highest accuracy)

Mistral Document AI · MISTRAL_API_KEY · docs-only — One OCR model with pinned versions; annotations give it an extract mode.

Model Modes List price / page (parse · ocr · extract) Notes
mistral/ocr-latest parse *, ocr *, extract * $0.004 · $0.004 · $0.005 Mistral OCR, latest alias (mistral-ocr-latest; currently OCR 4.1)
mistral/ocr-4-1 parse, ocr, extract $0.004 · $0.004 · $0.005 Mistral OCR 4.1 pinned (mistral-ocr-4-1; blocks + block confidence scores)
mistral/ocr-4-0 parse, ocr, extract $0.004 · $0.004 · $0.005 Mistral OCR 4.0 pinned (mistral-ocr-4-0; paragraph blocks, no block confidence)
mistral/ocr-2512 parse, ocr, extract $0.002 · $0.002 · $0.003 Mistral OCR 3 pinned (mistral-ocr-2512; cheaper, no paragraph blocks)

Azure AI Document Intelligence · AZURE_DOCUMENT_INTELLIGENCE_KEY (+ AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT) · docs-only — Prebuilt Document Intelligence models: read for OCR, layout for structure, fixed schemas for extraction.

Model Modes List price / page (parse · ocr · extract) Notes
azure/read ocr * — · $0.0015 · — Azure prebuilt-read: native OCR, words/lines with confidence (cheapest)
azure/layout parse *, ocr $0.01 · $0.01 · — Azure prebuilt-layout: markdown, paragraphs with roles, tables, polygons
azure/invoice extract * — · — · $0.01 Azure prebuilt-invoice: fixed invoice schema
azure/receipt extract — · — · $0.01 Azure prebuilt-receipt: fixed receipt schema
azure/id_document extract — · — · $0.01 Azure prebuilt-idDocument: fixed ID document schema
azure/tax_us_w2 extract — · — · $0.01 Azure prebuilt-tax.us.w2: fixed W-2 schema
azure/custom parse, ocr, extract $0.03 · $0.03 · $0.03 Azure custom model, id via provider_options.model_id

AWS Textract · AWS_ACCESS_KEY_ID (+ AWS_SECRET_ACCESS_KEY, AWS_REGION) · docs-only — One AWS API, four feature sets; SigV4-signed, no SDK.

Model Modes List price / page (parse · ocr · extract) Notes
textract/detect-text ocr * — · $0.0015 · — Textract DetectDocumentText: raw OCR, lines + words with boxes (cheapest)
textract/layout parse *, ocr $0.015 · $0.015 · — Textract AnalyzeDocument LAYOUT + TABLES: reading-order markdown + tables
textract/queries extract * — · — · $0.015 Textract AnalyzeDocument QUERIES: one natural-language query per schema field
textract/forms extract — · — · $0.05 Textract AnalyzeDocument FORMS: key-value pairs matched to schema fields

Google Gemini · GEMINI_API_KEY · docs-only — Vision-LLM transcription. Prices are per-page estimates (~1500 in + ~800 out tokens); cost_usd uses them until the response reports real usage.

Model Modes List price / page (parse · ocr · extract) Notes
gemini/2.5-flash parse *, ocr *, extract * $0.00245 · $0.00245 · $0.0012 Gemini 2.5 Flash: vision-LLM transcription, the price/quality default
gemini/2.5-pro parse, ocr, extract $0.009875 · $0.009875 · $0.004875 Gemini 2.5 Pro: highest accuracy, ~4x the cost of Flash
gemini/2.5-flash-lite parse, ocr, extract $0.00047 · $0.00047 · $0.00027 Gemini 2.5 Flash-Lite: cheapest and fastest, clean documents
gemini/3.5-flash parse, ocr, extract $0.00945 · $0.00945 · $0.00495 Gemini 3.5 Flash: frontier Flash generation (GA 2026-05-19)
gemini/3.5-flash-lite parse, ocr, extract $0.00245 · $0.00245 · $0.0012 Gemini 3.5 Flash-Lite: low-latency 3.x tier (GA 2026-07-21)
gemini/3.8-flash parse, ocr, extract $0.004125 · $0.004125 · $0.00225 Gemini 3.8 Flash: newest Flash model (GA 2026-09-02, introductory pricing)

OpenAI · OPENAI_API_KEY · docs-only — Vision transcription through the Responses API. Per-page prices are estimates, as for Gemini.

Model Modes List price / page (parse · ocr · extract) Notes
openai/gpt-5.6-luna parse *, ocr *, extract * $0.00114 · $0.00114 · $0.00114 OpenAI Responses API, gpt-5.6-luna (cheapest current vision model)
openai/gpt-5.6-terra parse, ocr, extract $0.0114 · $0.0114 · $0.0114 OpenAI Responses API, gpt-5.6-terra (balanced capability/price)
openai/gpt-5.6-sol parse, ocr, extract $0.02 · $0.02 · $0.02 OpenAI Responses API, gpt-5.6-sol (flagship GPT-5.6)
openai/gpt-6-astra parse, ocr, extract $0.05 · $0.05 · $0.05 OpenAI Responses API, gpt-6-astra (most capable, most expensive)

Anthropic (Claude) · ANTHROPIC_API_KEY · docs-only — Vision transcription through the Messages API. Per-page prices are estimates.

Model Modes List price / page (parse · ocr · extract) Notes
anthropic/claude-sonnet-5 parse *, ocr *, extract * $0.01 · $0.01 · $0.01 Claude Messages API, claude-sonnet-5 (balanced vision transcription)
anthropic/claude-haiku-4-5 parse, ocr, extract $0.005 · $0.005 · $0.005 Claude Messages API, claude-haiku-4-5 (cheapest, 200K context)
anthropic/claude-opus-5 parse, ocr, extract $0.025 · $0.025 · $0.025 Claude Messages API, claude-opus-5 (highest accuracy)

Mathpix · MATHPIX_APP_KEY (+ MATHPIX_APP_ID) · docs-only — Maths-first OCR: Mathpix Markdown with LaTeX, line and word polygons.

Model Modes List price / page (parse · ocr · extract) Notes
mathpix/pdf parse *, ocr * $0.005 · $0.005 · — Mathpix v3/pdf document OCR (image inputs auto-routed to v3/text); MMD + line polygons
mathpix/text parse, ocr $0.002 · $0.002 · — Mathpix v3/text single-image OCR (line + word polygons, per-image billing)

Datalab (Marker) · DATALAB_API_KEY · docs-only — Marker's hosted API: three quality modes at two price points.

Model Modes List price / page (parse · ocr · extract) Notes
datalab/fast parse, ocr $0.004 · $0.004 · — Datalab Convert mode=fast (lowest latency, digital-native documents)
datalab/balanced parse *, ocr * $0.004 · $0.004 · — Datalab Convert mode=balanced (Datalab's recommended default)
datalab/accurate parse, ocr $0.01 · $0.01 · — Datalab Convert mode=accurate (scans, dense layouts, complex tables)

Unstructured · UNSTRUCTURED_API_KEY · docs-only — The partitioner behind many RAG pipelines; flat per-page price across strategies.

Model Modes List price / page (parse · ocr · extract) Notes
unstructured/hi_res parse *, ocr * $0.015 · $0.015 · — Unstructured strategy=hi_res (layout model + OCR; coordinates, table HTML, confidence)
unstructured/fast parse, ocr $0.015 · $0.015 · — Unstructured strategy=fast (text-layer extraction, no OCR, rejects images)
unstructured/auto parse, ocr $0.015 · $0.015 · — Unstructured strategy=auto (routes each page to fast / hi_res / VLM)

Upstage Document Parse · UPSTAGE_API_KEY · docs-only — Korean/English layout parsing to HTML or markdown.

Model Modes List price / page (parse · ocr · extract) Notes
upstage/document-parse parse *, ocr * $0.01 · $0.01 · — Upstage Document Parse (layout to HTML/Markdown; 100 pages sync, 1000 async)
upstage/document-parse-nightly parse, ocr $0.01 · $0.01 · — Upstage Document Parse nightly build (newest layout model, may change without notice)

Landing AI (Agentic Document Extraction) · LANDINGAI_API_KEY · docs-only — Agentic Document Extraction (DPT-2): parse, then extract against your schema.

Model Modes List price / page (parse · ocr · extract) Notes
landingai/dpt-2 parse *, ocr *, extract * $0.03 · $0.03 · $0.04 Landing AI ADE DPT-2; extract runs parse then /v1/ade/extract with the JSON schema

Google Cloud Document AI · GOOGLE_DOCUMENTAI_ACCESS_TOKEN (+ GOOGLE_DOCUMENTAI_PROJECT, _LOCATION, _PROCESSOR_ID) · docs-only — Processor-based: pick the processor id in provider_options; OCR, layout, forms and prebuilt extractors.

Model Modes List price / page (parse · ocr · extract) Notes
google_documentai/ocr parse *, ocr * $0.0015 · $0.0015 · — Document OCR processor: native text + word boxes (processor id from config)
google_documentai/layout parse, ocr $0.01 · $0.01 · — Layout Parser processor: documentLayout blocks (headings, tables, lists)
google_documentai/form parse, ocr, extract * $0.03 · $0.03 · $0.03 Form Parser processor: paragraphs + tables, entities as extraction
google_documentai/prebuilt parse, ocr, extract $0.03 · $0.03 · $0.03 Prebuilt or custom extractor (invoice, W2, ...): entities as extraction

Self-hosted engines · no key · $0/page — out-of-process: a local binary or a server you run. Tesseract and Docling are verified locally; PaddleOCR is docs-only.

Model Modes List price / page (parse · ocr · extract) Notes
tesseract/default ocr * (native), parse * $0 · $0 · — local tesseract binary (TESSERACT_CMD), PDFs via pdftoppm; word/line boxes + confidences, no layout model; verified locally
docling/default parse *, ocr * $0 · $0 · — your docling-serve (DOCLING_BASE_URL): layout, tables, OCR; verified locally
paddleocr/default ocr * (native), parse * (PP-StructureV3) $0 · $0 · — your PaddleOCR serving (PADDLEOCR_BASE_URL); docs-only

Keep your Reducto / Extend / LlamaParse code#

Already integrated with a vendor? Ask for its shape and PuffinParse renders the response into that vendor's own JSON, whatever provider actually ran the call — switch the model string, keep your parser:

doc = puffinparse.parse("invoice.pdf", model="extend/parse_light", output_format="reducto")
block = doc["result"]["chunks"][0]["blocks"][0]            # Reducto's shape, Extend's engine
draw(block["bbox"]["left"], block["bbox"]["top"], block["type"])

output_format accepts reducto, extend, llamaparse, or puffinparse (the unified shape, and the default). It returns a plain dict instead of a dataclass, and works on parse, extract, their async variants, the Router, and the CLI (--output-format <vendor> with --format json). An unknown name raises BadRequestError before any network call.

What is guaranteed is structural fidelity — key set and nesting, one chunk/page per page, the content strings, the vendor's own block vocabulary and coordinate units, the billed page count — not byte equality with what the vendor would have returned. Fields PuffinParse does not model (presigned URLs, studio links, billing breakdowns, OCR word layers) are null or empty, and a few block types are lossy. docs/COMPAT.md enumerates all of it, per format; examples/switch_provider_keep_format.py is a runnable version of the above.

Provider-specific options#

Anything the common request doesn't cover is passed through verbatim:

puffinparse.parse("doc.pdf", model="reducto/standard",
              provider_options={"settings": {"return_ocr_data": True}, "async": True})
puffinparse.parse("doc.pdf", model="extend/parse_performance",
              provider_options={"blockOptions": {"figures": {"enabled": False}}})
puffinparse.ocr("doc.pdf", model="llamaparse/agentic",
            provider_options={"take_screenshot": True, "version": "2026-08-19"})

include_raw=True attaches the provider's original payload as resp.raw.

Router: fallbacks and load balancing#

router = puffinparse.Router(
    ["reducto/standard", "llamaparse/agentic", "extend/parse_light"],
    mode="parse",                # the mode every model must serve; "ocr" and "extract" too
    strategy="ordered",          # or "round_robin"
)
resp = router.parse("doc.pdf")   # falls back on provider / rate-limit / timeout / network errors
router.stats()                   # per-model successes, failures, latency, cost, pages

A router is bound to one mode: Router([...], mode="parse").ocr(...) raises InputError rather than quietly changing the answer's shape, and every model is validated against the mode when the router is built. Auth, bad-request and input errors never trigger a fallback.

Errors#

All errors derive from puffinparse.PuffinParseError and carry provider, status_code, job_id and retryable:

AuthenticationError, RateLimitError, BadRequestError, ProviderError, TimeoutError, UnsupportedModelError, InputError, NetworkError.

UnsupportedModelError also covers a model asked to do a mode it does not serve; the message names the mode and the models that do serve it.

Callbacks and logging#

Callbacks fire for every mode, with that mode's response object:

puffinparse.success_callback.append(lambda r: print(r.model, r.usage.pages, r.cost_usd))
puffinparse.failure_callback.append(lambda e: print("failed:", e))
puffinparse.init_logging("debug")    # Rust core tracing on stderr (or PUFFINPARSE_LOG=debug for the CLI)

CLI#

puffinparse providers                                             # models, modes, prices, key status
puffinparse providers --mode ocr                                  # only models that serve a mode
puffinparse parse invoice.pdf -m extend/parse_light               # markdown to stdout
puffinparse parse scan.png -m llamaparse/agentic -f json --raw    # full unified JSON (+ provider payload)
puffinparse parse doc.pdf -m extend/parse_light -f json \
    --output-format reducto                                   # ... in Reducto's response shape
puffinparse parse doc.pdf -m reducto/r-1 --pages 1-3 -f text
puffinparse ocr scan.png -m reducto/standard                      # plain text to stdout
puffinparse ocr scan.png -f json                                  # TextResponse: text + lines + words
puffinparse extract invoice.pdf -s schema.json --citations        # JSON object from a schema
puffinparse extract invoice.pdf -s '{"type":"object"}'            # inline schema also works
puffinparse extract invoice.pdf -s schema.json --output-format extend   # Extend's extract_run shape
puffinparse providers --json | jq '.output_formats'               # the vendor shapes this build renders

Gateway server#

Run PuffinParse as one HTTP endpoint so applications never hold provider keys:

puffinparse serve --config puffinparse.toml       # or: docker run ghcr.io/ajinkyashejul/puffinparse (docs/SERVER.md)
curl -H "Authorization: Bearer $TEAM_KEY" -F file=@invoice.pdf -F model=invoices \
     http://localhost:4000/v1/parse

puffinparse.toml defines aliases (invoices = [reducto/standard, extend/parse_performance] with ordered or round-robin fallback), provider keys as env: references, and virtual keys with model allow-lists, monthly USD budgets and per-minute limits. /v1/models, /v1/usage, /health and Prometheus /metrics are built in; request logs are JSON lines that never contain document content or secrets. Reference: docs/SERVER.md, sample: examples/server/puffinparse.toml.

Rust#

use puffinparse_core::{extract, ocr, parse, DocumentRequest, ExtractRequest};

let doc = parse(DocumentRequest::from_path("invoice.pdf").model("reducto/standard")).await?;
println!("{} pages, ${:.4}\n{}", doc.usage.pages, doc.cost_usd.unwrap_or(0.0), doc.markdown);

let text = ocr(DocumentRequest::from_path("scan.png").model("llamaparse/fast")).await?;
println!("{} lines on page 1", text.pages[0].lines.len());

let req = ExtractRequest::new(DocumentRequest::from_path("invoice.pdf").model("..."), schema);
let data = extract(req.citations(true)).await?.data;

Benchmark#

PuffinParse ships an open, reproducible benchmark. Ground truth is exact by construction (the documents are rendered from the same source as the truth files), metrics are deterministic text comparisons, and every run records the dataset hash, model, latency and cost.

python benchmark/generate_synthetic.py                       # regenerate the dataset (byte-reproducible)
puffinparse bench run --dataset benchmark/datasets/synthetic-v1 \
    --models reducto/standard extend/parse_performance llamaparse/cost_effective
puffinparse bench report benchmark/results/*.json > benchmark/LEADERBOARD.md
puffinparse bench score prediction.md truth.md                   # metrics for one pair, no network

Metrics (after NFKC + markdown stripping + whitespace collapsing, case-insensitive by default):

  • Overall = 100 × mean(char_similarity), where char_similarity = 1 − levenshtein / max(len)
  • CER, WER, word F1 (bag-of-words precision / recall)
  • Order: Kendall-τ-style agreement of shared line order (reading order)
  • Table: character similarity restricted to markdown table rows
  • Latency p50 / p95 and ms per page; $/1k pages from the price table

The current leaderboard is in benchmark/LEADERBOARD.md, and every document, output, diff and rule check is browsable at puffinparse.com/benchmark-results. Datasets: synthetic-v1 (exact truth by construction) and the headline combined-v3, which adds subsets of ParseBench, olmOCR-bench, OmniDocBench and DP-Bench converted by benchmark/adapters/ at pinned revisions. Older combined-v1 and combined-v2 runs are kept for comparison. Long runs are safe: --dry-run and --max-cost show and cap the spend before any call, and --resume continues an interrupted run.

How it maps providers#

Reducto Extend LlamaParse
Upload POST /upload → reducto:// id (URLs passed directly) POST /files/upload → file_… (URLs passed directly) multipart file or input_url
Parse POST /parse (sync) or /parse_async + GET /job/{id} POST /parse_runs + GET /parse_runs/{id} POST /api/v1/parsing/upload + GET …/job/{id} + …/result/json
Pages chunk_mode=page, blocks grouped by bbox.page page chunks, blocks by metadata.page.number pages[].md / items[]
Boxes already normalised divided by metadata.page.width/height bBox divided by page width/height
Usage usage.num_pages, usage.credits metrics.pageCount, usage.credits job_metadata.job_pages

Full details, including the exact wire formats verified against live responses, are in docs/SPEC.md.

Project layout#

crates/puffinparse-core     Rust library: types, providers, router, pricing, benchmark metrics
crates/puffinparse-cli      `puffinparse` binary
crates/puffinparse-python   PyO3 extension (puffinparse._core)
crates/puffinparse-node     napi-rs addon for the Node.js SDK
crates/puffinparse-server   HTTP gateway behind `puffinparse serve`
python/puffinparse          Python package (typed public API)
js/                     Node.js / TypeScript package
benchmark/              dataset generator, adapters, datasets, results, leaderboard, viewer
docs/SPEC.md            specification

Roadmap#

Planned work is tracked in GitHub Issues. Open items:

  • Packaging: publish 0.1.0 to PyPI, npm and crates.io with prebuilt wheels, addons and CLI binaries (#9).
  • Providers: live-verify the docs-only providers (#10), including PaddleOCR against a real server (#17).
  • Benchmark: add tesseract/default (#11) and docling/default with a hardware note (#12) to combined-v3; regenerate the leaderboard in one command (#13); score table-cell neighbour relations exactly (#15); a READoc long-document track (#16).
  • SDK: output_format="mistral" for code written against Mistral OCR responses (#14).
  • Playground: upload a document and compare up to three models side by side (#18).

Support#

PuffinParse is built in the open by one person. If it saved you time, a star on GitHub helps other developers find it. Reports of a wrong score, a missing provider or a confusing doc help even more: open an issue.

Contributing#

See CONTRIBUTING.md. Adding a provider is one Rust file plus a fixture test; the checklist is in the new provider issue template. Issues labelled good first issue are a good place to start, and AGENTS.md summarises the build commands and conventions for coding agents.

License#

MIT. See LICENSE.