# PuffinParse > PuffinParse is one API for every OCR and document-parsing provider (Reducto, Extend, LlamaParse, Mistral, Azure, Textract, Tesseract, Docling, ...). A Rust core with typed Python and TypeScript SDKs, a CLI, a self-hosted gateway, and an open benchmark that ranks providers on accuracy, latency and cost. Switching provider is a one-string change. Source: https://github.com/ajinkyashejul/puffinparse (MIT). The documentation lives under `/docs/`; every page below is also available as HTML at the same URL without the trailing `index.md`, and that HTML URL returns the Markdown when requested with `Accept: text/markdown`. ## When to use PuffinParse - You call hosted document-parsing / OCR providers and want one API for all of them: switching provider is a one-string change (`/`). - You want normalized output: the same markdown, typed blocks, bounding boxes, usage, cost and typed errors whichever provider produced them. - You want to compare providers on your own documents, or read the open benchmark (accuracy, latency, cost per 1,000 pages) before choosing one. ## When not to use it - You only ever use one provider and need features of its native SDK that the unified request does not expose (provider-specific options are passed through, but not every feature maps). - You need fully offline, high-accuracy layout parsing and no hosted provider: PuffinParse can drive local Tesseract and self-hosted Docling, but the accuracy comes from those engines, so using them (or another local parser) directly is simpler. ## For agents - [PuffinParse for agents](/docs/agents/index.md): install, model strings, choosing a model from the benchmark, minimal Python / TypeScript / CLI examples, key-handling rules and an onboarding prompt. The same page is served at [/agents.md](/agents.md). ## Benchmark data - [/benchmark-results/data/index.json](/benchmark-results/data/index.json): every benchmark run. Top-level `datasets` (name, version, licence, document count, categories) and `runs`; each run has `run_id`, `created_at`, `dataset`, `categories`, `file` (the full run) and `models`, where each model has a `summary` with `overall`, the text and table metrics, `latency_p50_ms`, `latency_p95_ms`, `cost_per_1k_pages_usd` and `by_category`. Prefer the newest `combined-v*` run. - `/benchmark-results/data/runs/.json`: one full run, including per-document scores, latency and cost for every model. - `/benchmark-results/data/outputs///.md`: the markdown each model produced for each document. ## Docs - [PuffinParse](/docs/index.md): Overview, install, usage, the model table and how PuffinParse maps each provider. - [Getting started](/docs/getting-started/index.md): Install, set a provider key, and make your first call from Python, the CLI or Rust. - [PuffinParse for agents](/docs/agents/index.md): For coding agents: install, model strings, choosing a model from the benchmark, minimal examples, key handling rules and a copyable onboarding prompt. - [Python SDK](/docs/python/index.md): The complete public Python API: the parse / ocr / extract modes, Router, response dataclasses, exceptions, callbacks, pricing and scoring helpers. - [TypeScript / Node.js SDK](/docs/typescript/index.md): The puffinparse npm package: async parse / ocr / extract on the Rust core, Router, camelCase response types, typed errors, pricing and scoring helpers. - [CLI](/docs/cli/index.md): Every puffinparse subcommand and flag: parse, ocr, extract, providers, and bench run / report / score. - [Gateway server](/docs/gateway/index.md): puffinparse serve: one HTTP endpoint with aliases, fallbacks, virtual keys, budgets, rate limits, JSON logs and Prometheus metrics. - [Rust](/docs/rust/index.md): Using the puffinparse-core crate: modes, the DocumentRequest builder, ParseResponse / TextResponse / ExtractResponse, Router, errors and the bench module. - [Providers](/docs/providers/index.md): Provider matrix: upload method, sync vs async, page and bbox handling, credits and shared behaviour. - [Reducto](/docs/providers/reducto/index.md): Reducto: endpoints, request flow, response mapping, errors, gotchas and provider_options. - [Extend](/docs/providers/extend/index.md): Extend: endpoints, request flow, response mapping, errors, gotchas and provider_options. - [LlamaParse](/docs/providers/llamaparse/index.md): LlamaParse: endpoints, request flow, response mapping, errors, gotchas and provider_options. - [Mistral Document AI](/docs/providers/mistral/index.md): Mistral OCR: single-call parse, paragraph blocks, document annotations for extraction. - [Azure Document Intelligence](/docs/providers/azure/index.md): Azure prebuilt-read / layout / invoice and friends via the 2024-11-30 REST API. - [AWS Textract](/docs/providers/textract/index.md): DetectDocumentText, AnalyzeDocument LAYOUT/TABLES, QUERIES and FORMS with in-house SigV4. - [Google Gemini](/docs/providers/gemini/index.md): Vision-LLM transcription and schema extraction with structured output; no geometry. - [OpenAI](/docs/providers/openai/index.md): Responses API vision transcription and strict-schema extraction. - [Anthropic (Claude)](/docs/providers/anthropic/index.md): Messages API transcription and tool-forced structured extraction. - [Mathpix](/docs/providers/mathpix/index.md): STEM and handwriting OCR to Mathpix Markdown with line and word polygons. - [Datalab (Marker)](/docs/providers/datalab/index.md): Hosted Marker Convert API: fast, balanced and accurate modes. - [Unstructured](/docs/providers/unstructured/index.md): Partition API elements with coordinates and table HTML. - [Upstage Document Parse](/docs/providers/upstage/index.md): Layout parsing to HTML/Markdown with element coordinates. - [Landing AI ADE](/docs/providers/landingai/index.md): Agentic Document Extraction: parse with groundings and schema extraction. - [Google Document AI](/docs/providers/google-documentai/index.md): Document OCR, Layout Parser, Form Parser and prebuilt processors. - [Tesseract (local)](/docs/providers/tesseract/index.md): Local Tesseract 5 via the tesseract binary: native OCR with word/line boxes and confidences, PDFs via pdftoppm. No key, $0. - [Docling (self-hosted)](/docs/providers/docling/index.md): Docling on your own docling-serve: layout parsing, tables and OCR over the v1 async API. No key, $0. - [PaddleOCR (self-hosted)](/docs/providers/paddleocr/index.md): PaddleOCR / PaddleX serving: OCR pipeline and PP-StructureV3 layout parsing. No key, $0. - [Benchmark](/docs/benchmark/index.md): How the open benchmark works: principles, metric definitions, how to run it, datasets and caveats. - [Leaderboard](/docs/benchmark/leaderboard/index.md): Current results: accuracy, latency and cost per 1,000 pages for every model, plus a per-category breakdown. - [Vendor benchmarks](/docs/benchmark/vendor-benchmarks/index.md): Survey of public OCR benchmarks and vendor-published results, with adapter notes for the combined dataset. - [Academic benchmarks](/docs/benchmark/academic-benchmarks/index.md): Survey of olmOCR-bench, OmniDocBench, DP-Bench, READoc and others: what they measure, licences at pinned revisions, how they map to PuffinParse. - [Dataset adapters](/docs/benchmark/adapters/index.md): How public benchmarks are converted into PuffinParse manifests and rules, what is skipped and why. - [Specification](/docs/project/spec/index.md): The product and architecture specification: goals, unified request/response, errors, router, provider mappings, benchmark design. - [Native-format compatibility](/docs/project/compat/index.md): Get responses in Reducto, Extend or LlamaParse native shape from any provider. - [Decisions](/docs/project/decisions/index.md): Architecture decision records — what was decided, why, and what it rules out. - [Contributing](/docs/project/contributing/index.md): Setup, the checks to run, how to add a provider or a dataset, and the PR checklist. - [Security](/docs/project/security/index.md): Supported versions, how to report a vulnerability, and how PuffinParse handles API keys. ## Optional - [Changelog](/docs/project/changelog/index.md): Keep-a-Changelog history of user-visible changes. - [llms-full.txt](/llms-full.txt): every page on this site concatenated, in nav order.