parse
Layout-aware markdown plus typed blocks with normalised boxes. For RAG chunks, tables and document structure.
{
"markdown": "# Invoice …",
"pages": [{"blocks": [
{"type": "table",
"bbox": [0.08, 0.41, 0.92, 0.63]}
]}]
}
Rust core · Python & TypeScript SDKs · gateway · open benchmark
Three modes — parse, ocr and extract — across 18 providers and 60 models. One request, one response shape, one set of errors, and an open benchmark that ranks every provider on accuracy, latency and cost.
pip install puffinparse
import puffinparse
doc = puffinparse.parse("invoice.pdf", model="reducto/standard")
print(doc.markdown, doc.pages[0].blocks[0].bbox)
print(doc.usage.pages, doc.cost_usd)
{
"markdown": "# ACME Industries\n\n## Invoice …",
"pages": [{"blocks": [
{"type": "title",
"bbox": [0.08, 0.07, 0.62, 0.11]},
{"type": "table",
"bbox": [0.08, 0.41, 0.92, 0.63]}
]}],
"usage": {"pages": 2},
"latency_ms": 2848,
"cost_usd": 0.0300
}
Any provider in, the same typed response out.
Switch in one line
Every provider has its own upload flow, polling loop, block vocabulary, coordinate system and billing unit. PuffinParse hides all of it. Pick a tab: only the highlighted model string changes.
import puffinparse
doc = puffinparse.parse("invoice.pdf", model="reducto/r-1")
doc.markdown # same unified markdown
doc.pages[0].blocks # same typed blocks, same normalised boxes
doc.usage.pages # same billed page count
doc.cost_usd # same field, priced from the model stringimport puffinparse
doc = puffinparse.parse("invoice.pdf", model="extend/parse_performance")
doc.markdown # same unified markdown
doc.pages[0].blocks # same typed blocks, same normalised boxes
doc.usage.pages # same billed page count
doc.cost_usd # same field, priced from the model stringimport puffinparse
doc = puffinparse.parse("invoice.pdf", model="mistral/ocr-latest")
doc.markdown # same unified markdown
doc.pages[0].blocks # same typed blocks, same normalised boxes
doc.usage.pages # same billed page count
doc.cost_usd # same field, priced from the model stringimport puffinparse
doc = puffinparse.parse("invoice.pdf", model="gemini/2.5-flash")
doc.markdown # same unified markdown
doc.pages[0].blocks # same typed blocks, same normalised boxes
doc.usage.pages # same billed page count
doc.cost_usd # same field, priced from the model stringProviders are only interchangeable within a mode. A model that cannot
serve the mode you asked for raises UnsupportedModelError before any network
call — never a quietly different response shape.
The registry
Read straight out of the Rust registry at build time, so this grid cannot drift from the code — 5 of them verified against live API responses, the rest implemented from the official reference with fixture-backed tests.
Provider reference → exactly what PuffinParse sends, what comes back, and how every field is mapped.
Three products, not one
A call picks one mode, and the mode decides the response type. Providers are only interchangeable within a mode.
Layout-aware markdown plus typed blocks with normalised boxes. For RAG chunks, tables and document structure.
{
"markdown": "# Invoice …",
"pages": [{"blocks": [
{"type": "table",
"bbox": [0.08, 0.41, 0.92, 0.63]}
]}]
}
Plain text in reading order with line and word boxes. For search indexes, redaction and overlays.
{
"text": "ACME Industries …",
"pages": [{
"lines": [{"text": "Total due",
"bbox": [0.6, 0.8, 0.8, 0.83]}],
"words": [{"text": "Total", "bbox": […]}]
}]
}
A JSON object shaped by your schema, with per-field confidence and citations back to the page. For invoices, forms, anything with fields.
{
"data": {"invoice_number": "INV-4182",
"total": 1280.5},
"fields": {"/total": {"confidence": 0.98}},
"citations": {"/total": [{"page": 2,
"bbox": […]}]}
}
puffinparse.list_models("extract") or
puffinparse providers --mode extract lists the models that serve a mode — it is a
registry fact, not a guess.
Native-format compatibility
Ask for a vendor's own JSON shape and PuffinParse renders the unified response into it — whatever provider actually produced it. Key set, nesting, block vocabulary, coordinate convention and billed page count, all in the vendor's units.
Your parser is loyal to a vendor. Your code doesn't have to be.
How the compatibility layer works → including
exactly what is guaranteed and what is rendered as null.
doc = puffinparse.parse("invoice.pdf",
model="gemini/2.5-flash",
output_format="reducto")
# -> Reducto's own parse JSON, from Gemini
doc.raw["result"]["chunks"][0]["blocks"][0]["bbox"]["left"]
Open benchmark
Ground truth is exact by construction — the documents are rendered from the same source as the truth files. The metrics are deterministic text comparisons with no LLM judge, provider result caches are disabled, and every per-document output is committed next to the run so you can verify any number yourself.
| # | Model | Overall | p50 latency | $/1k pages |
|---|---|---|---|---|
| 1 | llamaparse/cost_effective | 84.35 | 9,452 ms | $3.75 |
| 2 | llamaparse/agentic | 83.49 | 14,127 ms | $12.50 |
| 3 | reducto/r-1 | 82.40 | 3,459 ms | $10.00 |
| 4 | reducto/standard | 79.78 | 3,019 ms | $15.00 |
| 5 | extend/parse_performance | 77.43 | 21,965 ms | $25.00 |
| 6 | extend/parse_light | 76.46 | 32,038 ms | $6.25 |
combined-v3 v3.0.0 · 199 documents (synthetic, ParseBench, olmOCR-bench and OmniDocBench pages, each scored by its own truth) · higher Overall is better
puffinparse bench run --dataset benchmark/datasets/combined-v2 \
--models reducto/r-1 extend/parse_light llamaparse/cost_effective --dry-run
puffinparse bench report benchmark/results/2026-09-24-combined-v2.json
The results viewer opens every run document by document: the input, the ground truth, each model's raw output, a word-level diff and the commands to reproduce that exact score.
Results viewer → · Full leaderboard → · methodology and caveats → · raw result files on GitHub →
For agents
Every page on this site is also served as plain markdown, linked from the
HTML with <link rel="alternate" type="text/markdown">. No JavaScript is
needed to read anything. There is no MCP server yet — the markdown endpoints are the
interface.
/llms.txtthe site map, one line per
page/llms-full.txtevery page
concatenated, in nav order/docs/<page>/index.mdthe raw
markdown behind any page/docs/project/spec/the contract both
implementations follow