PuffinParse Specification#
One API for every OCR / document-parsing provider. Rust core, Python SDK, CLI, and an open benchmark that ranks providers on accuracy, latency and cost.
Status: v0.1 — providers: Reducto, Extend, LlamaParse.
1. Goals and non-goals#
Goals#
- Single call, any provider.
puffinparse.parse("invoice.pdf", model="reducto/standard")returns the sameParseResponseshape whether the backend is Reducto, Extend, LlamaParse, or anything added later. Switching provider is a one-string change, within a mode (§3.1). - Fast and lite. The core is a Rust library (
puffinparse-core) with a small dependency set. Python only wraps it (PyO3). No provider SDKs are vendored; every provider is talked to over plain HTTPS withreqwest. - Cost tracking. Every response carries
usage(pages, provider credits) and a computedcost_usdfrom an embedded, overridable pricing table. - Reliability primitives. Retries with backoff, per-call timeouts, and a
Routerwith ordered fallbacks and simple load balancing. - Open benchmark. A reproducible harness + datasets + metrics that ranks providers, with results committed to the repo and published as a leaderboard.
- Open-source standards. MIT license, CI (fmt, clippy, tests, Python lint + tests), semver, CHANGELOG, contributor docs, typed Python API, docstrings.
Non-goals (v0.1)#
- A hosted service. (A self-hosted gateway now exists:
puffinparse serve, §14.) - Bundling or running OCR models in-process. Local and self-hosted engines are supported only
as providers that call out to them — the
tesseractbinary viatokio::process, a docling-serve or PaddleOCR serving endpoint over HTTP (§8.4) — never through C bindings or an embedded runtime. - A parsing-only scope. v0.1 specifies and wires three modes end to end (§3.1); which models
serve
extractis a registry fact reported bylist_models("extract"), and the three original providers serveparseandocronly. Classification and other vendor products remain out of scope.
2. Architecture#
┌────────────────────────────────────────────────────────────────────┐
│ Python SDK (python/puffinparse) CLI (crates/puffinparse-cli) │
│ parse / ocr / extract (+ a*) puffinparse parse | ocr | extract │
│ Router(models, mode=...) puffinparse providers | bench │
└───────────────┬────────────────────────────────┬───────────────────┘
│ PyO3 (crates/puffinparse-python) │
┌───────────────▼────────────────────────────────▼───────────────────┐
│ puffinparse-core (Rust) │
│ ├─ modes parse → Parse / ocr → Text / extract → Extract │
│ ├─ types DocumentRequest / ParseResponse / TextResponse / │
│ │ ExtractResponse / Page / Block / Line / Word │
│ ├─ providers trait Provider { reducto, extend, llamaparse } │
│ ├─ router fallbacks, retries, strategy (ordered/round-robin) │
│ ├─ pricing embedded per-mode price table → cost_usd │
│ ├─ input path | bytes | url → DocumentInput │
│ └─ bench text normalisation + metrics (CER, WER, similarity)│
└────────────────────────────────────────────────────────────────────┘
Crates:
| Crate | Purpose |
|---|---|
crates/puffinparse-core |
Library. All provider logic, types, router, pricing, benchmark metrics. #![forbid(unsafe_code)]. |
crates/puffinparse-cli |
puffinparse binary: parse, ocr, extract, providers, bench run, bench report. |
crates/puffinparse-python |
PyO3 extension module puffinparse._core, built with maturin. |
crates/puffinparse-server |
HTTP gateway behind puffinparse serve (axum): aliases, virtual keys, budgets, metrics. §14, docs/SERVER.md. |
python/puffinparse |
Pure-Python public API, dataclasses, callbacks, typing. |
benchmark/ |
Datasets, manifests, ground truth, results, leaderboard generator. |
3. Modes and model naming#
3.1 Modes#
Document-AI vendors sell three different products, and they are not interchangeable. PuffinParse makes that explicit: every call names a mode, and the mode decides both which providers can serve it and what comes back.
| Mode | Entry point | Response | What it is for |
|---|---|---|---|
parse |
puffinparse.parse / puffinparse_core::parse |
ParseResponse (§5.1) |
layout-aware markdown + typed blocks: RAG chunks, tables, structure |
ocr |
puffinparse.ocr / puffinparse_core::ocr |
TextResponse (§5.2) |
plain text with line/word boxes: search, redaction, overlays |
extract |
puffinparse.extract / puffinparse_core::extract |
ExtractResponse (§5.3) |
a JSON object shaped by a schema, with per-field citations |
Providers are swappable only within a mode. A one-string provider switch is only honest
between models that do the same job: a markdown parse and a schema extraction are not
substitutes for each other, and a router that silently fell back from one to the other would
change the shape of the answer. So each model in the registry declares modes: &[Mode], and
ModelRef::parse_for(model, mode) rejects a model that does not serve the requested mode —
before any network call, with a message naming the mode and the models that do serve it.
Mode is a Rust enum (Mode::{Parse, Ocr, Extract}) with FromStr ("ocr"/"text",
"extract"/"extraction", case-insensitive), as_str, and Mode::ALL. In Python it is the
string literal type puffinparse.Mode = Literal["parse", "ocr", "extract"]. list_models(mode)
(list_models_for in Rust) lists the models for one mode; with no mode it lists all of them.
A provider that has no native OCR endpoint serves ocr from its own parse output (the default
Provider::ocr implementation): lines come from block text, words from lines, and the response
is tagged metadata["puffinparse_derived_from"] = "parse" so callers can tell native OCR geometry
from derived geometry. Nothing is derived across any other mode pair.
3.2 Model naming#
Like LiteLLM, the model string selects provider and model: "<provider>/<model>".
| Provider | Models (v0.1) | Modes | Maps to |
|---|---|---|---|
reducto |
reducto/standard (default), reducto/r-1, reducto/agentic |
parse, ocr | default settings / settings.model="r-1" / enhance.agentic=[{scope:text},{scope:table}] |
extend |
extend/parse_performance (default), extend/parse_light, extend/parse_auto |
parse, ocr | config.engine |
llamaparse |
llamaparse/fast, llamaparse/cost_effective (default), llamaparse/agentic, llamaparse/agentic_plus |
parse, ocr | tier form field (+ version=latest) |
The three providers above serve parse and ocr only. list_models("extract") reports which
models (if any) serve extraction; calling extract with a parse-only model raises
UnsupportedModelError before any network call.
Aliases: llama, llama_parse, llamacloud → llamaparse. Matching is case-insensitive.
model="reducto" (no slash) selects the provider's default model for the mode being called.
Unknown providers or models raise UnsupportedModelError before any network call, as does a
known model asked for a mode it does not declare.
Provider-specific knobs that do not fit the common request are passed through
provider_options (a JSON object) and merged into the provider request body
verbatim. This is the escape hatch; it never changes the response shape.
4. Unified request#
4.1 The common document request#
Every mode takes the same document request; parse and ocr take nothing else.
puffinparse.parse(
input, # str path | pathlib.Path | bytes | "https://..." URL
model: str = "reducto", # "<provider>/<model>", must support this mode
*,
filename: str | None = None, # required when input is bytes
pages: str | None = None, # "1-3,7" 1-based page selection (best effort per provider)
language: str | None = None, # BCP-47 hint, forwarded if provider supports it
output: Literal["markdown", "text"] = "markdown", # preferred `content` of blocks (parse only)
output_format: str | None = None, # "reducto" | "extend" | "llamaparse": return that vendor's
# native JSON shape instead of the unified response
# (docs/COMPAT.md, ADR-13); None = unified
provider_options: dict | None = None,
include_raw: bool = False, # attach the provider's raw JSON to response.raw
timeout: float = 300.0, # seconds, whole call including polling
max_retries: int = 2, # on 429 / 5xx / network errors, exponential backoff
api_key: str | None = None, # overrides env var
base_url: str | None = None, # overrides provider base URL
metadata: dict | None = None # echoed back, useful for callbacks/logging
) -> ParseResponse
aparse(...) is the async def equivalent. In Rust this is DocumentRequest, a builder over
the same fields. DocumentRequest also carries webhook_url: Option<String>, used only when the
request is submitted as a job (§15; Python submit(..., webhook_url=...)); parse ignores it.
4.2 ocr#
puffinparse.ocr(input, model="reducto", *, ...same keywords, minus `output`...) -> TextResponse
ocr always returns plain text, so output does not apply; aocr(...) is the async form.
4.3 extract#
puffinparse.extract(
input,
schema: dict, # JSON Schema (draft 2020-12 subset) object for the result
*,
model: str = "reducto", # must support the `extract` mode
instructions: str | None = None, # extra natural-language guidance, forwarded if supported
citations: bool = False, # ask for per-field page/box/source-text citations
...same common keywords as §4.1...
) -> ExtractResponse
aextract(...) is the async form. In Rust this is ExtractRequest { document: DocumentRequest
(flattened when serialised), schema, instructions, citations }. A schema that is not a JSON
object is rejected before any network call (TypeError in Python, InputError in the core).
4.4 Input handling#
Input handling (DocumentInput) is identical in all three modes:
- Path → read bytes, sniff MIME from extension (
mime_guess), upload. - Bytes → require
filename(used for MIME + provider upload). - URL (
http(s)://) → passed to the provider as a remote URL when the provider supports it (all three do); otherwise downloaded and uploaded.
Supported document types are whatever the provider accepts; PuffinParse does not pre-validate beyond a non-empty body.
5. Unified responses#
One response type per mode. All three carry the same envelope — id, provider, model,
provider_job_id, usage, cost_usd, latency_ms, created_at, metadata, raw — and
differ only in the payload.
5.1 parse → ParseResponse#
@dataclass
class ParseResponse:
id: str # puffinparse-generated uuid
provider: str # "reducto"
model: str # "reducto/standard"
provider_job_id: str | None
pages: list[Page]
markdown: str # whole document, pages joined by "\n\n"
text: str # plain text
usage: Usage
cost_usd: float | None # None if pricing unknown
latency_ms: int # wall-clock for the whole call, incl. polling
created_at: str # RFC 3339
metadata: dict
raw: Any | None # provider payload if include_raw
@dataclass
class Page:
page_number: int # 1-based
width: float | None # points/pixels if provider reports it
height: float | None
markdown: str
text: str
blocks: list[Block]
@dataclass
class Block:
type: BlockType # text | title | section_header | list | table | figure |
# header | footer | footnote | caption | formula | other
content: str # markdown (tables as markdown/HTML per provider)
text: str | None # plain text if provider gives a separate one
bbox: BBox | None # normalised 0..1 {x0, y0, x1, y1}, origin top-left
confidence: float | None # 0..1 if provider reports one
page_number: int
@dataclass
class Usage:
pages: int # pages billed/processed
credits: float | None # provider-native credit units, if any
provider_cost_usd: float | None # if the provider reports $ directly
Rules:
markdown/textat document level are derived from pages in order.- A provider that only returns per-chunk (not per-page) content gets pages
reconstructed from block page numbers; if that is impossible, a single page
page_number=1is emitted andUsage.pagesstill reflects the billed count. - Block types are mapped from each provider's vocabulary (see §8). Unknown →
other. bboxis normalised so consumers can draw overlays without knowing the page size. Providers reporting absolute coordinates are divided by page dims.
5.2 ocr → TextResponse#
@dataclass
class TextResponse:
id: str
provider: str
model: str
provider_job_id: str | None
pages: list[TextPage]
text: str # whole document, pages joined by "\n\n"
usage: Usage
cost_usd: float | None
latency_ms: int
created_at: str
metadata: dict # "puffinparse_derived_from": "parse" when derived (§3.1)
raw: Any | None
@dataclass
class TextPage:
page_number: int # 1-based
width: float | None
height: float | None
text: str # plain text in reading order, lines separated by "\n"
lines: list[Line]
words: list[Word]
@dataclass
class Line: # and Word, identical shape
text: str
bbox: BBox | None # normalised 0..1, origin top-left
confidence: float | None
No markdown, no block types: ocr is recognition, not layout analysis. When the result is
derived from a parse (§3.1), lines carry their block's box and words carry none.
5.3 extract → ExtractResponse#
@dataclass
class ExtractResponse:
id: str
provider: str
model: str
provider_job_id: str | None
data: Any # the extracted object, shaped by the request schema
fields: dict[str, FieldInfo] # keyed by JSON pointer into `data`, e.g. "/invoice/total"
usage: Usage
cost_usd: float | None
latency_ms: int
created_at: str
metadata: dict
raw: Any | None
@dataclass
class FieldInfo:
confidence: float | None
citations: list[Citation]
@dataclass
class Citation:
page_number: int # 1-based
bbox: BBox | None # normalised 0..1, origin top-left
text: str | None # source text the value was read from, if reported
fields is empty when the provider reports no per-field metadata; citations is only populated
when the request asked for them and the provider supports them. ExtractResponse.field_info(p)
and .citations(p) are Python conveniences over the pointer map.
6. Errors#
All errors derive from puffinparse.PuffinParseError:
| Error | When |
|---|---|
AuthenticationError |
401/403 from provider or missing API key |
RateLimitError |
429 (retried first; raised after max_retries) |
BadRequestError |
4xx other than auth/rate limit |
ProviderError |
5xx or provider-reported job failure |
TimeoutError |
overall timeout (upload + poll) exceeded |
UnsupportedModelError |
bad model string, or a model that does not serve the requested mode |
InputError |
unreadable file, bytes without filename, empty body, bad mode name, router asked for another mode |
Every error carries provider, status_code (if any), message, and
request_id/job_id when available.
Mode errors are raised before any network call. UnsupportedModelError from a mode mismatch
names the offending model, the modes it does serve, and the models that serve the mode you
asked for.
7. Router#
router = puffinparse.Router(
models=["reducto/standard", "llamaparse/agentic", "extend/parse_light"],
mode="parse", # "parse" (default) | "ocr" | "extract"
strategy="ordered", # "ordered" (fallback order) | "round_robin"
fallback_on=("ProviderError", "RateLimitError", "TimeoutError", "NetworkError"),
)
resp = router.parse("doc.pdf") # tries each in turn / rotates
A router is bound to one mode. Every model is validated against it at construction
(UnsupportedModelError otherwise), and calling a method for a different mode — say
Router([...], mode="parse").ocr(...) — raises InputError rather than answering with a
different shape. router.mode reports it; parse/ocr/extract (and aparse/aocr/
aextract) are the per-mode calls, each taking the same arguments as the module-level function
minus model.
Semantics: ordered → try models[0], on a fallback-eligible error move on.
round_robin → rotate the starting index per call, then fallback in order.
The router records per-model success/failure counts and average latency,
exposed as router.stats(). Auth/BadRequest/Input errors never trigger fallback.
When a fallback served the call, the response metadata carries puffinparse_fallback_index and
puffinparse_fallback_from_error.
8. Provider mapping#
All three mappings below were verified against live API responses on 2026-09-11;
the captured payloads live in crates/puffinparse-core/tests/fixtures/ and drive unit tests.
8.1 Reducto#
- Base:
https://platform.reducto.ai, headerAuthorization: Bearer <key>(REDUCTO_API_KEY,REDUCTO_BASE_URLoverride). - Flow:
POST /upload(multipartfile) →{file_id: "reducto://…"}; URLs are passed directly. ThenPOST /parse(sync, 900 s ceiling) with the v3 body{input, retrieval:{chunking:{chunk_mode:"page"}}, formatting:{table_output_format:"md"}, settings:{…}}. Withprovider_options.async = true:POST /parse_async→{job_id}→ pollGET /job/{id}(Pending→Completed|Failed); the parse payload isjob.result. pages→settings.page_range = [{start,end}](1-based).- Response:
result.typeis"full"(chunks[]) or"url"(fetchresult.url; the body is the sameFullResultobject).chunks[].blocks[]carrytype(values contain spaces, e.g."Section Header"),bbox{left,top,width,height,page,original_page}already normalised to 0..1,content,confidence("high"|"low"),granular_confidence.parse_confidence.usage.num_pages,usage.credits(null on per-product pricing accounts). - Block types:
Title→title,Section Header→section_header,Text/Key Value/Comment→text,List Item→list,Table→table,Figure→figure,Header→header,Footer→footer,Footnote→footnote,Caption→caption,Formula→formula, elseother. - Pages: with
chunk_mode=pageeach chunk is one page; page number is taken from the chunk's blocks'bbox.page(chunks themselves have no page field). Chunkcontentis kept as the page markdown. - Errors:
{"error":{"code","name","message"},"detail"},422Pydantic arrays, a bare nginx HTML403when the header is missing, and a non-standard442for password-protected files.
8.2 Extend#
- Base:
https://api.extend.ai(EXTEND_API_KEY,EXTEND_BASE_URL), headersAuthorization: Bearer <key>and the mandatoryx-extend-api-version: 2026-02-09.provider_options.workspace_idsetsx-extend-workspace-idfor org-scoped keys. - Flow:
POST /files/upload(multipart) →{id: "file_…"}; URLs are passed asfile:{url,name}. ThenPOST /parse_runs(async) → pollGET /parse_runs/{id}untilstatus ∈ {PROCESSED, FAILED}. (The syncPOST /parsehas a 5-minute hard limit, so PuffinParse always uses runs.) - Body:
{file, config:{target:"markdown", chunkingStrategy:{type:"page"}, engine, blockOptions:{tables:{targetFormat:"markdown"}}, advancedOptions:{pageRanges}}}.provider_optionskeystarget,chunkingStrategy,engine,engineVersion,blockOptions,advancedOptionsare merged intoconfig; others (metadata,dataRetention) at top level. - Response: the run object itself:
output.chunks[](type:"page",content,metadata.pageRange) withblocks[](type,content,metadata.page{number,width,height},metadata.avgOcrConfidence,boundingBox{left,top,right,bottom}in page pixels);metrics.pageCount,usage.credits.responseType=urlresults are fetched fromoutputUrl. - Block types:
heading→title,section_heading→section_header,text/key_value→text,table/table_head/table_cell→table,figure→figure,formula→formula,header,footer, elseother(page_number,barcode). - Errors:
{code, message, requestId, retryable}; failed runs carryfailureReason.
8.3 LlamaParse#
- Base:
https://api.cloud.llamaindex.ai(LLAMA_API_KEY,LLAMA_BASE_URL; EU:https://api.cloud.eu.llamaindex.ai), headerAuthorization: Bearer llx-…. - Flow:
POST /api/v1/parsing/upload(multipartfileorinput_url; form fieldstier,version=latest,language,target_pages(0-based, converted frompages), plus anyprovider_optionsas extra form fields) →{id, status}; pollGET /api/v1/parsing/job/{id}untilSUCCESS | PARTIAL_SUCCESS | ERROR | CANCELLED; thenGET /api/v1/parsing/job/{id}/result/json. - Response:
pages[].{page (1-based), text, md, items[], width, height}; items havetype(headingwithlvl,text,table),md,value,bBox{x,y,w,h,confidence}in page units.job_metadata.job_pages;job_credits_usageis0until billing settles and is therefore only reported when positive. - Block types:
headinglvl 1 →title, other headings →section_header,text→text,table→table, elseother. - Errors: FastAPI
{"detail": "…"}/{"detail": [ValidationError]}.
8.4 Self-hosted engines (Tesseract, Docling, PaddleOCR)#
Listed in model::SELF_HOSTED; ProviderInfo::self_hosted() is true, no API key is required
(env_var is empty, or names an optional key), prices are 0.0 with source: "self-hosted", and
puffinparse providers shows local in the Key column (self_hosted / key_required in --json
and in Python providers()).
tesseract/default:tesseract <image> stdout ... tsv, PDFs rasterised withpdftoppm;ocris native (TSV words/lines, confidences),parseis onetextblock per Tesseract paragraph.docling/default: docling-servePOST /v1/convert/source/async→ poll →GET /v1/result; DoclingDocument items mapped to blocks, bottom-left boxes flipped to top-left.paddleocr/default: PaddleX servingPOST /ocr(ocr) andPOST /layout-parsing(PP-StructureV3, parse).
Details, errors and limits: docs/providers/{tesseract,docling,paddleocr}.md.
9. Pricing#
crates/puffinparse-core/pricing.json (embedded via include_str!) maps
"<provider>/<model>" → { "parse": float, "ocr": float, "extract": float, "source": url,
"updated": date } — a per-page price per mode, each optional, since vendors price parsing,
OCR and extraction differently. cost_usd = usage.pages * price_per_page(model, mode) unless the
provider reports credits with a known credit price, in which case credits * per_credit_usd is
used. A mode with no price yields cost_usd = None.
Users can override one mode at a time with
puffinparse.set_pricing({"reducto/standard": 0.01}, "parse"), and ask for an estimate with
puffinparse.estimate_cost("reducto/standard", 1000, "ocr").
Prices are best-effort public list prices; the benchmark reports them as such.
10. Benchmark#
10.1 Principles#
- Reproducible: every run records provider, model, dataset hash, options, timestamp, and the raw provider outputs (optionally, gitignored).
- Machine-checkable ground truth: text-based metrics that don't need an LLM. An optional LLM-judge is a plug-in, never required for the leaderboard.
- Three axes: accuracy, latency (p50/p95 per page), cost per 1k pages.
- Open datasets only: synthetic documents generated by the repo (so the
ground truth is exact) plus adapters for public sets that users download
themselves. Implemented adapters: ParseBench (LlamaIndex), olmOCR-bench (AI2),
OmniDocBench (OpenDataLab, index only: research-only licence), DP-Bench (Upstage, MIT;
folded into
combined-v3). Planned: RealDocBench (Extend) and LongExtractBench (Reducto / micro_1) once public. Licences and mappings:docs/benchmarks/academic-benchmarks.md,docs/benchmarks/adapters.md. Vendor-published benchmarks are each won by their publisher; running all of them through one harness with one scoring pipeline is the point of the meta-benchmark. Each adapter isbenchmark/adapters/<name>.py, emits amanifest.json, and records the upstream version/commit so results stay tied to a dataset revision. The combined leaderboard reports per-dataset scores and an unweighted mean across datasets.
10.2 Dataset format#
benchmark/datasets/<name>/
manifest.json # {name, version, description, license, documents:[…]}
docs/<id>.<ext> # input file
truth/<id>.md # expected markdown (or .txt for text-only docs)
manifest.documents[]: {id, file, truth, pages, tags:[...], category}.
Categories in the built-in synthetic-v1 set: plain, invoice, table,
two_column, headings, noisy_scan, low_res, multipage, skewed, dense, faded,
receipt, complex_table. The generator (benchmark/generate_synthetic.py) is seeded and
byte-reproducible; the truth is produced from the same source the pixels are rendered from.
10.3 Metrics (computed in Rust, puffinparse_core::bench)#
Given predicted P and truth T after normalisation (NFKC, strip markdown syntax and HTML
tags, decode HTML entities, straighten quotes/dashes, collapse whitespace, lowercase for the
case_insensitive variant):
char_similarity = 1 - levenshtein(P, T) / max(|P|, |T|)(the primary metric of a plain transcript document)cer = levenshtein(P, T) / |T|wer = word_levenshtein(P_words, T_words) / |T_words|word_recall= fraction of truth word tokens present in prediction (bag-of-words)word_precision,word_f1table_score(when the truth has a table): char_similarity restricted to table rows. Tables are markdown pipe tables or HTML<table>s (thead/tbody, th/td,colspan/rowspanrepeated into every slot they cover, entities decoded), each reduced to rows of normalised cells; a row is its cells joined by a space.puffinparse_core::bench::tables.teds_grid(when the truth has a table): TEDS (Zhong et al. 2020) computed on the grid treetable > row > cell—1 - TED / max(|Tp|, |Tt|), Zhang–Shasha tree edit distance, unit insert/delete, cell rename = normalised Levenshtein of the contents. It is TEDS withoutthead/tbodynodes and span attributes, because the ground truth is markdown. Each truth table is matched to its best predicted table; extra predicted tables are not penalised.order_score: Kendall-τ–like agreement of the order of lines shared by both texts
Per document all metrics are recorded, plus the document's headline: table_score for
table-only documents, the rule pass rate for kind: rules, char_similarity otherwise.
Aggregate = mean over docs (failures count 0), plus per category. Summary.headline = mean of the
document headlines (0–1), Overall = 100 * headline; Summary.char_similarity is the plain
mean of the documents' char_similarity (a rule document has no transcript, so its
char_similarity field carries the pass rate).
Rule matching (score_rules) normalises both sides with the run's options, then drops every
space adjacent to punctuation, so tokenised rule text ((this " agreement ")) matches the printed
page. bag_of_sentences counts a sentence as present when a window of the prediction is ≥ 0.8
similar (BAG_SENTENCE_MIN_SIMILARITY); the rule passes when at least threshold of them are
(default BAG_DEFAULT_THRESHOLD = 0.8). A rule's optional max_diffs (olmOCR-bench) allows that
many Levenshtein edits in present / absent / order / table_cell matching.
The scorer is versioned: puffinparse_core::bench::SCORER_VERSION (currently 2) is written to every
result as scorer_version and bumped whenever a change would move a committed score.
10.4 Runner and outputs#
puffinparse bench run --dataset benchmark/datasets/synthetic-v1 \
--models reducto/standard extend/parse_performance llamaparse/agentic \
--out benchmark/results/<date>-synthetic-v1.json
puffinparse bench report benchmark/results/*.json --format markdown > benchmark/LEADERBOARD.md
Result JSON: {run_id, created_at, puffinparse_version, scorer_version, rescored_at?,
dataset:{name, version, documents, sha256}, normalize, models:[{model, docs:[{id, category, kind,
table_only, pages, metrics, headline, latency_ms, cost_usd, error}], summary:{documents, failed,
headline, char_similarity, cer, wer, word_f1, order_score, table_score, teds_grid, rule_pass_rate,
overall, latency_p50_ms, latency_p95_ms, latency_per_page_ms, total_pages, total_cost_usd,
cost_per_1k_pages_usd, by_category}}]}. The sha256 covers the manifest plus every input, truth
and rule file, so a result is tied to an exact dataset revision.
Consumers (the viewer in benchmark/site/, bench report) rank on summary.headline (0–1) when
present, else overall / 100, and show a document's headline when present. Compatibility with
older files: a file without scorer_version is scorer v1; its summary.headline is read as
overall / 100, teds_grid is absent, and — unlike v2 — its summary.char_similarity held the
headline (table-only documents contributed table_score). puffinparse bench rescore <result.json>
--outputs <dir> [--dataset <dir>] [--out <path>] re-scores a run from its saved per-document
outputs with the current scorer, without network access: metrics, headlines and summaries are
recomputed, kind/table_only/category are refreshed from the manifest, latency, cost, pages
and errors are kept as measured, scorer_version and rescored_at are set, and the dataset
sha256 is updated (with a warning) if the dataset changed since the run. A missing output for a
successful document is an error unless --keep-missing is given, which keeps that document's
recorded scores and records how many in rescore_kept_docs (used for sources whose outputs are
not committed, e.g. research-only datasets). A successful call that returned only whitespace is
flagged empty_output: true and counted in summary.empty_outputs; it is scored, not failed.
Each document record also carries audit fields (all optional, so older result files still load):
provider_job_id (the provider's id for the call — Reducto job id, Extend parse run id,
LlamaParse job id — taken from ParseResponse.provider_job_id, or from Error.job_id for a
failure), cache_hit (true if the provider reported a result-cache hit via a
<provider>_cache_hit metadata key, false if caches were disabled for the run, null unknown;
always written), attempts (calls the runner issued, >1 only with --retries; retries inside the
HTTP client are not counted), started_at (RFC 3339, first attempt) and, for failures,
error_kind (the ErrorKind serde name, "input" for an unreadable truth/rule file) next to the
error message.
Resilience: while running, every finished (model, document) record is appended and flushed to
<out>.partial.jsonl (first line {"type":"header", run_id, created_at, dataset_sha256,
normalize}, then {"type":"doc", model, doc} per call). The final JSON is written from those
records and the log is then deleted. bench run --resume reads the result JSON and/or the partial
log, refuses them if the dataset sha256 or normalisation differs, keeps the original run_id,
and calls only the pairs without a successful record. A leftover log without --resume is an
error, never silently overwritten. --dry-run prints the plan (calls, manifest pages, list-price
estimate per model) without network access; --max-cost <usd> aborts before the first call when
the estimate exceeds it (or when a planned model has no list price). The run ends with one line:
calls made, resumed, failed, total cost, this invocation's cost and wall time.
LEADERBOARD.md is regenerated from committed results and links to each run.
Document kinds: a manifest document is kind: "transcript" (default; truth markdown, scored by
the text metrics) or kind: "rules" (a rules file of machine-checkable assertions — present,
absent, order, table_cell, bag_of_sentences — scored by puffinparse_core::bench::score_rules,
reported as rule_pass_rate / rules_passed / rules_total in Metrics and rule_pass_rate in
Summary). Documents tagged table-only are headlined by table_score. Result JSON documents
carry kind and table_only; rule files are included in the dataset sha256.
11. Python SDK details#
python/puffinparse/__init__.pyexports the three modes —parse/aparse,ocr/aocr,extract/aextract— plusRouter, the response dataclasses (ParseResponse,Page,Block,BBox,Usage,TextResponse,TextPage,Line,Word,ExtractResponse,FieldInfo,Citation), theModeliteral ("parse" | "ocr" | "extract"), errors,set_pricing,estimate_cost,list_models,resolve_model,providers,modesand the callback lists.- Mode arguments are plain strings everywhere (
list_models("ocr"),Router([...], mode="ocr"),set_pricing({...}, "ocr"),estimate_cost(m, 1000, "ocr")), typed aspuffinparse.Mode. - Callbacks:
puffinparse.success_callback: list[Callable[[Response], None]]whereResponseis the union of the three response types, andpuffinparse.failure_callback— both fire for every mode, sync and async; awaitables returned by a callback are awaited (on the caller's loop fora*calls, on a private loop otherwise). - Bytes never cross the FFI boundary as base64: the Python layer passes the document
as a separate
bytesargument and the extension buildsDocumentInput::Bytes.extractfollows the same convention — the request dict isExtractRequestwith the document fields flattened into it. - Errors from a mode mismatch are enriched in Python: an
UnsupportedModelErrorfromextract/ocrnames the mode and the models that serve it (list_models(mode)), so a user who passes a parse-only model toextractis told what to use instead. Passing a non-dictschemaraisesTypeErrorbefore any FFI call. - Jobs (§15):
submit/asubmit→Job(dataclass mirroringJobHandle),retrieve/aretrieveandhandle_webhook/ahandle_webhook→Job | ParseResponse(puffinparse.JobResult), implemented inpython/puffinparse/jobs.pyover_core.submit,_core.retrieveand_core.parse_webhook. puffinparse.score,normalize_text,markdown_to_textexpose the benchmark metrics.- Logging:
PUFFINPARSE_LOG=debugenables tracing in the core; Python useslogging.getLogger("puffinparse"). - Typing: fully typed,
py.typedshipped; dataclasses mirror the Rust structs 1:1. - Build: maturin,
abi3-py39wheels,pip install puffinparse.
12. CLI#
puffinparse parse <file|url> [--model reducto/standard] [--format markdown|text|json] [--raw]
puffinparse ocr <file|url> [--model reducto/standard] [--format text|json]
puffinparse extract <file|url> --schema <file.json|inline JSON> [--instructions TEXT] [--citations]
puffinparse providers [--mode parse|ocr|extract] [--json] # models, modes, per-mode pricing, key status
puffinparse bench run|report|score|rescore # see §10 (parse mode; rescore is offline)
puffinparse serve [--config puffinparse.toml] [--host H] [--port P] # HTTP gateway, see §14
One subcommand per mode; --model must name a model that serves that subcommand's mode, and a
bare provider name resolves to its default model for that mode. All three share the common
options (--pages, --language, --options, --timeout, --max-retries, --api-key,
--base-url, --raw).
parseprints markdown (default), plain text, or the wholeParseResponseas JSON.ocrprints the plain text (default) or the wholeTextResponseas JSON, which includes per-pagelines[]andwords[]with boxes.extractalways prints theExtractResponseas JSON.--schemais either a path to a JSON file or inline JSON starting with{.providerslists every model with the modes it serves and its price in each mode;--modefilters the table to one mode and shows a single price column.
Non-JSON output prints a one-line summary to stderr ([model] N page(s) in T ms, est. $X).
Exit code 0 on success, 1 on provider error, 2 on usage/config error (including a model that
does not serve the requested mode).
13. Quality bar#
- Rust:
cargo fmt --check,cargo clippy -D warnings, unit tests with fixtures for each provider's parser, no network in tests (live tests are#[ignore]and run only when keys are present). - Python:
ruff,mypy --stricton the package,pytestwith a fake core for unit tests; live tests skipped without keys. - CI: GitHub Actions on push/PR (Linux; wheels build matrix on tags).
- Versioning: semver, single workspace version,
CHANGELOG.md(Keep a Changelog).
14. Gateway server#
puffinparse serve (crate puffinparse-server) exposes the three modes over HTTP for clients that should
not hold provider keys. Operator reference: SERVER.md. Contract:
- Endpoints.
POST /v1/parse | /v1/ocr | /v1/extracttake the §4 fields (model,pages,language,output,output_format,provider_options,include_raw,timeout,max_retries,metadata;schema/instructions/citationsfor extract) plusfallbacks: [str], as JSON (document_url, or base64document+filename) or multipart (filepart + the same fields).api_key,base_urland local paths are rejected. They return the §5 response JSON unchanged, or the vendor shape foroutput_format(parse, extract). AlsoGET /v1/models,GET /v1/usage,GET /health,GET /metrics(Prometheus text). - Jobs (§15).
POST /v1/jobstakes the/v1/parsebody pluswebhook_urland returns202 {id, object: "job", status: "pending", job: JobHandle}(withoutbase_url). Auth, allow-lists, aliases, budget pre-check andrpmapply as for/v1/parse, but an alias submits to its first target andfallbacksis rejected (no fallback for jobs).GET /v1/jobs/{id}checks the provider once:{id, object, status: pending|succeeded|failed, model, provider, provider_job_id, submitted_at, result?, error?},resultin the submit-timeoutput_format(or?output_format=),errorthe error object below.idis opaque and bound to the submitting key (other keys get404 not_found; the master key sees all). Cost is charged to that key once, when the job is first observed succeeded; polls count towardrpm, not the budget. Handles (no secrets) live in memory and thestate_fileforjob_retention_hours. OptionalPOST /v1/webhooks/{provider}([webhooks] enabled, sharedsecretas?token=orx-puffinparse-webhook-secret, else 404) resolves a provider webhook body with coreparse_webhook(+ one retrieve when needed) against a stored job and settles it the same way. - Config. One TOML file:
[server],master_key,[providers.<name>](api_key,base_url),[[models]]aliases (name,targets,strategy,fallback_on, with the §7 semantics and per-target credential overrides),[[keys]]virtual keys (id,key,modelsallow-list withprovider/*wildcards,monthly_budget_usd,rpm). Secrets may beenv:VAR. No master key and no keys means auth is off. - Accounting. Spend = response
cost_usd, per key per UTC calendar month, checked before each call (402once spent ≥ budget);rpmis a sliding 60 s window (429+Retry-After). State is in memory, optionally persisted to a JSONstate_file. - Errors. One body shape,
{"error": {type, message, provider, provider_status, job_id, request_id}}.ErrorKind→ HTTP: input / bad_request / unsupported_model → 400, rate_limit → 429, timeout → 504, provider / network / authentication → 502 (provider credentials are the operator's). Gateway-own types:unauthorized401,budget_exceeded402,model_not_allowed403,not_found404,payload_too_large413,key_rate_limited429. The provider's message is passed through. - Logs. One JSON line per request:
ts, request_id, key_id, method, path, mode, model, served_model, provider, fallback_index, pages, cost_usd, latency_ms, status, error_type, provider_status, plusjob_id, job_statuson jobs API lines. Never document content or URLs, provider error text, or any secret. Metrics label jobs trafficmode="job_submit" | "job_retrieve" | "webhook"and countpuffinparse_jobs_total{event}.
15. Asynchronous jobs and webhooks#
parse blocks until the provider is done (polling job-queue providers internally). The jobs API
splits that call in two so the caller owns the waiting — for long documents, large batches, or
pipelines driven by provider webhooks. An SDK cannot receive a webhook, so PuffinParse exposes the
primitives and a parser for webhook bodies; your web handler does the receiving. parse mode only.
puffinparse_core::submit_parse(DocumentRequest) -> Result<JobHandle>
puffinparse_core::retrieve_parse(&JobHandle) -> Result<JobStatus> // env credentials
puffinparse_core::retrieve_parse_with(&JobHandle, &RetrieveOptions) -> Result<JobStatus>
puffinparse_core::parse_webhook(model, &serde_json::Value) -> Result<WebhookEvent> // no network
puffinparse_core::resolve_webhook(model, &Value, &RetrieveOptions) -> Result<JobStatus>
job = puffinparse.submit("big.pdf", model="reducto/standard", webhook_url="https://…/hook") # -> Job
puffinparse.retrieve(job) # -> Job (still running) | ParseResponse; raises on failure
puffinparse.handle_webhook(body, model="reducto") # -> Job | ParseResponse; raises on failure
# asubmit / aretrieve / ahandle_webhook are the async equivalents
const job = await submit('big.pdf', { model: 'reducto/standard', webhookUrl }) // -> Job (camelCase JobHandle)
await retrieve(job, { apiKey?, outputFormat? }) // -> the same Job | ParseResponse; rejects on failure
await handleWebhook(body, { model: 'reducto' }) // -> Job | ParseResponse; rejects on failure
Over HTTP the gateway exposes the same pair as POST /v1/jobs / GET /v1/jobs/{id} (§14).
JobHandle (Job in Python): provider, model (qualified), job_id (the provider's id),
submitted_at (RFC 3339), output, include_raw, base_url, provider_state, metadata.
It never contains a secret: API keys are resolved again at retrieve time (env var or
RetrieveOptions.api_key / retrieve(api_key=…)), and provider_state holds only the
non-secret options a later call needs (Extend workspace_id, responseType). Serialise it with
serde / Job.to_dict() to store it or hand it to another process.
JobStatus: Pending | Succeeded(ParseResponse) | Failed(Error). A succeeded job is
normalised exactly like parse (qualified model, cost, request metadata echoed, raw only with
include_raw); latency_ms is the time since submission. A provider-side failure is
Ok(Failed(e)) with e.job_id set — Err means the status check itself failed (auth, network).
Python raises the typed exception for Failed.
Webhooks. webhook_url maps to each provider's own per-job setting; parse_webhook reads
the body the provider POSTs:
| Provider | Submit / retrieve | webhook_url → |
Webhook body → WebhookStatus |
|---|---|---|---|
reducto |
POST /parse_async / GET /job/{id} |
async.webhook = {"mode": "direct", "url": …} |
{"status", "job_id"}: Completed/Failed → Finished (retrieve for result or reason), else Pending |
extend |
POST /parse_runs / GET /parse_runs/{id} |
rejected (input error): Extend only has workspace webhook endpoints |
{"eventType": "parse_run.*", "payload": parse_run_status}: PROCESSED → Finished (or Succeeded if a full run with output), FAILED → Failed (reason + message), else Pending |
llamaparse |
POST /api/v1/parsing/upload / GET /api/v1/parsing/job/{id} (+ result/json) |
multipart field webhook_url |
the webhook_url result push {"txt","md","json":[pages]} → Succeeded; a LlamaCloud event {"event_type": "parse.*", "data": {"job_id"}} → Finished / Pending |
WebhookStatus::Finished means "terminal, but the body carries neither the result nor the error
detail"; resolve_webhook (Python handle_webhook) then makes one retrieve. Verifying webhook
authenticity (Extend and LlamaCloud HMAC signatures, a secret in Reducto async.metadata) is the
caller's job and must happen before the body is trusted. Other providers return
unsupported_model from submit_parse.