PuffinParse docs
/ llms.txt GitHub

PuffinParse Specification#

One API for every OCR / document-parsing provider. Rust core, Python SDK, CLI, and an open benchmark that ranks providers on accuracy, latency and cost.

Status: v0.1 — providers: Reducto, Extend, LlamaParse.


1. Goals and non-goals#

Goals#

  1. Single call, any provider. puffinparse.parse("invoice.pdf", model="reducto/standard") returns the same ParseResponse shape whether the backend is Reducto, Extend, LlamaParse, or anything added later. Switching provider is a one-string change, within a mode (§3.1).
  2. Fast and lite. The core is a Rust library (puffinparse-core) with a small dependency set. Python only wraps it (PyO3). No provider SDKs are vendored; every provider is talked to over plain HTTPS with reqwest.
  3. Cost tracking. Every response carries usage (pages, provider credits) and a computed cost_usd from an embedded, overridable pricing table.
  4. Reliability primitives. Retries with backoff, per-call timeouts, and a Router with ordered fallbacks and simple load balancing.
  5. Open benchmark. A reproducible harness + datasets + metrics that ranks providers, with results committed to the repo and published as a leaderboard.
  6. Open-source standards. MIT license, CI (fmt, clippy, tests, Python lint + tests), semver, CHANGELOG, contributor docs, typed Python API, docstrings.

Non-goals (v0.1)#

  • A hosted service. (A self-hosted gateway now exists: puffinparse serve, §14.)
  • Bundling or running OCR models in-process. Local and self-hosted engines are supported only as providers that call out to them — the tesseract binary via tokio::process, a docling-serve or PaddleOCR serving endpoint over HTTP (§8.4) — never through C bindings or an embedded runtime.
  • A parsing-only scope. v0.1 specifies and wires three modes end to end (§3.1); which models serve extract is a registry fact reported by list_models("extract"), and the three original providers serve parse and ocr only. Classification and other vendor products remain out of scope.

2. Architecture#

┌────────────────────────────────────────────────────────────────────┐
│  Python SDK (python/puffinparse)         CLI (crates/puffinparse-cli)      │
│  parse / ocr / extract (+ a*)        puffinparse parse | ocr | extract │
│  Router(models, mode=...)            puffinparse providers | bench     │
└───────────────┬────────────────────────────────┬───────────────────┘
                │ PyO3 (crates/puffinparse-python)   │
┌───────────────▼────────────────────────────────▼───────────────────┐
│  puffinparse-core (Rust)                                               │
│  ├─ modes        parse → Parse / ocr → Text / extract → Extract     │
│  ├─ types        DocumentRequest / ParseResponse / TextResponse /   │
│  │               ExtractResponse / Page / Block / Line / Word       │
│  ├─ providers    trait Provider     { reducto, extend, llamaparse } │
│  ├─ router       fallbacks, retries, strategy (ordered/round-robin) │
│  ├─ pricing      embedded per-mode price table → cost_usd           │
│  ├─ input        path | bytes | url  → DocumentInput                │
│  └─ bench        text normalisation + metrics (CER, WER, similarity)│
└────────────────────────────────────────────────────────────────────┘

Crates:

Crate Purpose
crates/puffinparse-core Library. All provider logic, types, router, pricing, benchmark metrics. #![forbid(unsafe_code)].
crates/puffinparse-cli puffinparse binary: parse, ocr, extract, providers, bench run, bench report.
crates/puffinparse-python PyO3 extension module puffinparse._core, built with maturin.
crates/puffinparse-server HTTP gateway behind puffinparse serve (axum): aliases, virtual keys, budgets, metrics. §14, docs/SERVER.md.
python/puffinparse Pure-Python public API, dataclasses, callbacks, typing.
benchmark/ Datasets, manifests, ground truth, results, leaderboard generator.

3. Modes and model naming#

3.1 Modes#

Document-AI vendors sell three different products, and they are not interchangeable. PuffinParse makes that explicit: every call names a mode, and the mode decides both which providers can serve it and what comes back.

Mode Entry point Response What it is for
parse puffinparse.parse / puffinparse_core::parse ParseResponse (§5.1) layout-aware markdown + typed blocks: RAG chunks, tables, structure
ocr puffinparse.ocr / puffinparse_core::ocr TextResponse (§5.2) plain text with line/word boxes: search, redaction, overlays
extract puffinparse.extract / puffinparse_core::extract ExtractResponse (§5.3) a JSON object shaped by a schema, with per-field citations

Providers are swappable only within a mode. A one-string provider switch is only honest between models that do the same job: a markdown parse and a schema extraction are not substitutes for each other, and a router that silently fell back from one to the other would change the shape of the answer. So each model in the registry declares modes: &[Mode], and ModelRef::parse_for(model, mode) rejects a model that does not serve the requested mode — before any network call, with a message naming the mode and the models that do serve it.

Mode is a Rust enum (Mode::{Parse, Ocr, Extract}) with FromStr ("ocr"/"text", "extract"/"extraction", case-insensitive), as_str, and Mode::ALL. In Python it is the string literal type puffinparse.Mode = Literal["parse", "ocr", "extract"]. list_models(mode) (list_models_for in Rust) lists the models for one mode; with no mode it lists all of them.

A provider that has no native OCR endpoint serves ocr from its own parse output (the default Provider::ocr implementation): lines come from block text, words from lines, and the response is tagged metadata["puffinparse_derived_from"] = "parse" so callers can tell native OCR geometry from derived geometry. Nothing is derived across any other mode pair.

3.2 Model naming#

Like LiteLLM, the model string selects provider and model: "<provider>/<model>".

Provider Models (v0.1) Modes Maps to
reducto reducto/standard (default), reducto/r-1, reducto/agentic parse, ocr default settings / settings.model="r-1" / enhance.agentic=[{scope:text},{scope:table}]
extend extend/parse_performance (default), extend/parse_light, extend/parse_auto parse, ocr config.engine
llamaparse llamaparse/fast, llamaparse/cost_effective (default), llamaparse/agentic, llamaparse/agentic_plus parse, ocr tier form field (+ version=latest)

The three providers above serve parse and ocr only. list_models("extract") reports which models (if any) serve extraction; calling extract with a parse-only model raises UnsupportedModelError before any network call.

Aliases: llama, llama_parse, llamacloud → llamaparse. Matching is case-insensitive.

model="reducto" (no slash) selects the provider's default model for the mode being called. Unknown providers or models raise UnsupportedModelError before any network call, as does a known model asked for a mode it does not declare.

Provider-specific knobs that do not fit the common request are passed through provider_options (a JSON object) and merged into the provider request body verbatim. This is the escape hatch; it never changes the response shape.


4. Unified request#

4.1 The common document request#

Every mode takes the same document request; parse and ocr take nothing else.

puffinparse.parse(
    input,                       # str path | pathlib.Path | bytes | "https://..." URL
    model: str = "reducto",      # "<provider>/<model>", must support this mode
    *,
    filename: str | None = None, # required when input is bytes
    pages: str | None = None,    # "1-3,7" 1-based page selection (best effort per provider)
    language: str | None = None, # BCP-47 hint, forwarded if provider supports it
    output: Literal["markdown", "text"] = "markdown",   # preferred `content` of blocks (parse only)
    output_format: str | None = None,   # "reducto" | "extend" | "llamaparse": return that vendor's
                                        # native JSON shape instead of the unified response
                                        # (docs/COMPAT.md, ADR-13); None = unified
    provider_options: dict | None = None,
    include_raw: bool = False,   # attach the provider's raw JSON to response.raw
    timeout: float = 300.0,      # seconds, whole call including polling
    max_retries: int = 2,        # on 429 / 5xx / network errors, exponential backoff
    api_key: str | None = None,  # overrides env var
    base_url: str | None = None, # overrides provider base URL
    metadata: dict | None = None # echoed back, useful for callbacks/logging
) -> ParseResponse

aparse(...) is the async def equivalent. In Rust this is DocumentRequest, a builder over the same fields. DocumentRequest also carries webhook_url: Option<String>, used only when the request is submitted as a job (§15; Python submit(..., webhook_url=...)); parse ignores it.

4.2 ocr#

puffinparse.ocr(input, model="reducto", *, ...same keywords, minus `output`...) -> TextResponse

ocr always returns plain text, so output does not apply; aocr(...) is the async form.

4.3 extract#

puffinparse.extract(
    input,
    schema: dict,                # JSON Schema (draft 2020-12 subset) object for the result
    *,
    model: str = "reducto",      # must support the `extract` mode
    instructions: str | None = None,  # extra natural-language guidance, forwarded if supported
    citations: bool = False,     # ask for per-field page/box/source-text citations
    ...same common keywords as §4.1...
) -> ExtractResponse

aextract(...) is the async form. In Rust this is ExtractRequest { document: DocumentRequest (flattened when serialised), schema, instructions, citations }. A schema that is not a JSON object is rejected before any network call (TypeError in Python, InputError in the core).

4.4 Input handling#

Input handling (DocumentInput) is identical in all three modes:

  • Path → read bytes, sniff MIME from extension (mime_guess), upload.
  • Bytes → require filename (used for MIME + provider upload).
  • URL (http(s)://) → passed to the provider as a remote URL when the provider supports it (all three do); otherwise downloaded and uploaded.

Supported document types are whatever the provider accepts; PuffinParse does not pre-validate beyond a non-empty body.


5. Unified responses#

One response type per mode. All three carry the same envelope — id, provider, model, provider_job_id, usage, cost_usd, latency_ms, created_at, metadata, raw — and differ only in the payload.

5.1 parse → ParseResponse#

@dataclass
class ParseResponse:
    id: str                    # puffinparse-generated uuid
    provider: str              # "reducto"
    model: str                 # "reducto/standard"
    provider_job_id: str | None
    pages: list[Page]
    markdown: str              # whole document, pages joined by "\n\n"
    text: str                  # plain text
    usage: Usage
    cost_usd: float | None     # None if pricing unknown
    latency_ms: int            # wall-clock for the whole call, incl. polling
    created_at: str            # RFC 3339
    metadata: dict
    raw: Any | None            # provider payload if include_raw

@dataclass
class Page:
    page_number: int           # 1-based
    width: float | None        # points/pixels if provider reports it
    height: float | None
    markdown: str
    text: str
    blocks: list[Block]

@dataclass
class Block:
    type: BlockType            # text | title | section_header | list | table | figure |
                               # header | footer | footnote | caption | formula | other
    content: str               # markdown (tables as markdown/HTML per provider)
    text: str | None           # plain text if provider gives a separate one
    bbox: BBox | None          # normalised 0..1 {x0, y0, x1, y1}, origin top-left
    confidence: float | None   # 0..1 if provider reports one
    page_number: int

@dataclass
class Usage:
    pages: int                 # pages billed/processed
    credits: float | None      # provider-native credit units, if any
    provider_cost_usd: float | None   # if the provider reports $ directly

Rules:

  • markdown/text at document level are derived from pages in order.
  • A provider that only returns per-chunk (not per-page) content gets pages reconstructed from block page numbers; if that is impossible, a single page page_number=1 is emitted and Usage.pages still reflects the billed count.
  • Block types are mapped from each provider's vocabulary (see §8). Unknown → other.
  • bbox is normalised so consumers can draw overlays without knowing the page size. Providers reporting absolute coordinates are divided by page dims.

5.2 ocr → TextResponse#

@dataclass
class TextResponse:
    id: str
    provider: str
    model: str
    provider_job_id: str | None
    pages: list[TextPage]
    text: str                  # whole document, pages joined by "\n\n"
    usage: Usage
    cost_usd: float | None
    latency_ms: int
    created_at: str
    metadata: dict             # "puffinparse_derived_from": "parse" when derived (§3.1)
    raw: Any | None

@dataclass
class TextPage:
    page_number: int           # 1-based
    width: float | None
    height: float | None
    text: str                  # plain text in reading order, lines separated by "\n"
    lines: list[Line]
    words: list[Word]

@dataclass
class Line:                    # and Word, identical shape
    text: str
    bbox: BBox | None          # normalised 0..1, origin top-left
    confidence: float | None

No markdown, no block types: ocr is recognition, not layout analysis. When the result is derived from a parse (§3.1), lines carry their block's box and words carry none.

5.3 extract → ExtractResponse#

@dataclass
class ExtractResponse:
    id: str
    provider: str
    model: str
    provider_job_id: str | None
    data: Any                  # the extracted object, shaped by the request schema
    fields: dict[str, FieldInfo]   # keyed by JSON pointer into `data`, e.g. "/invoice/total"
    usage: Usage
    cost_usd: float | None
    latency_ms: int
    created_at: str
    metadata: dict
    raw: Any | None

@dataclass
class FieldInfo:
    confidence: float | None
    citations: list[Citation]

@dataclass
class Citation:
    page_number: int           # 1-based
    bbox: BBox | None          # normalised 0..1, origin top-left
    text: str | None           # source text the value was read from, if reported

fields is empty when the provider reports no per-field metadata; citations is only populated when the request asked for them and the provider supports them. ExtractResponse.field_info(p) and .citations(p) are Python conveniences over the pointer map.


6. Errors#

All errors derive from puffinparse.PuffinParseError:

Error When
AuthenticationError 401/403 from provider or missing API key
RateLimitError 429 (retried first; raised after max_retries)
BadRequestError 4xx other than auth/rate limit
ProviderError 5xx or provider-reported job failure
TimeoutError overall timeout (upload + poll) exceeded
UnsupportedModelError bad model string, or a model that does not serve the requested mode
InputError unreadable file, bytes without filename, empty body, bad mode name, router asked for another mode

Every error carries provider, status_code (if any), message, and request_id/job_id when available.

Mode errors are raised before any network call. UnsupportedModelError from a mode mismatch names the offending model, the modes it does serve, and the models that serve the mode you asked for.


7. Router#

router = puffinparse.Router(
    models=["reducto/standard", "llamaparse/agentic", "extend/parse_light"],
    mode="parse",                # "parse" (default) | "ocr" | "extract"
    strategy="ordered",          # "ordered" (fallback order) | "round_robin"
    fallback_on=("ProviderError", "RateLimitError", "TimeoutError", "NetworkError"),
)
resp = router.parse("doc.pdf")   # tries each in turn / rotates

A router is bound to one mode. Every model is validated against it at construction (UnsupportedModelError otherwise), and calling a method for a different mode — say Router([...], mode="parse").ocr(...) — raises InputError rather than answering with a different shape. router.mode reports it; parse/ocr/extract (and aparse/aocr/ aextract) are the per-mode calls, each taking the same arguments as the module-level function minus model.

Semantics: ordered → try models[0], on a fallback-eligible error move on. round_robin → rotate the starting index per call, then fallback in order. The router records per-model success/failure counts and average latency, exposed as router.stats(). Auth/BadRequest/Input errors never trigger fallback. When a fallback served the call, the response metadata carries puffinparse_fallback_index and puffinparse_fallback_from_error.


8. Provider mapping#

All three mappings below were verified against live API responses on 2026-09-11; the captured payloads live in crates/puffinparse-core/tests/fixtures/ and drive unit tests.

8.1 Reducto#

  • Base: https://platform.reducto.ai, header Authorization: Bearer <key> (REDUCTO_API_KEY, REDUCTO_BASE_URL override).
  • Flow: POST /upload (multipart file) → {file_id: "reducto://…"}; URLs are passed directly. Then POST /parse (sync, 900 s ceiling) with the v3 body {input, retrieval:{chunking:{chunk_mode:"page"}}, formatting:{table_output_format:"md"}, settings:{…}}. With provider_options.async = true: POST /parse_async → {job_id} → poll GET /job/{id} (Pending → Completed | Failed); the parse payload is job.result.
  • pages → settings.page_range = [{start,end}] (1-based).
  • Response: result.type is "full" (chunks[]) or "url" (fetch result.url; the body is the same FullResult object). chunks[].blocks[] carry type (values contain spaces, e.g. "Section Header"), bbox{left,top,width,height,page,original_page} already normalised to 0..1, content, confidence ("high"|"low"), granular_confidence.parse_confidence. usage.num_pages, usage.credits (null on per-product pricing accounts).
  • Block types: Title→title, Section Header→section_header, Text/Key Value/Comment→text, List Item→list, Table→table, Figure→figure, Header→header, Footer→footer, Footnote→footnote, Caption→caption, Formula→formula, else other.
  • Pages: with chunk_mode=page each chunk is one page; page number is taken from the chunk's blocks' bbox.page (chunks themselves have no page field). Chunk content is kept as the page markdown.
  • Errors: {"error":{"code","name","message"},"detail"}, 422 Pydantic arrays, a bare nginx HTML 403 when the header is missing, and a non-standard 442 for password-protected files.

8.2 Extend#

  • Base: https://api.extend.ai (EXTEND_API_KEY, EXTEND_BASE_URL), headers Authorization: Bearer <key> and the mandatory x-extend-api-version: 2026-02-09. provider_options.workspace_id sets x-extend-workspace-id for org-scoped keys.
  • Flow: POST /files/upload (multipart) → {id: "file_…"}; URLs are passed as file:{url,name}. Then POST /parse_runs (async) → poll GET /parse_runs/{id} until status ∈ {PROCESSED, FAILED}. (The sync POST /parse has a 5-minute hard limit, so PuffinParse always uses runs.)
  • Body: {file, config:{target:"markdown", chunkingStrategy:{type:"page"}, engine, blockOptions:{tables:{targetFormat:"markdown"}}, advancedOptions:{pageRanges}}}. provider_options keys target, chunkingStrategy, engine, engineVersion, blockOptions, advancedOptions are merged into config; others (metadata, dataRetention) at top level.
  • Response: the run object itself: output.chunks[] (type:"page", content, metadata.pageRange) with blocks[] (type, content, metadata.page{number,width,height}, metadata.avgOcrConfidence, boundingBox{left,top,right,bottom} in page pixels); metrics.pageCount, usage.credits. responseType=url results are fetched from outputUrl.
  • Block types: heading→title, section_heading→section_header, text/key_value→text, table/table_head/table_cell→table, figure→figure, formula→formula, header, footer, else other (page_number, barcode).
  • Errors: {code, message, requestId, retryable}; failed runs carry failureReason.

8.3 LlamaParse#

  • Base: https://api.cloud.llamaindex.ai (LLAMA_API_KEY, LLAMA_BASE_URL; EU: https://api.cloud.eu.llamaindex.ai), header Authorization: Bearer llx-….
  • Flow: POST /api/v1/parsing/upload (multipart file or input_url; form fields tier, version=latest, language, target_pages (0-based, converted from pages), plus any provider_options as extra form fields) → {id, status}; poll GET /api/v1/parsing/job/{id} until SUCCESS | PARTIAL_SUCCESS | ERROR | CANCELLED; then GET /api/v1/parsing/job/{id}/result/json.
  • Response: pages[].{page (1-based), text, md, items[], width, height}; items have type (heading with lvl, text, table), md, value, bBox{x,y,w,h,confidence} in page units. job_metadata.job_pages; job_credits_usage is 0 until billing settles and is therefore only reported when positive.
  • Block types: heading lvl 1 → title, other headings → section_header, text→text, table→table, else other.
  • Errors: FastAPI {"detail": "…"} / {"detail": [ValidationError]}.

8.4 Self-hosted engines (Tesseract, Docling, PaddleOCR)#

Listed in model::SELF_HOSTED; ProviderInfo::self_hosted() is true, no API key is required (env_var is empty, or names an optional key), prices are 0.0 with source: "self-hosted", and puffinparse providers shows local in the Key column (self_hosted / key_required in --json and in Python providers()).

  • tesseract/default: tesseract <image> stdout ... tsv, PDFs rasterised with pdftoppm; ocr is native (TSV words/lines, confidences), parse is one text block per Tesseract paragraph.
  • docling/default: docling-serve POST /v1/convert/source/async → poll → GET /v1/result; DoclingDocument items mapped to blocks, bottom-left boxes flipped to top-left.
  • paddleocr/default: PaddleX serving POST /ocr (ocr) and POST /layout-parsing (PP-StructureV3, parse).

Details, errors and limits: docs/providers/{tesseract,docling,paddleocr}.md.


9. Pricing#

crates/puffinparse-core/pricing.json (embedded via include_str!) maps "<provider>/<model>" → { "parse": float, "ocr": float, "extract": float, "source": url, "updated": date } — a per-page price per mode, each optional, since vendors price parsing, OCR and extraction differently. cost_usd = usage.pages * price_per_page(model, mode) unless the provider reports credits with a known credit price, in which case credits * per_credit_usd is used. A mode with no price yields cost_usd = None. Users can override one mode at a time with puffinparse.set_pricing({"reducto/standard": 0.01}, "parse"), and ask for an estimate with puffinparse.estimate_cost("reducto/standard", 1000, "ocr"). Prices are best-effort public list prices; the benchmark reports them as such.


10. Benchmark#

10.1 Principles#

  1. Reproducible: every run records provider, model, dataset hash, options, timestamp, and the raw provider outputs (optionally, gitignored).
  2. Machine-checkable ground truth: text-based metrics that don't need an LLM. An optional LLM-judge is a plug-in, never required for the leaderboard.
  3. Three axes: accuracy, latency (p50/p95 per page), cost per 1k pages.
  4. Open datasets only: synthetic documents generated by the repo (so the ground truth is exact) plus adapters for public sets that users download themselves. Implemented adapters: ParseBench (LlamaIndex), olmOCR-bench (AI2), OmniDocBench (OpenDataLab, index only: research-only licence), DP-Bench (Upstage, MIT; folded into combined-v3). Planned: RealDocBench (Extend) and LongExtractBench (Reducto / micro_1) once public. Licences and mappings: docs/benchmarks/academic-benchmarks.md, docs/benchmarks/adapters.md. Vendor-published benchmarks are each won by their publisher; running all of them through one harness with one scoring pipeline is the point of the meta-benchmark. Each adapter is benchmark/adapters/<name>.py, emits a manifest.json, and records the upstream version/commit so results stay tied to a dataset revision. The combined leaderboard reports per-dataset scores and an unweighted mean across datasets.

10.2 Dataset format#

benchmark/datasets/<name>/
  manifest.json          # {name, version, description, license, documents:[…]}
  docs/<id>.<ext>        # input file
  truth/<id>.md          # expected markdown (or .txt for text-only docs)

manifest.documents[]: {id, file, truth, pages, tags:[...], category}. Categories in the built-in synthetic-v1 set: plain, invoice, table, two_column, headings, noisy_scan, low_res, multipage, skewed, dense, faded, receipt, complex_table. The generator (benchmark/generate_synthetic.py) is seeded and byte-reproducible; the truth is produced from the same source the pixels are rendered from.

10.3 Metrics (computed in Rust, puffinparse_core::bench)#

Given predicted P and truth T after normalisation (NFKC, strip markdown syntax and HTML tags, decode HTML entities, straighten quotes/dashes, collapse whitespace, lowercase for the case_insensitive variant):

  • char_similarity = 1 - levenshtein(P, T) / max(|P|, |T|) (the primary metric of a plain transcript document)
  • cer = levenshtein(P, T) / |T|
  • wer = word_levenshtein(P_words, T_words) / |T_words|
  • word_recall = fraction of truth word tokens present in prediction (bag-of-words)
  • word_precision, word_f1
  • table_score (when the truth has a table): char_similarity restricted to table rows. Tables are markdown pipe tables or HTML <table>s (thead/tbody, th/td, colspan/rowspan repeated into every slot they cover, entities decoded), each reduced to rows of normalised cells; a row is its cells joined by a space. puffinparse_core::bench::tables.
  • teds_grid (when the truth has a table): TEDS (Zhong et al. 2020) computed on the grid tree table > row > cell — 1 - TED / max(|Tp|, |Tt|), Zhang–Shasha tree edit distance, unit insert/delete, cell rename = normalised Levenshtein of the contents. It is TEDS without thead/tbody nodes and span attributes, because the ground truth is markdown. Each truth table is matched to its best predicted table; extra predicted tables are not penalised.
  • order_score: Kendall-τ–like agreement of the order of lines shared by both texts

Per document all metrics are recorded, plus the document's headline: table_score for table-only documents, the rule pass rate for kind: rules, char_similarity otherwise. Aggregate = mean over docs (failures count 0), plus per category. Summary.headline = mean of the document headlines (0–1), Overall = 100 * headline; Summary.char_similarity is the plain mean of the documents' char_similarity (a rule document has no transcript, so its char_similarity field carries the pass rate).

Rule matching (score_rules) normalises both sides with the run's options, then drops every space adjacent to punctuation, so tokenised rule text ((this " agreement ")) matches the printed page. bag_of_sentences counts a sentence as present when a window of the prediction is ≥ 0.8 similar (BAG_SENTENCE_MIN_SIMILARITY); the rule passes when at least threshold of them are (default BAG_DEFAULT_THRESHOLD = 0.8). A rule's optional max_diffs (olmOCR-bench) allows that many Levenshtein edits in present / absent / order / table_cell matching.

The scorer is versioned: puffinparse_core::bench::SCORER_VERSION (currently 2) is written to every result as scorer_version and bumped whenever a change would move a committed score.

10.4 Runner and outputs#

puffinparse bench run --dataset benchmark/datasets/synthetic-v1 \
    --models reducto/standard extend/parse_performance llamaparse/agentic \
    --out benchmark/results/<date>-synthetic-v1.json
puffinparse bench report benchmark/results/*.json --format markdown > benchmark/LEADERBOARD.md

Result JSON: {run_id, created_at, puffinparse_version, scorer_version, rescored_at?, dataset:{name, version, documents, sha256}, normalize, models:[{model, docs:[{id, category, kind, table_only, pages, metrics, headline, latency_ms, cost_usd, error}], summary:{documents, failed, headline, char_similarity, cer, wer, word_f1, order_score, table_score, teds_grid, rule_pass_rate, overall, latency_p50_ms, latency_p95_ms, latency_per_page_ms, total_pages, total_cost_usd, cost_per_1k_pages_usd, by_category}}]}. The sha256 covers the manifest plus every input, truth and rule file, so a result is tied to an exact dataset revision.

Consumers (the viewer in benchmark/site/, bench report) rank on summary.headline (0–1) when present, else overall / 100, and show a document's headline when present. Compatibility with older files: a file without scorer_version is scorer v1; its summary.headline is read as overall / 100, teds_grid is absent, and — unlike v2 — its summary.char_similarity held the headline (table-only documents contributed table_score). puffinparse bench rescore <result.json> --outputs <dir> [--dataset <dir>] [--out <path>] re-scores a run from its saved per-document outputs with the current scorer, without network access: metrics, headlines and summaries are recomputed, kind/table_only/category are refreshed from the manifest, latency, cost, pages and errors are kept as measured, scorer_version and rescored_at are set, and the dataset sha256 is updated (with a warning) if the dataset changed since the run. A missing output for a successful document is an error unless --keep-missing is given, which keeps that document's recorded scores and records how many in rescore_kept_docs (used for sources whose outputs are not committed, e.g. research-only datasets). A successful call that returned only whitespace is flagged empty_output: true and counted in summary.empty_outputs; it is scored, not failed.

Each document record also carries audit fields (all optional, so older result files still load): provider_job_id (the provider's id for the call — Reducto job id, Extend parse run id, LlamaParse job id — taken from ParseResponse.provider_job_id, or from Error.job_id for a failure), cache_hit (true if the provider reported a result-cache hit via a <provider>_cache_hit metadata key, false if caches were disabled for the run, null unknown; always written), attempts (calls the runner issued, >1 only with --retries; retries inside the HTTP client are not counted), started_at (RFC 3339, first attempt) and, for failures, error_kind (the ErrorKind serde name, "input" for an unreadable truth/rule file) next to the error message.

Resilience: while running, every finished (model, document) record is appended and flushed to <out>.partial.jsonl (first line {"type":"header", run_id, created_at, dataset_sha256, normalize}, then {"type":"doc", model, doc} per call). The final JSON is written from those records and the log is then deleted. bench run --resume reads the result JSON and/or the partial log, refuses them if the dataset sha256 or normalisation differs, keeps the original run_id, and calls only the pairs without a successful record. A leftover log without --resume is an error, never silently overwritten. --dry-run prints the plan (calls, manifest pages, list-price estimate per model) without network access; --max-cost <usd> aborts before the first call when the estimate exceeds it (or when a planned model has no list price). The run ends with one line: calls made, resumed, failed, total cost, this invocation's cost and wall time.

LEADERBOARD.md is regenerated from committed results and links to each run.

Document kinds: a manifest document is kind: "transcript" (default; truth markdown, scored by the text metrics) or kind: "rules" (a rules file of machine-checkable assertions — present, absent, order, table_cell, bag_of_sentences — scored by puffinparse_core::bench::score_rules, reported as rule_pass_rate / rules_passed / rules_total in Metrics and rule_pass_rate in Summary). Documents tagged table-only are headlined by table_score. Result JSON documents carry kind and table_only; rule files are included in the dataset sha256.


11. Python SDK details#

  • python/puffinparse/__init__.py exports the three modes — parse/aparse, ocr/aocr, extract/aextract — plus Router, the response dataclasses (ParseResponse, Page, Block, BBox, Usage, TextResponse, TextPage, Line, Word, ExtractResponse, FieldInfo, Citation), the Mode literal ("parse" | "ocr" | "extract"), errors, set_pricing, estimate_cost, list_models, resolve_model, providers, modes and the callback lists.
  • Mode arguments are plain strings everywhere (list_models("ocr"), Router([...], mode="ocr"), set_pricing({...}, "ocr"), estimate_cost(m, 1000, "ocr")), typed as puffinparse.Mode.
  • Callbacks: puffinparse.success_callback: list[Callable[[Response], None]] where Response is the union of the three response types, and puffinparse.failure_callback — both fire for every mode, sync and async; awaitables returned by a callback are awaited (on the caller's loop for a* calls, on a private loop otherwise).
  • Bytes never cross the FFI boundary as base64: the Python layer passes the document as a separate bytes argument and the extension builds DocumentInput::Bytes. extract follows the same convention — the request dict is ExtractRequest with the document fields flattened into it.
  • Errors from a mode mismatch are enriched in Python: an UnsupportedModelError from extract/ocr names the mode and the models that serve it (list_models(mode)), so a user who passes a parse-only model to extract is told what to use instead. Passing a non-dict schema raises TypeError before any FFI call.
  • Jobs (§15): submit/asubmit → Job (dataclass mirroring JobHandle), retrieve/aretrieve and handle_webhook/ahandle_webhook → Job | ParseResponse (puffinparse.JobResult), implemented in python/puffinparse/jobs.py over _core.submit, _core.retrieve and _core.parse_webhook.
  • puffinparse.score, normalize_text, markdown_to_text expose the benchmark metrics.
  • Logging: PUFFINPARSE_LOG=debug enables tracing in the core; Python uses logging.getLogger("puffinparse").
  • Typing: fully typed, py.typed shipped; dataclasses mirror the Rust structs 1:1.
  • Build: maturin, abi3-py39 wheels, pip install puffinparse.

12. CLI#

puffinparse parse   <file|url> [--model reducto/standard] [--format markdown|text|json] [--raw]
puffinparse ocr     <file|url> [--model reducto/standard] [--format text|json]
puffinparse extract <file|url> --schema <file.json|inline JSON> [--instructions TEXT] [--citations]
puffinparse providers [--mode parse|ocr|extract] [--json]   # models, modes, per-mode pricing, key status
puffinparse bench run|report|score|rescore                  # see §10 (parse mode; rescore is offline)
puffinparse serve [--config puffinparse.toml] [--host H] [--port P]   # HTTP gateway, see §14

One subcommand per mode; --model must name a model that serves that subcommand's mode, and a bare provider name resolves to its default model for that mode. All three share the common options (--pages, --language, --options, --timeout, --max-retries, --api-key, --base-url, --raw).

  • parse prints markdown (default), plain text, or the whole ParseResponse as JSON.
  • ocr prints the plain text (default) or the whole TextResponse as JSON, which includes per-page lines[] and words[] with boxes.
  • extract always prints the ExtractResponse as JSON. --schema is either a path to a JSON file or inline JSON starting with {.
  • providers lists every model with the modes it serves and its price in each mode; --mode filters the table to one mode and shows a single price column.

Non-JSON output prints a one-line summary to stderr ([model] N page(s) in T ms, est. $X). Exit code 0 on success, 1 on provider error, 2 on usage/config error (including a model that does not serve the requested mode).


13. Quality bar#

  • Rust: cargo fmt --check, cargo clippy -D warnings, unit tests with fixtures for each provider's parser, no network in tests (live tests are #[ignore] and run only when keys are present).
  • Python: ruff, mypy --strict on the package, pytest with a fake core for unit tests; live tests skipped without keys.
  • CI: GitHub Actions on push/PR (Linux; wheels build matrix on tags).
  • Versioning: semver, single workspace version, CHANGELOG.md (Keep a Changelog).

14. Gateway server#

puffinparse serve (crate puffinparse-server) exposes the three modes over HTTP for clients that should not hold provider keys. Operator reference: SERVER.md. Contract:

  • Endpoints. POST /v1/parse | /v1/ocr | /v1/extract take the §4 fields (model, pages, language, output, output_format, provider_options, include_raw, timeout, max_retries, metadata; schema / instructions / citations for extract) plus fallbacks: [str], as JSON (document_url, or base64 document + filename) or multipart (file part + the same fields). api_key, base_url and local paths are rejected. They return the §5 response JSON unchanged, or the vendor shape for output_format (parse, extract). Also GET /v1/models, GET /v1/usage, GET /health, GET /metrics (Prometheus text).
  • Jobs (§15). POST /v1/jobs takes the /v1/parse body plus webhook_url and returns 202 {id, object: "job", status: "pending", job: JobHandle} (without base_url). Auth, allow-lists, aliases, budget pre-check and rpm apply as for /v1/parse, but an alias submits to its first target and fallbacks is rejected (no fallback for jobs). GET /v1/jobs/{id} checks the provider once: {id, object, status: pending|succeeded|failed, model, provider, provider_job_id, submitted_at, result?, error?}, result in the submit-time output_format (or ?output_format=), error the error object below. id is opaque and bound to the submitting key (other keys get 404 not_found; the master key sees all). Cost is charged to that key once, when the job is first observed succeeded; polls count toward rpm, not the budget. Handles (no secrets) live in memory and the state_file for job_retention_hours. Optional POST /v1/webhooks/{provider} ([webhooks] enabled, shared secret as ?token= or x-puffinparse-webhook-secret, else 404) resolves a provider webhook body with core parse_webhook (+ one retrieve when needed) against a stored job and settles it the same way.
  • Config. One TOML file: [server], master_key, [providers.<name>] (api_key, base_url), [[models]] aliases (name, targets, strategy, fallback_on, with the §7 semantics and per-target credential overrides), [[keys]] virtual keys (id, key, models allow-list with provider/* wildcards, monthly_budget_usd, rpm). Secrets may be env:VAR. No master key and no keys means auth is off.
  • Accounting. Spend = response cost_usd, per key per UTC calendar month, checked before each call (402 once spent ≥ budget); rpm is a sliding 60 s window (429 + Retry-After). State is in memory, optionally persisted to a JSON state_file.
  • Errors. One body shape, {"error": {type, message, provider, provider_status, job_id, request_id}}. ErrorKind → HTTP: input / bad_request / unsupported_model → 400, rate_limit → 429, timeout → 504, provider / network / authentication → 502 (provider credentials are the operator's). Gateway-own types: unauthorized 401, budget_exceeded 402, model_not_allowed 403, not_found 404, payload_too_large 413, key_rate_limited 429. The provider's message is passed through.
  • Logs. One JSON line per request: ts, request_id, key_id, method, path, mode, model, served_model, provider, fallback_index, pages, cost_usd, latency_ms, status, error_type, provider_status, plus job_id, job_status on jobs API lines. Never document content or URLs, provider error text, or any secret. Metrics label jobs traffic mode="job_submit" | "job_retrieve" | "webhook" and count puffinparse_jobs_total{event}.

15. Asynchronous jobs and webhooks#

parse blocks until the provider is done (polling job-queue providers internally). The jobs API splits that call in two so the caller owns the waiting — for long documents, large batches, or pipelines driven by provider webhooks. An SDK cannot receive a webhook, so PuffinParse exposes the primitives and a parser for webhook bodies; your web handler does the receiving. parse mode only.

puffinparse_core::submit_parse(DocumentRequest) -> Result<JobHandle>
puffinparse_core::retrieve_parse(&JobHandle) -> Result<JobStatus>            // env credentials
puffinparse_core::retrieve_parse_with(&JobHandle, &RetrieveOptions) -> Result<JobStatus>
puffinparse_core::parse_webhook(model, &serde_json::Value) -> Result<WebhookEvent>   // no network
puffinparse_core::resolve_webhook(model, &Value, &RetrieveOptions) -> Result<JobStatus>
job = puffinparse.submit("big.pdf", model="reducto/standard", webhook_url="https://…/hook")  # -> Job
puffinparse.retrieve(job)                    # -> Job (still running) | ParseResponse; raises on failure
puffinparse.handle_webhook(body, model="reducto")   # -> Job | ParseResponse; raises on failure
# asubmit / aretrieve / ahandle_webhook are the async equivalents
const job = await submit('big.pdf', { model: 'reducto/standard', webhookUrl })  // -> Job (camelCase JobHandle)
await retrieve(job, { apiKey?, outputFormat? })   // -> the same Job | ParseResponse; rejects on failure
await handleWebhook(body, { model: 'reducto' })   // -> Job | ParseResponse; rejects on failure

Over HTTP the gateway exposes the same pair as POST /v1/jobs / GET /v1/jobs/{id} (§14).

JobHandle (Job in Python): provider, model (qualified), job_id (the provider's id), submitted_at (RFC 3339), output, include_raw, base_url, provider_state, metadata. It never contains a secret: API keys are resolved again at retrieve time (env var or RetrieveOptions.api_key / retrieve(api_key=…)), and provider_state holds only the non-secret options a later call needs (Extend workspace_id, responseType). Serialise it with serde / Job.to_dict() to store it or hand it to another process.

JobStatus: Pending | Succeeded(ParseResponse) | Failed(Error). A succeeded job is normalised exactly like parse (qualified model, cost, request metadata echoed, raw only with include_raw); latency_ms is the time since submission. A provider-side failure is Ok(Failed(e)) with e.job_id set — Err means the status check itself failed (auth, network). Python raises the typed exception for Failed.

Webhooks. webhook_url maps to each provider's own per-job setting; parse_webhook reads the body the provider POSTs:

Provider Submit / retrieve webhook_url → Webhook body → WebhookStatus
reducto POST /parse_async / GET /job/{id} async.webhook = {"mode": "direct", "url": …} {"status", "job_id"}: Completed/Failed → Finished (retrieve for result or reason), else Pending
extend POST /parse_runs / GET /parse_runs/{id} rejected (input error): Extend only has workspace webhook endpoints {"eventType": "parse_run.*", "payload": parse_run_status}: PROCESSED → Finished (or Succeeded if a full run with output), FAILED → Failed (reason + message), else Pending
llamaparse POST /api/v1/parsing/upload / GET /api/v1/parsing/job/{id} (+ result/json) multipart field webhook_url the webhook_url result push {"txt","md","json":[pages]} → Succeeded; a LlamaCloud event {"event_type": "parse.*", "data": {"job_id"}} → Finished / Pending

WebhookStatus::Finished means "terminal, but the body carries neither the result nor the error detail"; resolve_webhook (Python handle_webhook) then makes one retrieve. Verifying webhook authenticity (Extend and LlamaCloud HMAC signatures, a secret in Reducto async.metadata) is the caller's job and must happen before the body is trusted. Other providers return unsupported_model from submit_parse.