PuffinParse docs
/ llms.txt GitHub

Python SDK#

pip install puffinparse — Python 3.9+, fully typed (py.typed, mypy --strict clean). The package is a thin, typed wrapper over the Rust core (puffinparse._core); no provider logic lives in Python.

Everything below is the public surface exported from puffinparse.__all__.

import puffinparse

puffinparse.__version__          # version of the compiled core
puffinparse.modes()              # ["parse", "ocr", "extract"]

Modes#

Every call names a mode, and the mode decides the response type. Providers can be swapped freely within a mode; a model that does not serve the mode you asked for raises UnsupportedModelError before any network call.

Mode Call Async Returns Use it for
parse puffinparse.parse puffinparse.aparse ParseResponse Layout-aware markdown and typed blocks with boxes.
ocr puffinparse.ocr puffinparse.aocr TextResponse Plain text with line and word boxes — search indexes, redaction, overlays.
extract puffinparse.extract puffinparse.aextract ExtractResponse A JSON object shaped by your schema, with per-field citations.

Mode is Literal["parse", "ocr", "extract"] and MODES is the same tuple as a constant.

doc  = puffinparse.parse("invoice.pdf", model="reducto/standard")     # markdown + blocks
text = puffinparse.ocr("scan.png", model="llamaparse/fast")           # plain text + boxes
data = puffinparse.extract("invoice.pdf", schema, model="reducto/extract")

parse#

def parse(
    input: DocumentLike,
    model: str = "reducto",
    *,
    output_format: Optional[str] = None,
    filename: Optional[str] = None,
    pages: Optional[str] = None,
    language: Optional[str] = None,
    output: Literal["markdown", "text"] = "markdown",
    provider_options: Optional[dict[str, Any]] = None,
    include_raw: bool = False,
    timeout: float = 300.0,
    max_retries: int = 2,
    api_key: Optional[str] = None,
    base_url: Optional[str] = None,
    metadata: Optional[dict[str, Any]] = None,
) -> ParseResponse            # or dict[str, Any] when output_format is given

DocumentLike is str | os.PathLike[str] | bytes | bytearray | memoryview.

Parameter Meaning
input A file path, an http(s):// URL, or raw bytes (then filename is required).
model "<provider>/<model>", e.g. "reducto/standard". A bare provider name selects its default model for this mode.
filename Required when input is bytes; used to infer the document type.
pages 1-based page selection such as "1-3,7", forwarded best-effort.
language Language hint (ISO 639-1) when the provider supports it.
output Preferred block content: "markdown" (default) or "text".
output_format Return a vendor's own JSON dict instead of the dataclass: "reducto", "extend", "llamaparse", or "puffinparse" / None for the unified shape. See Native-format output.
provider_options Provider-specific options merged verbatim into the provider request body.
include_raw Attach the provider's raw payload as response.raw.
timeout Whole-call deadline in seconds (upload + polling + download).
max_retries Retries on 429 / 5xx / network errors with exponential backoff and full jitter.
api_key Override the API key (otherwise read from REDUCTO_API_KEY, EXTEND_API_KEY, LLAMA_API_KEY).
base_url Override the provider base URL.
metadata Free-form dict echoed back in response.metadata.

Passing bytes without filename raises InputError before any network call.

resp = puffinparse.parse("contract.pdf", model="extend/parse_performance")
resp = puffinparse.parse("https://example.com/doc.pdf", model="reducto/r-1", pages="1-2")
resp = puffinparse.parse(open("scan.png", "rb").read(), filename="scan.png", model="llamaparse/agentic")

aparse is the same call with await. It runs on the Rust runtime, so the event loop is never blocked.

resp = await puffinparse.aparse("contract.pdf", model="llamaparse/cost_effective")

ocr#

def ocr(
    input: DocumentLike,
    model: str = "reducto",
    *,
    filename: Optional[str] = None,
    pages: Optional[str] = None,
    language: Optional[str] = None,
    provider_options: Optional[dict[str, Any]] = None,
    include_raw: bool = False,
    timeout: float = 300.0,
    max_retries: int = 2,
    api_key: Optional[str] = None,
    base_url: Optional[str] = None,
    metadata: Optional[dict[str, Any]] = None,
) -> TextResponse

Same parameters as parse minus output (the mode implies plain text). Use it when you want text and geometry rather than document structure; for markdown, tables and block types use parse.

Providers without a native OCR endpoint serve this mode from their parse output. The response then carries metadata["puffinparse_derived_from"] == "parse", so you can tell the difference.

text = puffinparse.ocr("scan.png", model="reducto/r-1")
text.text                                    # whole document
for line in text.pages[0].lines:
    line.text, line.bbox, line.confidence
text = await puffinparse.aocr("scan.png")        # async

extract#

def extract(
    input: DocumentLike,
    schema: dict[str, Any],
    *,
    output_format: Optional[str] = None,
    model: str = "reducto",
    instructions: Optional[str] = None,
    citations: bool = False,
    filename: Optional[str] = None,
    pages: Optional[str] = None,
    language: Optional[str] = None,
    provider_options: Optional[dict[str, Any]] = None,
    include_raw: bool = False,
    timeout: float = 300.0,
    max_retries: int = 2,
    api_key: Optional[str] = None,
    base_url: Optional[str] = None,
    metadata: Optional[dict[str, Any]] = None,
) -> ExtractResponse          # or dict[str, Any] when output_format is given
Parameter Meaning
schema A JSON Schema object describing the fields you want.
model Must support extract; a parse-only model raises UnsupportedModelError before any network call.
instructions Optional natural-language guidance, forwarded to providers that accept it.
citations Ask for per-field citations (page, box, source text) where the provider supports them.
output_format As for parse, rendering the vendor's extract envelope. Best effort — see docs/COMPAT.md §7.

Everything else matches parse. aextract is the async variant.

schema = {
    "type": "object",
    "properties": {
        "vendor": {"type": "string"},
        "total": {"type": "number"},
    },
}

resp = puffinparse.extract("invoice.pdf", schema, model="reducto/extract",
                       instructions="Totals are inclusive of tax.", citations=True)

resp.data["total"]                      # the object your schema asked for
resp.citations("/total")                # [Citation(page_number=1, bbox=..., text="...")]
resp.field_info("/total").confidence

Router#

class Router:
    def __init__(
        self,
        models: list[str],
        *,
        mode: Mode = "parse",
        strategy: Literal["ordered", "round_robin"] = "ordered",
        fallback_on: Optional[list[str]] = None,
    ) -> None

A router is bound to one mode at construction: every model must support it, and calling a method for another mode raises InputError. That keeps fallbacks honest — a parse-only model can never quietly answer an extraction.

Member Type Description
models list[str] (property) The canonicalised model list.
mode Mode (property) The mode this router serves.
plan() list[str] The order models would be tried for the next call. Advances the round-robin cursor.
stats() dict[str, dict[str, Any]] Per model: successes, failures, total_latency_ms, total_cost_usd, total_pages.
parse(input, **kw) / aparse ParseResponse Same keyword arguments as puffinparse.parse except model, output_format included.
ocr(input, **kw) / aocr TextResponse Same, minus output.
extract(input, schema, *, output_format, instructions, citations, **kw) / aextract ExtractResponse Same as puffinparse.extract except model.

fallback_on is a list of error-kind names; the default is provider, rate_limit, timeout, network. Authentication, bad-request, unsupported-model and input errors never trigger a fallback. Unknown keyword arguments raise TypeError and the message lists what is accepted.

router = puffinparse.Router(["reducto/standard", "llamaparse/agentic"], strategy="round_robin")
resp = router.parse("doc.pdf", pages="1-5", timeout=120)
router.stats()["reducto/standard"]["successes"]

text_router = puffinparse.Router(["reducto/r-1", "extend/parse_light"], mode="ocr")
text_router.ocr("scan.png").text

Native-format output (output_format)#

PuffinParse normalises every provider to one response shape, which is the right default — and a migration cost if you are already integrated with a vendor. output_format removes it: ask for a vendor's shape and the response is rendered into that vendor's own JSON, whatever provider actually produced it.

doc = puffinparse.parse("invoice.pdf", model="extend/parse_light", output_format="reducto")

for chunk in doc["result"]["chunks"]:          # Reducto's shape, Extend's engine
    for block in chunk["blocks"]:
        draw(block["bbox"]["left"], block["bbox"]["top"], block["type"])
Value Shape
None (default), "puffinparse", "unified" PuffinParse's own response — a dataclass for None, the same JSON as a dict for "puffinparse".
"reducto" Reducto POST /parse response (response_type: "parse").
"extend" Extend parse_run object (GET /parse_runs/{id}).
"llamaparse" ("llama", "llama_parse") LlamaParse …/result/json payload.
  • Available on parse, aparse, extract, aextract, Router.parse / aparse / extract / aextract, and the CLI (--output-format, with --format json).
  • Names are case-insensitive and -/_ are interchangeable; puffinparse.output_formats() lists them. An unknown value raises BadRequestError before any network call.
  • The return type follows the argument: None gives the dataclass, a string gives a dict. Both are @overload-typed, so a type checker knows which one it is.
  • Callbacks always receive the dataclass, whatever the caller asked for — logging and cost tracking are unaffected.
  • output_format is independent of output: output picks markdown vs plain text inside block content, output_format picks the JSON envelope around it.

What is guaranteed is structural fidelity, not semantic identity: the key set and nesting, one chunk/page per unified page, the content strings, the vendor's own block vocabulary and coordinate units, and the billed page count. Not guaranteed: byte equality with what the vendor would have returned, fields PuffinParse does not model (they are rendered as null / [], never invented), or vendor-specific enrichments. Extract-mode rendering is explicitly best effort. docs/COMPAT.md lists every always-null field, the lossy block-type mappings and the coordinate conversions, per format.

puffinparse.output_formats()      # ["puffinparse", "reducto", "extend", "llamaparse"]

# same call, both shapes
doc = puffinparse.parse("invoice.pdf", model="reducto/standard")                       # ParseResponse
raw = puffinparse.parse("invoice.pdf", model="reducto/standard", output_format="extend")
raw["object"], raw["status"], raw["metrics"]["pageCount"]

router = puffinparse.Router(["reducto/standard", "extend/parse_light"])
router.parse("doc.pdf", output_format="llamaparse")["pages"][0]["items"]

Response types#

All response dataclasses mirror the Rust structs in puffinparse-core 1:1 and live in puffinparse.types. Response is the union ParseResponse | TextResponse | ExtractResponse.

Every response carries the same envelope: id, provider, model, provider_job_id, usage, cost_usd, latency_ms, created_at, metadata, raw, plus to_dict() and a from_dict() classmethod. cost_usd is pages × per_page_usd for the mode from the embedded price table (no provider reports dollars directly), and raw is None unless the call passed include_raw=True.

ParseResponse#

@dataclass
class ParseResponse:
    id: str
    provider: str
    model: str
    pages: list[Page]
    markdown: str
    text: str
    usage: Usage
    latency_ms: int
    created_at: str
    provider_job_id: Optional[str] = None
    cost_usd: Optional[float] = None
    metadata: dict[str, Any] = field(default_factory=dict)
    raw: Any = None
Member Description
num_pages (property) len(self.pages).
blocks (property) Every block from every page, in reading order.
tables (property) Blocks whose type is "table".
__str__ Returns markdown.

Page and Block#

@dataclass
class Page:
    page_number: int
    markdown: str
    text: str
    blocks: list[Block] = field(default_factory=list)
    width: Optional[float] = None
    height: Optional[float] = None

@dataclass
class Block:
    type: BlockType
    content: str
    page_number: int
    text: Optional[str] = None
    bbox: Optional[BBox] = None
    confidence: Optional[float] = None

Page.blocks_of(*types) filters by block type; Page.tables is shorthand for blocks_of("table"). width / height are None for providers that do not report page dimensions (Reducto).

BlockType is a Literal of "text", "title", "section_header", "list", "table", "figure", "header", "footer", "footnote", "caption", "formula", "other".

TextResponse, TextPage, Line, Word#

@dataclass
class TextResponse:
    id: str
    provider: str
    model: str
    pages: list[TextPage]
    text: str
    usage: Usage
    latency_ms: int
    created_at: str
    provider_job_id: Optional[str] = None
    cost_usd: Optional[float] = None
    metadata: dict[str, Any] = field(default_factory=dict)
    raw: Any = None

@dataclass
class TextPage:
    page_number: int
    text: str
    lines: list[Line] = field(default_factory=list)
    words: list[Word] = field(default_factory=list)
    width: Optional[float] = None
    height: Optional[float] = None

Line and Word are both { text: str, bbox: Optional[BBox], confidence: Optional[float] }. TextResponse.lines and .words flatten across pages, num_pages is the page count, and __str__ returns text.

ExtractResponse, FieldInfo, Citation#

@dataclass
class ExtractResponse:
    id: str
    provider: str
    model: str
    data: Any
    usage: Usage
    latency_ms: int
    created_at: str
    provider_job_id: Optional[str] = None
    fields: dict[str, FieldInfo] = field(default_factory=dict)
    cost_usd: Optional[float] = None
    metadata: dict[str, Any] = field(default_factory=dict)
    raw: Any = None

fields is keyed by JSON pointer into data ("/invoice/total"). field_info(pointer) returns the FieldInfo or None; citations(pointer) returns its citation list, or an empty list when the provider reported none. __str__ returns str(data).

@dataclass
class FieldInfo:
    confidence: Optional[float] = None
    citations: list[Citation] = field(default_factory=list)

@dataclass
class Citation:
    page_number: int                 # 1-based
    bbox: Optional[BBox] = None
    text: Optional[str] = None       # source text the value was read from

BBox and Usage#

@dataclass(frozen=True)
class BBox:
    x0: float
    y0: float
    x1: float
    y1: float

@dataclass
class Usage:
    pages: int = 0
    credits: Optional[float] = None
    provider_cost_usd: Optional[float] = None

Boxes are normalised to 0..1 relative to page size, origin top-left. BBox.width and .height are properties; to_pixels(page_width, page_height) returns absolute (x0, y0, x1, y1).

Metrics#

Returned by puffinparse.score.

@dataclass
class Metrics:
    char_similarity: float
    cer: float
    wer: float
    word_recall: float
    word_precision: float
    word_f1: float
    pred_chars: int
    truth_chars: int
    order_score: Optional[float] = None
    table_score: Optional[float] = None

Exceptions#

Every error raised by PuffinParse derives from PuffinParseError.

PuffinParseError                 kind = "error"
├── AuthenticationError      "authentication_error"    401/403, or no API key configured
├── RateLimitError           "rate_limit_error"        429 after retries were exhausted
├── BadRequestError          "bad_request_error"       other 4xx — the request itself is wrong
├── ProviderError            "provider_error"          5xx, malformed payload, failed job
├── TimeoutError             "timeout_error"           the whole-call deadline was exceeded
├── UnsupportedModelError    "unsupported_model_error" unknown model, or one that cannot serve the mode
├── InputError               "input_error"             unreadable input, bytes without filename
└── NetworkError             "network_error"           network / TLS / DNS failure after retries

Every instance carries:

Attribute Type Description
message str The provider's message, never swallowed.
kind str (class attribute) Stable machine-readable kind, as above.
provider Optional[str] Provider the call was routed to.
status_code Optional[int] HTTP status when there was one.
job_id Optional[str] Provider job / run id when there was one.
retryable bool Whether a retry could plausibly succeed.
to_dict() dict All of the above as a plain dict.
try:
    resp = puffinparse.parse("doc.pdf", model="reducto/standard")
except puffinparse.RateLimitError as e:
    print(e.provider, e.status_code, e.retryable)
except puffinparse.PuffinParseError as e:
    print(e.to_dict())

puffinparse.exceptions.from_core(exc) converts a raw puffinparse._core.CoreError into the typed exception; the SDK applies it for you.

Callbacks#

Two module-level lists, called after every call in every mode, including Router calls. Callbacks may be plain functions or coroutines; exceptions raised inside a callback are logged to the puffinparse logger and never propagate to the caller.

puffinparse.success_callback: list[Callable[[Response], None | Awaitable[None]]]
puffinparse.failure_callback: list[Callable[[PuffinParseError], None | Awaitable[None]]]
puffinparse.success_callback.append(lambda r: print(r.model, r.usage.pages, r.cost_usd))
puffinparse.failure_callback.append(lambda e: print("failed:", e))

Models and pricing#

puffinparse.modes() -> list[str]
puffinparse.output_formats() -> list[str]              # vendor shapes output_format accepts
puffinparse.list_models(mode: Optional[Mode] = None) -> list[str]
puffinparse.providers() -> list[dict[str, Any]]        # name, env var, base URL, docs, models + modes
puffinparse.resolve_model(model: str, mode: Optional[Mode] = None) -> str
puffinparse.pricing() -> dict[str, dict[str, Any]]     # model -> per-mode $/page, source, updated
puffinparse.set_pricing(prices: dict[str, float], mode: Mode = "parse") -> None
puffinparse.reset_pricing() -> None
puffinparse.estimate_cost(model: str, pages: int, mode: Mode = "parse") -> Optional[float]
puffinparse.list_models("extract")                     # only models that serve extract
puffinparse.resolve_model("reducto", "ocr")            # provider default *for that mode*
puffinparse.set_pricing({"reducto/standard": 0.012})   # your negotiated parse rate
puffinparse.estimate_cost("reducto/standard", 1000)    # 12.0
puffinparse.reset_pricing()

resolve_model raises UnsupportedModelError for an unknown string, or for a model that does not serve the requested mode — which makes it a cheap validator for user input. estimate_cost returns None when a model has no published price for that mode.

Scoring helpers#

The benchmark metrics are exposed directly — deterministic, offline, no LLM judge.

puffinparse.score(
    prediction: str,
    truth: str,
    *,
    case_insensitive: bool = True,
    strip_markdown: bool = True,
    strip_punctuation: bool = False,
) -> Metrics

puffinparse.normalize_text(text, *, case_insensitive=True, strip_markdown=True,
                       strip_punctuation=False) -> str

puffinparse.markdown_to_text(markdown: str) -> str
m = puffinparse.score(resp.markdown, open("truth.md").read())
print(m.char_similarity, m.cer, m.wer, m.word_f1, m.order_score, m.table_score)

See Benchmark for what each metric means.

Logging#

puffinparse.init_logging(level: str = "info") -> None

Enables the Rust core's tracing output on stderr; "debug" shows every HTTP step. The CLI uses the PUFFINPARSE_LOG environment variable for the same thing. Python-side messages (such as a raising callback) go to the standard logging logger named puffinparse.

See also#