Python SDK#
pip install puffinparse — Python 3.9+, fully typed (py.typed, mypy --strict clean). The package is
a thin, typed wrapper over the Rust core (puffinparse._core); no provider logic lives in Python.
Everything below is the public surface exported from puffinparse.__all__.
import puffinparse
puffinparse.__version__ # version of the compiled core
puffinparse.modes() # ["parse", "ocr", "extract"]
Modes#
Every call names a mode, and the mode decides the response type. Providers can be swapped
freely within a mode; a model that does not serve the mode you asked for raises
UnsupportedModelError before any network call.
| Mode | Call | Async | Returns | Use it for |
|---|---|---|---|---|
parse |
puffinparse.parse |
puffinparse.aparse |
ParseResponse |
Layout-aware markdown and typed blocks with boxes. |
ocr |
puffinparse.ocr |
puffinparse.aocr |
TextResponse |
Plain text with line and word boxes — search indexes, redaction, overlays. |
extract |
puffinparse.extract |
puffinparse.aextract |
ExtractResponse |
A JSON object shaped by your schema, with per-field citations. |
Mode is Literal["parse", "ocr", "extract"] and MODES is the same tuple as a constant.
doc = puffinparse.parse("invoice.pdf", model="reducto/standard") # markdown + blocks
text = puffinparse.ocr("scan.png", model="llamaparse/fast") # plain text + boxes
data = puffinparse.extract("invoice.pdf", schema, model="reducto/extract")
parse#
def parse(
input: DocumentLike,
model: str = "reducto",
*,
output_format: Optional[str] = None,
filename: Optional[str] = None,
pages: Optional[str] = None,
language: Optional[str] = None,
output: Literal["markdown", "text"] = "markdown",
provider_options: Optional[dict[str, Any]] = None,
include_raw: bool = False,
timeout: float = 300.0,
max_retries: int = 2,
api_key: Optional[str] = None,
base_url: Optional[str] = None,
metadata: Optional[dict[str, Any]] = None,
) -> ParseResponse # or dict[str, Any] when output_format is given
DocumentLike is str | os.PathLike[str] | bytes | bytearray | memoryview.
| Parameter | Meaning |
|---|---|
input |
A file path, an http(s):// URL, or raw bytes (then filename is required). |
model |
"<provider>/<model>", e.g. "reducto/standard". A bare provider name selects its default model for this mode. |
filename |
Required when input is bytes; used to infer the document type. |
pages |
1-based page selection such as "1-3,7", forwarded best-effort. |
language |
Language hint (ISO 639-1) when the provider supports it. |
output |
Preferred block content: "markdown" (default) or "text". |
output_format |
Return a vendor's own JSON dict instead of the dataclass: "reducto", "extend", "llamaparse", or "puffinparse" / None for the unified shape. See Native-format output. |
provider_options |
Provider-specific options merged verbatim into the provider request body. |
include_raw |
Attach the provider's raw payload as response.raw. |
timeout |
Whole-call deadline in seconds (upload + polling + download). |
max_retries |
Retries on 429 / 5xx / network errors with exponential backoff and full jitter. |
api_key |
Override the API key (otherwise read from REDUCTO_API_KEY, EXTEND_API_KEY, LLAMA_API_KEY). |
base_url |
Override the provider base URL. |
metadata |
Free-form dict echoed back in response.metadata. |
Passing bytes without filename raises InputError before any network call.
resp = puffinparse.parse("contract.pdf", model="extend/parse_performance")
resp = puffinparse.parse("https://example.com/doc.pdf", model="reducto/r-1", pages="1-2")
resp = puffinparse.parse(open("scan.png", "rb").read(), filename="scan.png", model="llamaparse/agentic")
aparse is the same call with await. It runs on the Rust runtime, so the event loop is never
blocked.
resp = await puffinparse.aparse("contract.pdf", model="llamaparse/cost_effective")
ocr#
def ocr(
input: DocumentLike,
model: str = "reducto",
*,
filename: Optional[str] = None,
pages: Optional[str] = None,
language: Optional[str] = None,
provider_options: Optional[dict[str, Any]] = None,
include_raw: bool = False,
timeout: float = 300.0,
max_retries: int = 2,
api_key: Optional[str] = None,
base_url: Optional[str] = None,
metadata: Optional[dict[str, Any]] = None,
) -> TextResponse
Same parameters as parse minus output (the mode implies plain text). Use it when you want text
and geometry rather than document structure; for markdown, tables and block types use parse.
Providers without a native OCR endpoint serve this mode from their parse output. The response then
carries metadata["puffinparse_derived_from"] == "parse", so you can tell the difference.
text = puffinparse.ocr("scan.png", model="reducto/r-1")
text.text # whole document
for line in text.pages[0].lines:
line.text, line.bbox, line.confidence
text = await puffinparse.aocr("scan.png") # async
extract#
def extract(
input: DocumentLike,
schema: dict[str, Any],
*,
output_format: Optional[str] = None,
model: str = "reducto",
instructions: Optional[str] = None,
citations: bool = False,
filename: Optional[str] = None,
pages: Optional[str] = None,
language: Optional[str] = None,
provider_options: Optional[dict[str, Any]] = None,
include_raw: bool = False,
timeout: float = 300.0,
max_retries: int = 2,
api_key: Optional[str] = None,
base_url: Optional[str] = None,
metadata: Optional[dict[str, Any]] = None,
) -> ExtractResponse # or dict[str, Any] when output_format is given
| Parameter | Meaning |
|---|---|
schema |
A JSON Schema object describing the fields you want. |
model |
Must support extract; a parse-only model raises UnsupportedModelError before any network call. |
instructions |
Optional natural-language guidance, forwarded to providers that accept it. |
citations |
Ask for per-field citations (page, box, source text) where the provider supports them. |
output_format |
As for parse, rendering the vendor's extract envelope. Best effort — see docs/COMPAT.md §7. |
Everything else matches parse. aextract is the async variant.
schema = {
"type": "object",
"properties": {
"vendor": {"type": "string"},
"total": {"type": "number"},
},
}
resp = puffinparse.extract("invoice.pdf", schema, model="reducto/extract",
instructions="Totals are inclusive of tax.", citations=True)
resp.data["total"] # the object your schema asked for
resp.citations("/total") # [Citation(page_number=1, bbox=..., text="...")]
resp.field_info("/total").confidence
Router#
class Router:
def __init__(
self,
models: list[str],
*,
mode: Mode = "parse",
strategy: Literal["ordered", "round_robin"] = "ordered",
fallback_on: Optional[list[str]] = None,
) -> None
A router is bound to one mode at construction: every model must support it, and calling a method
for another mode raises InputError. That keeps fallbacks honest — a parse-only model can never
quietly answer an extraction.
| Member | Type | Description |
|---|---|---|
models |
list[str] (property) |
The canonicalised model list. |
mode |
Mode (property) |
The mode this router serves. |
plan() |
list[str] |
The order models would be tried for the next call. Advances the round-robin cursor. |
stats() |
dict[str, dict[str, Any]] |
Per model: successes, failures, total_latency_ms, total_cost_usd, total_pages. |
parse(input, **kw) / aparse |
ParseResponse |
Same keyword arguments as puffinparse.parse except model, output_format included. |
ocr(input, **kw) / aocr |
TextResponse |
Same, minus output. |
extract(input, schema, *, output_format, instructions, citations, **kw) / aextract |
ExtractResponse |
Same as puffinparse.extract except model. |
fallback_on is a list of error-kind names; the default is provider, rate_limit, timeout,
network. Authentication, bad-request, unsupported-model and input errors never trigger a
fallback. Unknown keyword arguments raise TypeError and the message lists what is accepted.
router = puffinparse.Router(["reducto/standard", "llamaparse/agentic"], strategy="round_robin")
resp = router.parse("doc.pdf", pages="1-5", timeout=120)
router.stats()["reducto/standard"]["successes"]
text_router = puffinparse.Router(["reducto/r-1", "extend/parse_light"], mode="ocr")
text_router.ocr("scan.png").text
Native-format output (output_format)#
PuffinParse normalises every provider to one response shape, which is the right default — and a
migration cost if you are already integrated with a vendor. output_format removes it: ask for a
vendor's shape and the response is rendered into that vendor's own JSON, whatever provider
actually produced it.
doc = puffinparse.parse("invoice.pdf", model="extend/parse_light", output_format="reducto")
for chunk in doc["result"]["chunks"]: # Reducto's shape, Extend's engine
for block in chunk["blocks"]:
draw(block["bbox"]["left"], block["bbox"]["top"], block["type"])
| Value | Shape |
|---|---|
None (default), "puffinparse", "unified" |
PuffinParse's own response — a dataclass for None, the same JSON as a dict for "puffinparse". |
"reducto" |
Reducto POST /parse response (response_type: "parse"). |
"extend" |
Extend parse_run object (GET /parse_runs/{id}). |
"llamaparse" ("llama", "llama_parse") |
LlamaParse …/result/json payload. |
- Available on
parse,aparse,extract,aextract,Router.parse/aparse/extract/aextract, and the CLI (--output-format, with--format json). - Names are case-insensitive and
-/_are interchangeable;puffinparse.output_formats()lists them. An unknown value raisesBadRequestErrorbefore any network call. - The return type follows the argument:
Nonegives the dataclass, a string gives adict. Both are@overload-typed, so a type checker knows which one it is. - Callbacks always receive the dataclass, whatever the caller asked for — logging and cost tracking are unaffected.
output_formatis independent ofoutput:outputpicks markdown vs plain text inside block content,output_formatpicks the JSON envelope around it.
What is guaranteed is structural fidelity, not semantic identity: the key set and nesting, one
chunk/page per unified page, the content strings, the vendor's own block vocabulary and coordinate
units, and the billed page count. Not guaranteed: byte equality with what the vendor would have
returned, fields PuffinParse does not model (they are rendered as null / [], never invented), or
vendor-specific enrichments. Extract-mode rendering is explicitly best effort.
docs/COMPAT.md lists every
always-null field, the lossy block-type mappings and the coordinate conversions, per format.
puffinparse.output_formats() # ["puffinparse", "reducto", "extend", "llamaparse"]
# same call, both shapes
doc = puffinparse.parse("invoice.pdf", model="reducto/standard") # ParseResponse
raw = puffinparse.parse("invoice.pdf", model="reducto/standard", output_format="extend")
raw["object"], raw["status"], raw["metrics"]["pageCount"]
router = puffinparse.Router(["reducto/standard", "extend/parse_light"])
router.parse("doc.pdf", output_format="llamaparse")["pages"][0]["items"]
Response types#
All response dataclasses mirror the Rust structs in puffinparse-core 1:1 and live in puffinparse.types.
Response is the union ParseResponse | TextResponse | ExtractResponse.
Every response carries the same envelope: id, provider, model, provider_job_id, usage,
cost_usd, latency_ms, created_at, metadata, raw, plus to_dict() and a from_dict()
classmethod. cost_usd is pages × per_page_usd for the mode from the embedded price table (no
provider reports dollars directly), and raw is None unless the call passed include_raw=True.
ParseResponse#
@dataclass
class ParseResponse:
id: str
provider: str
model: str
pages: list[Page]
markdown: str
text: str
usage: Usage
latency_ms: int
created_at: str
provider_job_id: Optional[str] = None
cost_usd: Optional[float] = None
metadata: dict[str, Any] = field(default_factory=dict)
raw: Any = None
| Member | Description |
|---|---|
num_pages (property) |
len(self.pages). |
blocks (property) |
Every block from every page, in reading order. |
tables (property) |
Blocks whose type is "table". |
__str__ |
Returns markdown. |
Page and Block#
@dataclass
class Page:
page_number: int
markdown: str
text: str
blocks: list[Block] = field(default_factory=list)
width: Optional[float] = None
height: Optional[float] = None
@dataclass
class Block:
type: BlockType
content: str
page_number: int
text: Optional[str] = None
bbox: Optional[BBox] = None
confidence: Optional[float] = None
Page.blocks_of(*types) filters by block type; Page.tables is shorthand for
blocks_of("table"). width / height are None for providers that do not report page
dimensions (Reducto).
BlockType is a Literal of "text", "title", "section_header", "list", "table",
"figure", "header", "footer", "footnote", "caption", "formula", "other".
TextResponse, TextPage, Line, Word#
@dataclass
class TextResponse:
id: str
provider: str
model: str
pages: list[TextPage]
text: str
usage: Usage
latency_ms: int
created_at: str
provider_job_id: Optional[str] = None
cost_usd: Optional[float] = None
metadata: dict[str, Any] = field(default_factory=dict)
raw: Any = None
@dataclass
class TextPage:
page_number: int
text: str
lines: list[Line] = field(default_factory=list)
words: list[Word] = field(default_factory=list)
width: Optional[float] = None
height: Optional[float] = None
Line and Word are both { text: str, bbox: Optional[BBox], confidence: Optional[float] }.
TextResponse.lines and .words flatten across pages, num_pages is the page count, and
__str__ returns text.
ExtractResponse, FieldInfo, Citation#
@dataclass
class ExtractResponse:
id: str
provider: str
model: str
data: Any
usage: Usage
latency_ms: int
created_at: str
provider_job_id: Optional[str] = None
fields: dict[str, FieldInfo] = field(default_factory=dict)
cost_usd: Optional[float] = None
metadata: dict[str, Any] = field(default_factory=dict)
raw: Any = None
fields is keyed by JSON pointer into data ("/invoice/total").
field_info(pointer) returns the FieldInfo or None; citations(pointer) returns its citation
list, or an empty list when the provider reported none. __str__ returns str(data).
@dataclass
class FieldInfo:
confidence: Optional[float] = None
citations: list[Citation] = field(default_factory=list)
@dataclass
class Citation:
page_number: int # 1-based
bbox: Optional[BBox] = None
text: Optional[str] = None # source text the value was read from
BBox and Usage#
@dataclass(frozen=True)
class BBox:
x0: float
y0: float
x1: float
y1: float
@dataclass
class Usage:
pages: int = 0
credits: Optional[float] = None
provider_cost_usd: Optional[float] = None
Boxes are normalised to 0..1 relative to page size, origin top-left. BBox.width and .height are
properties; to_pixels(page_width, page_height) returns absolute (x0, y0, x1, y1).
Metrics#
Returned by puffinparse.score.
@dataclass
class Metrics:
char_similarity: float
cer: float
wer: float
word_recall: float
word_precision: float
word_f1: float
pred_chars: int
truth_chars: int
order_score: Optional[float] = None
table_score: Optional[float] = None
Exceptions#
Every error raised by PuffinParse derives from PuffinParseError.
PuffinParseError kind = "error"
├── AuthenticationError "authentication_error" 401/403, or no API key configured
├── RateLimitError "rate_limit_error" 429 after retries were exhausted
├── BadRequestError "bad_request_error" other 4xx — the request itself is wrong
├── ProviderError "provider_error" 5xx, malformed payload, failed job
├── TimeoutError "timeout_error" the whole-call deadline was exceeded
├── UnsupportedModelError "unsupported_model_error" unknown model, or one that cannot serve the mode
├── InputError "input_error" unreadable input, bytes without filename
└── NetworkError "network_error" network / TLS / DNS failure after retries
Every instance carries:
| Attribute | Type | Description |
|---|---|---|
message |
str |
The provider's message, never swallowed. |
kind |
str (class attribute) |
Stable machine-readable kind, as above. |
provider |
Optional[str] |
Provider the call was routed to. |
status_code |
Optional[int] |
HTTP status when there was one. |
job_id |
Optional[str] |
Provider job / run id when there was one. |
retryable |
bool |
Whether a retry could plausibly succeed. |
to_dict() |
dict |
All of the above as a plain dict. |
try:
resp = puffinparse.parse("doc.pdf", model="reducto/standard")
except puffinparse.RateLimitError as e:
print(e.provider, e.status_code, e.retryable)
except puffinparse.PuffinParseError as e:
print(e.to_dict())
puffinparse.exceptions.from_core(exc) converts a raw puffinparse._core.CoreError into the typed
exception; the SDK applies it for you.
Callbacks#
Two module-level lists, called after every call in every mode, including Router calls. Callbacks
may be plain functions or coroutines; exceptions raised inside a callback are logged to the
puffinparse logger and never propagate to the caller.
puffinparse.success_callback: list[Callable[[Response], None | Awaitable[None]]]
puffinparse.failure_callback: list[Callable[[PuffinParseError], None | Awaitable[None]]]
puffinparse.success_callback.append(lambda r: print(r.model, r.usage.pages, r.cost_usd))
puffinparse.failure_callback.append(lambda e: print("failed:", e))
Models and pricing#
puffinparse.modes() -> list[str]
puffinparse.output_formats() -> list[str] # vendor shapes output_format accepts
puffinparse.list_models(mode: Optional[Mode] = None) -> list[str]
puffinparse.providers() -> list[dict[str, Any]] # name, env var, base URL, docs, models + modes
puffinparse.resolve_model(model: str, mode: Optional[Mode] = None) -> str
puffinparse.pricing() -> dict[str, dict[str, Any]] # model -> per-mode $/page, source, updated
puffinparse.set_pricing(prices: dict[str, float], mode: Mode = "parse") -> None
puffinparse.reset_pricing() -> None
puffinparse.estimate_cost(model: str, pages: int, mode: Mode = "parse") -> Optional[float]
puffinparse.list_models("extract") # only models that serve extract
puffinparse.resolve_model("reducto", "ocr") # provider default *for that mode*
puffinparse.set_pricing({"reducto/standard": 0.012}) # your negotiated parse rate
puffinparse.estimate_cost("reducto/standard", 1000) # 12.0
puffinparse.reset_pricing()
resolve_model raises UnsupportedModelError for an unknown string, or for a model that does not
serve the requested mode — which makes it a cheap validator for user input. estimate_cost returns
None when a model has no published price for that mode.
Scoring helpers#
The benchmark metrics are exposed directly — deterministic, offline, no LLM judge.
puffinparse.score(
prediction: str,
truth: str,
*,
case_insensitive: bool = True,
strip_markdown: bool = True,
strip_punctuation: bool = False,
) -> Metrics
puffinparse.normalize_text(text, *, case_insensitive=True, strip_markdown=True,
strip_punctuation=False) -> str
puffinparse.markdown_to_text(markdown: str) -> str
m = puffinparse.score(resp.markdown, open("truth.md").read())
print(m.char_similarity, m.cer, m.wer, m.word_f1, m.order_score, m.table_score)
See Benchmark for what each metric means.
Logging#
puffinparse.init_logging(level: str = "info") -> None
Enables the Rust core's tracing output on stderr; "debug" shows every HTTP step. The CLI uses
the PUFFINPARSE_LOG environment variable for the same thing. Python-side messages (such as a raising
callback) go to the standard logging logger named puffinparse.
See also#
- Getting started
- Specification — the unified request/response contract these types implement
- Providers — per-provider mapping, supported modes and
provider_options