Getting started#
PuffinParse gives you one call for every OCR / document-parsing provider. Install it, set one key, and parse a document in under a minute.
1. Install#
Python#
pip install puffinparse
Python 3.9+. The wheel bundles the Rust core — there is no toolchain to install and no provider SDK to add.
CLI#
The puffinparse binary is built from the Rust workspace:
cargo install --git https://github.com/ajinkyashejul/puffinparse puffinparse-cli
# or, from a clone:
cargo build --release -p puffinparse-cli # ./target/release/puffinparse
Rust#
[dependencies]
puffinparse-core = "0.1"
tokio = { version = "1", features = ["rt-multi-thread", "macros"] }
From source#
git clone https://github.com/ajinkyashejul/puffinparse && cd puffinparse
python -m venv .venv && . .venv/bin/activate
pip install maturin && maturin develop --release # builds puffinparse._core into the venv
cargo build --release -p puffinparse-cli
2. Set a key#
Each provider reads its own environment variable. Set only the ones you use.
| Provider | Environment variable | Get a key |
|---|---|---|
| Reducto | REDUCTO_API_KEY |
reducto.ai |
| Extend | EXTEND_API_KEY |
extend.ai |
| LlamaParse | LLAMA_API_KEY (starts with llx-) |
cloud.llamaindex.ai |
export REDUCTO_API_KEY=...
export EXTEND_API_KEY=...
export LLAMA_API_KEY=llx-...
Keys can also be passed per call (api_key=... / --api-key), and base URLs overridden with
REDUCTO_BASE_URL, EXTEND_BASE_URL, LLAMA_BASE_URL. See .env.example.
Check what is configured:
puffinparse providers
3. First call#
Python#
import puffinparse
resp = puffinparse.parse("invoice.pdf", model="reducto/standard")
print(resp.markdown) # unified markdown, identical shape for every provider
print(resp.usage.pages, resp.cost_usd, resp.latency_ms)
print(resp.pages[0].blocks[0].type, resp.pages[0].blocks[0].bbox)
Async is the same call with await:
resp = await puffinparse.aparse("invoice.pdf", model="llamaparse/cost_effective")
CLI#
puffinparse parse invoice.pdf -m extend/parse_light # markdown on stdout
puffinparse parse scan.png -m llamaparse/agentic -f json --raw # unified JSON + provider payload
Rust#
use puffinparse_core::{parse, DocumentRequest};
#[tokio::main]
async fn main() -> puffinparse_core::Result<()> {
let resp = parse(DocumentRequest::from_path("invoice.pdf").model("reducto/standard")).await?;
println!("{} pages, ${:.4}", resp.usage.pages, resp.cost_usd.unwrap_or(0.0));
println!("{}", resp.markdown);
Ok(())
}
4. Pick a mode#
parse is one of three modes. Each has its own response type, and providers can only be swapped
within a mode.
| Mode | Python | CLI | You get |
|---|---|---|---|
parse |
puffinparse.parse(...) |
puffinparse parse doc.pdf |
Layout-aware markdown + typed blocks with boxes. |
ocr |
puffinparse.ocr(...) |
puffinparse ocr scan.png |
Plain text with line and word boxes. |
extract |
puffinparse.extract(..., schema) |
puffinparse extract doc.pdf -s schema.json |
A JSON object shaped by your schema, with citations. |
text = puffinparse.ocr("scan.png", model="reducto/r-1")
text.text, text.pages[0].lines[0].bbox
data = puffinparse.extract(
"invoice.pdf",
{"type": "object", "properties": {"total": {"type": "number"}}},
model="reducto/standard",
citations=True,
)
data.data["total"], data.citations("/total")
puffinparse providers --mode extract lists the models that serve a mode; asking a model for a mode it
does not support raises UnsupportedModelError before any network call.
5. Switch providers#
The model string is the only thing that changes. "<provider>/<model>", like LiteLLM; a bare
provider name selects its default model.
puffinparse.parse("doc.pdf", model="reducto/r-1")
puffinparse.parse("doc.pdf", model="extend/parse_performance")
puffinparse.parse("doc.pdf", model="llamaparse/agentic_plus")
puffinparse.parse("doc.pdf", model="reducto") # -> reducto's default parse model
Unknown providers or models raise UnsupportedModelError before any network call.
6. Add fallbacks#
router = puffinparse.Router(
["reducto/standard", "llamaparse/agentic", "extend/parse_light"],
mode="parse", # a router is bound to one mode
strategy="ordered", # or "round_robin"
)
resp = router.parse("doc.pdf")
router.stats() # successes, failures, latency, cost, pages per model
Provider, rate-limit, timeout and network errors fall through to the next model. Auth, bad-request and input errors never do — they would fail on every provider.
Where to next#
- Python SDK — every function, dataclass and exception.
- CLI — every subcommand and flag.
- Rust —
puffinparse-corecrate usage. - Providers — exactly what PuffinParse sends and how the response is mapped.
- Benchmark — how the leaderboard is produced, and its caveats.
- Specification — the contract the implementations follow.