Rust#
puffinparse-core is the whole implementation: unified types, every provider, the router, pricing and
the benchmark metrics. The Python SDK and the CLI are thin wrappers over it.
#![forbid(unsafe_code)], no vendor SDK crates — every provider is spoken to over plain HTTPS with
reqwest and tokio.
API reference: docs.rs/puffinparse-core (published on the first crates.io release).
Until then, cargo doc -p puffinparse-core --open from a clone.
Install#
[dependencies]
puffinparse-core = "0.1"
tokio = { version = "1", features = ["rt-multi-thread", "macros"] }
From the repository while it is pre-release:
puffinparse-core = { git = "https://github.com/ajinkyashejul/puffinparse" }
First call#
use puffinparse_core::{parse, DocumentRequest};
#[tokio::main]
async fn main() -> puffinparse_core::Result<()> {
let resp = parse(DocumentRequest::from_path("invoice.pdf").model("reducto/standard")).await?;
println!("{} pages, ${:.4}", resp.usage.pages, resp.cost_usd.unwrap_or(0.0));
println!("{}", resp.markdown);
Ok(())
}
Each entry point resolves the model, builds the provider, times the call, fills in provider,
model, latency_ms and cost_usd, merges request metadata into the response, and drops raw
unless include_raw was set.
Modes#
A call has a mode, and each mode has its own request and response type. Providers can only be swapped within a mode — a model that does not support the requested mode is rejected before any network call.
| Mode | Entry point | Request | Response | What you get |
|---|---|---|---|---|
Mode::Parse |
parse |
DocumentRequest |
ParseResponse |
Layout-aware markdown and typed blocks with boxes. |
Mode::Ocr |
ocr |
DocumentRequest |
TextResponse |
Plain text with line and word boxes, no layout semantics. |
Mode::Extract |
extract |
ExtractRequest |
ExtractResponse |
A JSON object shaped by your schema, with per-field citations. |
pub async fn parse(request: DocumentRequest) -> Result<ParseResponse>;
pub async fn ocr(request: DocumentRequest) -> Result<TextResponse>;
pub async fn extract(request: ExtractRequest) -> Result<ExtractResponse>;
/// Blocking wrapper around `parse`; builds a small current-thread runtime per call.
pub fn parse_blocking(request: DocumentRequest) -> Result<ParseResponse>;
Mode is Parse | Ocr | Extract, with Mode::ALL, as_str() and a Display impl.
DocumentRequest builder#
Every setter takes self and returns Self, so requests chain.
use puffinparse_core::{DocumentRequest, OutputFormat};
let req = DocumentRequest::from_path("doc.pdf")
.model("extend/parse_performance")
.pages("1-3,7")
.language("en")
.output(OutputFormat::Markdown)
.provider_options(serde_json::json!({ "blockOptions": { "figures": { "enabled": false } } }))
.include_raw(true)
.timeout_secs(120.0)
.max_retries(3)
.api_key(std::env::var("EXTEND_API_KEY").unwrap())
.base_url("https://api.extend.ai");
Constructors#
| Constructor | Input |
|---|---|
DocumentRequest::from_path(path) |
Local file. |
DocumentRequest::from_bytes(data, filename) |
In-memory bytes; the filename infers the type. |
DocumentRequest::from_url(url) |
Remote document, passed to the provider where supported. |
DocumentRequest::from_str_input(s) |
URL if s starts with http:// / https://, else a path. |
DocumentRequest::new(DocumentInput) |
The general form. |
DocumentInput is an enum of Path { path }, Bytes { data, filename } and Url { url }, with
filename(), mime_type() and describe() helpers.
Fields and defaults#
| Field | Type | Default |
|---|---|---|
input |
DocumentInput |
— |
model |
String |
"reducto" |
pages |
Option<String> |
None |
language |
Option<String> |
None |
output |
OutputFormat |
Markdown |
provider_options |
Option<serde_json::Value> |
None |
include_raw |
bool |
false |
timeout_secs |
f64 |
300.0 |
max_retries |
u32 |
2 |
api_key |
Option<String> |
None (falls back to the provider's env var) |
base_url |
Option<String> |
None (falls back to <PROVIDER>_BASE_URL, then the built-in) |
metadata |
BTreeMap<String, Value> |
empty |
req.option("key") reads a single key back out of provider_options.
ParseResponse#
pub struct ParseResponse {
pub id: String,
pub provider: String,
pub model: String,
pub pages: Vec<Page>,
pub markdown: String,
pub text: String,
pub usage: Usage,
pub latency_ms: u64,
pub created_at: String,
pub provider_job_id: Option<String>,
pub cost_usd: Option<f64>,
pub metadata: BTreeMap<String, serde_json::Value>,
pub raw: Option<serde_json::Value>,
}
Page { page_number, markdown, text, blocks, width, height },
Block { type, content, page_number, text, bbox, confidence },
BBox { x0, y0, x1, y1 } normalised 0..1 with a top-left origin, and
Usage { pages, credits, provider_cost_usd }.
Everything derives Serialize / Deserialize, so a response round-trips through JSON unchanged —
that is exactly what puffinparse parse -f json prints.
Helpers in types: ParseResponse::from_pages, page_count, join_pages, pages_from_blocks,
strip_html_tags, markdown_to_text.
TextResponse (ocr mode)#
Same envelope — id, provider, model, provider_job_id, usage, cost_usd, latency_ms,
created_at, metadata, raw — with text for the whole document and pages: Vec<TextPage>.
pub struct TextPage {
pub page_number: u32,
pub text: String,
pub lines: Vec<Line>,
pub words: Vec<Word>,
pub width: Option<f64>,
pub height: Option<f64>,
}
Line and Word are both { text, bbox: Option<BBox>, confidence: Option<f64> }.
ExtractRequest / ExtractResponse (extract mode)#
use puffinparse_core::{extract, DocumentRequest, ExtractRequest};
let req = ExtractRequest::new(
DocumentRequest::from_path("invoice.pdf").model("reducto/standard"),
serde_json::json!({
"type": "object",
"properties": { "total": { "type": "number" }, "vendor": { "type": "string" } },
}),
)
.instructions("Totals are inclusive of tax.")
.citations(true);
let resp = extract(req).await?;
println!("{}", resp.data); // the object your schema describes
for (pointer, info) in &resp.fields { // "/total" -> confidence + citations
println!("{pointer}: {:?} {:?}", info.confidence, info.citations);
}
ExtractRequest flattens a DocumentRequest (so every builder setter above still applies) and adds
schema (a JSON Schema 2020-12 subset), optional instructions, and citations: bool.
ExtractResponse carries data, plus fields: BTreeMap<String, FieldInfo> keyed by JSON pointer,
where FieldInfo { confidence: Option<f64>, citations: Vec<Citation> } and
Citation { page_number, bbox: Option<BBox>, text: Option<String> }.
Router#
use puffinparse_core::{DocumentRequest, Mode, Router, RouterConfig, Strategy};
let router = Router::new(
RouterConfig::new(vec!["reducto/standard".into(), "llamaparse/agentic".into()])
.mode(Mode::Parse), // every model must support this mode
)?;
let resp = router.parse(&DocumentRequest::from_path("doc.pdf")).await?;
for (model, s) in router.stats() {
println!("{model}: {} ok, {} failed, avg {:?} ms", s.successes, s.failures, s.avg_latency_ms());
}
RouterConfig::new(models) fills in Mode::Parse, Strategy::Ordered and the default
fallback_on: Provider, RateLimit, Timeout, Network. Set strategy to
Strategy::RoundRobin to rotate the starting model per call. The request's own model is ignored
— the router chooses. router.models() returns the canonicalised list, router.mode() the mode it
is pinned to, and router.plan() the order the next call will try (advancing the round-robin
cursor). Strategy parses from a string ("ordered" / "fallback", "round_robin" /
"roundrobin").
The router has one method per mode: parse, ocr and extract. Calling one whose mode does not
match the router's configured mode is an error.
ModelStats carries successes, failures, total_latency_ms, total_cost_usd, total_pages
and avg_latency_ms().
Errors#
pub enum ErrorKind {
Authentication, RateLimit, BadRequest, Provider,
Timeout, UnsupportedModel, Input, Network,
}
Error carries the kind, the provider message, and optional provider, status_code and
job_id. type Result<T> = std::result::Result<T, Error>. The Python exception hierarchy is a
1:1 mirror of these kinds.
match parse(req).await {
Ok(resp) => println!("{}", resp.markdown),
Err(e) if e.kind == puffinparse_core::ErrorKind::RateLimit => { /* back off */ }
Err(e) => eprintln!("{e}"),
}
Models and pricing#
use puffinparse_core::{list_models, list_models_for, model_info, Mode, ModelRef, PROVIDERS};
for p in PROVIDERS { // name, display_name, env_var, base_url, docs, models
for m in p.models {
println!("{} {:?} {}", m.qualified(), m.modes, if m.default { "*" } else { "" });
}
}
list_models(); // every "<provider>/<model>"
list_models_for(Mode::Extract); // only the models that serve this mode
model_info("reducto", "r-1"); // -> Option<&ModelInfo>
let ModelRef { provider, model } = ModelRef::parse("reducto")?; // -> reducto / standard
let checked = ModelRef::parse_for("reducto", Mode::Ocr)?; // also checks the mode
let usd = puffinparse_core::pricing::estimate_cost("reducto/standard", Mode::Parse, 12);
let table = puffinparse_core::pricing::all_prices();
ModelInfo is { provider, model, description, default, modes }; default marks the provider's
default model within each mode it supports. ModelRef::parse / parse_for are the validation
gate: unknown providers, unknown models, or a model that does not serve the requested mode produce
ErrorKind::UnsupportedModel before any network call. Aliases llama, llama_parse and
llamacloud resolve to llamaparse.
Benchmark module#
puffinparse_core::bench is the deterministic scoring used by puffinparse bench and
puffinparse.score — no LLM judge, no network.
use puffinparse_core::bench::{normalize, score, summarize, Metrics, NormalizeOptions, Summary};
let opts = NormalizeOptions {
case_insensitive: true,
strip_markdown: true,
strip_punctuation: false,
};
let m: Metrics = score(&prediction, &truth, opts);
println!("{:.3} char sim, {:.3} CER, {:.3} WER", m.char_similarity, m.cer, m.wer);
let s: Summary = summarize(&[Some(m), None]); // None counts as a failed document
println!("overall {:.2} over {} docs ({} failed)", s.overall, s.documents, s.failed);
| Item | Description |
|---|---|
normalize(s, opts) |
NFKC, markdown stripped, quotes/dashes straightened, whitespace collapsed. |
score(pred, truth, opts) |
Metrics { char_similarity, cer, wer, word_recall, word_precision, word_f1, order_score, table_score, pred_chars, truth_chars }. |
summarize(&[Option<Metrics>]) |
Summary { documents, failed, char_similarity, cer, wer, word_f1, order_score, table_score, overall }. |
levenshtein(a, b) |
The generic edit distance used underneath. |
See Benchmark for the definition of each metric.
Module map#
| Module | Contents |
|---|---|
types |
Mode, DocumentRequest, ParseResponse, TextResponse, TextPage, Line, Word, ExtractRequest, ExtractResponse, Citation, FieldInfo, Page, Block, BBox, Usage, DocumentInput, OutputFormat. |
error |
Error, ErrorKind, Result. |
model |
ModelRef, ModelInfo, ProviderInfo, PROVIDERS, list_models, list_models_for, model_info. |
pricing |
Embedded, overridable per-page price table. |
http |
Shared client: retry/backoff with jitter, deadline, polling helper. |
provider |
Provider trait plus key / base-URL / multipart helpers. |
providers |
reducto, extend, llamaparse, and build(). |
router |
Router (one method per mode), RouterConfig, Strategy, ModelStats. |
bench |
Normalisation, metrics, summaries. |
util |
deep_merge, page-range parsing. |
Adding a provider#
One file under crates/puffinparse-core/src/providers/, implementing Provider, registered in
providers/mod.rs and model::PROVIDERS, plus pricing.json entries, a fixture-backed
normalisation test with a real redacted payload, and a page under Providers.
The full checklist is in Contributing.
See also#
- Specification — the contract all three surfaces implement
- CLI · Python SDK