PuffinParse docs
/ llms.txt GitHub

Decision log#

Architecture decision records, newest last. One entry per decision that someone might otherwise re-litigate. Format: context → decision → consequences. Add an entry when you change direction; do not edit old entries, supersede them.

ADR-1: Rust core with a thin Python wrapper#

Context. The brief asks for "fastest and litest" and a Python SDK. Provider calls are HTTP + JSON; the benchmark needs Levenshtein over long documents.

Decision. All provider logic, normalisation, routing, pricing and metrics live in the liteocr-core crate. Python (python/liteocr) is a typed wrapper over a PyO3 module (liteocr._core) and holds no provider logic. The CLI is a separate crate on the same core.

Consequences. One implementation to test per provider; a future gateway server and a TypeScript SDK reuse the core. Cost: Python contributors need a Rust toolchain (maturin develop), and wheels must be built per platform (handled by release.yml).

ADR-2: "<provider>/<model>" model strings, LiteLLM style#

Decision. The model string selects provider and quality tier (reducto/r-1, llamaparse/agentic). A bare provider name selects its default. Provider-specific knobs go in provider_options and are merged verbatim into the provider request.

Consequences. Switching provider is a one-string change; the response shape never depends on options. The registry in crates/liteocr-core/src/model.rs is the only place models are declared; pricing is keyed by the same string.

ADR-3: Normalised, top-left-origin bounding boxes#

Decision. Every block bbox is {x0, y0, x1, y1} in 0..=1 of the page. Reducto already reports normalised boxes; Extend and LlamaParse boxes are divided by the page dimensions they report. Boxes are None when dimensions are unknown.

Consequences. Consumers can draw overlays without knowing page size or DPI. Precision of Extend's DPI-scaled coordinates is preserved because the division uses the same units.

ADR-4: Extend always via /parse_runs; Reducto sync by default#

Decision. Extend's sync /parse has a 5-minute hard limit and its docs call it "for onboarding"; LiteOCR always creates a run and polls. Reducto's sync /parse allows 15 minutes and returns inline, so it is the default; provider_options.async = true switches to /parse_async + /job/{id}.

Consequences. One code path per provider is exercised by default; both are covered by the same unified deadline (timeout), so polling never exceeds the caller's budget.

ADR-5: Page reconstruction rules#

Decision. Pages are the unit of the unified response. Reducto is asked for chunk_mode=page and pages are derived from blocks[].bbox.page; Extend uses page chunks and metadata.page.number; LlamaParse's JSON result is already per page. Chunk-level markdown from the provider is preferred over re-joining blocks, so the page markdown is what the provider itself renders.

Consequences. Benchmarks compare provider-rendered markdown, not a LiteOCR re-rendering.

ADR-6: Cost is computed from a list-price table, not from provider credits#

Decision. cost_usd = pages × per_page_usd from pricing.json, overridable at runtime. Provider credits are passed through in usage.credits but not converted: Reducto credits are null on new pricing plans, Extend credits depend on the plan's $/credit, and LlamaParse reports 0 credits until billing settles.

Consequences. Costs are comparable across providers and stable, but are list prices; the benchmark labels them as such.

ADR-7: Benchmark ground truth is exact by construction; metrics are deterministic#

Decision. synthetic-v1 renders documents from the same source as the truth markdown, so there is no annotation noise. Metrics are text-only (char similarity, CER, WER, word F1, order, table) computed in Rust. No LLM judge is required for the leaderboard.

Consequences. Fully reproducible and cheap to run, but synthetic documents are cleaner than real scans; the combined open dataset (see TASKS) is the answer to that, not a looser scorer.

ADR-8: Benchmark runs disable provider result caches#

Context. LlamaParse caches results for 48 hours; a re-run returned in ~450 ms instead of ~9 s and would have misreported latency.

Decision. bench run passes cache-busting options per provider (do_not_cache + invalidate_cache for LlamaParse) unless --allow-cache is given.

Consequences. Latency in the leaderboard reflects real processing. Repeated runs cost real credits.

ADR-9: (withdrawn)#

Repository housekeeping that no longer applies.

ADR-10: Combined open benchmark instead of a new vendor benchmark#

Context. Each vendor publishes a benchmark it wins (ParseBench, RealDocBench, LongExtractBench).

Decision. LiteOCR will not author a competing "vendor-neutral" opinion benchmark. It will run every public benchmark through one harness with one scoring pipeline and publish per-dataset and combined scores, with every input, truth and output inspectable online. Datasets whose license permits redistribution are vendored in the repo; others are downloaded by an adapter at run time with the upstream revision pinned in the manifest.

Consequences. The site is the verification surface; the CLI is the reproduction surface. Some benchmarks (olmOCR-bench) need their own scorer implemented rather than transcript similarity.

ADR-11: Modes — providers are only interchangeable within a mode#

Context. "OCR provider" covers different products: plain text recognition (Textract DetectDocumentText, Azure Read, Google Document OCR), layout-aware parsing to markdown and typed blocks (Reducto, Extend, LlamaParse, Mistral OCR, Datalab, Unstructured, Upstage, Landing AI), and schema-driven structured extraction (Reducto Extract, Extend Extract, LlamaExtract, Azure prebuilt models, Textract Queries/Forms, vision LLMs with structured output). Swapping a parse model for an extract model is not a like-for-like switch.

Decision. The core defines Mode::{Parse, Ocr, Extract} with its own request/response type per mode (DocumentRequest → ParseResponse, DocumentRequest → TextResponse, ExtractRequest → ExtractResponse) and its own entry point (parse, ocr, extract). Every registry model declares the modes it supports; resolution (ModelRef::parse_for) and the Router reject models outside the requested mode. Pricing is per model and mode. A layout provider serves ocr by deriving lines/words from its parse output (marked with liteocr_derived_from = "parse"), so plain-text callers can still use it, but native OCR endpoints override that when they exist.

Consequences. The Python API becomes liteocr.parse / liteocr.ocr / liteocr.extract (the pre-release liteocr.ocr that meant parse is renamed; nothing was published). Vision LLMs (Gemini, OpenAI, Anthropic) fit as parse/extract models without geometry, which the response makes explicit by returning blocks without boxes. Adding a fourth mode (classify, split) is additive: a new enum variant, types, trait method and entry point.

ADR-12: Provider fan-out rules#

Decision. Each provider lives in one file (crates/liteocr-core/src/providers/<name>.rs), is declared up front in providers/mod.rs, and is implemented independently against the Provider trait. Shared registration points (model.rs registry, pricing.json, build(), README table, .env.example) are edited only by the integrator after a provider lands, from snippets in the provider's hand-off. Providers without a key in the environment are implemented from official docs with fixtures built from documented responses and #[ignore] live tests; the task board records which ones have been verified live.

Consequences. Many providers can be built in parallel without merge conflicts; the cost is a short integration step per provider and, for unverified providers, a "docs-only" label until someone with a key runs the live tests.

ADR-13: Native-format compatibility is a renderer over the unified response#

Context. The unified response is the product, but it is also the migration cost: a team already parsing Reducto's result.chunks[].blocks[].bbox.left cannot try another provider without rewriting the code that reads the result. The one thing that would make switching free is getting answers back in the shape they already parse.

Decision. Add crates/liteocr-core/src/compat/: a pure, infallible renderer render_parse(&ParseResponse, Format) -> serde_json::Value (plus a best-effort render_extract) with one module per vendor shape — Format::{Liteocr, Reducto, Extend, LlamaParse}. It is applied after a call, not inside it: parse/ocr/extract keep returning the unified structs, DocumentRequest.output_format only records and validates the caller's choice (validate_output_format), and the SDK/CLI call ParseResponse::to_format(&str) on the result. So routing, retries, pricing, fallbacks and every existing test are untouched.

What is promised is structural fidelity, not semantic identity: the vendor's key set, nesting, chunk/page and block counts, content strings, block-type vocabulary and coordinate units. Fields LiteOCR does not model are rendered null/empty and enumerated in docs/COMPAT.md, never invented. Extend and LlamaParse need a page size for their unit boxes; when the source provider reports none (Reducto, vision LLMs) the renderer assumes a 1000×1000 page and, for Extend, records metadata.liteocr_synthetic_page_dims = true in the run's free-form metadata map.

The claim is enforced rather than asserted: for each provider, compat/roundtrip.rs runs fixture → provider::normalize → render_parse(same format) and compares against the original fixture with a skeleton_diff (key sets, counts, contents, types up to the provider's own forward mapping, boxes within 1e-6, billed pages), plus cross-format and no-geometry cases. The tests live in the crate because the normalize functions are pub(crate).

Consequences. Adding a fourth shape is one module plus one enum variant. Lossy edges are real and documented: type mappings are not injective (Extend's key_value returns as text), the Reducto render uses a reduced vocabulary (footnote/caption/formula/other → Text), and render_extract is explicitly weaker than the parse path until extract fixtures exist for all three providers. Because the renderer only reads the unified types, any provider added later gets all three native shapes for free — and any unified field a new provider cannot fill shows up as a null in someone's vendor-shaped payload, which is the honest outcome.

ADR-14: The Node.js SDK is a napi-rs addon over the same core#

Context. A TypeScript SDK was on the roadmap once the Python surface stabilised. The options were a WASM build, a wrapper around the CLI, or a native N-API addon.

Decision. crates/liteocr-node exposes the core through napi 3; the js/ package is a thin layer. The addon only converts values: requests and responses cross as the core's serde JSON, errors as a prefixed JSON payload (LITEOCR_CORE_ERROR:<json>) that the JS layer rebuilds into typed LiteOCRError subclasses. The JS layer converts to camelCase with explicit per-type converters; data, metadata and raw are never renamed, and index.d.ts is hand-written against SPEC §5 (the napi-generated native.d.ts is internal). The crate is a normal workspace member: napi's dyn-symbols keeps cargo test --workspace free of Node. It uses deny(unsafe_code) because napi's macro expansion is incompatible with forbid, and declares rust-version = "1.88" (napi 3) while the rest of the workspace stays at 1.80. timeout is in seconds, as in the other SDKs, and LiteOCRError.kind uses the ErrorKind serde values.

Consequences. One implementation behind three surfaces (Python, Node, CLI). No provider logic in JS. Prebuilt binaries per platform are required for npm install without a Rust toolchain; the release matrix builds them but publishing is not wired yet.

ADR-15: The gateway is a thin axum layer over the core, TOML-configured, with no database#

Context. LiteLLM's proxy is what teams actually deploy: one endpoint, central keys, budgets and logs. LiteOCR needed the same without growing a second implementation of providers.

Decision. crates/liteocr-server (axum + tower-http, which are HTTP frameworks, not provider SDKs) calls liteocr_core::{parse, ocr, extract} and runs its own fallback loop, because the core Router cannot give each target its own credentials or base URL. Configuration is TOML (already idiomatic in Rust, lighter than YAML). Usage and budget state live in memory with an optional JSON state file; a database is out of scope. Clients may not send api_key or base_url, so they cannot redirect the gateway's credentials, and local file paths are refused. Provider credential failures map to 502, not 401, because the caller's own key was valid.

Consequences. Single binary, no infrastructure to run. Budgets can be overshot by requests in flight, and key changes need a restart. A database, HTTP key management, async job endpoints and metrics auth are follow-ups, not blockers.

ADR-16: Vendor only what the licence permits; index the rest#

Context. The combined benchmark (ADR-10) pulls in public datasets whose licences differ: olmOCR-bench is ODC-BY-1.0, OmniDocBench has no licence and is marked research-only / non-commercial.

Decision. A dataset is vendored into the repo (with attribution) only when its licence permits redistribution. Otherwise the repo holds a manifest with upstream paths, a pinned revision and image/truth hashes, and the adapter materialises the data locally. Combined datasets are versioned and never rewritten once results exist (combined-v2 supersedes combined-v1 for new runs). Upstream tests that cannot be expressed faithfully in the shared rule schema are skipped and counted, never weakened silently; the one relaxed mapping (olmOCR "left/right of" → same row) is counted as relaxed.

Consequences. Anyone can reproduce every score, but index-only sources need a fetch step before a run (their documents carry a fetch-required tag). Stats files make the coverage of each conversion auditable.

ADR-17: The benchmark results viewer stays vanilla JS#

Context. TASKS listed an open choice for benchmark/site/ between a React app on Extend UI (PDF viewer and layout overlays out of the box) and the zero-build vanilla viewer. The viewer has to be where every benchmark claim can be checked: page rendering, side-by-side outputs, diffs, rule checklists, bbox overlays, charts, deep links.

Decision. Keep static HTML/CSS/ES2018 in benchmark/site/src/: no framework, no bundler, no npm install; the build stays stdlib Python (Pillow optional). The one third-party runtime dependency is pdf.js, loaded lazily from cdnjs at a pinned version with SRI, only when a PDF is opened, with the build-time PNG as fallback. Overlays and charts are inline SVG. The per-document rule checklist uses a JS port of liteocr-core's score_rules, and every rules page compares its count with the recorded Rust score and flags any disagreement.

Consequences. Vercel and Pages build the site with one uv run command and nothing to audit. The data contract (data/index.json, data/runs/, data/outputs/…/<doc>.{md,json}) is independent of the front-end, so this can be revisited without touching the build. The JS port must follow scorer changes in bench.rs; the mismatch badge makes drift visible.

ADR-18: Webhooks are exposed as primitives, not received#

Context. TASKS asked for "webhooks instead of polling". An SDK cannot host an HTTP endpoint, and providers differ: Reducto and LlamaParse accept a per-job webhook URL, Extend only has workspace-level webhook endpoints.

Decision. LiteOCR offers submit_parse / retrieve_parse plus a pure parse_webhook / resolve_webhook that normalises a provider's webhook body into JobStatus (doing one retrieve when the body only names the job). JobHandle is serialisable and secret-free; credentials are resolved again at retrieve time. DocumentRequest.webhook_url maps to per-job provider webhooks where they exist and is rejected with an input error where they don't. Only parse mode has jobs for now.

Consequences. Users wire LiteOCR into their own web handler; the gateway (ADR-15) can later add job endpoints on top of the same primitives. Webhook body shapes come from vendor docs until real deliveries are captured.

ADR-19: Local and self-hosted engines are out-of-process providers#

Context. SPEC §1 listed local models as a v0.1 non-goal, but LiteOCR was unusable without a paid key and the benchmark had no open baseline.

Decision. Support local and self-hosted engines only as providers that call out of process: a CLI binary via tokio::process (Tesseract, with pdftoppm for PDFs) or an HTTP server the user runs (docling-serve, PaddleOCR/PaddleX serving). No C bindings, FFI, embedded runtimes or model weights ship in LiteOCR. They are listed in model::SELF_HOSTED, need no key, and are priced at 0.0 with source "self-hosted".

Consequences. #![forbid(unsafe_code)] and the no-SDK rule still hold. Benchmark latency for these engines depends on the user's hardware, so leaderboard rows need a hardware note. Live tests need a binary or a server rather than a key.

ADR-20: The scorer is versioned; committed runs are re-scored offline#

Context. The first combined run exposed scorer defects (HTML tables scored 0, tokenised punctuation in rule text, a dead bag_of_sentences threshold) that moved the leaderboard by up to four points. Re-calling providers to fix a scoring bug would cost money and change latency and provider versions at the same time.

Decision. Any scorer change that moves scores bumps SCORER_VERSION in liteocr-core and re-scores committed runs from their saved outputs with liteocr bench rescore, never by re-calling providers; measured latency and cost are kept. Result JSON records scorer_version (absent = 1) and rescored_at. The viewer's JS rule checker follows every scorer version and flags any document where it disagrees with the recorded Rust score. The bag_of_sentences threshold of 0.8 is justified by sentences the reference extraction fused together. The table structure metric is TEDS on the row/cell grid and is named teds_grid, not TEDS, because the truth has no header/body or span structure.

Consequences. Leaderboards stay comparable across scorer fixes at no cost, and a reader can see which scorer produced a number. Saved outputs are part of every committed run (already required).

ADR-21: Gateway jobs are owned, not routed#

Context. ADR-18 left gateway job endpoints for later; ADR-15 gave the gateway fallbacks and per-key budgets.

Decision. /v1/jobs submits to exactly one deployment (an alias's first target, or the next in a round-robin rotation) with no fallback: failures only surface at retrieve time, and retrying would mean resubmitting. Jobs are stored under opaque gateway ids bound to the submitting key (the master key can read all, as with /v1/usage), without secrets; provider credentials are resolved from config on every poll. Cost is charged once, on the first observed success, under the usage lock. Provider webhooks are received only when the operator opts in with a shared secret.

Consequences. Clients poll the gateway, never the provider. Providers can reuse job ids (LlamaParse returns the cached job for an identical upload), so one webhook may settle several gateway jobs. Vendor HMAC verification and automatic webhook registration remain open.

ADR-22: One design language; the viewer renders PDF pages at build time#

Context. The docs, landing page and viewer each had their own palette (blue in the docs, vermilion on the landing page), and the viewer showed raw ids, eleven metric tiles per document and full-table heatmaps; PDF pages depended on pdf.js from a CDN and often did not render.

Decision. website/assets/tokens.css is the single source of colour, type, space and shape, described in docs/DESIGN.md; every surface links it first and styles only through its custom properties. The viewer follows the language's principles (one number per view, evidence one click away, human names with raw ids on request) and renders PDF pages to WebP at build time with pypdfium2 (a pip wheel that works on Vercel); pdf.js remains only as a lazy fallback. Layout-box overlays are off by default.

Consequences. A palette or type change is one edit. The site build needs pypdfium2 (added to vercel.json and pages.yml); without it the build falls back to Pillow page-1 extraction.

ADR-23: The product is renamed PuffinParse#

Context. liteocr on PyPI belongs to an unrelated OCR engine, several GitHub projects already use the name, and "OCR" undersells a tool whose modes are parse, OCR and extract. The owner's preference, LiteParse, is a LlamaIndex product. About 110 names were checked against domains, PyPI/npm/crates.io and web collisions.

Decision. The product, crates (puffinparse-{core,cli,python,node,server}), Python package (puffinparse, native module puffinparse._core), Node package, CLI binary, environment variables (PUFFINPARSE_*), metadata keys (puffinparse_*), error base class (PuffinParseError), the native output format value ("puffinparse") and the site (puffinparse.vercel.app) all use the new name. PuffinParse had every checked domain (.com, .dev, .ai, .io) and every registry name free and no product collision; the puffin's black, white and orange match the existing palette.

Consequences. Committed benchmark results, recorded provider fixtures (which contain "LiteOCR" in document text), the CHANGELOG history and earlier ADRs keep the old name. Result files written before the rename carry liteocr_version, which the CLI and site builder read as an alias. liteocr.vercel.app keeps serving the same project.

ADR-24: A puffin mark and mascot; brand imagery only in brand moments#

Context. After the rename (ADR-23) the product had no logo: each surface improvised a different mark (a scan line, an orange square, plain text), and the favicon was still the pre-ADR-22 blue. The puffin's black, white and orange already matched the palette.

Decision. Two brand assets. The mark (website/assets/mark.svg): a puffin head in a rounded-square tile, the beak carrying two stripes read as parsed lines; it is the favicon and sits beside the wordmark on every surface. The puffin (website/assets/puffin.svg): a mascot holding three document fish, used only in brand moments (landing hero, 404, empty states, social card, README), never beside data in the tools. Both are flat SVG in existing tokens plus three brand tokens (--mark-tile, --mascot-body, --mascot-wing) that lift in dark mode. A committed social card (og.png) is the og:image for every page. Alongside, two table rules tightened: headers are sentence case, and bold marks a best value only when it is unique as displayed.

Consequences. "No decorative imagery" now reads "none in the tools". Inline copies of the mark live in website/build.py (MARK) and the viewer's index.html; the social card must be re-rendered (website/og/render.py, Playwright) when the card, mark or mascot changes. The launch videos still show the text wordmark until they are re-rendered.

Amended 2026-09-25: the mark is C2 of the explored set (kept after comparing six alternatives), and the mascot became the waving puffin with three pages in its beak (instead of three document fish), which says "documents" more directly. docs/DESIGN.md gained a Brand motion section.