Decision log#
Architecture decision records, newest last. One entry per decision that someone might otherwise re-litigate. Format: context → decision → consequences. Add an entry when you change direction; do not edit old entries, supersede them.
ADR-1: Rust core with a thin Python wrapper#
Context. The brief asks for "fastest and litest" and a Python SDK. Provider calls are HTTP + JSON; the benchmark needs Levenshtein over long documents.
Decision. All provider logic, normalisation, routing, pricing and metrics live in the
liteocr-core crate. Python (python/liteocr) is a typed wrapper over a PyO3 module
(liteocr._core) and holds no provider logic. The CLI is a separate crate on the same core.
Consequences. One implementation to test per provider; a future gateway server and a
TypeScript SDK reuse the core. Cost: Python contributors need a Rust toolchain
(maturin develop), and wheels must be built per platform (handled by release.yml).
ADR-2: "<provider>/<model>" model strings, LiteLLM style#
Decision. The model string selects provider and quality tier (reducto/r-1,
llamaparse/agentic). A bare provider name selects its default. Provider-specific knobs go in
provider_options and are merged verbatim into the provider request.
Consequences. Switching provider is a one-string change; the response shape never depends
on options. The registry in crates/liteocr-core/src/model.rs is the only place models are
declared; pricing is keyed by the same string.
ADR-3: Normalised, top-left-origin bounding boxes#
Decision. Every block bbox is {x0, y0, x1, y1} in 0..=1 of the page. Reducto already
reports normalised boxes; Extend and LlamaParse boxes are divided by the page dimensions they
report. Boxes are None when dimensions are unknown.
Consequences. Consumers can draw overlays without knowing page size or DPI. Precision of Extend's DPI-scaled coordinates is preserved because the division uses the same units.
ADR-4: Extend always via /parse_runs; Reducto sync by default#
Decision. Extend's sync /parse has a 5-minute hard limit and its docs call it
"for onboarding"; LiteOCR always creates a run and polls. Reducto's sync /parse allows 15
minutes and returns inline, so it is the default; provider_options.async = true switches to
/parse_async + /job/{id}.
Consequences. One code path per provider is exercised by default; both are covered by the
same unified deadline (timeout), so polling never exceeds the caller's budget.
ADR-5: Page reconstruction rules#
Decision. Pages are the unit of the unified response. Reducto is asked for chunk_mode=page
and pages are derived from blocks[].bbox.page; Extend uses page chunks and
metadata.page.number; LlamaParse's JSON result is already per page. Chunk-level markdown from
the provider is preferred over re-joining blocks, so the page markdown is what the provider
itself renders.
Consequences. Benchmarks compare provider-rendered markdown, not a LiteOCR re-rendering.
ADR-6: Cost is computed from a list-price table, not from provider credits#
Decision. cost_usd = pages × per_page_usd from pricing.json, overridable at runtime.
Provider credits are passed through in usage.credits but not converted: Reducto credits are
null on new pricing plans, Extend credits depend on the plan's $/credit, and LlamaParse reports 0
credits until billing settles.
Consequences. Costs are comparable across providers and stable, but are list prices; the benchmark labels them as such.
ADR-7: Benchmark ground truth is exact by construction; metrics are deterministic#
Decision. synthetic-v1 renders documents from the same source as the truth markdown, so
there is no annotation noise. Metrics are text-only (char similarity, CER, WER, word F1, order,
table) computed in Rust. No LLM judge is required for the leaderboard.
Consequences. Fully reproducible and cheap to run, but synthetic documents are cleaner than real scans; the combined open dataset (see TASKS) is the answer to that, not a looser scorer.
ADR-8: Benchmark runs disable provider result caches#
Context. LlamaParse caches results for 48 hours; a re-run returned in ~450 ms instead of ~9 s and would have misreported latency.
Decision. bench run passes cache-busting options per provider (do_not_cache +
invalidate_cache for LlamaParse) unless --allow-cache is given.
Consequences. Latency in the leaderboard reflects real processing. Repeated runs cost real credits.
ADR-9: (withdrawn)#
Repository housekeeping that no longer applies.
ADR-10: Combined open benchmark instead of a new vendor benchmark#
Context. Each vendor publishes a benchmark it wins (ParseBench, RealDocBench, LongExtractBench).
Decision. LiteOCR will not author a competing "vendor-neutral" opinion benchmark. It will run every public benchmark through one harness with one scoring pipeline and publish per-dataset and combined scores, with every input, truth and output inspectable online. Datasets whose license permits redistribution are vendored in the repo; others are downloaded by an adapter at run time with the upstream revision pinned in the manifest.
Consequences. The site is the verification surface; the CLI is the reproduction surface. Some benchmarks (olmOCR-bench) need their own scorer implemented rather than transcript similarity.
ADR-11: Modes — providers are only interchangeable within a mode#
Context. "OCR provider" covers different products: plain text recognition (Textract DetectDocumentText, Azure Read, Google Document OCR), layout-aware parsing to markdown and typed blocks (Reducto, Extend, LlamaParse, Mistral OCR, Datalab, Unstructured, Upstage, Landing AI), and schema-driven structured extraction (Reducto Extract, Extend Extract, LlamaExtract, Azure prebuilt models, Textract Queries/Forms, vision LLMs with structured output). Swapping a parse model for an extract model is not a like-for-like switch.
Decision. The core defines Mode::{Parse, Ocr, Extract} with its own request/response
type per mode (DocumentRequest → ParseResponse, DocumentRequest → TextResponse,
ExtractRequest → ExtractResponse) and its own entry point (parse, ocr, extract).
Every registry model declares the modes it supports; resolution (ModelRef::parse_for) and
the Router reject models outside the requested mode. Pricing is per model and mode.
A layout provider serves ocr by deriving lines/words from its parse output (marked with
liteocr_derived_from = "parse"), so plain-text callers can still use it, but native OCR
endpoints override that when they exist.
Consequences. The Python API becomes liteocr.parse / liteocr.ocr / liteocr.extract
(the pre-release liteocr.ocr that meant parse is renamed; nothing was published). Vision
LLMs (Gemini, OpenAI, Anthropic) fit as parse/extract models without geometry, which the
response makes explicit by returning blocks without boxes. Adding a fourth mode (classify,
split) is additive: a new enum variant, types, trait method and entry point.
ADR-12: Provider fan-out rules#
Decision. Each provider lives in one file (crates/liteocr-core/src/providers/<name>.rs),
is declared up front in providers/mod.rs, and is implemented independently against the
Provider trait. Shared registration points (model.rs registry, pricing.json,
build(), README table, .env.example) are edited only by the integrator after a provider
lands, from snippets in the provider's hand-off. Providers without a key in the environment are
implemented from official docs with fixtures built from documented responses and #[ignore]
live tests; the task board records which ones have been verified live.
Consequences. Many providers can be built in parallel without merge conflicts; the cost is a short integration step per provider and, for unverified providers, a "docs-only" label until someone with a key runs the live tests.
ADR-13: Native-format compatibility is a renderer over the unified response#
Context. The unified response is the product, but it is also the migration cost: a team already
parsing Reducto's result.chunks[].blocks[].bbox.left cannot try another provider without
rewriting the code that reads the result. The one thing that would make switching free is getting
answers back in the shape they already parse.
Decision. Add crates/liteocr-core/src/compat/: a pure, infallible renderer
render_parse(&ParseResponse, Format) -> serde_json::Value (plus a best-effort render_extract)
with one module per vendor shape — Format::{Liteocr, Reducto, Extend, LlamaParse}. It is applied
after a call, not inside it: parse/ocr/extract keep returning the unified structs,
DocumentRequest.output_format only records and validates the caller's choice
(validate_output_format), and the SDK/CLI call ParseResponse::to_format(&str) on the result.
So routing, retries, pricing, fallbacks and every existing test are untouched.
What is promised is structural fidelity, not semantic identity: the vendor's key set, nesting,
chunk/page and block counts, content strings, block-type vocabulary and coordinate units. Fields
LiteOCR does not model are rendered null/empty and enumerated in docs/COMPAT.md, never invented.
Extend and LlamaParse need a page size for their unit boxes; when the source provider reports none
(Reducto, vision LLMs) the renderer assumes a 1000×1000 page and, for Extend, records
metadata.liteocr_synthetic_page_dims = true in the run's free-form metadata map.
The claim is enforced rather than asserted: for each provider, compat/roundtrip.rs runs
fixture → provider::normalize → render_parse(same format) and compares against the original
fixture with a skeleton_diff (key sets, counts, contents, types up to the provider's own forward
mapping, boxes within 1e-6, billed pages), plus cross-format and no-geometry cases. The tests live
in the crate because the normalize functions are pub(crate).
Consequences. Adding a fourth shape is one module plus one enum variant. Lossy edges are real
and documented: type mappings are not injective (Extend's key_value returns as text), the
Reducto render uses a reduced vocabulary (footnote/caption/formula/other → Text), and
render_extract is explicitly weaker than the parse path until extract fixtures exist for all
three providers. Because the renderer only reads the unified types, any provider added later gets
all three native shapes for free — and any unified field a new provider cannot fill shows up as a
null in someone's vendor-shaped payload, which is the honest outcome.
ADR-14: The Node.js SDK is a napi-rs addon over the same core#
Context. A TypeScript SDK was on the roadmap once the Python surface stabilised. The options were a WASM build, a wrapper around the CLI, or a native N-API addon.
Decision. crates/liteocr-node exposes the core through napi 3; the js/ package is a thin
layer. The addon only converts values: requests and responses cross as the core's serde JSON,
errors as a prefixed JSON payload (LITEOCR_CORE_ERROR:<json>) that the JS layer rebuilds into
typed LiteOCRError subclasses. The JS layer converts to camelCase with explicit per-type
converters; data, metadata and raw are never renamed, and index.d.ts is hand-written
against SPEC §5 (the napi-generated native.d.ts is internal). The crate is a normal workspace
member: napi's dyn-symbols keeps cargo test --workspace free of Node. It uses
deny(unsafe_code) because napi's macro expansion is incompatible with forbid, and declares
rust-version = "1.88" (napi 3) while the rest of the workspace stays at 1.80. timeout is in
seconds, as in the other SDKs, and LiteOCRError.kind uses the ErrorKind serde values.
Consequences. One implementation behind three surfaces (Python, Node, CLI). No provider logic
in JS. Prebuilt binaries per platform are required for npm install without a Rust toolchain;
the release matrix builds them but publishing is not wired yet.
ADR-15: The gateway is a thin axum layer over the core, TOML-configured, with no database#
Context. LiteLLM's proxy is what teams actually deploy: one endpoint, central keys, budgets and logs. LiteOCR needed the same without growing a second implementation of providers.
Decision. crates/liteocr-server (axum + tower-http, which are HTTP frameworks, not provider
SDKs) calls liteocr_core::{parse, ocr, extract} and runs its own fallback loop, because the core
Router cannot give each target its own credentials or base URL. Configuration is TOML (already
idiomatic in Rust, lighter than YAML). Usage and budget state live in memory with an optional
JSON state file; a database is out of scope. Clients may not send api_key or base_url, so they
cannot redirect the gateway's credentials, and local file paths are refused. Provider credential
failures map to 502, not 401, because the caller's own key was valid.
Consequences. Single binary, no infrastructure to run. Budgets can be overshot by requests in flight, and key changes need a restart. A database, HTTP key management, async job endpoints and metrics auth are follow-ups, not blockers.
ADR-16: Vendor only what the licence permits; index the rest#
Context. The combined benchmark (ADR-10) pulls in public datasets whose licences differ: olmOCR-bench is ODC-BY-1.0, OmniDocBench has no licence and is marked research-only / non-commercial.
Decision. A dataset is vendored into the repo (with attribution) only when its licence permits
redistribution. Otherwise the repo holds a manifest with upstream paths, a pinned revision and
image/truth hashes, and the adapter materialises the data locally. Combined datasets are
versioned and never rewritten once results exist (combined-v2 supersedes combined-v1 for new
runs). Upstream tests that cannot be expressed faithfully in the shared rule schema are skipped
and counted, never weakened silently; the one relaxed mapping (olmOCR "left/right of" → same row)
is counted as relaxed.
Consequences. Anyone can reproduce every score, but index-only sources need a fetch step before
a run (their documents carry a fetch-required tag). Stats files make the coverage of each
conversion auditable.
ADR-17: The benchmark results viewer stays vanilla JS#
Context. TASKS listed an open choice for benchmark/site/ between a React app on Extend UI
(PDF viewer and layout overlays out of the box) and the zero-build vanilla viewer. The viewer has
to be where every benchmark claim can be checked: page rendering, side-by-side outputs, diffs,
rule checklists, bbox overlays, charts, deep links.
Decision. Keep static HTML/CSS/ES2018 in benchmark/site/src/: no framework, no bundler, no
npm install; the build stays stdlib Python (Pillow optional). The one third-party runtime
dependency is pdf.js, loaded lazily from cdnjs at a pinned version with SRI, only when a PDF is
opened, with the build-time PNG as fallback. Overlays and charts are inline SVG. The per-document
rule checklist uses a JS port of liteocr-core's score_rules, and every rules page compares its
count with the recorded Rust score and flags any disagreement.
Consequences. Vercel and Pages build the site with one uv run command and nothing to audit.
The data contract (data/index.json, data/runs/, data/outputs/…/<doc>.{md,json}) is
independent of the front-end, so this can be revisited without touching the build. The JS port
must follow scorer changes in bench.rs; the mismatch badge makes drift visible.
ADR-18: Webhooks are exposed as primitives, not received#
Context. TASKS asked for "webhooks instead of polling". An SDK cannot host an HTTP endpoint, and providers differ: Reducto and LlamaParse accept a per-job webhook URL, Extend only has workspace-level webhook endpoints.
Decision. LiteOCR offers submit_parse / retrieve_parse plus a pure parse_webhook /
resolve_webhook that normalises a provider's webhook body into JobStatus (doing one retrieve
when the body only names the job). JobHandle is serialisable and secret-free; credentials are
resolved again at retrieve time. DocumentRequest.webhook_url maps to per-job provider webhooks
where they exist and is rejected with an input error where they don't. Only parse mode has jobs
for now.
Consequences. Users wire LiteOCR into their own web handler; the gateway (ADR-15) can later add job endpoints on top of the same primitives. Webhook body shapes come from vendor docs until real deliveries are captured.
ADR-19: Local and self-hosted engines are out-of-process providers#
Context. SPEC §1 listed local models as a v0.1 non-goal, but LiteOCR was unusable without a paid key and the benchmark had no open baseline.
Decision. Support local and self-hosted engines only as providers that call out of process: a
CLI binary via tokio::process (Tesseract, with pdftoppm for PDFs) or an HTTP server the user
runs (docling-serve, PaddleOCR/PaddleX serving). No C bindings, FFI, embedded runtimes or model
weights ship in LiteOCR. They are listed in model::SELF_HOSTED, need no key, and are priced at
0.0 with source "self-hosted".
Consequences. #![forbid(unsafe_code)] and the no-SDK rule still hold. Benchmark latency for
these engines depends on the user's hardware, so leaderboard rows need a hardware note. Live tests
need a binary or a server rather than a key.
ADR-20: The scorer is versioned; committed runs are re-scored offline#
Context. The first combined run exposed scorer defects (HTML tables scored 0, tokenised
punctuation in rule text, a dead bag_of_sentences threshold) that moved the leaderboard by up to
four points. Re-calling providers to fix a scoring bug would cost money and change latency and
provider versions at the same time.
Decision. Any scorer change that moves scores bumps SCORER_VERSION in liteocr-core and
re-scores committed runs from their saved outputs with liteocr bench rescore, never by
re-calling providers; measured latency and cost are kept. Result JSON records scorer_version
(absent = 1) and rescored_at. The viewer's JS rule checker follows every scorer version and
flags any document where it disagrees with the recorded Rust score. The bag_of_sentences
threshold of 0.8 is justified by sentences the reference extraction fused together. The table
structure metric is TEDS on the row/cell grid and is named teds_grid, not TEDS, because the
truth has no header/body or span structure.
Consequences. Leaderboards stay comparable across scorer fixes at no cost, and a reader can see which scorer produced a number. Saved outputs are part of every committed run (already required).
ADR-21: Gateway jobs are owned, not routed#
Context. ADR-18 left gateway job endpoints for later; ADR-15 gave the gateway fallbacks and per-key budgets.
Decision. /v1/jobs submits to exactly one deployment (an alias's first target, or the next in
a round-robin rotation) with no fallback: failures only surface at retrieve time, and retrying
would mean resubmitting. Jobs are stored under opaque gateway ids bound to the submitting key (the
master key can read all, as with /v1/usage), without secrets; provider credentials are resolved
from config on every poll. Cost is charged once, on the first observed success, under the usage
lock. Provider webhooks are received only when the operator opts in with a shared secret.
Consequences. Clients poll the gateway, never the provider. Providers can reuse job ids (LlamaParse returns the cached job for an identical upload), so one webhook may settle several gateway jobs. Vendor HMAC verification and automatic webhook registration remain open.
ADR-22: One design language; the viewer renders PDF pages at build time#
Context. The docs, landing page and viewer each had their own palette (blue in the docs, vermilion on the landing page), and the viewer showed raw ids, eleven metric tiles per document and full-table heatmaps; PDF pages depended on pdf.js from a CDN and often did not render.
Decision. website/assets/tokens.css is the single source of colour, type, space and shape,
described in docs/DESIGN.md; every surface links it first and styles only through its custom
properties. The viewer follows the language's principles (one number per view, evidence one click
away, human names with raw ids on request) and renders PDF pages to WebP at build time with
pypdfium2 (a pip wheel that works on Vercel); pdf.js remains only as a lazy fallback. Layout-box
overlays are off by default.
Consequences. A palette or type change is one edit. The site build needs pypdfium2 (added to
vercel.json and pages.yml); without it the build falls back to Pillow page-1 extraction.
ADR-23: The product is renamed PuffinParse#
Context. liteocr on PyPI belongs to an unrelated OCR engine, several GitHub projects already
use the name, and "OCR" undersells a tool whose modes are parse, OCR and extract. The owner's
preference, LiteParse, is a LlamaIndex product. About 110 names were checked against domains,
PyPI/npm/crates.io and web collisions.
Decision. The product, crates (puffinparse-{core,cli,python,node,server}), Python package
(puffinparse, native module puffinparse._core), Node package, CLI binary, environment variables
(PUFFINPARSE_*), metadata keys (puffinparse_*), error base class (PuffinParseError), the
native output format value ("puffinparse") and the site (puffinparse.vercel.app) all use the
new name. PuffinParse had every checked domain (.com, .dev, .ai, .io) and every registry name free
and no product collision; the puffin's black, white and orange match the existing palette.
Consequences. Committed benchmark results, recorded provider fixtures (which contain
"LiteOCR" in document text), the CHANGELOG history and earlier ADRs keep the old name. Result files
written before the rename carry liteocr_version, which the CLI and site builder read as an alias.
liteocr.vercel.app keeps serving the same project.
ADR-24: A puffin mark and mascot; brand imagery only in brand moments#
Context. After the rename (ADR-23) the product had no logo: each surface improvised a different mark (a scan line, an orange square, plain text), and the favicon was still the pre-ADR-22 blue. The puffin's black, white and orange already matched the palette.
Decision. Two brand assets. The mark (website/assets/mark.svg): a puffin head in a
rounded-square tile, the beak carrying two stripes read as parsed lines; it is the favicon and sits
beside the wordmark on every surface. The puffin (website/assets/puffin.svg): a mascot holding
three document fish, used only in brand moments (landing hero, 404, empty states, social card,
README), never beside data in the tools. Both are flat SVG in existing tokens plus three brand
tokens (--mark-tile, --mascot-body, --mascot-wing) that lift in dark mode. A committed social
card (og.png) is the og:image for every page. Alongside, two table rules tightened: headers are
sentence case, and bold marks a best value only when it is unique as displayed.
Consequences. "No decorative imagery" now reads "none in the tools". Inline copies of the mark
live in website/build.py (MARK) and the viewer's index.html; the social card must be
re-rendered (website/og/render.py, Playwright) when the card, mark or mascot changes. The launch
videos still show the text wordmark until they are re-rendered.
Amended 2026-09-25: the mark is C2 of the explored set (kept after comparing six alternatives), and
the mascot became the waving puffin with three pages in its beak (instead of three document fish),
which says "documents" more directly. docs/DESIGN.md gained a Brand motion section.