PuffinParse docs
/ llms.txt GitHub

Changelog#

All notable changes to this project are documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning. The workspace crates (puffinparse-core, puffinparse-cli, puffinparse-python, …) and the Python package puffinparse share a single version. Entries before the rename say LiteOCR.

Unreleased#

Added#

  • THIRD_PARTY_NOTICES.md (generated by scripts/third_party_notices.py, which also fails on a copyleft-only dependency and writes the full licence texts with --full), an Acknowledgements section in the README, BibTeX citations and licence files for the benchmark datasets, provenance notes for the test fixtures, and source credits for the TEDS, olmOCR and OmniDocBench logic the scorer and adapters mirror.

Changed#

  • The results viewer build never copies documents tagged fetch-required (OmniDocBench, research-only), nor provider outputs for them, even in a clone that fetched them locally.

0.1.2 - 2026-10-08#

Fixed#

  • Release versions now match across all manifests (js/package-lock.json still said 0.1.0), so the release pipeline publishes again; 0.1.1 was tagged but never published.

Added#

  • Agent-friendly website: docs URLs return their Markdown for Accept: text/markdown (with Vary: Accept, and a Markdown 404 for unknown docs paths); one schema.org JSON-LD block per page (SoftwareApplication, WebSite with a docs-search SearchAction backed by /docs/?q=, TechArticle, and licensed Dataset entries on the benchmark viewer); robots.txt names the major AI crawlers explicitly; llms.txt gains when-to-use guidance and the benchmark JSON endpoints; and a new /docs/agents/ page, also served as /agents.md, with a copyable onboarding prompt.

0.1.1 - 2026-10-08 [not published]#

Fixed#

  • The source distribution now includes the LICENSE file its metadata names; PyPI rejected the 0.1.0 sdist for this, so 0.1.0 shipped wheels only.

0.1.0 - 2026-10-08#

First public release (PyPI wheels, CLI binaries, gateway image).

Changed#

  • Provider verification labels are consistent across the README, the provider reference, the landing page and the docs site: live-verified (Reducto, Extend, LlamaParse), verified locally (Tesseract, Docling) and docs-only (the rest, implemented from documentation and tested against fixtures, not yet run live; tracked in issue #10). Tesseract and Docling were previously counted as live-verified on the landing page.
  • README roadmap now lists the open GitHub issues; added issue templates for benchmark and dataset suggestions, a shorter pull-request template, AGENTS.md, and private vulnerability reporting in SECURITY.md.

  • Benchmark and docs tables use sentence-case headers, and a column's best value is bold only when it is unique as displayed (ties are no longer bolded).

  • Renamed to PuffinParse (ADR-23). Crates puffinparse-*, Python package puffinparse (import puffinparse), Node package and CLI binary puffinparse, environment variables PUFFINPARSE_* (was LITEOCR_*), metadata keys puffinparse_*, PuffinParseError, and output_format="puffinparse" for the native shape. The site moves to puffinparse.vercel.app (liteocr.vercel.app still works). Old result files with liteocr_version still load.

Added#

  • Release pipeline (docs/RELEASING.md). A v* tag publishes abi3 wheels and an sdist to PyPI, puffinparse plus five prebuilt platform packages to npm, puffinparse-core/-server/-cli to crates.io, the gateway image to ghcr.io/ajinkyashejul/puffinparse, and CLI archives with SHA256SUMS to GitHub Releases, all through trusted publishing (no stored registry tokens). npm and crates.io are switched on per repository variable once their one-time setup is done.

  • Tesseract baseline in combined-v3. tesseract/default (Tesseract 5.5.1, one OpenMP thread per process) was added to the 2026-09-25 run with bench run --resume: 199 documents, 0 failures, Overall 56.36 (dpbench 86.77, olmocr 41.67, omnidocbench 35.70, parsebench 20.73, synthetic 97.98). The six API rows are unchanged. Licence audit notes in the DP-Bench README (5 of the 40 vendored pages are Upstage's own documents, kept under its MIT declaration) and the olmOCR README (redistribution with attribution complies with ODC-BY and AI2's guidelines).

  • Puffin brand (ADR-24, docs/DESIGN.md Brand): a puffin mark (favicon and header logo on the landing page, docs and benchmark viewer) and a waving puffin mascot with three pages in its beak, shown on the landing hero, the 404 page and the README. Every page now has an Open Graph / Twitter social card (website/assets/og.png). The landing demo shows a response on arrival instead of starting blank.

  • First combined-v3 results (6 API models × 199 documents, adds a 40-page DP-Bench subset; 0 failures, $14.65): llamaparse/cost_effective leads (84.35) ahead of llamaparse/agentic (83.49) and reducto/r-1 (82.40). It is the new headline in benchmark/LEADERBOARD.md and the viewer; combined-v2 keeps the Tesseract baseline row.

  • Design language (docs/DESIGN.md, website/assets/tokens.css): paper-and-ink neutrals, one scan accent, verdict and proof-mark colours, layout-box hues, a system-font type scale, the bounding-box and scan-line signatures; every text token passes WCAG AA in both themes. The docs, landing page and benchmark viewer all style through it.

  • Benchmark viewer redesign. Human document titles ("Headers & footers 3") with raw ids in Details; one primary number per view with secondary metrics, methodology and reproduce commands in disclosures; a compact leaderboard with per-source columns and an All metrics toggle; a calmer score-vs-cost chart with non-overlapping labels; plain-English check rows; a documents list with mean score and best model. PDF pages are rendered at build time with pypdfium2, so olmOCR-bench and ParseBench pages show without pdf.js.
  • First combined-v2 results (7 models × 159 documents, 0 failures, $11.74): llamaparse/cost_effective leads (83.79) ahead of llamaparse/agentic (83.05) and reducto/r-1 (81.32); tesseract/default is the free baseline (48.68, 97.94 on synthetic). benchmark/LEADERBOARD.md now has one section per dataset with combined-v2 as the headline. The viewer shows research-only documents (OmniDocBench) as scores only, with the command to fetch the data, instead of broken links.
  • Jobs everywhere. Gateway: POST /v1/jobs (parse body + webhook_url, 202 with a job id) and GET /v1/jobs/{id} (pending / succeeded / failed, output_format on retrieval); job ids are bound to the submitting key, cost is charged once when the job first succeeds, handles persist in state_file (server.job_retention_hours), and an opt-in POST /v1/webhooks/{provider} receiver ([webhooks] enabled, shared secret) settles jobs from provider webhook bodies. New liteocr_jobs_total metric and job_id/job_status log fields. Node SDK: submit, retrieve, handleWebhook with a typed Job.
  • DP-Bench (Upstage, MIT) adapter: 40 committed single-page PDFs with reading-order transcript truth (python -m benchmark.adapters dpbench, pinned at 24702c61), and combined-v3 = combined-v2 + DP-Bench (199 documents). Headers and footers stay in the DP-Bench truth because DP-Bench's own NID scores them; each source keeps its publisher's conventions.
  • tesseract/default runs tesseract with OMP_THREAD_LIMIT=1 unless set, so parallel pages no longer oversubscribe the CPU (70 s → 0.7 s for a small page on 4 cores).
  • Benchmark results flag a successful call that returned no text as empty_output: true (counted in summary.empty_outputs, shown as (+N empty) in the leaderboard), and bench rescore --keep-missing keeps the recorded scores of documents whose outputs are not committed.
  • Scorer v2 (scorer_version: 2 in result JSON). HTML <table> output is scored like markdown tables (it used to score 0, which ranked Reducto r-1 last); a new teds_grid metric (TEDS on the row/cell grid); rule matching ignores spaces next to punctuation (ParseBench rule text is tokenised); bag_of_sentences matches sentences fuzzily (0.8) and passes at 0.8 (1.0 was dead signal); olmOCR max_diffs tolerances are honoured; HTML entities are decoded; and results carry an explicit summary.headline and per-document headline, so char_similarity means literal character similarity again. Python and Node Metrics gain teds_grid / tedsGrid, and the viewer's rule checker follows the same rules.
  • liteocr bench rescore re-scores a run from its saved outputs with the current scorer, no network calls, keeping measured latency and cost. Both committed runs are re-scored and benchmark/LEADERBOARD.md regenerated: on combined-v1, reducto/r-1 moves from last (87.19) to second (91.37); llamaparse/agentic stays first (93.08).
  • docs/benchmarks/findings.md: why four models returned nothing for parsebench/text_multicolumns_2col (a Form XObject with a ±2^1023 /BBox that 32-bit renderers clip to empty; fixing the file fixes both Reducto and Extend, no option does).
  • Self-hosted engines, no key, $0/page. tesseract/default (local tesseract binary, PDFs via pdftoppm; native OCR with word/line boxes and confidences), docling/default (docling-serve v1 async API; layout, tables, OCR) and paddleocr/default (PaddleOCR/PaddleX serving: /ocr and PP-StructureV3 /layout-parsing; docs-only). Tesseract and Docling are live-verified locally. liteocr providers shows local for them; --json and Python providers() include self_hosted.
  • README, docs site and landing page cover the TypeScript SDK, the gateway, jobs/webhooks, self-hosted engines and the new benchmark sources; the docs navigation gains the gateway, the self-hosted provider pages, the academic-benchmark survey and the adapter notes.
  • Async jobs API and webhooks. Split a parse into submit and retrieve so the caller owns the waiting (long documents, batches, webhook-driven pipelines). Rust: submit_parse, retrieve_parse / retrieve_parse_with, parse_webhook, resolve_webhook, JobHandle (serialisable, never holds a key), JobStatus, and DocumentRequest.webhook_url. Python: liteocr.submit / asubmit, retrieve / aretrieve, handle_webhook / ahandle_webhook and liteocr.Job. Covers Reducto, Extend and LlamaParse (live-verified); webhook_url maps to Reducto async.webhook and the LlamaParse webhook_url field, and is rejected for Extend, which only has workspace-level webhooks. SPEC §15.
  • Benchmark viewer: verify every claim. A per-document inspector at #/<run>/<model>/<doc> shows the page itself (pdf.js for PDFs, pinned with SRI, PNG fallback), every model's output side by side, a word diff against the truth, and a pass/fail rule checklist for rules documents that is checked against the recorded score. Layout-box overlays by block type appear when a run saved unified responses. The leaderboard ranks by summary.headline when present and adds p90 latency, a score-vs-cost/latency scatter with the Pareto frontier, and per-source and source × category breakdowns. Every view is a shareable link, with keyboard navigation (j/k, m, 1–4, d, o, ?), the product site's design and theme switch, and a mobile layout.
  • liteocr bench run --save-outputs also writes the unified response as <doc>.json next to <doc>.md, which the viewer uses for its layout overlay.
  • More public benchmarks in the combined dataset. olmocr (40 AI2 olmOCR-bench PDFs, 205 rules, ODC-BY-1.0) and omnidocbench (40 pages across 10 document types, English and Chinese; index only — images and truth are fetched at a pinned revision by python -m benchmark.adapters omnidocbench because the data is research-only), joined with synthetic-v1 and ParseBench as combined-v2 (159 documents; combined-v1 is unchanged). Upstream tests that cannot be expressed faithfully (math, baseline, positional absences, vertical table neighbours) are skipped and counted in conversion-stats.json, never dropped silently. A dataset self-check test proves every converted rule is satisfiable. Survey of academic benchmarks with licences at pinned revisions: docs/benchmarks/academic-benchmarks.md.
  • Gateway server: liteocr serve (crates/liteocr-server). An HTTP gateway in front of every provider, the LiteOCR equivalent of the LiteLLM proxy: POST /v1/parse|ocr|extract (JSON or multipart, output_format, fallbacks), GET /v1/models, /v1/usage, /health and Prometheus /metrics. A liteocr.toml config defines model aliases with ordered or round-robin fallback, provider keys as env: references, and virtual keys with model allow-lists, monthly USD budgets and per-minute rate limits. Every request is logged as one JSON line (never document content, provider error text or secrets), and all errors share one JSON body. Clients cannot override api_key/base_url or read local files. Ships a multi-stage distroless Dockerfile and examples/server/liteocr.toml; see docs/SERVER.md.
  • Node.js / TypeScript SDK. The liteocr npm package in js/ runs on a napi-rs addon over the same Rust core (crates/liteocr-node): async parse / ocr / extract (with fallbacks), Router, camelCase typed responses (index.d.ts), LiteOCRError subclasses mapped from the core ErrorKind, and the pricing, model and scoring helpers. CI builds the addon and runs the typecheck and node:test suite; the release workflow builds prebuilt .node binaries as artifacts (npm publishing not wired yet, so build from source for now). Docs at /docs/typescript/.
  • liteocr bench run --resume: every finished call is appended to <out>.partial.jsonl, so an interrupted or partly failed run continues where it stopped and only re-runs missing or failed (model, document) pairs. --dry-run prints the plan (calls, pages, list-price estimate) without calling any provider; --max-cost <usd> aborts before the first call when the estimate is higher; --retries N re-issues documents after retryable errors (off by default, since a retried job may be billed twice). Result documents record provider_job_id, cache_hit, attempts, started_at and error_kind, and the run ends with a summary line (calls, failures, cost, wall time).
  • Native-format compatibility (output_format). A response can be rendered in a provider's own JSON shape instead of the unified one, so an integration already written against Reducto, Extend or LlamaParse can switch the underlying provider without rewriting its parsing code: ParseResponse::to_format("reducto") in Rust (liteocr_core::compat::{Format, render_parse, render_extract}), with DocumentRequest.output_format carrying and validating the choice. What is guaranteed is structural fidelity — key set, chunk/page and block counts, content strings, block-type vocabulary, coordinate units and billed pages — not byte equality; the always-null fields and the lossy type mappings are enumerated in docs/COMPAT.md, and each provider's fixture is round-tripped through its own renderer in the test suite. Extract-mode rendering is best effort. See ADR-13.
  • output_format in the Python SDK and the CLI. liteocr.parse(..., output_format="reducto") (and aparse, extract, aextract, plus the matching Router methods) returns the vendor's own JSON as a dict instead of the dataclass; output_format=None (the default) or "liteocr" keeps the unified shape. The value is validated in the core before any network call — an unknown name raises BadRequestError listing liteocr | reducto | extend | llamaparse — the rendering happens in Rust (liteocr._core.render_parse / render_extract), and success callbacks still receive the dataclass. Overloads type the return (None → dataclass, str → dict), liteocr.output_formats() lists the accepted values, and examples/switch_provider_keep_format.py shows a provider swap with the parsing code untouched. The CLI gains --output-format <vendor> on parse (with --format json; ignored with a warning on stderr otherwise) and on extract, and liteocr providers --json now emits {"providers": [...], "output_formats": [...]} instead of a bare array.

Changed#

  • Internal: the vision-LLM providers (Gemini, OpenAI, Anthropic) share providers/vlm.rs; behaviour and request bodies are unchanged.
  • Adapter HTML→markdown table conversion no longer doubles backslashes, so LaTeX in table cells reaches the scorer as a parser would print it.
  • Modes. Every call now names a mode — parse (markdown + typed blocks), ocr (plain text with line/word boxes) or extract (a JSON object from a schema, with per-field confidence and citations) — and providers can only be swapped within a mode. The Python SDK exposes parse/aparse, ocr/aocr (now plain text, not markdown) and extract/aextract, the new TextResponse / ExtractResponse dataclasses, a mode-bound Router(models, mode=...), and mode arguments on list_models, resolve_model, set_pricing and estimate_cost; pricing is now per page per mode. The CLI gains liteocr ocr and liteocr extract, and liteocr providers shows modes, per-mode prices and a --mode filter. The old liteocr.ocr (which returned markdown) is now liteocr.parse, and OcrResponse is now ParseResponse. Which models serve which mode is reported by liteocr.list_models(mode) and liteocr providers --mode <mode>.
  • README model table. The "Model names" section is regenerated from the registry and pricing.json: every provider group carries its env var and verification status (live-verified vs docs-only), every model its modes, the default per mode and the list price per page for parse · ocr · extract.

Fixed#

  • Extend provider_options={"responseType": "url"} is now sent as the GET /parse_runs/{id} query parameter; it used to go in the request body, so the presigned-output path never triggered. Reducto and Extend url-typed results are now covered by live-captured fixtures and loopback tests.

Initial development - 2026-09-11#

Initial release.

Added#

  • Rust core (liteocr-core) — a unified request/response contract for document parsing: OcrRequest accepting a path, bytes or URL, and an OcrResponse with pages, blocks (normalised block types and 0..1 bounding boxes), document-level markdown/text, usage, cost_usd and latency_ms, identical across every provider. #![forbid(unsafe_code)].
  • Providers — Reducto, Extend and LlamaParse, each talked to over plain HTTPS with no vendored SDK, addressed by "<provider>/<model>" model strings (reducto/standard, extend/parse_performance, llamaparse/cost_effective, …) with a per-provider default and a provider_options pass-through escape hatch.
  • Router — ordered fallbacks and round-robin load balancing across models, with per-model success/failure and latency stats; only provider, rate-limit and timeout errors trigger fallback.
  • Pricing and cost tracking — an embedded, overridable pricing.json price table mapping model → per-page (or per-credit) USD, used to compute cost_usd on every response.
  • Reliability primitives — retries with exponential backoff on 429/5xx and network errors, per-call timeouts, and whole-call deadlines that cover upload plus polling.
  • Benchmark metrics — character similarity (the primary score), CER, WER, word-level F1/recall, reading-order agreement and a table-restricted score, all computed in Rust over NFKC-normalised text.
  • CLI (liteocr) — parse (markdown/text/json output, optional raw provider payload), providers (models, pricing and API-key status), and bench run / bench report / bench score for running the benchmark, scoring results and generating the leaderboard.
  • Python SDK (liteocr) — ocr(), the async aocr(), Router, typed dataclasses (OcrResponse, Page, Block, BBox, Usage), the full error hierarchy, set_pricing(), list_models(), and success/failure callbacks. Fully typed, ships py.typed, distributed as abi3 wheels for CPython 3.9+.
  • synthetic-v1 benchmark dataset (v1.1.0) — 39 deterministically generated documents across 13 categories (plain, headings, invoice, table, two-column, noisy scan, low resolution, multi-page PDF, skewed, dense, faded, receipt, complex table) with exact markdown ground truth, plus the generator that produces them byte-for-byte.
  • First leaderboard (benchmark/LEADERBOARD.md) from a run over seven models across the three providers; bench run disables provider result caches by default (--allow-cache to opt out) so latency reflects real work.