Contributing to PuffinParse#
Thanks for helping build PuffinParse — one API for every OCR / document-parsing provider. This document covers local setup, the checks CI runs, and the contributions we get asked about most: verifying or adding a provider and adding a benchmark dataset. If you use a coding agent, point it at AGENTS.md, which summarises the same rules.
Good first issues#
Issues labelled
good first issue
are scoped to one area and list a "done when" condition. Comment on the issue to
claim it so two people don't do the same work. Another useful first contribution
if you have a key for one of the docs-only providers (Mistral, Azure,
Textract, Gemini, OpenAI, Anthropic, Mathpix, Datalab, Unstructured, Upstage,
Landing AI, Google Document AI, PaddleOCR) is running its live tests and
reporting what differs; see
issue #10.
Questions and ideas that are not yet issues go to
Discussions.
By participating you agree to the Code of Conduct. PuffinParse is MIT licensed; contributions are accepted under the same license.
1. Setup#
Prerequisites:
- Rust stable (≥ 1.80 — the workspace
rust-version), via rustup. - Python 3.9+ (3.9 is the minimum we build abi3 wheels for).
- A C toolchain (whatever
ccyour platform ships) for the native deps. - Node.js 18+ only if you work on the TypeScript SDK in
js/.
git clone https://github.com/ajinkyashejul/puffinparse
cd puffinparse
# Rust side
cargo build --workspace
# Python side — use a virtualenv; maturin installs into the active one.
python -m venv .venv && source .venv/bin/activate
pip install maturin ruff mypy pytest
maturin develop # builds crates/puffinparse-python, installs `puffinparse`
maturin develop compiles the PyO3 extension (puffinparse._core) and links it
against the pure-Python package in python/puffinparse, so edits to the Python
sources take effect immediately; re-run it after changing any Rust code.
Use maturin develop --release when you care about speed (benchmarks, large
documents) — debug builds of the core are slow.
Provider keys go in a .env (copy .env.example) or in your shell:
export REDUCTO_API_KEY=...
export EXTEND_API_KEY=...
export LLAMA_API_KEY=... # LlamaCloud / LlamaParse, starts with llx-
You do not need any keys to contribute. Every test that talks to a live
provider skips itself when the relevant key is absent, and CI sets no secrets.
Live tests are opt-in: PUFFINPARSE_LIVE_TESTS=1 pytest python/tests -q for
Python and cargo test -p puffinparse-core -- --ignored for Rust. Never commit
a .env file or paste a key into an issue, log or fixture.
Debugging#
Set PUFFINPARSE_LOG to turn on the core's tracing output — request URLs, retry
and backoff decisions, poll loops, router fallbacks:
PUFFINPARSE_LOG=debug puffinparse parse invoice.pdf --model reducto/standard
PUFFINPARSE_LOG=puffinparse_core::providers=trace pytest python/tests -q
On the Python side the same events also surface through
logging.getLogger("puffinparse").
2. Tests, lint, formatting#
Run everything the CI runs with make lint test (and make test-node for the
TypeScript SDK), or piecemeal:
# Tests
cargo test --workspace
pytest python/tests -q
cd js && npm ci && npm run build:debug && npm test # Node SDK
# Lint / format
cargo fmt --all # or `cargo fmt --all --check` to only verify
cargo clippy --workspace --all-targets -- -D warnings
ruff check python/ benchmark/ examples/
ruff format --check python/ benchmark/ examples/
mypy python/puffinparse
Rules of the road:
cargo fmtoutput is authoritative; do not hand-format around it.- Clippy warnings are errors in CI. Prefer fixing over
#[allow]; if an#[allow]is genuinely right, put a one-line comment saying why. - No network in unit tests. Provider parsing is tested against recorded
JSON in
crates/puffinparse-core/tests/fixtures/. Live tests are#[ignore]d and/or key-gated. puffinparse-coreis#![forbid(unsafe_code)]. Keep it that way.- The Python package is fully typed and
mypy --strict-clean; new public API needs annotations and a docstring.
3. Adding a provider#
This is the highest-value contribution. A provider is a single file
implementing one trait. Check for an existing
new provider issue first,
or open one from the template so we can agree on model naming before you write
code.
- Implement the trait. Create
crates/puffinparse-core/src/providers/<name>.rsand implementOcrProvider(seecrates/puffinparse-core/src/provider.rs). Use the shared HTTP helpers insrc/http.rsso you inherit retries, backoff, deadlines and error classification — do not build your ownreqwest::Client, and do not vendor a provider SDK. Map the provider's response onto the unifiedOcrResponse/Page/Block/Usagetypes insrc/types.rs: - normalise
bboxto 0..1 with a top-left origin; - map the provider's block vocabulary onto
BlockType, unknown →other; - fill
Usage.pageswith the billed page count; - map provider errors onto the
Errorvariants insrc/error.rs(401/403 → auth, 429 → rate limit, 5xx / failed job → provider error); onlyProviderError,RateLimitErrorandTimeoutErrorare fallback-eligible in the router, so classify carefully. - Register it. Add the module and a
matcharm tobuild()incrates/puffinparse-core/src/providers/mod.rs. - Declare its models. Add a
ProviderInfoentry toPROVIDERSincrates/puffinparse-core/src/model.rs:name,display_name,env_var,base_url,docs, and oneModelInfoper mode with exactly onedefault: true. Model strings are"<provider>/<model>"; keep them short, lowercase and stable — they are public API. - Add pricing. Add
"<provider>/<model>"entries tocrates/puffinparse-core/src/pricing.jsonwithper_page_usd, asourceURL pointing at the public pricing page, and theupdateddate. Public list prices only. - Add a fixture + normalisation test. Save one real (redacted) response as
crates/puffinparse-core/tests/fixtures/<name>_<endpoint>.jsonand add a unit test that parses it and asserts the normalised output: page count, block types, a bbox inside 0..1,usage.pages, and that the document-levelmarkdownis the pages joined in order. Scrub keys, job ids, customer names and anything else non-public from the fixture. - Document it. Add
docs/providers/<name>.md(same sections as the existing pages) with a status banner, a row indocs/providers/README.md, the API-key env var (and any*_BASE_URLoverride) in.env.example, the model table inREADME.md, and a line in the## [Unreleased]section ofCHANGELOG.md. Label the provider live-verified only if its live tests passed against the real API; otherwise it is docs-only. - Benchmark it. If you have keys, run the benchmark with
--save-outputs, commit the result JSON underbenchmark/results/and the per-document outputs underbenchmark/results/outputs/<run_id>/, and add a section tobenchmark/LEADERBOARD.md(it is stitched per dataset by hand until #13 lands).
Verifying a docs-only provider#
Run its #[ignore]d live tests with your key
(cargo test -p puffinparse-core <provider> -- --ignored --nocapture), fix any
wire-format differences, replace the hand-built fixture with a redacted real
response, and change the label to live-verified in docs/providers/README.md,
the provider page's banner and README.md. One PR per provider.
The Python SDK needs no changes: it forwards whatever model string the core accepts.
Local and self-hosted engines#
Engines that run on the user's machine or their own server (tesseract,
docling, paddleocr) follow the same steps, with these differences:
- No API key. Add the provider name to
SELF_HOSTEDinmodel.rs; setenv_varto""(or to an optional key, as Docling does) andbase_urlto the local default.puffinparse providersthen showslocalin the Key column instead of a missing-key cross. Read the base URL from<NAME>_BASE_URLviaprovider::resolve_base_url. - Price 0.
pricing.jsongets0.0for each mode with"source": "self-hosted (...)". - Local binaries are run with
tokio::process(never C bindings or FFI), with the call's deadline andkill_on_drop. A missing binary must produce an error that names the binary and how to install it or point at it (TESSERACT_CMD). Shared helpers (download a URL input, base64, file-type sniffing, a self-cleaning scratch directory) are incrates/puffinparse-core/src/providers/local.rs. - Fixtures. Capture a real output from a local install where you can (the
Tesseract TSV and docling-serve fixtures are real); otherwise shape it from
the server's documented schema and mark the doc page docs-only. The
#[ignore]d live test needs the binary or server rather than a key. - Document how to install or start the engine in
docs/providers/<name>.md.
4. Adding a benchmark dataset#
Datasets live in benchmark/datasets/<name>/ and follow SPEC §10.2:
benchmark/datasets/<name>/
manifest.json # {name, version, description, license, documents:[…]}
docs/<id>.<ext> # the input file
truth/<id>.md # expected markdown (.txt for text-only documents)
Each entry in manifest.documents[] is:
{
"id": "invoice_001",
"file": "docs/invoice_001.png",
"truth": "truth/invoice_001.md",
"pages": 1,
"category": "invoice",
"tags": ["clean", "table"]
}
idis unique within the dataset and is what shows up in results JSON.categorygroups documents for per-category scores. The built-insynthetic-v1categories areplain,invoice,table,two_column,noisy_scan,handwriting_like,low_res,rotated,headings,multipage.tagsare free-form; documents taggedtableadditionally get atable_score.licensein the manifest is required and must permit redistribution. Only open datasets are committed. For a public set that cannot be redistributed (olmOCR-bench, OmniDocBench), contribute a downloader/adapter that produces this layout locally instead of the files themselves. The same goes for outputs: per-document outputs of research-only datasets (OmniDocBench) are not committed, only their scores.
Ground truth must be exact — that is why the built-in set is generated:
make dataset (python benchmark/generate_synthetic.py) renders documents
deterministically from the same source text it writes to truth/. Prefer
extending the generator over hand-writing truth files.
Verify with a cheap model before proposing the dataset:
cargo run -p puffinparse-cli --release -- bench run \
--dataset benchmark/datasets/<name> \
--models llamaparse/cost_effective \
--out benchmark/results/$(date +%F)-<name>.json
5. Commit messages and pull requests#
Write commit subjects in the imperative mood, ≤ 72 characters, no trailing
period — Add Mistral OCR provider, not Added… / Adds….
Conventional Commits prefixes
(feat:, fix:, docs:, perf:, refactor:, test:, chore:) are
welcome but optional. Explain the why in the body; wrap it at 72 columns.
Keep one logical change per commit and rebase rather than merge main.
Before opening a PR:
- [ ]
cargo fmt --all --checkis clean - [ ]
cargo clippy --workspace --all-targets -- -D warningsis clean - [ ]
cargo test --workspacepasses - [ ]
ruff check/ruff format --check/mypy python/puffinparseare clean - [ ]
pytest python/tests -qpasses - [ ] New behaviour has a test (fixture-based, no network)
- [ ] Docs updated —
README.md,docs/SPEC.mdif the contract changed,.env.examplefor a new key - [ ]
CHANGELOG.md## [Unreleased]has an entry - [ ] Public API changes are semver-appropriate and noted in the PR description
Fill in the PR template (summary, test plan, checklist). Small, focused PRs get reviewed fastest. If you are planning something large — a new crate, a change to the unified response shape, a new metric — open an issue first so we can agree on the design.
6. Releasing (maintainers)#
- Update
CHANGELOG.md: move## [Unreleased]items under a new## [x.y.z] - YYYY-MM-DDheading. - Bump
workspace.package.versionin the rootCargo.toml, runcargo check --workspaceto refreshCargo.lock, and commit. - Tag
vx.y.zand push the tag..github/workflows/release.ymlbuilds wheels + an sdist + CLI archives, publishes to PyPI via trusted publishing, and creates the GitHub Release.