Benchmark adapters#
How public OCR benchmarks become PuffinParse datasets.
ADR-10 says PuffinParse does not author a competing benchmark: it runs every
public benchmark through one harness. An adapter is the piece that makes that true — it
fetches one upstream benchmark at a pinned revision and rewrites it into
benchmark/datasets/<name>/manifest.json. Datasets whose license permits redistribution are
vendored into the repository; the rest stay an index plus a sha256, fetched on demand.
Everything lives in benchmark/adapters/:
| File | What it is |
|---|---|
base.py |
the framework: Adapter, the manifest model, the rule schema, shared helpers |
parsebench.py |
LlamaIndex ParseBench → benchmark/datasets/parsebench/ |
olmocr.py |
AI2 olmOCR-bench → benchmark/datasets/olmocr/ (kind: rules) |
omnidocbench.py |
OpenDataLab OmniDocBench → benchmark/datasets/omnidocbench/ (kind: transcript, index only, fetched at run time) |
dpbench.py |
Upstage DP-Bench → benchmark/datasets/dpbench/ (kind: transcript, 40 PDFs committed) |
combined.py |
union of the datasets → benchmark/datasets/combined-v1/ (adapter combined, frozen), benchmark/datasets/combined-v2/ (adapter combined-v2) and benchmark/datasets/combined-v3/ (adapter combined-v3: v2 + dpbench) |
__main__.py |
the python -m benchmark.adapters CLI |
CLI#
python -m benchmark.adapters --list
python -m benchmark.adapters <name> [--limit N] [--out DIR] [--cache DIR] [--seed N] [--no-download]
| Flag | Meaning |
|---|---|
--limit N |
cap the number of documents built (selection stays deterministic) |
--out DIR |
output dataset directory (default: the adapter's default_out) |
--cache DIR |
download cache (default: $PUFFINPARSE_BENCH_CACHE, else $XDG_CACHE_HOME/puffinparse/benchmarks, else ~/.cache/puffinparse/benchmarks) — never inside the repo |
--seed N |
selection seed, default 1234 |
--no-download |
build from an existing cache, never hit the network |
Downloads go over HTTPS through whatever proxy the environment configures. Behind the agent
proxy, point huggingface_hub at the CA bundle (the adapter sets these if they are unset) —
never disable TLS verification:
export REQUESTS_CA_BUNDLE=/root/.ccr/ca-bundle.crt SSL_CERT_FILE=/root/.ccr/ca-bundle.crt
The CLI prints a build summary: documents by kind, committed size, and the adapter's own counters (rules converted, rules skipped by upstream type, documents skipped and why).
Manifest extension#
The manifest stays exactly what benchmark/README.md and
docs/SPEC.md §10.2 describe:
{"name", "version", "description", "license",
"documents": [{"id", "file", "truth", "pages", "category", "tags"}]}
Adapters add optional keys. Every added key is additive and ignored by the current Rust
Manifest / ManifestDoc (serde ignores unknown fields by default), so old manifests keep
working and new manifests load in an unmodified CLI.
Document-level additions#
| Key | Type | Meaning |
|---|---|---|
kind |
"transcript" | "rules" |
how the document is scored. Default "transcript" — absent means transcript, which is what every pre-existing manifest is |
rules |
path | for kind: "rules", the assertion file, relative to the dataset directory |
source_id |
string | the upstream id, kept verbatim when id had to be slugified |
upstream_path |
string | the document's path inside the upstream repo, for on-demand fetch |
sha256 |
hex | the document file's hash |
license |
SPDX | per-document license, when a dataset mixes them |
attribution |
string | the citation this document must carry |
source_url |
string | where the page originally came from (olmOCR-bench records one URL per test) |
truth_sha256 |
hex | hash of a truth file that is generated locally and not committed (OmniDocBench) |
truth is always emitted, even for kind: "rules" documents, where it is the empty string.
That is not cosmetic: the Rust ManifestDoc declares pub truth: String without
#[serde(default)], so a missing key fails to deserialise the whole manifest.
Manifest-level additions#
generator, upstream ({repo_id, revision, url, repo_type}), attribution, notes, and —
for combined datasets — sources:
"sources": [{"name", "version", "license", "path", "documents", "manifest_sha256", "attribution"}]
kind: "transcript"#
The existing behaviour: truth is markdown, scored by puffinparse_core::bench with
char_similarity / cer / wer / word_f1 / order_score / table_score.
kind: "rules"#
truth is empty and rules points at a JSON list of machine-checkable assertions. One
schema covers every rule-based upstream benchmark we care about (ParseBench today,
olmOCR-bench next):
{
"id": "text_dense_baoutou_order_623",
"type": "present" | "absent" | "order" | "table_cell" | "bag_of_sentences",
"text": "…", // present / absent
"before": "…", "after": "…", // order
"cell": {"row_header": "…", "col_header": "…", "value": "…"}, // table_cell
"sentences": ["…"], // bag_of_sentences
"threshold": 0.8, // bag_of_sentences: required pass fraction
"case_sensitive": false,
"max_diffs": 0, // optional: upstream fuzzy allowance (olmOCR-bench)
"source": "text_dense__baoutou_order_623"
}
| Type | Passes when the parsed markdown… |
|---|---|
present |
contains text (after the run's normalisation) |
absent |
does not contain text |
order |
contains both before and after, and the first occurrence of before precedes the first occurrence of after (with max_diffs > 0: some fuzzy match of before starts before some fuzzy match of after, as olmOCR does) |
table_cell |
has a table (markdown pipe table or HTML <table>, spans repeated into every slot) with a row matching cell.row_header and a column matching cell.col_header whose cell equals cell.value |
bag_of_sentences |
contains at least threshold (fraction, default 0.8) of sentences; a sentence counts when some window of the output is ≥ 0.8 similar to it (1 − edit distance / length) |
Matching normalisation (scorer v2): the run's normalize, then
every space adjacent to punctuation is dropped on both sides, so a tokenised rule such as
(this " agreement ") matches the printed (this "Agreement") and vice versa. Words still need
their spaces.
case_sensitive is always present and always explicit. source is the upstream rule id, so any
score can be pushed back to the publisher's own harness for cross-checking. max_diffs is
optional and records how many Levenshtein edits the upstream scorer tolerates. Since scorer v2 the
Rust scorer honours it: present / absent use a fuzzy substring search (Sellers' algorithm, like
upstream's fuzzysearch.find_near_matches(max_l_dist=max_diffs)), order uses the fuzzy rule
above, and table_cell tolerates max_diffs edits in the header matches and the value. 0 or
absent means exact matching after normalisation.
A document's rule score is passed / total; a dataset's is the mean over its rule documents.
That is deliberately the same shape as an accuracy in [0, 1], so it slots next to
char_similarity in the existing Summary.
ParseBench mapping#
Upstream: llamaindex/ParseBench,
pinned to commit 2805a1d940f95a203e0ae4b88be9934f7765b3fc, Apache-2.0, 2,078 single-page
documents and 169,011 assertions in five JSONL files.
Rule conversion, whole upstream dataset#
| Upstream file | Upstream type | Rules | → PuffinParse | Converted | Skipped |
|---|---|---|---|---|---|
table.jsonl |
expected_markdown |
503 | kind: transcript, HTML table → markdown |
503 | 0 |
text_content.jsonl |
missing_specific_word |
105,369 | present (rule.word) |
105,369 | 0 |
missing_specific_sentence |
18,768 | present (rule.sentence) |
18,768 | 0 | |
order |
13,087 | order (rule.before / rule.after) |
13,087 | 0 | |
missing_sentence_percent |
503 | bag_of_sentences (rule.bag_of_sentence keys, threshold: 0.8) |
503 | 0 | |
unexpected_sentence_percent |
503 | — | 0 | 503 | |
too_many_sentence_occurence_percent |
503 | — | 0 | 503 | |
missing_word_percent |
506 | — | 0 | 506 | |
unexpected_word_percent |
506 | — | 0 | 506 | |
too_many_word_occurence_percent |
506 | — | 0 | 506 | |
bag_of_digit_percent |
486 | — | 0 | 486 | |
is_header |
278 | — | 0 | 278 | |
is_footer |
307 | — | 0 | 307 | |
text_formatting.jsonl |
is_title, is_bold, is_italic, is_sup, is_sub, is_underline, is_strikeout, is_mark, is_latex, is_code_block, title_hierarchy_percent |
5,997 | — | 0 | 5,997 |
chart.jsonl |
chart_data_point |
4,864 | — | 0 | 4,864 |
layout.jsonl |
element bounding boxes | 16,325 | — | 0 | 16,325 |
| Total | 169,011 | 138,230 | 30,781 |
Why the skips:
unexpected_*/too_many_*/*_percentcounters are precision-direction assertions: "the output must not contain sentences the page does not have". Checking one needs the complete reference text, and ParseBench never publishes it for the text split — only bags. Abag_of_*compared against an incomplete reference would punish correct output.bag_of_digit_percentis a digit histogram, which nopresent/absentassertion expresses.is_header/is_footerassert a string's role on the page, not its presence.text_formatting.jsonlasserts style (bold, italic, superscript, LaTeX, heading level). Markdown carries some of this, but our normalisation strips it before scoring, so a rule about emphasis would be unscoreable. A futureformattingkind could pick these up.chart.jsonlasserts a value read off a chart;layout.jsonlasserts bounding boxes with IoA matching. Neither has a text-similarity or substring analogue.
Net: 138,230 of 169,011 upstream assertions (81.8%) are expressible, covering the table
and text_content dimensions — 1,009 of 2,078 pages (48.6%).
Committed subset#
The full conversion would vendor 517 MB of PDFs, so only a curated subset is committed
(benchmark/datasets/parsebench/, 5.7 MB total, 4.0 MB of PDFs):
| Docs | Rules | Notes | |
|---|---|---|---|
kind: transcript (table split) |
25 | — | 18 tagged merged-cells, 8 tagged hard, all tagged table-only |
kind: rules (text split) |
15 | 3,703 | present 3,303 · order 385 · bag_of_sentences 15 |
Selection is deterministic for a given --seed: table pages are taken at most one per source
document with roughly a third hard; text pages round-robin across all eight ParseBench
document-type buckets (simple, ocr, multicolumns, multilang, misc, dense, sparse,
handwritting). Documents over 400 KB, over the 11 MB budget, or with more than one page are
skipped and counted.
manifest.full.json in the same directory indexes all 1,009 convertible documents with their
upstream_path and is not runnable as-is; subset-committed.txt lists the 40 committed ids.
The table-only caveat#
ParseBench's table truth is the page's table, not the page. A parser returns the whole
page. Measured on apple_10_k_page1 with reducto/standard:
| Metric | Value | Reading |
|---|---|---|
table_score |
0.990 | the real signal — markdown table rows only |
char_similarity |
0.311 | meaningless here: the prediction also has the rest of the page |
cer / wer |
2.212 / 1.938 | same |
overall |
31.13 | derived from char_similarity, so also meaningless |
So: score table-only documents with table_score. Every such document carries the
table-only tag precisely so a scorer can pick the right primary metric. merged-cells marks
the 18 documents where colspan/rowspan had to be flattened by repeating cells — cell content
survives, table structure (ParseBench's GriTS) does not.
Framework reference (base.py)#
class Adapter(ABC):
name: str # CLI argument and dataset name
license: str # SPDX for the redistributable part
upstream: Upstream # repo_id + pinned revision + url
default_out: str # repo-relative output directory
def download(self, cache_dir: Path) -> Path: ...
def build(self, out_dir: Path, limit: int | None = None, seed: int = 1234) -> Manifest: ...
Shared helpers:
| Helper | What it does |
|---|---|
sha256_file, sha256_bytes |
streamed hashing for manifest provenance |
write_manifest, write_rules, write_json |
stable 2-space JSON with a trailing newline; returns the SHA-256 written |
slugify |
ASCII, lowercase, _-separated ids safe for filenames and --filter |
html_table_to_markdown |
HTML <table> → GitHub pipe table; flattens colspan/rowspan by repeating the cell into every position it covers, and reports merged_cells so the caller can add the merged-cells tag. Inline markup is dropped, <br> becomes a space, \| is escaped (backslashes are kept verbatim so LaTeX in cells survives), short rows are padded |
pdf_page_count |
pypdf when importable, otherwise a byte scan: the /Count of the root /Type /Pages node (also inside inflated object streams), falling back to distinct /Type /Page object numbers so an incrementally-updated PDF is not double counted |
default_cache_dir |
$PUFFINPARSE_BENCH_CACHE → $XDG_CACHE_HOME/puffinparse/benchmarks → ~/.cache/puffinparse/benchmarks |
pypdf is optional on purpose. It pulls in cryptography, whose Rust extension can raise a
pyo3 PanicException (a BaseException, not an Exception) on a mismatched wheel, so the
import is guarded broadly and cached; a broken optional dependency degrades to the byte scan
instead of aborting a build.
Adding an adapter#
- Create
benchmark/adapters/<name>.py. - Subclass
Adapter, setname,license,default_out,descriptionandupstreamwith a pinned commit — not a branch. - Implement
download(cache_dir)(fetch only what conversion needs; keep the big media trees for a per-document fetch) andbuild(out_dir, limit, seed). - Decorate the class with
@registerand import the module frombenchmark/adapters/__init__.py. - Use
Doc/Manifest/Ruleandwrite_manifest/write_rulesso output stays byte-stable, and record every skip withself.bump(...)so the summary is honest. - If the dataset belongs in the combined benchmark, add a new
SOURCES_V<n>andCombinedV<n>Adapterincombined.py(ascombined-v3did for DP-Bench). Never change the sources of an existing combined version: a result is only reproducible against the exact manifest it scored. Write aREADME.mdin the dataset directory with the license, and list the dataset inbenchmark/README.md. - Add the dataset to
crates/puffinparse-core/tests/benchmark_datasets.rs. Every rule must pass against a witness built from its own document, and every transcript must score 1.0 against itself.
Not adaptable into either kind, per the vendor benchmark survey:
- RealDoc-Bench: QA over a parse. It needs an LLM reader, which conflicts with the deterministic-metrics principle, and its source PDFs are not redistributable.
- RealDoc-Bench-Layout: bounding boxes, no text.
- LongExtractBench: schema-driven extraction, a different mode.
The academic benchmark survey covers olmOCR-bench, OmniDocBench, DP-Bench, READoc, the Nanonets IDP leaderboard, Fox, CC-OCR and OCRBench v2. DP-Bench is now adapted (below); READoc, as a long-document track, is the next candidate.
olmOCR-bench mapping#
Upstream: allenai/olmOCR-bench at
54a96a6fb6a2bd3b297e59869491db4d3625b711, ODC-BY-1.0. It has 1,403 single-page PDFs and
7,010 tests, plus 9 explicit baseline tests. The semantics were read from
olmocr/bench/tests.py at f7cfe4c22098b154c76b6ec950d1c0a464eecf8d. Every document is
kind: rules.
| Upstream test | → PuffinParse | Converted | Skipped | Note |
|---|---|---|---|---|
present |
present (upstream case_sensitive, default true) |
721 | 0 | |
absent |
absent |
622 | 201 | positional (first_n / last_n) absences are skipped: a whole-page absence would be the wrong assertion |
order |
order, case_sensitive: true |
1,061 | 0 | |
table |
table_cell |
636 | 384 | top_heading → col_header, left_heading → row_header; left/right → row_header (327, relaxed to "same row"); up/down skipped |
math |
— | 0 | 3,385 | KaTeX-rendered equivalence |
baseline |
— | 0 | 9 | plus 1,394 implicit per-PDF baseline tests the upstream runner adds |
| Total | 3,040 | 3,979 | 824 of 1,403 PDFs keep ≥ 1 rule |
All rule text has whitespace runs collapsed. Upstream's normalize_text does the same on both
sides, so nothing is lost, and it matters for table headers annotated across two lines.
What is committed:
- 40 documents, 8 from each split that has convertible tests: 6.0 MB of PDFs and 205 rules.
manifest.full.json, an index of all 824 convertible documents.conversion-stats.json, which records every count.
Documents whose rules are all absent carry the absent-only tag, because an empty parse
passes them. Details and caveats:
benchmark/datasets/olmocr/README.md.
OmniDocBench mapping#
Upstream: opendatalab/OmniDocBench
at aa1ee96d106dbe53d0ae59474d75c6e6d9b53fec. It has 1,651 page images and one annotation JSON
with blocks, reading order and text / LaTeX / HTML. Each page becomes a kind: transcript
document. The truth is built the way upstream's tools/json2md.py builds it:
- blocks are emitted in reading order, with
truncatedchains merged; - titles become
#headings, tables are converted from HTML to pipe tables, and display formulas stay as$$…$$; - headers, footers, page numbers, page footnotes,
abandonregions and figures are dropped, because OmniDocBench does not score them either.
Pages with *_mask regions (126) or less than 120 characters of truth (82) are skipped, which
leaves 1,443 convertible pages. category is the upstream data_source. Language, layout,
subset, special issues, has-table, has-formula and merged-cells are tags.
Licence. The card has no licence and says "research purposes only and not for commercial
use". So only manifest.json, with the image sha256 and truth_sha256, is committed. docs/
and truth/ are excluded from version control and produced by
python -m benchmark.adapters omnidocbench. Every document carries the fetch-required tag.
Details: benchmark/datasets/omnidocbench/README.md.
DP-Bench mapping#
Upstream: upstage/dp-bench at
24702c61a2fb13325534be664653bc6e60250d13 — data and evaluate.py pinned together. 200
single-page PDFs and dataset/reference.json, which lists every layout element of a page
(12 categories) in reading order (id), with content.text and, for tables, content.html.
Licence. The card front-matter says license: mit; the repository has no LICENSE file at
that revision. The adapter reads the card at the pinned revision and refuses to build if it no
longer says MIT. The 20 Upstage-internal pages are covered only by that declaration.
Truth. Each page becomes one kind: transcript document, following what upstream scores:
| Upstream category | → truth | Why |
|---|---|---|
Heading1 |
# heading |
|
Table |
pipe table from content.html (merged cells repeated, tag merged-cells) |
upstream scores tables with TEDS; table_score / teds_grid do the same job |
Equation |
LaTeX inside $$ … $$ (tags has-equation, has-formula) |
|
Figure, Chart |
dropped | upstream's NID ignores figure, table, chart (--ignore-classes-for-layout) |
Paragraph, Caption, List, Footnote, Index, Header, Footer |
content.text verbatim |
NID scores every one of them, so headers and footers stay — unlike OmniDocBench, whose own matching drops them |
No rules are emitted: a document has one kind, and the table truth inside the transcript is
already scored by table_score and teds_grid.
Selection. category is the page's dominant layout feature, in priority order table >
equation > chart > figure > index > list > text. Upstream has 42 / 18 / 43 / 38 / 10 /
9 / 40 pages in those buckets; the committed subset takes 10 / 5 / 6 / 5 / 3 / 4 / 7 (40 pages,
4.5 MB of PDFs, 4 with merged-cell tables). Pages over 600 KB or past the 8 MB budget are skipped
and counted. All 200 upstream pages convert (none skipped). Every element category present is a
tag (has-table, has-header, has-chart …). The reference has no per-page source (Library of
Congress / OER / Upstage) or domain, so neither can be tagged. Counts:
benchmark/datasets/dpbench/conversion-stats.json.
Not comparable to published DP-Bench numbers: upstream NID concatenates element text with
newlines removed (not replaced by spaces) and compares with rapidfuzz.fuzz.ratio, on text
only; PuffinParse scores a markdown transcript with its own normalisation, tables included.
Upstream also discards predicted text that lies inside a ground-truth figure or chart region
(--filter-by-gt-area). A transcript has no regions, so chart labels a parser transcribes count
as extra text here: tesseract/default scores 65 on chart pages and 99 on text pages.
Self-checks#
crates/puffinparse-core/tests/benchmark_datasets.rs keeps the adapters honest:
olmocr_rules_are_satisfiablebuilds a witness prediction from each document's own assertions: everypresenttext, everyorderpair, and eachtable_cellas a one-row table. Every rule, including theabsentones, must pass, so an adapter cannot emit a rule the scorer can never satisfy. It also requires an empty parse to fail every document that is notabsent-only.omnidocbench_truth_scores_itself_perfectlyscores each locally built truth against itself and checks that thehas-tabletag agrees withtable_score.dpbench_truth_scores_itself_perfectlyscores each committed truth against itself (1.0,table_scoreandteds_grid1.0 onhas-tablepages, tag agreement) and requires an empty parse to score 0.combined_v2_references_resolve/combined_v3_references_resolvecheck every committed../path ofcombined-v2/combined-v3.- Two
#[ignore]d helpers: PUFFINPARSE_SELFCHECK_RULES_DIRruns the witness check over any directory of rule files. On the full 824-document olmOCR conversion: 3,040 rules, 0 unsatisfiable.PUFFINPARSE_SELFCHECK_PRED_DIRscores real extractions, such aspdftotextoutput.
Licensing#
| Dataset | Redistributable? | What is committed |
|---|---|---|
synthetic-v1 |
yes, CC0-1.0 | everything |
parsebench |
yes, Apache-2.0 (publisher's terms) | 40 documents + truth + rules |
olmocr |
yes, ODC-BY-1.0 (attribution; AI2 Responsible Use Guidelines) | 40 PDFs + rules |
omnidocbench |
no: no licence; "research only, not for commercial use" | manifest index only (sha256, truth_sha256); fetched at run time |
dpbench |
yes, MIT (dataset card) | 40 PDFs + truth |
| RealDoc-Bench / LongExtractBench | annotations only | would be manifest + sha256 only |
Rules for any future adapter:
- Record the SPDX id in the manifest
license, and per document when a dataset mixes them. - Carry the publisher's citation in
attributionon the manifest and on each document, so a combined manifest never loses it. - Never vendor bytes whose license does not clearly allow it — index them with
upstream_path+sha256and fetch at run time. - Keep
upstream.revisiona commit hash. A result JSON is only reproducible if the input is.
Known gaps in the Rust side#
Both gaps are closed: ManifestDoc carries kind / rules, rule files are hashed into the dataset SHA, kind: rules documents are scored with puffinparse_core::bench::score_rules, and table-only documents are headlined by table_score (summarize_with).