PuffinParse docs
/ llms.txt GitHub

Academic and community OCR / document-parsing benchmarks: a survey for PuffinParse#

Researched 2026-09-24. Companion to the vendor benchmark review, which covers ParseBench, RealDoc-Bench and LongExtractBench. Every licence below was read from the dataset card on Hugging Face (https://huggingface.co/api/datasets/<id> plus the README at the pinned revision) or from the repository's LICENSE file. Revisions are the commits that were current when this was written. olmOCR-bench and OmniDocBench now have adapters (see adapters.md). For the others this page recommends what to do next.

PuffinParse scores two document kinds (SPEC §10): transcript (reference markdown, scored with char/word edit distance, reading order and table_score) and rules (machine-checkable present / absent / order / table_cell / bag_of_sentences assertions). "Maps to" below means into one of those two kinds.

Summary#

Benchmark Measures Size Licence (verified) Redistributable in PuffinParse (MIT)? GT format Maps to Recommendation
olmOCR-bench (AI2) PDF → markdown unit tests 1,403 PDFs, 7,010 tests ODC-BY-1.0 (card) Yes, with attribution JSONL unit tests rules: 3,040 of 7,019 tests (43%) Adapted: 40-doc subset committed
OmniDocBench (OpenDataLab) end-to-end page parsing: text, tables, formulas, reading order 1,651 pages, 10 doc types, EN + ZH none on the card; "research purposes only and not for commercial use" No, index only one JSON: blocks + order + text/LaTeX/HTML transcript (1,443 of 1,651 pages) Adapted: 40-page index, fetched at run time
DP-Bench (Upstage) element serialisation (NID), table structure (TEDS / TEDS-S) 200 single-page PDFs, 36 MB MIT (card) Yes (see the caveat below) reference.json: elements with category, coordinates, text/html/markdown transcript Adapted (benchmark/datasets/dpbench, in combined-v3)
READoc (ISCAS) realistic PDF → markdown on whole documents 2,233 docs (arXiv + GitHub) MIT (card) Ground truth yes; PDFs doubtful one markdown file per document transcript, multi-page Adapter for a long-document track; fetch PDFs at run time
Nanonets IDP leaderboard aggregator: olmOCR-bench, OmniDocBench, "IDP Core" (KIE, VQA, OCR, tables, classification) 6,406 IDP Core samples plus the two above harness MIT; datasets mixed Per dataset per-dataset — No adapter of its own; its two public page benchmarks are covered above
Fox (UCAS) fine-grained, region/line/colour-focused page OCR, EN + ZH 1 zip (focus_benchmark_test.zip) CC-BY-NC-SA-4.0 (card) No (non-commercial, share-alike) prompt → text pairs partly transcript Skip: focus-prompted task, NC licence
CC-OCR (Alibaba / Qwen) LMM literacy: scene text, multilingual, doc parsing, KIE 39 subsets, 7,058 images card says MIT; README says only "the source code is licensed under MIT" Unclear for the images TSV with base64 images doc_parsing track → transcript Maybe later, fetched at run time
OCRBench v2 LMM OCR QA across 31 scenarios 10,000 QA pairs, 1.1 GB MIT (card); images drawn from many source datasets Unclear for the images QA pairs + eval type — (QA, not page transcription) Skip

Pinned revisions:

Benchmark Data Revision Scorer / code Revision
olmOCR-bench HF allenai/olmOCR-bench 54a96a6fb6a2bd3b297e59869491db4d3625b711 github.com/allenai/olmocr (Apache-2.0), olmocr/bench/tests.py f7cfe4c22098b154c76b6ec950d1c0a464eecf8d
OmniDocBench HF opendatalab/OmniDocBench aa1ee96d106dbe53d0ae59474d75c6e6d9b53fec github.com/opendatalab/OmniDocBench (Apache-2.0) f133a71e9e91c3621c7ce8994200a7b394a06eb3
DP-Bench HF upstage/dp-bench (data and evaluate.py together) 24702c61a2fb13325534be664653bc6e60250d13 same repo same
READoc HF lazyc/READoc 782bf954e4d9a31ae2bcfe14ff59f1f4b2467592 github.com/icip-cas/READoc not pinned (not adapted)
Fox HF ucaslcl/Fox_benchmark_data d6c6f5202d61bf28c8cf9b7b777754dc9bbdc0e0 github.com/ucaslcl/Fox not pinned
CC-OCR HF wulipc/CC-OCR c64517e92179991d509776064174776700cdd5a2 github.com/AlibabaResearch/AdvancedLiterateMachinery (Benchmarks/CC-OCR) not pinned
OCRBench v2 HF ling99/OCRBench_v2 c7e7cdf23bdb6774661e9b0caf0d9935a42feb8b github.com/Yuliang-Liu/MultimodalOCR not pinned
IDP leaderboard — — github.com/NanoNets/idp-leaderboard-benchmarks (MIT) not pinned

1. olmOCR-bench (Allen Institute for AI): adapted#

  • Paper: arXiv:2502.18443 ("olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models"). Code: allenai/olmocr (Apache-2.0). Data: allenai/olmOCR-bench.

What it measures. Whether a PDF → markdown converter gets specific, checkable facts right. It does not compare the output against a reference transcript. Each test is a unit test: this sentence is present, this running header is absent, this paragraph precedes that one, this table cell sits under that heading, this equation renders the same as the reference.

Size. 1,403 single-page PDFs and 7,010 tests in seven JSONL files, one per source. The runner also adds one implicit baseline test per PDF. There are 9 explicit baseline tests, so the files hold 7,019 rows.

Split PDFs Tests Test types
arxiv_math 522 2,927 math
old_scans_math 36 458 math
table_tests 188 1,020 (+2 baseline) table
old_scans 98 526 present 279, absent 70, order 177
headers_footers 266 753 (+7 baseline) absent
multi_column 231 884 order
long_tiny_text 62 442 present

Licence. The card front-matter says license: odc-by, and the README says the dataset is "licensed under ODC-BY-1.0 … intended for research and educational use in accordance with AI2's Responsible Use Guidelines". ODC-BY allows redistribution with attribution. The underlying pages are third-party (arXiv, Library of Congress, Internet Archive, crawled PDFs), and every test carries the page's origin url. The adapter keeps that URL as source_url on each document.

Ground truth and scorer. One JSON object per test: {pdf, page, id, type, max_diffs, checked, url, …} plus type-specific keys: text, case_sensitive, first_n, last_n (present/absent); before, after (order); cell, up, down, left, right, top_heading, left_heading (table); math (LaTeX). Semantics, read from olmocr/bench/tests.py:

  • Both sides go through normalize_text: whitespace collapsed, **/__/*/_ emphasis stripped, NFC, and typographic quotes and dashes folded to ASCII.
  • present / absent use rapidfuzz.partial_ratio with threshold 1 - max_diffs/len(text). first_n / last_n restrict the search to the start or end of the output. present is case-sensitive by default.
  • order uses fuzzysearch.find_near_matches with max_l_dist = max_diffs and passes if any before match starts before any after match. It is case-sensitive.
  • table parses markdown and HTML tables, finds a cell by fuzz.ratio, then checks the immediate up/down/left/right neighbour or the column/row heading.
  • math renders both equations with KaTeX and compares the rendered symbol layout.
  • baseline fails on empty output, long repeated n-grams, or CJK and emoji characters.

The leaderboard averages pass rates per split, not per document.

Mapping to PuffinParse. Every test is already a rule, so documents are kind: rules:

Upstream → PuffinParse Converted Skipped Fidelity
present present 721 0 exact substring, stricter than fuzzy when max_diffs > 0
absent (no first_n/last_n) absent 622 — exact, looser than fuzzy when max_diffs > 0
absent with first_n/last_n — 0 201 a whole-page absence would fail a page number that also appears in the body
order order (case-sensitive) 1,061 0 exact; 94% of the multi_column tests have max_diffs > 0, so this is stricter than upstream
table with top_heading / left_heading / no relation table_cell with col_header / row_header 309 0 exact cell equality (upstream: fuzz.ratio ≥ max(0.5, …))
table with left / right table_cell with row_header = neighbour 327 0 relaxed: "immediately left/right of" becomes "in the same row as"
table with up / down — 0 384 the schema has no column-adjacency relation, and "value exists" would inflate scores
math — 0 3,385 KaTeX-rendered equivalence has no text-rule analogue
baseline — 0 9 (+1,394 implicit) repetition and charset heuristic
Total 3,040 3,979 824 of 1,403 PDFs keep at least one rule

max_diffs is kept on every converted rule, so a fuzzy scorer can honour it later.

Recommendation. Done. See benchmark/datasets/olmocr/README.md. Two scorer extensions would recover most of what is skipped or approximated: fuzzy matching driven by max_diffs, and up/down/left/right neighbour fields on table_cell. Math is deliberately out of scope until PuffinParse has a formula metric.

2. OmniDocBench (OpenDataLab / Shanghai AI Laboratory): adapted as an index#

  • Paper: arXiv:2412.07626. Code: opendatalab/OmniDocBench (Apache-2.0). Data: opendatalab/OmniDocBench, not gated.

What it measures. End-to-end parsing of diverse real pages: text (normalised edit distance), tables (TEDS), display formulas (CDM and edit distance) and reading order (edit distance over block order). It also has layout-detection and single-module tracks.

Size. 1,651 page images (PNG/JPG, over 1 GB; the first 1,000 alone are 1,019 MB) plus OmniDocBench.json (42 MB), which includes a 296-page hard subset added 2026-04-09. Page attributes:

  • data_source (10): book 276, PPT2PDF 253, academic_literature 215, exam_paper 193, colorful_textbook 159, newspaper 151, magazine 149, research_report 132, note 118, historical_document 5.
  • language: simplified_chinese 765, english 755, en_ch_mixed 116, traditional_chinese 13, other 2.
  • layout: single / double / three column, mixed, other.
  • special_issue: watermark, fuzzy_scan, colorful_background, table flags, and so on.
  • subset: v1.5, table_hard, layout_hard, equation_hard.

Licence. The card has no licence field (cardData is null in the API). The only terms are the Copyright Statement: "The PDFs are collected from public online channels and community user contributions. Content that is not allowed for distribution has been removed. The dataset is for research purposes only and not for commercial use." That grants no redistribution right and forbids commercial use, so PuffinParse, an MIT repository, does not vendor any of it: not the images, and not the truth derived from the annotations. The evaluation code is Apache-2.0.

Ground truth. Per page: page_info (image path, size, attributes) and layout_dets. Each block has a category_type (28 block classes), a polygon, a reading order, and text / latex / html. extra.relation holds truncated links, which join paragraphs split across columns, and parent_son links, which attach captions. Upstream's tools/json2md.py shows how to turn this into markdown.

Mapping to PuffinParse. One transcript document per page. Blocks are emitted in order, with truncated chains merged. Titles become # headings, tables become pipe tables converted from the HTML (merged cells flattened, tagged merged-cells), display formulas stay as $$…$$, and headers, footers, page numbers, page footnotes, abandon regions and figures are dropped. Those are the page furniture OmniDocBench itself does not score. Pages with *_mask regions (126) are skipped because their content is deliberately unannotated. So are pages whose truth is under 120 characters (82). That leaves 1,443 of 1,651 pages convertible. data_source is the category. Language, layout, subset, special issues and has-table / has-formula are tags.

What is lost: TEDS table structure (PuffinParse's table_score compares cell text row by row), CDM formula matching (LaTeX is compared as text), and OmniDocBench's block-level matching, which forgives reading-order differences. PuffinParse's char_similarity over the whole page does not.

Recommendation. Done as a fetch-at-run-time index. See benchmark/datasets/omnidocbench/README.md. If OpenDataLab ever publishes an explicit licence that permits redistribution, the same adapter can commit the files by deleting its .gitignore.

3. DP-Bench (Upstage): adapted#

  • Data, inference scripts and evaluate.py all live in HF upstage/dp-bench. There is no paper.

What it measures. Two things:

  • NID (normalised indel distance: edit distance without substitutions) over the text elements serialised in reading order. Tables, figures and charts are excluded.
  • TEDS / TEDS-S over the 55 tables.

Size. 200 single-page PDFs (36 MB, largest 4.3 MB): 90 from the Library of Congress, 90 from Open Educational Resources, 20 from Upstage internal documents. The documents contain 1,822 layout elements in 12 classes: Paragraph 804, Heading1 194, Footer 168, Caption 154, Header 101, List 91, Chart 67, Footnote 63, Equation 58, Figure 57, Table 55, Index 10.

Licence. The card front-matter says license: mit. The Library of Congress pages are largely public domain and OER pages are openly licensed. The 20 Upstage internal pages are covered only by the MIT declaration, so record that in the dataset README.

Ground truth. dataset/reference.json (1.5 MB) is {<pdf>: {elements: [{category, coordinates, id, page, content: {text, html, markdown}}]}}. Tables are HTML, equations LaTeX, and id is the reading order.

Mapping. A transcript per page: elements in id order, tables as pipe tables via html_table_to_markdown, and headers and footers either dropped (NID excludes nothing but tables, figures and charts, so they could stay) or tagged. This is structurally the same as the OmniDocBench adapter, so it is roughly a day's work, and the whole dataset (36 MB) is small enough that a 40–60 page subset commits in about 5 MB.

Status. Adapted by benchmark/adapters/dpbench.py: headers and footers are kept (NID scores them), figures and charts dropped, 40 pages committed (4.5 MB), folded into combined-v3. See adapters.md.

4. READoc (Institute of Software, CAS)#

  • Paper: arXiv:2409.05137. Code: icip-cas/READoc. Data: lazyc/READoc.

What it measures. Realistic document structured extraction: a whole PDF (arXiv papers, GitHub READMEs rendered to PDF) converted to one markdown file. It scores text, headings, tables, formulas and reading order after a standardisation and segmentation step (the "DSE Evaluation Suite").

Size. 2,233 documents: arxiv_ground_truth/ has 1,009 markdown files and github_ground_truth/ has 1,224. The PDFs come as arxiv.zip (1.19 GB), github.zip (0.59 GB) and zenodo.zip (1.45 GB).

Licence. The card says license: mit. The arXiv PDFs keep their authors' licences, most often arXiv's non-exclusive distribution licence, which does not let third parties redistribute. Treat the markdown truth as MIT and the PDFs as fetch-only.

Ground truth. One markdown file per document, e.g. arxiv_ground_truth/0705.4297.md (91 KB).

Mapping. transcript, multi-page, exactly PuffinParse's existing kind, but with documents of 10 to 40 pages. That makes it a latency and cost stress test as much as an accuracy one.

Recommendation. Worth an adapter as a separate long-document track (readoc-arxiv, about 30 documents). Ship the index plus truth and fetch the PDFs from the zips at run time. Selective extraction needs HTTP range reads of the zip central directory; downloading 1.2 GB for 30 documents is wasteful. Keep it out of combined-* because its per-document cost dwarfs the single-page sets.

5. Nanonets IDP leaderboard#

  • Site idp-leaderboard.org. Harness NanoNets/idp-leaderboard-benchmarks (MIT). Toolkit NanoNets/docext.

What it measures. It is an aggregator rather than a dataset. It runs olmOCR-bench (from allenai/olmOCR-bench), OmniDocBench (from opendatalab/OmniDocBench) and its own "IDP Core": key-information extraction, VQA, OCR, document classification, long-document processing, table extraction and confidence calibration. IDP Core has 6,406 samples drawn from existing datasets (Nanonets KIE, DocILE, handwritten forms, ChartQA, DocVQA, OCR handwriting and diacritics sets, table-extraction sets), loaded through docext from Hugging Face.

Licence. The harness is MIT. The IDP Core datasets keep their own terms, and several need an agreement or are research-only (DocILE, DocVQA). Not verified dataset by dataset here.

Recommendation. No adapter of its own. Its two page-parsing benchmarks are now PuffinParse adapters, so PuffinParse can report comparable per-benchmark numbers directly. The KIE and table tracks belong with a future extract-mode benchmark and would need a per-dataset licence audit.

6. Fox (UCAS / MEGVII)#

  • Data ucaslcl/Fox_benchmark_data, a single focus_benchmark_test.zip.

What it measures. Fine-grained, focus-prompted document understanding in English and Chinese: OCR of a region given a box, a line, or a colour highlight, plus page-level OCR and multi-page variants. Scoring uses edit distance, F1, BLEU and METEOR.

Licence. CC-BY-NC-SA-4.0 (card and tags). The non-commercial and share-alike terms are incompatible with vendoring into an MIT repository.

Recommendation. Skip. The task is prompt-conditioned, and PuffinParse's providers expose no "focus this box" input. Only the plain page-OCR slice would map to transcript, and the licence keeps it fetch-only.

7. CC-OCR (Alibaba / Qwen team)#

  • Paper: arXiv:2412.02210. Data wulipc/CC-OCR, the TSV version used by VLMEvalKit.

What it measures. Literacy of large multimodal models across four tracks: multi-scene text reading, multilingual text reading, document parsing (documents, tables, formulas; scored with normalised edit distance and TEDS), and key-information extraction. It has 39 subsets and 7,058 images, of which 41% come from real applications.

Licence. The card front-matter says mit, but the README's licence section says only that "the source code is licensed under the MIT License". The images' own terms are not stated.

Mapping. The doc_parsing subsets give an image and a reference markdown, HTML or LaTeX string, which maps to transcript, with tables via HTML → pipe table.

Recommendation. Possible later, as a fetch-at-run-time index like OmniDocBench, until the data licence is clarified. Low priority: the parsing track is small and mostly overlaps OmniDocBench.

8. OCRBench v2#

  • Paper: arXiv:2501.00321. Data ling99/OCRBench_v2 (parquet, 1.1 GB, 10,000 QA pairs). The first OCRBench (echo840/OCRBench) carries no licence on its card.

What it measures. LMM OCR ability as question answering over 31 scenarios and 23 tasks: text recognition, referring, spotting, key-information extraction, parsing, reasoning. Scoring is task-specific (exact match, ANLS, TEDS, IoU, and so on).

Licence. The card says license: mit. The images are collected from many public datasets that keep their own terms.

Recommendation. Skip. It is QA over images, not page transcription, and it conflicts with the "score parse output" design exactly as RealDoc-Bench's QA track does.


What changes in PuffinParse because of this survey#

  1. benchmark/adapters/olmocr.py and benchmark/adapters/omnidocbench.py exist, and combined-v2 includes both. combined-v1 is unchanged because it has committed results.
  2. Scorer work that would raise fidelity, in priority order:
  3. max_diffs fuzzy matching for present, absent, order and table_cell. The field is already in the rule files.
  4. Neighbour relations (up, down, left, right) on table_cell.
  5. HTML tables in predictions (in progress elsewhere).
  6. A structural table metric (TEDS) next to table_score.
  7. Next adapter: READoc as a long-document track (DP-Bench is done: combined-v3).