From 3d874e5b210c49f5027c09477ec5e379759475c3 Mon Sep 17 00:00:00 2001 From: hallelx2 Date: Thu, 17 Sep 2026 22:47:36 +0100 Subject: [PATCH] docs(bench): score pdfgrab against the field, and fix the ICDAR metric MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every number published here so far was measured against pdfplumber alone, using an aggregation that was never the competition's. score.py pooled every adjacency relation across the corpus and scored once. ICDAR 2013 averages precision/recall/F1 per document. The competition's own evaluator prints per-table figures and aggregates nothing; Namysl et al. reproduce the competition results as "per-document averages". Pooling weights a document by how many relations it happens to contain, which is a different question. The gap is not cosmetic: pdfgrab reads 0.442 per-document against 0.362 pooled, pdfplumber 0.458 against 0.370. So the citable figure was understated by ~0.08 and was answering the wrong question. compare.py now reports both and ranks on the per-document column. The three earlier evaluations keep their numbers and gain a note. A dated measurement records what was true on the day; rewriting it to match a later correction destroys the only thing that made it evidence. On the field itself, across ten systems on 125 documents: pdfgrab is fifth, at 0.442, statistically level with the pdfplumber it ports (0.458) — which is the parity claim holding rather than failing. It is roughly 10x faster than anything of comparable accuracy, 81ms against pdfplumber's 794ms and camelot lattice's 1554ms. Speed is the defensible claim; accuracy is not. No Go library beats it. coregx/gxpdf — the only other permissively licensed Go table extractor — scores 0.179, last of ten, and its cells come back as merged text blocks with per-glyph spacing unresolved. unidoc/unipdf has better output than either but is commercial-only since v5, so it cannot be benchmarked or depended on. camelot's stream flavour wins outright at 0.582, on 0.762 recall against our 0.422. That is a direct hit on the known weakness, and it is not simply "whitespace inference works" — pdfplumber's equivalent mode scores 0.248, second worst in the table. Filed separately. The harness gains a pluggable adapter registry, wall-clock and failure counts per system, and a gxpdf extractor. A library that is not installed is reported as skipped, never scored as zero — the first run of this benchmark had tabula silently returning 0.000 because JAVA_HOME was unset, which would have published a config fault as a capability measurement. Both full runs agree to every decimal place on all ten systems. --- README.md | 19 +- bench/README.md | 1 + bench/icdar2013/compare.py | 160 ++++++++++ bench/icdar2013/systems.py | 296 ++++++++++++++++++ docs/README.md | 1 + .../2026-08-02-icdar2013-table-structure.md | 8 + ...026-08-02-strategy-auto-negative-result.md | 8 + ...-08-03-hybrid-ceiling-oracle-boundaries.md | 8 + ...-field-comparison-and-metric-correction.md | 174 ++++++++++ 9 files changed, 668 insertions(+), 7 deletions(-) create mode 100644 bench/icdar2013/compare.py create mode 100644 bench/icdar2013/systems.py create mode 100644 docs/evaluations/2026-09-17-field-comparison-and-metric-correction.md diff --git a/README.md b/README.md index 9a772ca..f78a6c1 100644 --- a/README.md +++ b/README.md @@ -608,13 +608,18 @@ stdlib-only. **0.0000pt on both axes** — the golden envelope is now asserted at 0.01pt. See [the evaluation](docs/evaluations/2026-08-02-font-metrics-and-table-fidelity.md). -- `v0.5.x` — **table detection**. The ICDAR 2013 benchmark puts - end-to-end F1 at 0.362, level with pdfplumber's 0.370, and the - diagnostic is unambiguous: precision 0.865, recall 0.229. What we - extract is right; we miss three quarters of the tables, because the - `lines` strategy needs *intersecting* rulings and a horizontally-ruled - table produces none. See - [the evaluation](docs/evaluations/2026-08-02-icdar2013-table-structure.md). +- `v0.5.x` — **table detection**. Measured against ten systems on ICDAR + 2013, pdfgrab's end-to-end F1 is **0.442** (per-document, the + competition's protocol) — fifth of ten, level with pdfplumber's 0.458, + and roughly **10x faster** than anything of comparable accuracy + (81 ms/doc against pdfplumber's 794). The diagnostic is unambiguous: + precision 0.545, recall 0.422. What we extract is right; we miss most + of the tables, because the `lines` strategy needs *intersecting* + rulings and a horizontally-ruled table produces none. Given a correct + grid the same extractor reaches **0.935**, so detection is the whole + gap. camelot's `stream` flavour reaches 0.762 recall on this corpus and + is the rule-based result worth porting. See + [the field comparison](docs/evaluations/2026-09-17-field-comparison-and-metric-correction.md). - `v0.6.x` — performance pass: parser speed against pdfminer.six and pdfplumber on a representative corpus. diff --git a/bench/README.md b/bench/README.md index 20172ee..81a9dcd 100644 --- a/bench/README.md +++ b/bench/README.md @@ -13,6 +13,7 @@ Each harness downloads what it needs into a scratch directory. | Benchmark | Dataset | Measures | Report | | --- | --- | --- | --- | | [`icdar2013/`](icdar2013/) | ICDAR 2013 Table Competition (125 PDFs) | table detection + structure | [2026-08-02](../docs/evaluations/2026-08-02-icdar2013-table-structure.md) | +| [`icdar2013/compare.py`](icdar2013/compare.py) | same, 10 systems | pdfgrab vs the whole field | [2026-09-17](../docs/evaluations/2026-09-17-field-comparison-and-metric-correction.md) | | [`icdar2013/oracle.py`](icdar2013/oracle.py) | same, with ground-truth boundaries | the ceiling a layout model could reach | [2026-08-03](../docs/evaluations/2026-08-03-hybrid-ceiling-oracle-boundaries.md) | ## Why the numbers live in `docs/evaluations/` diff --git a/bench/icdar2013/compare.py b/bench/icdar2013/compare.py new file mode 100644 index 0000000..9736183 --- /dev/null +++ b/bench/icdar2013/compare.py @@ -0,0 +1,160 @@ +"""Score every available table-extraction system on ICDAR 2013. + + python bench/icdar2013/compare.py [--limit N] + +Where score.py answers "is pdfgrab as good as pdfplumber", this answers +the broader question: how does it stand against the field. Same corpus, +same metric, same process — the only thing that varies is the extractor. + +Reports accuracy AND wall-clock. A benchmark that reports only F1 hides +the trade a pipeline actually has to make. +""" + +from __future__ import annotations + +import argparse +import json +import os +import sys +from collections import Counter + +HERE = os.path.dirname(os.path.abspath(__file__)) +sys.path.insert(0, HERE) + +from score import gt_relations, norm, prf, relations_from_grid, score # noqa: E402 +from systems import Timing, build_adapters, timed # noqa: E402 + + +def relations(tables) -> Counter: + rels: Counter = Counter() + for grid in tables: + rels += relations_from_grid([[norm(c) for c in row] for row in grid]) + return rels + + +def find_pairs(root: str, limit: int = 0) -> list[tuple[str, str]]: + pairs = [] + for dirpath, _, files in os.walk(root): + for f in sorted(files): + if not f.endswith("-str.xml"): + continue + pdf = os.path.join(dirpath, f.replace("-str.xml", ".pdf")) + if os.path.exists(pdf): + pairs.append((pdf, os.path.join(dirpath, f))) + pairs.sort() + return pairs[:limit] if limit else pairs + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("root", help="corpus root") + ap.add_argument("exe", help="built pdfgrab extractor") + ap.add_argument("--limit", type=int, default=0) + ap.add_argument("--gxpdf", default="", help="built coregx/gxpdf extractor") + ap.add_argument("--json", default="", help="also write results here") + args = ap.parse_args() + + pairs = find_pairs(args.root, args.limit) + adapters = build_adapters(args.exe, args.gxpdf) + + active, skipped = [], [] + for a in adapters: + ok, why = a.available() + (active if ok else skipped).append((a, why)) + + print(f"corpus : {args.root}") + print(f"documents : {len(pairs)}") + print(f"systems : {len(active)} active, {len(skipped)} skipped\n") + + if skipped: + print("skipped (library not importable):") + for a, why in skipped: + print(f" {a.name:<26} pip install {a.install}") + print(f" {'':<26} {why}") + print() + + # Two aggregations, because they are different numbers and only one of + # them is the competition's. + # + # micro — pool every relation across the corpus, then score once. + # Weights a document by how many relations it has. + # macro — score each document, then average the per-document F1s. + # A one-table document counts as much as a five-table one. + # + # ICDAR 2013 specifies MACRO ("per-document averages"), so that is the + # number comparable to published results. Micro is reported alongside + # because it is the more natural read of "how many relations did we get + # right", and quoting one while the reader assumes the other is exactly + # how benchmark numbers get misused. + totals = {a.name: [0, 0, 0] for a, _ in active} + per_doc = {a.name: [] for a, _ in active} + timings = {a.name: Timing() for a, _ in active} + + for i, (pdf, xml) in enumerate(pairs, 1): + gt = gt_relations(xml) + for a, _ in active: + got = relations(timed(a.extract, pdf, timings[a.name])) + c, nd, ng = score(gt, got) + totals[a.name][0] += c + totals[a.name][1] += nd + totals[a.name][2] += ng + per_doc[a.name].append(prf(c, nd, ng)) + if i % 10 == 0: + print(f" ...{i}/{len(pairs)}", flush=True) + + rows = [] + for a, _ in active: + mp, mr, mf = prf(*totals[a.name]) + docs = per_doc[a.name] + n = len(docs) or 1 + Mp = sum(d[0] for d in docs) / n + Mr = sum(d[1] for d in docs) / n + Mf = sum(d[2] for d in docs) / n + t = timings[a.name] + rows.append({ + "system": a.name, + "version": a.version(), + "macro_precision": round(Mp, 3), + "macro_recall": round(Mr, 3), + "macro_f1": round(Mf, 3), + "micro_precision": round(mp, 3), + "micro_recall": round(mr, 3), + "micro_f1": round(mf, 3), + "mean_ms": round(t.mean_ms(), 1), + "p95_ms": round(t.p95_ms(), 1), + "failures": t.failures, + "note": a.note, + }) + rows.sort(key=lambda d: d["macro_f1"], reverse=True) + + w = max(len(r["system"]) for r in rows) + 2 + print(f"\n{'':<{w}} {'--- per-document (ICDAR) ---':^25} {'-- pooled --':^17}") + print(f"{'system':<{w}} {'prec':>7} {'recall':>7} {'F1':>8} " + f"{'prec':>7} {'F1':>8} {'ms/doc':>9} {'p95 ms':>9} {'fails':>6}") + print("-" * (w + 60)) + for r in rows: + print(f"{r['system']:<{w}} {r['macro_precision']:>7.3f} " + f"{r['macro_recall']:>7.3f} {r['macro_f1']:>8.3f} " + f"{r['micro_precision']:>7.3f} {r['micro_f1']:>8.3f} " + f"{r['mean_ms']:>9.1f} {r['p95_ms']:>9.1f} {r['failures']:>6d}") + + gtn = next(iter(totals.values()))[2] if totals else 0 + print(f"\nground-truth relations: {gtn}") + print("\nMetric: adjacency relations (Goebel et al.), END-TO-END — find the") + print("table AND grid it. NOT comparable to published structure-only scores,") + print("which are handed the table region.") + print() + print("Ranked on PER-DOCUMENT F1, which is the ICDAR 2013 protocol and the") + print("column to cite. Pooled F1 is shown too because it answers a different") + print("question (how many relations were right overall) and the two diverge") + print("whenever documents differ in size.") + + if args.json: + with open(args.json, "w") as fh: + json.dump({"documents": len(pairs), "results": rows}, fh, indent=2) + print(f"\nwrote {args.json}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/bench/icdar2013/systems.py b/bench/icdar2013/systems.py new file mode 100644 index 0000000..d6605de --- /dev/null +++ b/bench/icdar2013/systems.py @@ -0,0 +1,296 @@ +"""Table-extraction systems under comparison, as pluggable adapters. + +Every adapter answers the same question — "what tables are on this page, +as grids of cell text" — and the scorer turns those grids into adjacency +relations. That keeps the comparison honest: each library is scored on +the structure it reports, not on how it chose to serialise it. + +An adapter that cannot import its library is *unavailable*, not broken. +The suite reports which systems ran and which were skipped, so a partial +environment produces a partial table rather than a crash or, worse, a +silently missing row that reads as a zero. + +Contract for an extract function: + + extract(pdf_path: str) -> list[list[list[str]]] + +i.e. a list of tables, each a list of rows, each a list of cell strings. +Cells may be None or empty; the scorer normalises whitespace. + +Adding a system: write the function, add one Adapter to ADAPTERS. Do not +reach into the scorer. +""" + +from __future__ import annotations + +import importlib +import os +import subprocess +import time +from dataclasses import dataclass, field +from typing import Callable + +Grid = list[list[str]] +Tables = list[Grid] + + +@dataclass +class Adapter: + """One system under test.""" + + name: str + extract: Callable[[str], Tables] + + # Import name checked for availability, plus the pip target that + # provides it, so a skip message can tell the reader how to fix it. + module: str | None = None + install: str = "" + + # Free-text note carried into the report — a caveat that would + # otherwise be lost between running the benchmark and reading it. + note: str = "" + + # Was this system trained on data drawn from the same distribution as + # the test corpus? Every deep-learning table model is trained on + # PubTables-1M / FinTabNet / PubTabNet, which overlap this benchmark's + # document population; rule-based extractors are trained on nothing. + # Comparing the two without saying so flatters the learned systems. + trained_on_distribution: bool = False + + # Was the table region handed to the system? Everything here is + # end-to-end (it must find the table itself), but the flag exists so an + # oracle-boundary row can sit in the same table without being mistaken + # for a comparable one. A missed table donates ALL of its relations to + # false negatives, so this single bit moves F1 by 2-3x. + region_given: bool = False + + def available(self) -> tuple[bool, str]: + if self.module is None: + return True, "" + try: + importlib.import_module(self.module) + return True, "" + except Exception as e: + return False, f"{type(e).__name__}: {e}" + + def version(self) -> str: + if self.module is None: + return "" + try: + m = importlib.import_module(self.module) + return str(getattr(m, "__version__", "") or getattr(m, "VERSION", "") or "") + except Exception: + return "" + + +@dataclass +class Timing: + """Wall-clock per system, accumulated across documents. + + Worth measuring alongside accuracy. A library that is 3% more accurate + and 40x slower is not obviously the better choice for a pipeline that + ingests documents on a request path, and a benchmark that reports only + F1 hides that trade entirely. + """ + + seconds: float = 0.0 + docs: int = 0 + failures: int = 0 + per_doc: list[float] = field(default_factory=list) + + def add(self, dt: float) -> None: + self.seconds += dt + self.docs += 1 + self.per_doc.append(dt) + + def mean_ms(self) -> float: + return 1000.0 * self.seconds / self.docs if self.docs else 0.0 + + def p95_ms(self) -> float: + if not self.per_doc: + return 0.0 + ordered = sorted(self.per_doc) + idx = min(len(ordered) - 1, int(0.95 * len(ordered))) + return 1000.0 * ordered[idx] + + +def timed(fn: Callable[[str], Tables], pdf: str, t: Timing) -> Tables: + """Run fn, recording wall-clock and swallowing per-document failures. + + A library that throws on one malformed PDF should score zero for that + document, not abort the whole run — but the failure is counted and + reported, because "extracted nothing" and "crashed" are different + facts about a library. + """ + start = time.perf_counter() + try: + out = fn(pdf) + except Exception: + t.failures += 1 + out = [] + t.add(time.perf_counter() - start) + return out + + +# --- adapters --------------------------------------------------------- + + +def pdfgrab(exe: str, strategy: str, merge: bool = False) -> Callable[[str], Tables]: + """pdfgrab, via the benchmark's Go extractor binary.""" + + def run(pdf: str) -> Tables: + import json + + cmd = [exe, "-strategy", strategy] + if merge: + cmd.append("-merge") + cmd.append(pdf) + out = subprocess.run(cmd, capture_output=True, timeout=120).stdout + return [t["rows"] for t in json.loads(out or b"[]")] + + return run + + +def gxpdf_tables(exe: str) -> Callable[[str], Tables]: + """coregx/gxpdf, via a sibling Go extractor binary. + + The one direct competitor pdfgrab has inside Go: MIT, pure Go (no CGo), + and the only permissively-licensed Go library that claims table + extraction. Its own docs claim "100% accuracy on bank statements", + which is a narrow enough claim to be worth testing on a general corpus. + + Built separately rather than linked, so a panic or a hang in a + third-party library cannot take the harness down with it. + """ + + def run(pdf: str) -> Tables: + import json + + out = subprocess.run([exe, pdf], capture_output=True, timeout=120).stdout + return [t["rows"] for t in json.loads(out or b"[]")] + + return run + + +def pdfplumber_tables(strategy: str) -> Callable[[str], Tables]: + def run(pdf: str) -> Tables: + import pdfplumber + + settings = {"vertical_strategy": strategy, "horizontal_strategy": strategy} + tables: Tables = [] + with pdfplumber.open(pdf) as doc: + for page in doc.pages: + tables.extend(page.extract_tables(settings)) + return tables + + return run + + +def pymupdf_tables(pdf: str) -> Tables: + """PyMuPDF's find_tables, added in 1.23. + + Its own strategy selection is internal, so there is no per-axis knob + to match against the others — it is scored as the library ships. + """ + try: + import pymupdf + except ImportError: # <1.24 shipped only the legacy name + import fitz as pymupdf + + tables: Tables = [] + with pymupdf.open(pdf) as doc: + for page in doc: + for tbl in page.find_tables().tables: + tables.append(tbl.extract()) + return tables + + +def camelot_tables(flavor: str) -> Callable[[str], Tables]: + """Camelot. 'lattice' needs ruled cells; 'stream' infers from whitespace. + + Camelot reads a page range rather than a document, so 'all' is passed + explicitly — its default is page 1 only, which would quietly score a + multi-page document on its first page and look like a recall problem. + """ + + def run(pdf: str) -> Tables: + import camelot + + tables: Tables = [] + for t in camelot.read_pdf(pdf, pages="all", flavor=flavor, suppress_stdout=True): + tables.append([[str(c) for c in row] for row in t.df.values.tolist()]) + return tables + + return run + + +def tabula_tables(lattice: bool) -> Callable[[str], Tables]: + """tabula-py, the Java tabula wrapper. Needs a JVM on PATH.""" + + def run(pdf: str) -> Tables: + import tabula + + dfs = tabula.read_pdf( + pdf, pages="all", lattice=lattice, stream=not lattice, + multiple_tables=True, silent=True, + ) + tables: Tables = [] + for df in dfs: + # tabula promotes the first row to a header; put it back, or + # every table silently loses its header row and with it the + # vertical relations that row participates in. + header = [str(c) for c in df.columns.tolist()] + rows = [[str(c) for c in row] for row in df.values.tolist()] + if any(not h.startswith("Unnamed") for h in header): + rows.insert(0, header) + tables.append(rows) + return tables + + return run + + +def build_adapters(exe: str, gx_exe: str = "") -> list[Adapter]: + """The comparison set. + + `exe` is the built pdfgrab extractor; `gx_exe` the gxpdf one, which is + optional because it is a third-party Go module the harness should not + require. + """ + adapters = [ + Adapter("pdfgrab (lines)", pdfgrab(exe, "lines"), + note="the library under test"), + Adapter("pdfgrab (auto)", pdfgrab(exe, "auto"), + note="one-axis-ruled support, opt-in"), + Adapter("pdfplumber (lines)", pdfplumber_tables("lines"), + module="pdfplumber", install="pdfplumber", + note="the implementation pdfgrab is a port of"), + Adapter("pdfplumber (text)", pdfplumber_tables("text"), + module="pdfplumber", install="pdfplumber", + note="whitespace-inferred; high recall, low precision"), + Adapter("PyMuPDF find_tables", pymupdf_tables, + module="pymupdf", install="pymupdf"), + # camelot 2.0 dropped Ghostscript for pdfium and the [base] extra + # no longer exists — pip warns and installs the bare package. + Adapter("camelot (lattice)", camelot_tables("lattice"), + module="camelot", install="camelot-py", + note="ruled cells; pdfium backend since 1.0"), + Adapter("camelot (stream)", camelot_tables("stream"), + module="camelot", install="camelot-py", + note="whitespace-inferred"), + Adapter("tabula (lattice)", tabula_tables(True), + module="tabula", install="tabula-py", + note="DORMANT: last release 2024-10; JVM required"), + Adapter("tabula (stream)", tabula_tables(False), + module="tabula", install="tabula-py", + note="DORMANT: last release 2024-10; JVM required"), + ] + + if gx_exe and os.path.exists(gx_exe): + # Inserted right after pdfgrab: Go-vs-Go is the comparison that + # decides whether pdfgrab is worth maintaining at all, so it should + # sit next to it in the output rather than at the bottom. + adapters.insert(2, Adapter( + "gxpdf (Go)", gxpdf_tables(gx_exe), + note="the only other permissively-licensed Go table extractor")) + + return adapters diff --git a/docs/README.md b/docs/README.md index 7a5d080..81b16fb 100644 --- a/docs/README.md +++ b/docs/README.md @@ -20,6 +20,7 @@ release history in [`CHANGELOG.md`](../CHANGELOG.md). | date | subject | headline | | --- | --- | --- | +| [2026-09-17](evaluations/2026-09-17-field-comparison-and-metric-correction.md) | the field, and a metric correction | **per-doc F1 0.442**, 5th of 10, ~10x faster than anything comparable. No Go library comes close. Earlier figures were pooled, not the competition's metric | | [2026-08-03](evaluations/2026-08-03-hybrid-ceiling-oracle-boundaries.md) | hybrid ceiling with oracle boundaries | **0.362 → 0.935.** Given a correct grid, extraction is near-perfect — structure is the whole gap | | [2026-08-02](evaluations/2026-08-02-icdar2013-table-structure.md) | ICDAR 2013 table detection + structure | F1 0.362 end-to-end; the bottleneck is **detection**, not cell accuracy | | [2026-08-02](evaluations/2026-08-02-strategy-auto-negative-result.md) | `StrategyAuto` for one-axis-ruled tables | **negative result** — detection 22%→18% missed, but F1 0.362→0.358. Shipped opt-in only | diff --git a/docs/evaluations/2026-08-02-icdar2013-table-structure.md b/docs/evaluations/2026-08-02-icdar2013-table-structure.md index 0a88220..36fea1f 100644 --- a/docs/evaluations/2026-08-02-icdar2013-table-structure.md +++ b/docs/evaluations/2026-08-02-icdar2013-table-structure.md @@ -6,6 +6,14 @@ **Dataset:** ICDAR 2013 Table Competition, Smock's corrected edition — 125 born-digital PDFs, 39,524 ground-truth adjacency relations **Reference:** pdfplumber 0.11.9 +> **Metric note, added 2026-09-17.** The F1 figures below are **pooled** — +> every adjacency relation across the corpus scored in one batch. The ICDAR +> 2013 protocol averages **per document**, which puts pdfgrab at **0.442** and +> pdfplumber at **0.458** on the same data. These numbers are left as they were +> measured; see +> [2026-09-17](2026-09-17-field-comparison-and-metric-correction.md) for the +> corrected metric and a comparison against the full field. + ## Result | system | precision | recall | F1 | diff --git a/docs/evaluations/2026-08-02-strategy-auto-negative-result.md b/docs/evaluations/2026-08-02-strategy-auto-negative-result.md index 662b23b..8bdd3dd 100644 --- a/docs/evaluations/2026-08-02-strategy-auto-negative-result.md +++ b/docs/evaluations/2026-08-02-strategy-auto-negative-result.md @@ -6,6 +6,14 @@ **Verdict:** shipped as **opt-in**; does **not** become a default. The hypothesis it tested is disproved. +> **Metric note, added 2026-09-17.** The F1 figures below are **pooled** — +> every adjacency relation across the corpus scored in one batch. The ICDAR +> 2013 protocol averages **per document**, which puts pdfgrab at **0.442** and +> pdfplumber at **0.458** on the same data. These numbers are left as they were +> measured; see +> [2026-09-17](2026-09-17-field-comparison-and-metric-correction.md) for the +> corrected metric and a comparison against the full field. + ## Hypothesis The [ICDAR 2013 evaluation](2026-08-02-icdar2013-table-structure.md) found diff --git a/docs/evaluations/2026-08-03-hybrid-ceiling-oracle-boundaries.md b/docs/evaluations/2026-08-03-hybrid-ceiling-oracle-boundaries.md index 042f6eb..85e0771 100644 --- a/docs/evaluations/2026-08-03-hybrid-ceiling-oracle-boundaries.md +++ b/docs/evaluations/2026-08-03-hybrid-ceiling-oracle-boundaries.md @@ -5,6 +5,14 @@ **Harness:** [`bench/icdar2013/oracle.py`](../../bench/icdar2013/oracle.py) **Question:** if a layout model supplied correct rows and columns, how good would extraction be? That number decides whether the model is worth deploying. +> **Metric note, added 2026-09-17.** The F1 figures below are **pooled** — +> every adjacency relation across the corpus scored in one batch. The ICDAR +> 2013 protocol averages **per document**, which puts pdfgrab at **0.442** and +> pdfplumber at **0.458** on the same data. These numbers are left as they were +> measured; see +> [2026-09-17](2026-09-17-field-comparison-and-metric-correction.md) for the +> corrected metric and a comparison against the full field. + ## Result | system | precision | recall | F1 | diff --git a/docs/evaluations/2026-09-17-field-comparison-and-metric-correction.md b/docs/evaluations/2026-09-17-field-comparison-and-metric-correction.md new file mode 100644 index 0000000..750532b --- /dev/null +++ b/docs/evaluations/2026-09-17-field-comparison-and-metric-correction.md @@ -0,0 +1,174 @@ +# The field, and a metric we had been computing wrong + +**Date:** 2026-09-17 +**Harness:** [`bench/icdar2013/compare.py`](../../bench/icdar2013/compare.py) · [`systems.py`](../../bench/icdar2013/systems.py) +**Corpus:** ICDAR 2013 Table Competition, Smock-corrected edition — 125 PDFs, 39,524 ground-truth adjacency relations +**Questions:** where does pdfgrab sit against the whole field rather than against pdfplumber alone, and is any Go library better? + +## Correction first: our number was not the competition's metric + +Every pdfgrab figure published before today — the 0.362 that appears throughout +this repo — was computed by **pooling every adjacency relation across the whole +corpus and scoring once**. That is micro-averaging. + +ICDAR 2013 does not do that. It computes precision/recall/F1 **per document and +averages over documents**. Evidence, in order of authority: + +- The competition's own evaluator, `tamirhassan/dataset-tools` (Apache-2.0), + prints per-table precision/recall and performs **no aggregation at all** — it + emits the raw counts and leaves pooling to the caller. +- Namysl et al. (VISAPP 2022), reproducing the competition results, state the + aggregation explicitly: "*We report the precision, recall, and F1 score + (per-document averages) for the complete recognition process.*" + +The two differ by a meaningful margin, because a document with one large table +and a document with six small ones count equally under one scheme and very +unequally under the other: + +| | pooled *(what we published)* | per-document *(the competition's)* | +|---|---|---| +| pdfgrab (`lines`) | 0.362 | **0.442** | +| pdfplumber (`lines`) | 0.370 | **0.458** | + +**The citable end-to-end figure for pdfgrab is 0.442, not 0.362.** The harness +now reports both and ranks on the per-document column. + +The earlier evaluations are left as written. They record what was measured on +the day with the metric as it was then implemented; a note now points here. +Rewriting a dated measurement is worse than annotating it. + +## Results + +Ten systems, same corpus, same metric, same process. End-to-end: each system +must **find** the table and **grid** it. + +| System | Version | per-doc P | per-doc R | **per-doc F1** | pooled F1 | ms/doc | p95 ms | +|---|---|---|---|---|---|---|---| +| camelot (stream) | 2.0.0 | 0.514 | 0.762 | **0.582** | 0.716 | 300 | 899 | +| PyMuPDF `find_tables` | 1.28.2 | 0.560 | 0.472 | **0.485** | 0.392 | 730 | 2116 | +| camelot (lattice) | 2.0.0 | 0.520 | 0.453 | **0.467** | 0.396 | 1554 | 3237 | +| pdfplumber (lines) | 0.11.10 | 0.558 | 0.440 | **0.458** | 0.370 | 794 | 2066 | +| **pdfgrab (auto)** | — | 0.548 | 0.422 | **0.443** | 0.358 | **86** | 274 | +| **pdfgrab (lines)** | — | 0.545 | 0.422 | **0.442** | 0.362 | **81** | 247 | +| tabula (stream) | 2.10.0 | 0.387 | 0.460 | **0.397** | 0.437 | 1100 | 2680 | +| tabula (lattice) | 2.10.0 | 0.258 | 0.339 | **0.257** | 0.107 | 164 | 415 | +| pdfplumber (text) | 0.11.10 | 0.188 | 0.559 | **0.248** | 0.267 | 1456 | 4043 | +| **gxpdf (Go)** | v0.9.4 | 0.201 | 0.176 | **0.179** | 0.293 | 73 | 213 | + +No system recorded a failure on any document. + +## What it says + +### No Go library beats pdfgrab + +`coregx/gxpdf` (MIT, pure Go, v0.9.4, 2026-08-02) is the only other +permissively-licensed Go library that extracts tables. It scores **0.179** — +last of ten, 2.5x behind pdfgrab. + +The raw output shows why. On `eu-001.pdf` it returns cells like: + +``` +"N i t r og en ox id es (N Ox/N O2 ) 100 00 0 - - \nHy dro g en Cy anide..." +``` + +Whole text blocks merged into one cell, with the per-glyph spacing unresolved — +the class of error the AFM metrics work fixed here. Its README claims "100% +accuracy on bank statements", which is plausibly true and narrow: ruled +financial tables are the case `lattice`-style detection handles well. + +For completeness, the Go field: `unidoc/unipdf` v5 has the best output of any +Go library — `TextTable` with per-cell bbox *and* grid indices — but is +**commercial-licence-only** since v5 (the AGPL option was removed), so it cannot +be benchmarked without a key and cannot be depended on by an MIT project. +`klippa-app/go-pdfium` gives per-character boxes and, notably, runs +**CGo-free in WebAssembly mode** — but has no table layer. `pdfcpu` has no text +extraction at all. `ledongthuc/pdf`, `rsc.io/pdf` and `dslipak/pdf` give text +positions but no tables. + +### Against Python, pdfgrab is mid-pack — and that is the honest claim + +Fifth of ten. The gap to pdfplumber is 0.016, which is the port working as +intended: pdfgrab reproduces its ancestor's behaviour, including its ceiling. + +What pdfgrab wins is **throughput**: 81 ms/doc against pdfplumber's 794 and +camelot lattice's 1554 — roughly **10x faster than anything of comparable +accuracy**, and 4x faster than the system that beat it. Combined with a single +static binary, no Python runtime and no model download, that is a real and +defensible position. "More accurate than Python" is not. + +### camelot's `stream` is the result worth studying + +It wins outright, and it wins on **recall: 0.762 against pdfgrab's 0.422**. +That is a direct attack on the known weakness — `lines` requires intersecting +rulings, so booktabs-style and horizontally-ruled-only tables are invisible to +it, and 22% of documents yield no table at all. + +The instructive part is that this is *not* simply "whitespace inference beats +ruling detection": pdfplumber's equivalent whitespace mode (`text`) scores +**0.248**, the second-worst result in the table. Same idea, very different +implementation. camelot 2.0 rebuilt its backend on `playa-pdf` this year and its +stream flavour is doing something materially better than the algorithm pdfgrab +inherited. + +That makes it the highest-value thing to read next, and it is **rule-based** — +so any gain is portable to pure Go with no model, no network and no new +dependency. + +## Determinism + +The whole table was produced twice, from scratch, in two independent processes, +and the per-document F1 of all ten systems agreed to **every decimal place** — +delta 0.0000 on every row. + +That is worth stating rather than assuming. Every system here is rule-based and +runs at a fixed configuration, so there is no sampling to average out, but +"should be deterministic" and "was deterministic" are different claims and only +one of them is a measurement. A number nobody can reproduce is not a result. + +## Caveats that must travel with these numbers + +**End-to-end, not structure-only.** Published ICDAR 2013 figures of 0.85–0.95 +are structure-only: the system is handed the table region. Our own oracle +experiment measures that regime at **0.935**. The two are not comparable and +must never appear in the same column. For reference, the end-to-end +(`GT Border = N`) column of the competition reproduction runs +KYTHE 0.522 · pdf2table 0.585 · TABFIND 0.696 · Nurminen 0.837 · +FineReader 0.877. + +**Detection failure is scored without mercy.** A table missed at IoU < 0.5 +contributes *every one of its ground-truth relations* as a false negative, and a +spurious detection contributes every predicted relation as a false positive. +There is no partial credit for a nearly-right region. This is why a 0.935 +structure score and a 0.442 end-to-end score are consistent rather than +contradictory. + +**Corpus identity.** This is the Smock-corrected **125-PDF superset** +(competition set + the 2012 practice data), not the 67-PDF competition test set +that every historical number above was measured on. The corrected ground truth +is also a slightly more forgiving target — the same TATR checkpoint gains +~1.3 DAR points from the corrections alone. + +**Adjacency relations only sees non-blank cells,** and compares them by exact +string match after whitespace normalisation. It is blind to empty-cell +misalignment, and how an implementation decides a cell is "non-empty" moves the +score without any change in structure quality. + +**tabula needs `JAVA_HOME`.** Without it, jpype cannot find `libjvm.so`, every +call throws, and the system scores a clean 0.000 that looks like a measurement. +It is not one. The first run of this benchmark hit exactly that and was +discarded. + +## Reproduce + +```sh +pip install pdfplumber pymupdf camelot-py tabula-py jpype1 +export JAVA_HOME=... # or tabula silently scores zero +python bench/icdar2013/run.py # fetches the corpus, builds the extractor +python bench/icdar2013/compare.py \ + ~/.cache/pdfgrab-bench/ICDAR-2013-Table-Competition-Corrected \ + ~/.cache/pdfgrab-bench/bench-extract \ + --gxpdf ~/.cache/pdfgrab-bench/gx-extract +``` + +Adding a system is one `Adapter` in `systems.py`; a library that is not +installed is reported as skipped rather than scored as zero.