Skip to content

fix(pdf): bump pdf-inspector to 1.20.0 so RTL text extracts in logical order - #175

Open
shlomsh wants to merge 1 commit into
firecrawl:mainfrom
shlomsh:bump-pdf-inspector-1.20
Open

shlomsh wants to merge 1 commit into
firecrawl:mainfrom
shlomsh:bump-pdf-inspector-1.20

Conversation

@shlomsh

@shlomsh shlomsh commented Sep 15, 2026 •

Copy link
Copy Markdown

Closes #170.

pdf-inspector 1.14.2 emits right-to-left runs in visual order, so Hebrew, Arabic and Persian text comes out character-reversed (in my corpus, 52 of 77 Hebrew PDFs). The fix landed upstream in firecrawl/pdf-inspector#440 and #441 and shipped in pdf-inspector 1.16.0 on 2026-08-21, but @firecrawl/anydoc-wasm 0.2.4 (published after that) still locks 1.14.2, so a plain npm install reproduces the bug.

This bumps Cargo.toml to 1.20.0 (current release) and regenerates Cargo.lock with cargo update -p pdf-inspector; lopdf moves 0.42 → 0.45 as part of that, everything else is transitive.

Snapshots. One snapshot changes, pdf/text.pdf, and each line of it is an improvement:

Checked locally, matching ci.yml: cargo clippy --workspace --all-targets --all-features -- -D warnings clean, cargo test --locked green (286 + 1 + 9 tests), cargo check --workspace builds anydoc-wasm and anydoc-python on the new lock.

Happy to adjust the version pin if you'd rather take 1.16 as the floor than 1.20.0.


Summary by cubic

Bumps pdf-inspector from 1.14.2 to 1.20.0 so right-to-left text (Hebrew, Arabic, Persian) extracts in logical order instead of character-reversed, closing #170.

  • lopdf moves 0.42 → 0.45 as a transitive update.
  • The one snapshot change (pdf/text.pdf) improves each affected line: Persian now reads as written, the endnote marker becomes <sup>, and the table heading stays with its rows.
  • Published @firecrawl/anydoc-wasm 0.2.4 still pins 1.14.2, so this affects source builds only.

Written for commit 3863d52. Summary will update on new commits.

Review in cubic

…l order

pdf-inspector 1.14.2 emitted right-to-left runs in visual order, so Hebrew,
Arabic and Persian came out character-reversed. Fixed upstream in
pdf-inspector#440 and #441 (released in 1.16.0), but every published
anydoc-wasm still pinned 1.14.2. Closes firecrawl#170.

The only snapshot that moves is pdf/text.pdf, and every change is an
improvement: "Persian with ZWNJ" now reads as written, the endnote marker
becomes a proper <sup>, and the table heading is grouped with its rows.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 3 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Re-trigger cubic

@al-ashalash

Copy link
Copy Markdown

I measured the bump target on Arabic specifically: 1.20.0 is not far enough.

Setup. Windows x64, Node bindings from npm, one version per temp folder, same input for all. Sample: a three-line Arabic HTML page printed with headless Edge — nothing attached and nothing hosted; two commands reproduce it, happy to share them. Metrics against the known text: character error rate (CER), the same rate with all spaces removed, and the share of the expected words that survive intact. I also checked that processPdf — the process_pdf_mem path anydoc uses — returns the same text as extractPagesMarkdown on Arabic PDFs, and that 1.20.0's output for text.pdf matches this PR's snapshot byte for byte.

pdf-inspector extractPagesMarkdown on that page
1.14.2 (pinned here) CER 0.258 — word boundaries shifted, and one two-character glyph cluster transposed
1.16.0 – 1.21.0 CER 0.146, 0.000 with spaces removed, 32% of words intact — every letter correct, most word boundaries misplaced
1.22.0 CER 0.000 — identical to the expected text
1.22.1 / 1.23.0 / 1.24.0 CER 0.000

So on this sample firecrawl/pdf-inspector#440 and firecrawl/pdf-inspector#441 (1.16.0) fixed letter order but left the spacing, and the text only becomes usable at 1.22.0 — which includes firecrawl/pdf-inspector#552 (RTL lines read back through the Unicode Bidirectional Algorithm), likely the fix, though I haven't bisected.

Suggest targeting 1.23.0 or newer rather than 1.20.0 (and if you prefer a floor, 1.22.0 rather than 1.16). It costs nothing extra here: text.pdf converts identically on 1.20.0 through 1.24.0, so this PR's snapshot stays as is. 1.23.0 also adds firecrawl/pdf-inspector#561 (right-to-left runs already shown in reading order are read forwards), which matters for logically-stored producers.

Worth noting for the snapshot: the only Arabic-script text in the PDF fixtures is one Persian word in text.pdf, not stored as presentation forms — enough to catch letter order (it did, in this PR's diff), but with no word boundary it cannot catch the spacing bug that survived to 1.21.0. I can contribute two generated fixtures, one from a visually-stored producer and one from a logically-stored producer, as regression guards; happy to post them if useful.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

anydoc-wasm still ships pdf-inspector 1.14.2 — the RTL extraction fix (pdf-inspector#440) isn't included

2 participants