Skip to content

fix: table cells read via get_textbox — extract() gave 37.50 as 3750 after importing pymupdf4llm - #43

Merged
NameetP merged 1 commit into
mainfrom
fix/table-extract-decimals
Oct 2, 2026
Merged

NameetP merged 1 commit into
mainfrom
fix/table-extract-decimals

Conversation

@NameetP

@NameetP NameetP commented Oct 2, 2026

Copy link
Copy Markdown
Owner

The bug

pymupdf4llm's import calls pymupdf.TOOLS.unset_quad_corrections(True) for the whole process, and import pdfmux imports it. From then on, Table.extract() on a ruled table returns:

Cell extract() after the import Meaning after cleanup
37.50 "3750\n." 3750 (100×)
1,063.50 "106350\n, ." 106350
Card 0 "Card0" wrong text

It stays broken until something calls pymupdf4llm.to_markdown(), which happens to overwrite pymupdf.table.FLAGS. So the same page reads right or wrong depending on call order. A running-balance reconciliation can't catch it, because the 100× is uniform.

I reproduced it in isolation: a fresh process is correct, and the same file goes wrong right after import pymupdf4llm. Verified on PyMuPDF 1.27.1 / pymupdf4llm 0.3.4. Found while building the hosted extract_tables tool (#42).

Exposure in pdfmux

  • extractors/fast.py (--quality fast with tables): currently correct only because to_markdown() runs first on that path.
  • segment.py _detect_table_regions: broken whenever pymupdf4llm has been imported. Inside the pipeline its text only feeds a log line, but detect_segments is importable library code.
  • Users: no wrong user-facing pdfmux convert output was found. This removes the trap before a different call order hits it.

Fix

pdfmux/table_cells.py's table_cell_texts(page, table) reads each cell with page.get_textbox(cell_rect), which is correct in every global state. Both call sites use it.

Tests

tests/test_table_cells.py reproduces the post-import state and checks:

  • the helper reads cells exactly
  • the fast extractor's structured tables are exact
  • segment text is exact
  • a sentinel records that upstream is still broken, and skips itself if that's ever fixed

Both call-site tests fail on main; I checked by reverting the fix.

Full suite: 744 passed, 4 skipped. The MCP test files were excluded locally because this venv lacks mcp>=2; CI runs them.

Follow-up: #42's src/pdfmux/remote/tables.py has its own copy of this fix (_cell_texts). After both merge it should switch to the shared helper.

🤖 Generated with Claude Code

…" as 3750 after importing pymupdf4llm

pymupdf4llm's import calls pymupdf.TOOLS.unset_quad_corrections(True) process-wide; until a
to_markdown() call resets pymupdf.table.FLAGS, Table.extract() splits punctuation and drops spaces
("37.50" -> "3750\n.", "Card 0" -> "Card0"): a silent, uniform 100x that reconciliation can't see.

New pdfmux/table_cells.py table_cell_texts() reads each cell with page.get_textbox(cell rect),
correct in every global state. fast.py (correct today only by call order) and segment.py (broken
whenever pymupdf4llm was imported) now use it. 4 tests reproduce the exact post-import state; both
call-site tests fail on the old code. Full suite: 744 passed, 4 skipped (MCP test files excluded:
local venv lacks mcp>=2).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@NameetP
NameetP merged commit c899e7c into main Oct 2, 2026
7 checks passed
@NameetP
NameetP deleted the fix/table-extract-decimals branch October 2, 2026 19:39
NameetP added a commit that referenced this pull request Oct 2, 2026
…sheet)

- remote/tables.py drops its private _cell_texts copy for pdfmux.table_cells (merged in #43).
- outputs.py: wb.worksheets[0] instead of Optional wb.active (pyright reportOptionalSubscript).
- ci.yml: install the remote extra so test_remote_server can import httpx/starlette/openpyxl.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
NameetP added a commit that referenced this pull request Oct 6, 2026
…ed, nothing-skipped) (#42)

* remote: deterministic extract_tables engine (stitching, typed numbers/dates, balance reconciliation)

No models, no network, no paid fallback. 30 tests on generated PDFs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* remote: ChatGPT plugin server (pdfmux serve --remote) — fetch, sandbox, XLSX, kit

- extract_tables tool: file (openai/fileParams) or pdf_url → CSV inline (wrapped untrusted) +
  15-minute signed XLSX link + "rows verified N of M"; neutral copy, annotations.
- fetch.py: https-only, every resolved address must be global (blocks loopback, RFC1918,
  169.254.169.254 metadata, 100.64/10 Tailscale, v4-mapped v6), connected peer re-checked
  (rebinding), manual re-validated redirects, 25 MB streaming cap, %PDF- magic.
- sandbox.py + worker.py: parse in a separate process — docker run --network none --read-only
  --cap-drop ALL --user nobody with mem/cpu/pid limits in production, rlimits locally; timeout
  enforced from outside (SIGALRM can't stop C code). Front end never parses; worker never fetches.
- outputs.py: CSV/XLSX from typed values (numbers stay numbers, dates real dates, summary sheet).
- kit.py: Python port of chatgpt-plugin-kit (token trust, last-XFF IP, fail-closed limits,
  kill-switch file, banned copy, untrusted wrapper, JSONL logs without cell values).
- FIX: PyMuPDF 1.27 find_tables().extract() splits "37.50" into "3750\n." and drops spaces —
  a silent 100x on every amount. Cells are now read via page.get_textbox(cell rect). Regression
  test pinned. (Core pdfmux's two extract() call sites are a follow-up.)

Tests: 62 remote tests. Full suite: 803 passed; the only failures are the MCP test files,
which need mcp>=2 that this local venv lacks (pre-existing, unrelated).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* remote: sandbox image + make job files readable by the unprivileged container user

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* packaging: remote extra for the hosted ChatGPT plugin

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* examples: two-page sample bank statement (synthetic) for the remote extract_tables demo/tests

Force-added past the *.pdf ignore on purpose: it's a deliberate public fixture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* remote: use shared table_cell_texts; fix CI (remote extra, typed worksheet)

- remote/tables.py drops its private _cell_texts copy for pdfmux.table_cells (merged in #43).
- outputs.py: wb.worksheets[0] instead of Optional wb.active (pyright reportOptionalSubscript).
- ci.yml: install the remote extra so test_remote_server can import httpx/starlette/openpyxl.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* tests: import sibling test module the way bare pytest resolves it (CI runs pytest, not python -m)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant