Repository navigation
fix: table cells read via get_textbox — extract() gave 37.50 as 3750 after importing pymupdf4llm - #43
Merged
Merged
Conversation
…" as 3750 after importing pymupdf4llm
pymupdf4llm's import calls pymupdf.TOOLS.unset_quad_corrections(True) process-wide; until a
to_markdown() call resets pymupdf.table.FLAGS, Table.extract() splits punctuation and drops spaces
("37.50" -> "3750\n.", "Card 0" -> "Card0"): a silent, uniform 100x that reconciliation can't see.
New pdfmux/table_cells.py table_cell_texts() reads each cell with page.get_textbox(cell rect),
correct in every global state. fast.py (correct today only by call order) and segment.py (broken
whenever pymupdf4llm was imported) now use it. 4 tests reproduce the exact post-import state; both
call-site tests fail on the old code. Full suite: 744 passed, 4 skipped (MCP test files excluded:
local venv lacks mcp>=2).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
NameetP
added a commit
that referenced
this pull request
Oct 2, 2026
…sheet) - remote/tables.py drops its private _cell_texts copy for pdfmux.table_cells (merged in #43). - outputs.py: wb.worksheets[0] instead of Optional wb.active (pyright reportOptionalSubscript). - ci.yml: install the remote extra so test_remote_server can import httpx/starlette/openpyxl. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
NameetP
added a commit
that referenced
this pull request
Oct 6, 2026
…ed, nothing-skipped) (#42) * remote: deterministic extract_tables engine (stitching, typed numbers/dates, balance reconciliation) No models, no network, no paid fallback. 30 tests on generated PDFs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * remote: ChatGPT plugin server (pdfmux serve --remote) — fetch, sandbox, XLSX, kit - extract_tables tool: file (openai/fileParams) or pdf_url → CSV inline (wrapped untrusted) + 15-minute signed XLSX link + "rows verified N of M"; neutral copy, annotations. - fetch.py: https-only, every resolved address must be global (blocks loopback, RFC1918, 169.254.169.254 metadata, 100.64/10 Tailscale, v4-mapped v6), connected peer re-checked (rebinding), manual re-validated redirects, 25 MB streaming cap, %PDF- magic. - sandbox.py + worker.py: parse in a separate process — docker run --network none --read-only --cap-drop ALL --user nobody with mem/cpu/pid limits in production, rlimits locally; timeout enforced from outside (SIGALRM can't stop C code). Front end never parses; worker never fetches. - outputs.py: CSV/XLSX from typed values (numbers stay numbers, dates real dates, summary sheet). - kit.py: Python port of chatgpt-plugin-kit (token trust, last-XFF IP, fail-closed limits, kill-switch file, banned copy, untrusted wrapper, JSONL logs without cell values). - FIX: PyMuPDF 1.27 find_tables().extract() splits "37.50" into "3750\n." and drops spaces — a silent 100x on every amount. Cells are now read via page.get_textbox(cell rect). Regression test pinned. (Core pdfmux's two extract() call sites are a follow-up.) Tests: 62 remote tests. Full suite: 803 passed; the only failures are the MCP test files, which need mcp>=2 that this local venv lacks (pre-existing, unrelated). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * remote: sandbox image + make job files readable by the unprivileged container user Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * packaging: remote extra for the hosted ChatGPT plugin Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * examples: two-page sample bank statement (synthetic) for the remote extract_tables demo/tests Force-added past the *.pdf ignore on purpose: it's a deliberate public fixture. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * remote: use shared table_cell_texts; fix CI (remote extra, typed worksheet) - remote/tables.py drops its private _cell_texts copy for pdfmux.table_cells (merged in #43). - outputs.py: wb.worksheets[0] instead of Optional wb.active (pyright reportOptionalSubscript). - ci.yml: install the remote extra so test_remote_server can import httpx/starlette/openpyxl. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * tests: import sibling test module the way bare pytest resolves it (CI runs pytest, not python -m) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
pymupdf4llm's import callspymupdf.TOOLS.unset_quad_corrections(True)for the whole process, andimport pdfmuximports it. From then on,Table.extract()on a ruled table returns:extract()after the import37.50"3750\n."1,063.50"106350\n, ."Card 0"Card0"It stays broken until something calls
pymupdf4llm.to_markdown(), which happens to overwritepymupdf.table.FLAGS. So the same page reads right or wrong depending on call order. A running-balance reconciliation can't catch it, because the 100× is uniform.I reproduced it in isolation: a fresh process is correct, and the same file goes wrong right after
import pymupdf4llm. Verified on PyMuPDF 1.27.1 / pymupdf4llm 0.3.4. Found while building the hostedextract_tablestool (#42).Exposure in pdfmux
extractors/fast.py(--quality fastwith tables): currently correct only becauseto_markdown()runs first on that path.segment.py_detect_table_regions: broken whenever pymupdf4llm has been imported. Inside the pipeline its text only feeds a log line, butdetect_segmentsis importable library code.pdfmux convertoutput was found. This removes the trap before a different call order hits it.Fix
pdfmux/table_cells.py'stable_cell_texts(page, table)reads each cell withpage.get_textbox(cell_rect), which is correct in every global state. Both call sites use it.Tests
tests/test_table_cells.pyreproduces the post-import state and checks:Both call-site tests fail on
main; I checked by reverting the fix.Full suite: 744 passed, 4 skipped. The MCP test files were excluded locally because this venv lacks
mcp>=2; CI runs them.Follow-up: #42's
src/pdfmux/remote/tables.pyhas its own copy of this fix (_cell_texts). After both merge it should switch to the shared helper.🤖 Generated with Claude Code