fix(apa-48): ship the parse runtime in the worker image; harness faults are not bad documents - #115
Conversation
…ts are not bad documents The deployed worker could not parse a single document. Dockerfile.worker shipped the static Go binary on bare Alpine: no python3, no tools/. Every ingested document reached internal/parser/liteparse client.go exec of "python3 tools/parsers/liteparse/shim.py", failed at cmd.Start(), and was failTerminal()'d as a parse failure before extraction. Runtime restored in the image - Multi-stage build adds a pinned Python + LiteParse runtime. The wheel is resolved out of uv.lock at build time (musllinux, sha256-verified), so uv.lock stays the single source of truth and no pin is duplicated here. - The shim is copied to /app/tools/parsers/liteparse/shim.py, the exact path DefaultShimPath resolves to from WORKDIR. - Build context moves to the repo root: the shim and pyproject.toml/ uv.lock live outside apps/api. Go stage stays scoped to apps/api/. - New .dockerignore keeps .env, .git and caches out of the context. - Image: 48.3MB -> 132MB. Correctness of the artifact over size. Harness faults are not document faults - shim envelope code "harness" (unreadable input, failed vendor import, serialization failure) and an unresolvable interpreter or missing shim now map to liteparse.ErrRuntimeUnavailable, not parser.ErrParseFailure. - The worker routes that to the EXISTING OutcomeTransient (non-2xx/Nack), reusing the isPermanent/retryableStoreErr idiom. No new OutcomeKind, no new exception taxonomy, no new ExceptionCode. It still fails closed: no extraction, no evidence, no SUCCESS, and no FAILED document row blaming the customer's file. - Real document faults stay TERMINAL and still record the FAILED row. - The sentinel lives in the adapter, not internal/parser: that package is a sealed contract that must not import adapters (parser.go:304). The gate that should have caught this never gated anything - ci.yml ran the live shim BEFORE `Set up Python` / `uv sync`, so TestLiveShimSmoke skipped on every run. It now runs after Python setup with CLAIMOPS_REQUIRE_LITEPARSE=1, so a broken parser environment fails the job instead of skipping. The local skip is kept only for an absent optional venv. - The smoke asserts the full evidence surface, not just "it started": pages, tables, cells, bounding boxes, block IDs, block types, parser metadata and reading order, from a real corpus PDF (CASE-001). Reading order asserts the stream starts at the page top, deliberately NOT strict Y monotonicity, which observed vendor output does not satisfy. - New container smoke (tools/parsers/liteparse/container_smoke.sh) runs inside the built image and exercises the production relative shim path. Unchanged: worker entrypoint, internal/parser interface and error set, OCRConfidence/ConfidenceAvailable provenance, ADR-007, ADR-009.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Aparnap2
left a comment
There was a problem hiding this comment.
REVIEW RESULT: PASS / merge-ready, with one post-merge release check.
The P0 is correctly addressed at the deployed-artifact boundary, not just the host test boundary. The worker image now contains python3, the pinned LiteParse runtime, and the exact production shim path; the container smoke checks the real relative execution path and evidence payload. The parser runtime fault is correctly distinguished from document corruption and routed through existing OutcomeTransient semantics, avoiding a durable FAILED document classification for an environment defect.
The CI ordering fix is important and correctly makes the LiteParse smoke fail rather than skip when the runtime is broken. The parser/worker regression tests preserve the existing parser and OCR confidence contracts, and the eval surface is unchanged.
The observed non-monotonic vendor block Y-order is handled correctly: the smoke asserts only the property actually guaranteed (stream starts at page top) instead of inventing monotonic ordering semantics.
I agree with the +83.7 MB image increase as a consequence of keeping LiteParse in the production path for now; the breakdown shows the bulk is Python/vendor binaries. This should be treated as an operational cost to revisit with a managed OCR implementation, not as a reason to compromise the current parser contract in the P0.
One release requirement remains: because the repository's integration workflow does not build/deploy the Cloud Run worker image itself, the actual worker image must be rebuilt/redeployed from the new repo-root context (docker build -f apps/api/Dockerfile.worker .) and one real document must parse successfully before declaring the production outage closed. I cannot verify any external image-build pipeline from this repository alone.
Decision: merge-ready. Do not combine GCP Document AI migration or Python-domain cleanup with this P0. After merge/redeploy, proceed to PR #114 and then the real-LLM qualification matrix.
What changed
The deployed worker could not parse a single document.
Dockerfile.workershipped the static Go binary on bare Alpine — nopython3, notools/— so every document hitcmd.Start()failure ininternal/parser/liteparseand wasfailTerminal()'d before extraction.uv.lockat build time (musllinux, sha256-verified), souv.lockstays the single source of truth. Shim copied to/app/tools/parsers/liteparse/shim.py— the exact pathDefaultShimPathresolves to fromWORKDIR. Build context moves to repo root (shim +pyproject.toml/uv.locklive outsideapps/api); new.dockerignorekeeps.env/.git/caches out.harnessenvelope + unresolvable interpreter + missing shim map toliteparse.ErrRuntimeUnavailable, routed by the worker into the existingOutcomeTransient. No new OutcomeKind, no new exception taxonomy. Still fails closed: no extraction, no evidence, no FAILED document row blaming the customer's file. Real document faults stay TERMINAL.uv syncand skipped on every run. It now runs after Python setup withCLAIMOPS_REQUIRE_LITEPARSE=1, so a broken parser env fails the job. Smoke asserts full evidence (pages, tables, cells, boxes, block IDs, block types, metadata, reading order) from corpus PDF CASE-001. Newcontainer_smoke.shruns inside the built image via the production relative shim path.Unchanged: worker entrypoint,
internal/parserinterface + error set,OCRConfidence/ConfidenceAvailable, ADR-007, ADR-009,workflows/claim-investigation.yaml.Evidence
go test ./...41 pkgs,-raceon touched pkgs,pytest tests/unit17,ruffclean)docker build+ container smoke, 48.3MB -> 132MB)uid=10001(apprunner),WORKDIR=/app, liteparse 2.14.4TestAdapterConformance, eval-v1, ADR-009 OCR + APA-12 HITL)CLAIMOPS_REQUIRE_LITEPARSE=1FAILs when liteparse is removed, SKIPs only without the flagNot production-qualified on unit tests alone: the image must be rebuilt and redeployed, and CI must be observed green on the PR.
Blockers
None. One decision to review: the build context changed to repo root. It is forced (shim +
uv.lockare outsideapps/api) and fails loudly at build time if a caller still uses the old context, but any out-of-band pipeline that built this image needs the new command.Next step
Rebuild/redeploy
claimops-worker, confirm one document parses end-to-end, then merge.