From d47a73c4816a4fffd22b3aedbecf1e1e8f1c936c Mon Sep 17 00:00:00 2001 From: Aparna Pradhan Date: Mon, 28 Sep 2026 19:01:50 +0530 Subject: [PATCH] fix(apa-48): ship the parse runtime in the worker image; harness faults are not bad documents The deployed worker could not parse a single document. Dockerfile.worker shipped the static Go binary on bare Alpine: no python3, no tools/. Every ingested document reached internal/parser/liteparse client.go exec of "python3 tools/parsers/liteparse/shim.py", failed at cmd.Start(), and was failTerminal()'d as a parse failure before extraction. Runtime restored in the image - Multi-stage build adds a pinned Python + LiteParse runtime. The wheel is resolved out of uv.lock at build time (musllinux, sha256-verified), so uv.lock stays the single source of truth and no pin is duplicated here. - The shim is copied to /app/tools/parsers/liteparse/shim.py, the exact path DefaultShimPath resolves to from WORKDIR. - Build context moves to the repo root: the shim and pyproject.toml/ uv.lock live outside apps/api. Go stage stays scoped to apps/api/. - New .dockerignore keeps .env, .git and caches out of the context. - Image: 48.3MB -> 132MB. Correctness of the artifact over size. Harness faults are not document faults - shim envelope code "harness" (unreadable input, failed vendor import, serialization failure) and an unresolvable interpreter or missing shim now map to liteparse.ErrRuntimeUnavailable, not parser.ErrParseFailure. - The worker routes that to the EXISTING OutcomeTransient (non-2xx/Nack), reusing the isPermanent/retryableStoreErr idiom. No new OutcomeKind, no new exception taxonomy, no new ExceptionCode. It still fails closed: no extraction, no evidence, no SUCCESS, and no FAILED document row blaming the customer's file. - Real document faults stay TERMINAL and still record the FAILED row. - The sentinel lives in the adapter, not internal/parser: that package is a sealed contract that must not import adapters (parser.go:304). The gate that should have caught this never gated anything - ci.yml ran the live shim BEFORE `Set up Python` / `uv sync`, so TestLiveShimSmoke skipped on every run. It now runs after Python setup with CLAIMOPS_REQUIRE_LITEPARSE=1, so a broken parser environment fails the job instead of skipping. The local skip is kept only for an absent optional venv. - The smoke asserts the full evidence surface, not just "it started": pages, tables, cells, bounding boxes, block IDs, block types, parser metadata and reading order, from a real corpus PDF (CASE-001). Reading order asserts the stream starts at the page top, deliberately NOT strict Y monotonicity, which observed vendor output does not satisfy. - New container smoke (tools/parsers/liteparse/container_smoke.sh) runs inside the built image and exercises the production relative shim path. Unchanged: worker entrypoint, internal/parser interface and error set, OCRConfidence/ConfidenceAvailable provenance, ADR-007, ADR-009. --- .dockerignore | 44 +++++ .github/workflows/ci.yml | 37 +++- apps/api/Dockerfile.worker | 63 +++++- apps/api/internal/parser/liteparse/client.go | 13 +- apps/api/internal/parser/liteparse/errors.go | 44 ++++- .../internal/parser/liteparse/errors_test.go | 9 +- .../internal/parser/liteparse/live_test.go | 183 +++++++++++++++--- .../parser/liteparse/runtime_errors_test.go | 137 +++++++++++++ .../internal/worker/parser_runtime_test.go | 111 +++++++++++ apps/api/internal/worker/processor.go | 32 +++ tools/parsers/liteparse/container_smoke.sh | 168 ++++++++++++++++ 11 files changed, 800 insertions(+), 41 deletions(-) create mode 100644 .dockerignore create mode 100644 apps/api/internal/parser/liteparse/runtime_errors_test.go create mode 100644 apps/api/internal/worker/parser_runtime_test.go create mode 100755 tools/parsers/liteparse/container_smoke.sh diff --git a/.dockerignore b/.dockerignore new file mode 100644 index 0000000..7aaa511 --- /dev/null +++ b/.dockerignore @@ -0,0 +1,44 @@ +# Docker build context hygiene for the repo-root worker build. +# +# The worker image builds from the REPO ROOT (see apps/api/Dockerfile.worker): +# the Go module is apps/api, but the parse runtime needs the repo-root shim +# (tools/) and the pinned Python dependency source (pyproject.toml, uv.lock). +# +# Everything below is excluded so the build context stays small, secret-free +# and reproducible. The negative-pattern intent is deny-by-default: only what +# the worker image actually needs may enter the context. + +# VCS + local state +.git +.gitignore +.github + +# Secrets and environments. .env must never reach a build layer. +.env +.env.* +.venv +venv + +# Tooling caches +.mypy_cache +.ruff_cache +.pytest_cache +**/__pycache__ +**/*.pyc + +# Test + evaluation material. The container smoke MOUNTS the corpus +# read-only at run time; fixtures are never baked into the production image. +fixtures +tests +eval +mocks +docs +workflows +infra + +# Docs and non-build root files +prd.md +README.md +AGENTS.md +skills-lock.json +Makefile diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index fd7ea69..c71543a 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -6,9 +6,13 @@ # emulators are unreachable (see requirePool in # apps/api/internal/repository/postgres/postgres_test.go and the Makefile # header: "integration-gated tests skip gracefully ... exit 0"). -# - The LiteParse live smoke below skips when the `liteparse` Python package -# is absent (see liveShim in -# apps/api/internal/parser/liteparse/live_test.go). +# +# The LiteParse live smoke is the exception and no longer self-skips here: +# it runs AFTER `uv sync`, with CLAIMOPS_REQUIRE_LITEPARSE=1, so an absent +# or broken `liteparse` FAILS the job. Before APA-48 it ran before the +# Python setup and skipped on every run, which is how a worker image with +# no interpreter shipped green (see liveShim in +# apps/api/internal/parser/liteparse/live_test.go). # # Full parser benchmark (#31: Python + liteparse + 45 PDFs, ~1-2 min) is # intentionally NOT run here. The benchmark-smoke step runs only the gated @@ -63,13 +67,6 @@ jobs: # package list down to something that excludes them. run: go test ./... - - name: benchmark smoke (live shim only, skips without liteparse) - working-directory: apps/api - # Gated smoke from #31: real shim + mapping + Validate path against - # CASE-001, or a clean SKIP when liteparse is not installed. This is - # the smoke gate, not the fidelity benchmark. - run: go test ./internal/parser/liteparse/ -run TestLiveShimSmoke -count=1 -v - - name: Set up Python uses: actions/setup-python@v5 with: @@ -84,6 +81,26 @@ jobs: # `dev` dependency-group in pyproject.toml. run: uv sync --group dev + # APA-48: this gate used to run BEFORE `Set up Python` / `uv sync`, so + # liteparse was never importable and TestLiveShimSmoke SKIPPED on every + # run — a permanently green pipeline while every ingested document + # terminal-FAILED at parse in production. It now runs after the Python + # environment exists, and CLAIMOPS_REQUIRE_LITEPARSE=1 makes a broken + # parser environment a hard failure instead of a skip. + # + # Kept out of `go test ./...` on purpose: a skip-then-succeed interaction + # there would hide the gate behind an unrelated package failure. As its + # own step, a missing runtime fails THIS job loudly. + - name: benchmark smoke (live shim, fails if parser env is broken) + working-directory: apps/api + # Gated smoke from #31: real shim + mapping + Validate path against + # CASE-001, asserting pages, tables, bounding boxes, block IDs, block + # types, parser metadata and reading-order evidence. This is the + # smoke gate, not the fidelity benchmark. + env: + CLAIMOPS_REQUIRE_LITEPARSE: '1' + run: go test ./internal/parser/liteparse/ -run TestLiveShimSmoke -count=1 -v + - name: Python unit tests run: uv run pytest tests/unit -q diff --git a/apps/api/Dockerfile.worker b/apps/api/Dockerfile.worker index cace419..b889807 100644 --- a/apps/api/Dockerfile.worker +++ b/apps/api/Dockerfile.worker @@ -1,23 +1,74 @@ # ClaimOps Worker — production image (Cloud Run contract). # Serves POST /events/document-ingested for Pub/Sub push delivery. -# Build from the module root: docker build -f Dockerfile.worker . -# Alpine-only, squeezed: static binary on minimal Alpine, non-root, -# ca-certificates for TLS. No shell tools in the final layer. +# +# Build from the REPO ROOT (context changed in APA-48): +# docker build -f apps/api/Dockerfile.worker -t claimops-worker . +# +# The context is the repo root, not the module root, because the deployed +# worker needs two things that live outside apps/api: +# 1. tools/parsers/liteparse/shim.py — the parse shim, which +# internal/parser/liteparse resolves as DefaultShimPath relative to +# WORKDIR, i.e. /app/tools/parsers/liteparse/shim.py, and +# 2. pyproject.toml + uv.lock — the pinned LiteParse dependency source. +# Everything the Go stage needs is still scoped to apps/api/, so the Go +# build is byte-for-byte the same as before. +# +# APA-48: the image previously shipped the static Go binary on bare Alpine +# with no interpreter and no tools/, so internal/parser/liteparse +# client.go exec of "python3 tools/parsers/liteparse/shim.py" failed at +# cmd.Start() and EVERY ingested document died terminal-FAILED at parse. +# The parse runtime is now part of the artifact, and the container smoke +# (tools/parsers/liteparse/container_smoke.sh) gates it inside this image. + +# ---------- Stage 1: static Go binary ---------- FROM golang:1.27-alpine AS build RUN apk add --no-cache git ca-certificates WORKDIR /src -COPY go.mod go.sum ./ +COPY apps/api/go.mod apps/api/go.sum ./ RUN go mod download -COPY . . +COPY apps/api/ ./ RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags="-s -w" -o /out/worker ./cmd/worker \ && ls -la /out/worker +# ---------- Stage 2: pinned LiteParse runtime ---------- +# Isolated so no pip cache, no build tooling, and no other project +# dependency lands in the final image. Only the one wheel the parse shim +# imports is installed: the shim needs liteparse and nothing else, and the +# worker image stays "minimum" as the header requires. +FROM alpine:3.21 AS pybuild +RUN apk add --no-cache python3 py3-pip +COPY pyproject.toml uv.lock /src/ +# Resolve the wheel straight out of uv.lock instead of re-typing a URL or +# version here, so uv.lock stays the single source of truth and cannot +# drift from the image. musllinux is the Alpine (musl) build; cp310-abi3 +# is satisfied by Alpine's Python 3.12. +RUN python3 - <<'PY' +import pathlib, tomllib +lock = tomllib.loads(pathlib.Path("/src/uv.lock").read_text()) +wheels = [w for pkg in lock["package"] if pkg["name"] == "liteparse" for w in pkg["wheels"]] +match = [w for w in wheels if "musllinux_1_2_x86_64" in w["url"]] +if len(match) != 1: + raise SystemExit(f"uv.lock: expected exactly 1 musllinux liteparse wheel, found {len(match)}") +w = match[0] +pathlib.Path("/tmp/req.txt").write_text(f"{w['url']}#sha256={w['hash'].split(':', 1)[1]}\n") +print("locked:", w["url"].rsplit('/', 1)[-1], w["hash"]) +PY +RUN pip install --no-cache-dir --break-system-packages --target=/opt/lp -r /tmp/req.txt \ + && PYTHONPATH=/opt/lp python3 -c "import liteparse; print('build-stage liteparse import OK')" + +# ---------- Stage 3: the deployed artifact ---------- FROM alpine:3.21 -RUN apk add --no-cache ca-certificates \ +RUN apk add --no-cache ca-certificates python3 \ && adduser -D -H -u 10001 apprunner \ && rm -rf /var/cache/apk/* WORKDIR /app COPY --from=build /out/worker /app/worker +# Python site-packages only — the stdlib comes from the apk python3 package +# already installed above, so nothing from pybuild but the wheel is copied. +COPY --from=pybuild /opt/lp/ /usr/lib/python3.12/site-packages/ +# The shim at the exact path DefaultShimPath resolves to from WORKDIR /app. +# Copied from repo source so the image can never run a stale shim. +COPY tools/parsers/liteparse/shim.py /app/tools/parsers/liteparse/shim.py ENV PORT=8080 APP_ENV=prod PUSH_AUTH_MODE=oidc EXPOSE 8080 USER apprunner diff --git a/apps/api/internal/parser/liteparse/client.go b/apps/api/internal/parser/liteparse/client.go index f40e9a2..d688750 100644 --- a/apps/api/internal/parser/liteparse/client.go +++ b/apps/api/internal/parser/liteparse/client.go @@ -5,6 +5,7 @@ import ( "context" "fmt" "io" + "os" "os/exec" "time" ) @@ -54,6 +55,16 @@ func NewExecRunner(pythonBin, shimPath string, timeout time.Duration) *ExecRunne // Run implements Runner. func (r *ExecRunner) Run(ctx context.Context, pdfPath string) ([]byte, error) { + // Pre-flight the shim. A shim that is not in the image is a misbuilt + // artifact, not a bad document: without this the interpreter starts, + // exits 2 with "can't open file" on stderr, and mapExitError files it + // as a document fault. Checking first keeps the runtime class explicit + // and avoids paying for a subprocess that cannot succeed. + if _, err := os.Stat(r.shimPath); err != nil { + return nil, fmt.Errorf("%w: liteparse shim not readable at %q: %v", + ErrRuntimeUnavailable, r.shimPath, err) + } + runCtx := ctx cancel := context.CancelFunc(func() {}) if _, hasDeadline := ctx.Deadline(); !hasDeadline { @@ -70,7 +81,7 @@ func (r *ExecRunner) Run(ctx context.Context, pdfPath string) ([]byte, error) { cmd.Stderr = &stderr if err := cmd.Start(); err != nil { - return nil, wrapExec("start", err, nil, ctx) + return nil, wrapSpawn(err, ctx) } // Bound stdout: read at most MaxStdoutBytes+1 to detect overflow. limited := io.LimitReader(stdout, MaxStdoutBytes+1) diff --git a/apps/api/internal/parser/liteparse/errors.go b/apps/api/internal/parser/liteparse/errors.go index 54cb889..4934589 100644 --- a/apps/api/internal/parser/liteparse/errors.go +++ b/apps/api/internal/parser/liteparse/errors.go @@ -4,6 +4,7 @@ import ( "context" "errors" "fmt" + "os/exec" "strings" "claimops-api/internal/parser" @@ -14,6 +15,26 @@ import ( // failure, not a caller media error). var errShimOverflow = errors.New("liteparse: shim output overflow") +// ErrRuntimeUnavailable marks a LiteParse EXECUTION-ENVIRONMENT fault: the +// interpreter could not be resolved, the shim is not present at +// DefaultShimPath, or the vendor library failed to import inside the +// subprocess. The shim reports the in-process cases with envelope code +// "harness" (shim.py:99, 109, 148). +// +// APA-48: this is deliberately NOT parser.ErrParseFailure. ErrParseFailure +// is the DOCUMENT-fault class — the worker failTerminal()s it, durably +// recording the customer's file as unparseable. A misbuilt image or an +// uninstalled dependency is our own fault, says nothing about the +// document, and may heal on redelivery, so the worker routes this to the +// existing OutcomeTransient instead. Before the fix these conditions were +// terminal-FAILED: a worker image shipping without an interpreter burned +// every ingested document. +// +// The sentinel lives here, not in internal/parser, because that package is +// a sealed contract that must never import adapters (parser.go:304) and +// enumerates its error set; no other adapter can raise this condition. +var ErrRuntimeUnavailable = errors.New("liteparse: parser runtime unavailable") + // unsupportedMedia classifies a non-vendor-handled input. Returned BEFORE // spawning the shim so unsupported media never pays subprocess cost. func unsupportedMedia(mediaType string) error { @@ -55,6 +76,21 @@ func wrapExec(stage string, err error, stderr []byte, ctx context.Context) error return fmt.Errorf("%w: liteparse shim %s: %v", parser.ErrParseFailure, stage, err) } +// wrapSpawn maps a failure to LAUNCH the shim subprocess. An unresolvable +// interpreter (exec.ErrNotFound — the exact failure of an image built +// without python3) is an environment fault, not a document fault, so it +// gets ErrRuntimeUnavailable instead of ErrParseFailure. Cancellation +// still takes precedence. +func wrapSpawn(err error, ctx context.Context) error { + if ctx.Err() != nil { + return wrapCtx(ctx.Err()) + } + if errors.Is(err, exec.ErrNotFound) { + return fmt.Errorf("%w: %v", ErrRuntimeUnavailable, err) + } + return wrapExec("start", err, nil, ctx) +} + // mapExitError maps a non-zero shim exit (or signal kill) to a typed // error. ctx cancellation (including the timeout kill) maps to the ctx // error; anything else is a shim crash -> ErrParseFailure with a stderr @@ -83,9 +119,15 @@ func mapEnvelopeError(code, message string) error { // gate let through). Same taxonomy as the Go pre-spawn gate. return fmt.Errorf("%w: liteparse shim: %s", parser.ErrUnsupportedMediaType, truncate(message, 500)) - case "parse_failure", "timeout", "harness": + case "parse_failure", "timeout": return fmt.Errorf("%w: liteparse shim %s: %s", parser.ErrParseFailure, code, truncate(message, 500)) + case "harness": + // Runtime/environment fault reported from inside the shim (cannot + // read input, vendor import failed, payload serialization failed). + // Never a document fault — see ErrRuntimeUnavailable. + return fmt.Errorf("%w: liteparse shim %s: %s", + ErrRuntimeUnavailable, code, truncate(message, 500)) default: return fmt.Errorf("%w: liteparse shim unknown error code %q: %s", parser.ErrParseFailure, code, truncate(message, 500)) diff --git a/apps/api/internal/parser/liteparse/errors_test.go b/apps/api/internal/parser/liteparse/errors_test.go index b62516c..0f9f36b 100644 --- a/apps/api/internal/parser/liteparse/errors_test.go +++ b/apps/api/internal/parser/liteparse/errors_test.go @@ -12,7 +12,14 @@ func TestMapEnvelopeErrorTaxonomy(t *testing.T) { if err := mapEnvelopeError("unsupported_media", "x"); !errors.Is(err, parser.ErrUnsupportedMediaType) { t.Fatalf("unsupported_media = %v", err) } - for _, code := range []string{"parse_failure", "timeout", "harness", "bogus"} { + // APA-48: "harness" moved OUT of this list. It reports a runtime/ + // environment fault (shim cannot read input, vendor import failed, + // serialization failed), not a defect in the document, so it now maps to + // ErrRuntimeUnavailable; the worker routes that to the existing + // OutcomeTransient instead of failTerminal, which used to durably record + // a good customer's file as unparseable whenever our own image was broken. + // See runtime_errors_test.go for the dedicated coverage. + for _, code := range []string{"parse_failure", "timeout", "bogus"} { if err := mapEnvelopeError(code, "x"); !errors.Is(err, parser.ErrParseFailure) { t.Fatalf("%s = %v, want ErrParseFailure", code, err) } diff --git a/apps/api/internal/parser/liteparse/live_test.go b/apps/api/internal/parser/liteparse/live_test.go index cacf47f..b25ae35 100644 --- a/apps/api/internal/parser/liteparse/live_test.go +++ b/apps/api/internal/parser/liteparse/live_test.go @@ -2,6 +2,7 @@ package liteparse import ( "context" + "fmt" "os" "os/exec" "path/filepath" @@ -11,50 +12,187 @@ import ( "claimops-api/internal/parser" ) -// TestLiveShimSmoke runs the real shim + mapping + Validate path against -// a minimal generated PDF. Skipped when liteparse is not installed. -// This is a smoke gate, not a fidelity test: #31 owns corpus scoring. +// requireEnvEnv makes the live parser gate a hard failure instead of a skip +// when set to "1". CI sets it AFTER Python/uv setup, so a broken production +// parser environment fails the job instead of quietly skipping. Locally it +// stays unset so a contributor without the venv is not blocked; that skip is +// the only legitimate one (see liveShim). +const requireLiveEnv = "CLAIMOPS_REQUIRE_LITEPARSE" + +// TestLiveShimSmoke runs the real shim + mapping + Validate path against a +// real corpus PDF (CASE-001, the same bytes eval-v1 scores). It is a smoke +// gate, not a fidelity test: #31 owns corpus scoring. +// +// APA-48: this gate previously never gated anything. In CI it ran BEFORE +// `Set up Python` / `uv sync`, so liteparse was not importable and it +// SKIPPED on every run — a green pipeline while every document terminal- +// FAILED in production. It now runs after Python setup, and with +// CLAIMOPS_REQUIRE_LITEPARSE=1 a missing runtime is a FAILURE. The +// assertions below cover the full evidence surface, because "the shim +// started" is not the same claim as "a usable Parser result came back". func TestLiveShimSmoke(t *testing.T) { py, shim := liveShim(t) - content := []byte("%PDF-1.4\nlive smoke\n" + string(make([]byte, 600))) - in, err := parser.NewTrustedDocument("doc-live", "t1", "c1", "live.pdf", - "application/pdf", "abc123", int64(len(content)), content) - if err != nil { - t.Fatal(err) - } - a := New(NewExecRunner(py, shim, 60*time.Second), 60*time.Second) - // CASE-001 is a real corpus PDF; use its bytes for a faithful smoke. - realPDF, err := os.ReadFile(filepath.Join("..", "..", "..", "..", "..", "fixtures", "parser_eval", "v1", "CASE-001", "document.pdf")) + realPDF, err := os.ReadFile(corpusCase001(t)) if err != nil { + requireLive(t, fmt.Sprintf("corpus PDF unavailable: %v", err)) t.Skipf("corpus PDF unavailable: %v", err) } - in2, err := parser.NewTrustedDocument("doc-live-1", "t1", "c1", "bill.pdf", + in, err := parser.NewTrustedDocument("doc-live-1", "t1", "c1", "bill.pdf", "application/pdf", "abc123", int64(len(realPDF)), realPDF) if err != nil { t.Fatal(err) } - _ = in - doc, err := a.Parse(context.Background(), in2) + a := New(NewExecRunner(py, shim, 60*time.Second), 60*time.Second) + + doc, err := a.Parse(context.Background(), in) if err != nil { t.Fatalf("live parse: %v", err) } if err := doc.Validate(); err != nil { t.Fatalf("live artifact invalid: %v", err) } - if len(doc.Pages) == 0 || len(doc.Pages[0].Blocks) == 0 { - t.Fatal("live artifact has no content") + + // --- metadata: the artifact must stay traceable to parser + input --- + if doc.Metadata.ParserName != AdapterName { + t.Errorf("ParserName = %q, want %q", doc.Metadata.ParserName, AdapterName) + } + if doc.Metadata.ParserVersion != AdapterVersion { + t.Errorf("ParserVersion = %q, want %q", doc.Metadata.ParserVersion, AdapterVersion) + } + if doc.Metadata.SourceSHA256 != "abc123" || doc.Metadata.SourceMedia != "application/pdf" { + t.Errorf("provenance metadata = %+v, want the trusted input echoed", doc.Metadata) + } + if doc.Metadata.PageCount != len(doc.Pages) { + t.Errorf("PageCount = %d, want %d", doc.Metadata.PageCount, len(doc.Pages)) + } + + // --- pages: present, and numbered 1..N in order --- + if len(doc.Pages) == 0 { + t.Fatal("live artifact has no pages") + } + + blocks, tables, cells := 0, 0, 0 + sawHeading := false + for _, page := range doc.Pages { + blocks += len(page.Blocks) + tables += len(page.Tables) + + // --- block IDs: unique, and sequential in reading order --- + // Sequential IDs are the evidence that the vendor's stream order + // survived the conversion instead of being reshuffled. + firstY := 1.0 + anyY := false + for i, b := range page.Blocks { + if want := fmt.Sprintf("b%d", i+1); b.ID != want { + t.Errorf("page %d block[%d].ID = %q, want %q (reading order not preserved)", + page.Number, i, b.ID, want) + } + if b.Evidence.BlockID != b.ID { + t.Errorf("page %d block %q evidence BlockID = %q, want it to match", + page.Number, b.ID, b.Evidence.BlockID) + } + switch b.Type { + case parser.BlockHeading: + sawHeading = true + case parser.BlockText, parser.BlockListItem, parser.BlockTableCell, parser.BlockFigure: + default: + t.Errorf("page %d block %q has unknown type %q", page.Number, b.ID, b.Type) + } + // --- bounding boxes: present and normalized, never fabricated --- + if b.Evidence.Box == nil { + t.Errorf("page %d block %q has no bounding box (provenance must not be dropped)", page.Number, b.ID) + continue + } + for name, v := range map[string]float64{ + "x0": b.Evidence.Box.X0, "y0": b.Evidence.Box.Y0, + "x1": b.Evidence.Box.X1, "y1": b.Evidence.Box.Y1, + } { + if v < 0 || v > 1 { + t.Errorf("page %d block %q box %s = %v, want within [0,1]", page.Number, b.ID, name, v) + } + } + if b.Evidence.Box.X1 < b.Evidence.Box.X0 || b.Evidence.Box.Y1 < b.Evidence.Box.Y0 { + t.Errorf("page %d block %q box is inverted: %+v", page.Number, b.ID, b.Evidence.Box) + } + if !anyY || b.Evidence.Box.Y0 < firstY { + firstY, anyY = b.Evidence.Box.Y0, true + } + } + // --- reading-order evidence: the stream STARTS at the page top --- + // Deliberately not a monotonic-Y assertion: observed vendor output + // is not strictly monotonic (multi-column), and asserting it would + // test a guarantee the parser never made. + if anyY && page.Blocks[0].Evidence.Box != nil { + if got := page.Blocks[0].Evidence.Box.Y0; got > firstY+1e-9 { + t.Errorf("page %d first block Y0 = %v, want the topmost block (Y0 %v)", page.Number, got, firstY) + } + } + + // --- tables: reconstructed, with cell-level evidence --- + for _, tb := range page.Tables { + if tb.ID == "" { + t.Errorf("page %d table has blank id", page.Number) + } + if tb.Page != page.Number { + t.Errorf("table %q page = %d, want %d", tb.ID, tb.Page, page.Number) + } + for _, row := range tb.Rows { + for _, c := range row.Cells { + cells++ + if c.Evidence.Box == nil { + t.Errorf("page %d table %q cell has no bounding box", page.Number, tb.ID) + } + if c.RowSpan < 1 || c.ColSpan < 1 { + t.Errorf("page %d table %q cell spans = %dx%d, want >= 1x1", + page.Number, tb.ID, c.RowSpan, c.ColSpan) + } + } + } + } + } + + if tables == 0 || cells == 0 { + t.Errorf("live artifact reconstructed %d table(s) / %d cell(s); CASE-001 is a hospital bill with a line-item table", tables, cells) + } + if !sawHeading { + t.Error("live artifact has no heading blocks (block-type mapping lost)") + } + if blocks == 0 { + t.Error("live artifact has no content blocks") + } + t.Logf("live: %d pages, %d blocks, %d tables, %d cells", + len(doc.Pages), blocks, tables, cells) +} + +// corpusCase001 is the eval-v1 corpus document the live smoke parses. A +// real document, not a synthetic one: the smoke must exercise the same +// bytes production sees. +func corpusCase001(t *testing.T) string { + t.Helper() + return filepath.Join("..", "..", "..", "..", "..", "fixtures", "parser_eval", "v1", "CASE-001", "document.pdf") +} + +// requireLive turns a skip into a failure when CI demands a real parser +// environment, so a misbuilt production artifact cannot pass green. +func requireLive(t *testing.T, msg string) { + t.Helper() + if os.Getenv(requireLiveEnv) == "1" { + t.Fatalf("%s: %s is set, so the parser environment must be real; the live shim gate must not skip in CI", requireLiveEnv, msg) } - t.Logf("live: %d pages, %d blocks, %d tables", - len(doc.Pages), len(doc.Pages[0].Blocks), len(doc.Pages[0].Tables)) } -// liveShim resolves the venv interpreter + shim path, skipping when the -// vendor library is unavailable. +// liveShim resolves a working interpreter + shim path. Skips when the +// optional local Python environment is absent (the only legitimate skip: +// the venv is a contributor convenience, not part of the Go build). Under +// CLAIMOPS_REQUIRE_LITEPARSE=1 a missing runtime fails instead — that is +// the mode CI uses, so an uninstalled liteparse can never be mistaken for +// a healthy pipeline. func liveShim(t *testing.T) (py, shim string) { t.Helper() root := filepath.Join("..", "..", "..", "..", "..") shim = filepath.Join(root, "tools", "parsers", "liteparse", "shim.py") if _, err := os.Stat(shim); err != nil { + requireLive(t, fmt.Sprintf("shim unavailable: %v", err)) t.Skipf("shim unavailable: %v", err) } for _, cand := range []string{filepath.Join(root, ".venv", "bin", "python"), "python3"} { @@ -66,6 +204,7 @@ func liveShim(t *testing.T) (py, shim string) { return cand, abs } } - t.Skip("liteparse not installed") + requireLive(t, "liteparse is not importable by any candidate interpreter") + t.Skip("liteparse not installed (set CLAIMOPS_REQUIRE_LITEPARSE=1 to make this a failure)") return "", "" } diff --git a/apps/api/internal/parser/liteparse/runtime_errors_test.go b/apps/api/internal/parser/liteparse/runtime_errors_test.go new file mode 100644 index 0000000..c34e550 --- /dev/null +++ b/apps/api/internal/parser/liteparse/runtime_errors_test.go @@ -0,0 +1,137 @@ +package liteparse + +import ( + "context" + "errors" + "os" + "path/filepath" + "testing" + "time" + + "claimops-api/internal/parser" +) + +// APA-48: a runtime/environment fault is NOT a bad document. +// +// The deployed worker image shipped without an interpreter and without the +// shim, so every ingested document died terminal-FAILED at parse. The +// envelope code for exactly that condition ("harness", shim.py:99/109/148) +// was mapped into parser.ErrParseFailure, which is the document-fault +// class: the worker failTerminal()s it, persists a FAILED document row, and +// blames the customer's file for our broken image. These tests pin the +// separation so the two classes cannot collapse back together. + +func TestHarnessEnvelopeIsRuntimeNotDocumentFault(t *testing.T) { + err := mapEnvelopeError("harness", "liteparse not installed: No module named 'liteparse'") + if !errors.Is(err, ErrRuntimeUnavailable) { + t.Fatalf("harness envelope = %v, want ErrRuntimeUnavailable", err) + } + if errors.Is(err, parser.ErrParseFailure) { + t.Fatalf("harness envelope = %v, must NOT be classified as a document fault", err) + } + if errors.Is(err, parser.ErrUnsupportedMediaType) { + t.Fatalf("harness envelope = %v, must NOT be classified as a media fault", err) + } +} + +// A genuine document fault must keep its document classification. This +// guards against over-correcting every parse error into "runtime". +func TestDocumentFaultsKeepDocumentClassification(t *testing.T) { + cases := map[string]error{ + "parse_failure": parser.ErrParseFailure, + "timeout": parser.ErrParseFailure, + "unsupported_media": parser.ErrUnsupportedMediaType, + "weird": parser.ErrParseFailure, + } + for code, want := range cases { + err := mapEnvelopeError(code, "x") + if !errors.Is(err, want) { + t.Fatalf("%s = %v, want %v", code, err, want) + } + if errors.Is(err, ErrRuntimeUnavailable) { + t.Fatalf("%s = %v, must not be a runtime fault", code, err) + } + } +} + +// The exact production failure: DefaultPythonBin cannot be resolved +// because the image has no interpreter. exec.CommandContext fails at +// Start with exec.ErrNotFound. +func TestMissingInterpreterIsRuntimeFault(t *testing.T) { + shim := filepath.Join(t.TempDir(), "shim.py") + if err := os.WriteFile(shim, []byte("import sys\n"), 0o600); err != nil { + t.Fatal(err) + } + r := NewExecRunner("python3-not-installed-in-image", shim, 10*time.Second) + + _, err := r.Run(context.Background(), shim) + if err == nil { + t.Fatal("Run with a missing interpreter = nil, want error") + } + if !errors.Is(err, ErrRuntimeUnavailable) { + t.Fatalf("missing interpreter = %v, want ErrRuntimeUnavailable", err) + } + if errors.Is(err, parser.ErrParseFailure) { + t.Fatalf("missing interpreter = %v, must NOT be a document fault", err) + } +} + +// A missing shim is the same class of fault: the image is misbuilt, the +// document is fine. Detected before spawning so it cannot be misread as a +// shim crash (which would surface as a document fault via mapExitError). +func TestMissingShimIsRuntimeFault(t *testing.T) { + missing := filepath.Join(t.TempDir(), "not-shipped.py") + r := NewExecRunner("python3", missing, 10*time.Second) + + _, err := r.Run(context.Background(), filepath.Join(t.TempDir(), "doc.pdf")) + if err == nil { + t.Fatal("Run with a missing shim = nil, want error") + } + if !errors.Is(err, ErrRuntimeUnavailable) { + t.Fatalf("missing shim = %v, want ErrRuntimeUnavailable", err) + } + if errors.Is(err, parser.ErrParseFailure) { + t.Fatalf("missing shim = %v, must NOT be a document fault", err) + } +} + +// A shim that runs and reports a runtime fault (e.g. the dependency import +// blew up inside the subprocess) must also surface as runtime, never as a +// bad document. +func TestHarnessEnvelopeFromRunnerIsRuntimeFault(t *testing.T) { + dir := t.TempDir() + shim := filepath.Join(dir, "shim.py") + // Stands in for a real interpreter that cannot import the vendor lib. + body := "import json,sys\nprint(json.dumps({'ok':False,'error':{'code':'harness','message':'liteparse not installed'}}))\n" + if err := os.WriteFile(shim, []byte(body), 0o600); err != nil { + t.Fatal(err) + } + runner := NewExecRunner(pythonForTest(t), shim, 10*time.Second) + + raw, err := runner.Run(context.Background(), shim) + if err != nil { + t.Fatalf("Run = %v, want nil (exit 0 + envelope is the contract)", err) + } + _, err = Convert("doc-1", "sha", "application/pdf", raw) + if !errors.Is(err, ErrRuntimeUnavailable) { + t.Fatalf("Convert(harness envelope) = %v, want ErrRuntimeUnavailable", err) + } + if errors.Is(err, parser.ErrParseFailure) { + t.Fatalf("Convert(harness envelope) = %v, must NOT be a document fault", err) + } +} + +// pythonForTest resolves an interpreter for tests that actually spawn one. +func pythonForTest(t *testing.T) string { + t.Helper() + if py := os.Getenv("CLAIMOPS_TEST_PYTHON"); py != "" { + return py + } + for _, cand := range []string{"python3", "python"} { + if _, err := os.Stat(filepath.Join("/usr/bin", cand)); err == nil { + return cand + } + } + t.Skip("no interpreter available for subprocess test") + return "" +} diff --git a/apps/api/internal/worker/parser_runtime_test.go b/apps/api/internal/worker/parser_runtime_test.go new file mode 100644 index 0000000..42eb674 --- /dev/null +++ b/apps/api/internal/worker/parser_runtime_test.go @@ -0,0 +1,111 @@ +package worker + +import ( + "context" + "errors" + "fmt" + "testing" + + "claimops-api/internal/documents" + "claimops-api/internal/parser" + "claimops-api/internal/parser/liteparse" +) + +// APA-48: harness faults must not become failTerminal. +// +// Before this change a missing interpreter produced parser.ErrParseFailure, +// so processor.go:824 failTerminal()'d the document: it persisted a FAILED +// document row and acknowledged the message, permanently recording a good +// customer's file as unparseable because our own image was misbuilt. The +// correct routing reuses the EXISTING OutcomeTransient kind (the same +// non-2xx/Nack signal used for cancellation, store and load faults): the +// document is fine, the environment is not, and redelivery may heal it. +// +// fail closed, never open: still no extraction, no evidence, no SUCCESS. + +// errParser is a parser stub that always fails with the given error. +type errParser struct{ err error } + +func (p errParser) Name() string { return liteparse.AdapterName } +func (p errParser) Version() string { return liteparse.AdapterVersion } +func (p errParser) Parse(context.Context, parser.TrustedDocument) (parser.ParsedDocument, error) { + return parser.ParsedDocument{}, p.err +} + +func TestParserRuntimeFaultIsTransientNotBadDocument(t *testing.T) { + fetch, store, loader, checker := happyFixture() + p := NewProcessor(fetch, store, loader, checker) + // Exactly the production shape: the adapter wrapped a harness fault. + p.Parser = errParser{err: fmt.Errorf("liteparse shim harness: %w", liteparse.ErrRuntimeUnavailable)} + + out := p.Handle(context.Background(), goodEvent(t, "doc-runtime")) + + if out.Kind != OutcomeTransient { + t.Fatalf("runtime fault kind = %q, want %q (not failTerminal/bad-document)", + out.Kind, OutcomeTransient) + } + if out.Status != documents.StFailed { + t.Fatalf("runtime fault status = %q, want %q", out.Status, documents.StFailed) + } + if out.Extraction != documents.ExtractionNotAttempted { + t.Fatalf("runtime fault extraction = %q, want %q (fail closed)", + out.Extraction, documents.ExtractionNotAttempted) + } + if out.Err == nil { + t.Fatal("runtime fault Err = nil, want the cause surfaced") + } + if !errors.Is(out.Err, liteparse.ErrRuntimeUnavailable) { + t.Fatalf("runtime fault Err = %v, want it to wrap ErrRuntimeUnavailable", out.Err) + } + + // The critical assertion: the environment fault must not be recorded as + // a defective document. A FAILED row is a durable statement about the + // customer's file; there is no such statement here. + store.mu.Lock() + failedRows := 0 + for _, d := range store.docs { + if d.ID == "doc-runtime" { + failedRows++ + } + } + evRows := len(store.ev) + store.mu.Unlock() + if failedRows != 0 { + t.Fatalf("persisted %d document row(s) for doc-runtime; a runtime fault must not blame the document", failedRows) + } + if evRows != 0 { + t.Fatalf("persisted %d evidence row(s); a failed parse must produce no evidence", evRows) + } +} + +// Guard against over-correction: a real document fault stays terminal and +// still records the FAILED row, so bad documents are not silently retried +// forever. +func TestDocumentParseFaultStaysTerminal(t *testing.T) { + fetch, store, loader, checker := happyFixture() + p := NewProcessor(fetch, store, loader, checker) + p.Parser = errParser{err: fmt.Errorf("liteparse parse failed: %w", parser.ErrParseFailure)} + + out := p.Handle(context.Background(), goodEvent(t, "doc-badfile")) + + if out.Kind != OutcomeTerminal { + t.Fatalf("document fault kind = %q, want %q", out.Kind, OutcomeTerminal) + } + if out.Extraction != documents.ExtractionNotAttempted { + t.Fatalf("document fault extraction = %q, want %q", out.Extraction, documents.ExtractionNotAttempted) + } + if !errors.Is(out.Err, parser.ErrParseFailure) { + t.Fatalf("document fault Err = %v, want it to wrap ErrParseFailure", out.Err) + } + store.mu.Lock() + failedRows := 0 + for _, d := range store.docs { + if d.ID == "doc-badfile" { + failedRows++ + } + } + store.mu.Unlock() + if failedRows != 1 { + t.Fatalf("persisted %d FAILED row(s) for a bad document, want 1", failedRows) + } +} diff --git a/apps/api/internal/worker/processor.go b/apps/api/internal/worker/processor.go index 17c154e..3fc96e8 100644 --- a/apps/api/internal/worker/processor.go +++ b/apps/api/internal/worker/processor.go @@ -51,6 +51,7 @@ import ( "claimops-api/internal/metrics" "claimops-api/internal/observability" "claimops-api/internal/parser" + "claimops-api/internal/parser/liteparse" "claimops-api/internal/ports" "claimops-api/internal/sufficiency" "claimops-api/internal/verify" @@ -284,10 +285,24 @@ type documentIngestedEvent struct { // (ErrCorruptedBlob / ErrBlobTooLarge from the #16B integrity boundary). // errors.As/Is is the seam: fetch adapters alias and wrap these sentinels, // so detection survives the ContentFetcher interface boundary. +// isPermanent reports whether a fetch failure is a permanent integrity fault +// (TERMINAL) rather than a transport fault worth redelivery (TRANSIENT). func isPermanent(err error) bool { return errors.Is(err, ErrCorruptedBlob) || errors.Is(err, ErrBlobTooLarge) } +// isParserRuntimeErr reports whether a parse failure is a broken parser +// EXECUTION ENVIRONMENT (APA-48) rather than a defect in the document: a +// missing interpreter, a missing shim, or an uninstallable vendor +// dependency. These are our fault, not the customer's, and the document +// stays unproven — so the caller must NOT failTerminal (which acknowledges +// the delivery and persists a FAILED document row). It routes to the +// existing OutcomeTransient so redelivery can heal it once the artifact is +// rebuilt. Sibling of isPermanent: same shape, opposite verdict. +func isParserRuntimeErr(err error) bool { + return errors.Is(err, liteparse.ErrRuntimeUnavailable) +} + // retryableStoreErr reports whether a store/dependency error merits another // attempt (S3/F8). Integrity-constraint violations (pgconn class 23) are // data faults retry cannot heal: the repository layer already converts real @@ -821,6 +836,23 @@ func (p *Processor) runNewPipeline(ctx context.Context, tenant, claimID, blobKey if errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded) { return Outcome{DocumentID: docID, Status: documents.StFailed, Extraction: documents.ExtractionNotAttempted, Kind: OutcomeTransient, Attempts: 1, Err: err}, effective } + // APA-48: the parser's execution environment is broken (interpreter + // or shim missing, vendor dependency not importable). The document + // is unproven, not bad. Reuse the existing OutcomeTransient signal — + // non-2xx/Nack, so redelivery may heal it — instead of failTerminal, + // which would acknowledge the message AND persist a FAILED document + // row, durably blaming the customer's file for our misbuilt image. + // Fails closed: no extraction, no evidence, no SUCCESS. + if isParserRuntimeErr(err) { + return Outcome{ + DocumentID: docID, + Status: documents.StFailed, + Extraction: documents.ExtractionNotAttempted, + Kind: OutcomeTransient, + Attempts: 1, + Err: fmt.Errorf("worker: parse document: %w", err), + }, effective + } return failTerminal(documents.DocType(effective), fmt.Errorf("worker: parse document: %w", err)) } validateErr := parsed.Validate() diff --git a/tools/parsers/liteparse/container_smoke.sh b/tools/parsers/liteparse/container_smoke.sh new file mode 100755 index 0000000..cfedfac --- /dev/null +++ b/tools/parsers/liteparse/container_smoke.sh @@ -0,0 +1,168 @@ +#!/bin/sh +# Container-level parse smoke for the ClaimOps worker image (APA-48). +# +# Runs INSIDE the built worker image and proves the deployed artifact can +# actually execute the LiteParse shim, rather than proving it on the host +# where a developer venv hides a broken image. The image under test is the +# only place this truth lives: a worker that ships without an interpreter +# terminal-FAILs every document at parse time, and no host test can see it. +# +# Usage (from the repo root): +# docker run --rm \ +# -v "$PWD:/harness:ro" \ +# -v "$PWD/fixtures:/corpus:ro" \ +# --entrypoint sh claimops-worker:p0 \ +# /harness/tools/parsers/liteparse/container_smoke.sh \ +# /corpus/parser_eval/v1/CASE-001/document.pdf +# +# The corpus PDF is MOUNTED, never baked into the image: the production +# artifact must not carry test fixtures. Only the shim is expected to be +# present in the image itself, at the path internal/parser/liteparse +# resolves from WORKDIR (DefaultShimPath). +# +# Exit 0 = the parse path runs in this artifact. Non-zero at the FIRST +# failed gate, naming the gate, so a red build says which link broke. + +set -eu + +# Default to the path the Go adapter expects: DefaultShimPath resolved +# against WORKDIR /app. +WORKDIR="${WORKDIR:-/app}" +REL_SHIM="${REL_SHIM:-tools/parsers/liteparse/shim.py}" +SHIM_PATH="${SHIM_PATH:-$WORKDIR/$REL_SHIM}" +PDF_PATH="${1:-}" + +fail() { + echo "SMOKE FAIL [$1] $2" >&2 + exit 1 +} + +pass() { + echo " ok: $1" +} + +echo "== ClaimOps worker image parse smoke ==" +echo "shim: $SHIM_PATH" +echo "corpus: $PDF_PATH" + +# --- Gate 1: the shim is present at the path the Go code resolves ------- +[ -f "$SHIM_PATH" ] || fail shim-present "shim MISSING at $SHIM_PATH (Go resolves DefaultShimPath from WORKDIR /app; image ships no tools/)" +pass "shim present at $SHIM_PATH" + +# --- Gate 2: an interpreter is on PATH --------------------------------- +# internal/parser/liteparse DefaultPythonBin is "python3"; ExecRunner +# resolves it through PATH. Alpine without python3 exits 127 here, which is +# exactly the cmd.Start() failure that failTerminal()s every document. +command -v python3 >/dev/null 2>&1 || fail python-present "python3 NOT on PATH (DefaultPythonBin unresolvable -> every parse terminal-FAILs)" +pass "python3 on PATH: $(command -v python3)" + +# --- Gate 3: the interpreter runs the shim the way production does ------ +# Production (internal/parser/liteparse client.go) execs the shim by its +# DefaultShimPath RELATIVE to the image WORKDIR, not by absolute path: +# exec.CommandContext(ctx, "python3", "tools/parsers/liteparse/shim.py", tmp) +# So the relative resolution is exercised here too; an absolute-path-only +# check would miss a WORKDIR/image mismatch. +OUT="${TMPDIR:-/tmp}/shim-out.$$" +ERR="${TMPDIR:-/tmp}/shim-err.$$" +if ! (cd "$WORKDIR" && python3 "$REL_SHIM" "$PDF_PATH") >"$OUT" 2>"$ERR"; then + fail shim-executes "shim exited non-zero from WORKDIR=$WORKDIR: $(head -c 300 "$ERR")" +fi +pass "shim executed (exit 0) via the production relative path '$REL_SHIM' from WORKDIR $WORKDIR" + +# --- Gates 4-8: a real Parser payload with intact evidence ------------- +# Validated with the image's own python3 so this stays a single-artifact +# check. Each gate names one evidence class that must survive the shim. +python3 - "$OUT" <<'PYEOF' +import json, sys + +raw = open(sys.argv[1]).read() +env = json.loads(raw) +if not env.get("ok"): + code = (env.get("error") or {}).get("code") + msg = (env.get("error") or {}).get("message") + if code == "harness": + print("SMOKE FAIL [shim-runtime] shim reported harness fault: %s" % msg, file=sys.stderr) + else: + print("SMOKE FAIL [shim-ok] envelope ok:false code=%s: %s" % (code, msg), file=sys.stderr) + sys.exit(1) + +payload = env["payload"] +pages = payload.get("pages") or [] +if not pages: + print("SMOKE FAIL [pages] shim returned zero pages", file=sys.stderr) + sys.exit(1) + +total_blocks = 0 +total_tables = 0 +total_cells = 0 +kinds = set() +reading_order_ok = True +first_block_top = None +min_y = None + +for p in pages: + w, h = p["width"], p["height"] + if not w or not h: + print("SMOKE FAIL [page-dims] page %s has degenerate dims %sx%s" % (p["page_num"], w, h), file=sys.stderr) + sys.exit(1) + blocks = p.get("blocks") + if blocks is None: + print("SMOKE FAIL [blocks] page %s emitted no blocks" % p["page_num"], file=sys.stderr) + sys.exit(1) + for b in blocks: + total_blocks += 1 + kinds.add(b["kind"]) + if b["kind"] == "table": + total_tables += 1 + rows = ([b["header"]] if b.get("header") else []) + (b.get("rows") or []) + for r in rows: + for c in r: + total_cells += 1 + if not c.get("bbox"): + print("SMOKE FAIL [bbox] table cell has no bbox", file=sys.stderr) + sys.exit(1) + continue + bb = b.get("bbox") + if not bb: + print("SMOKE FAIL [bbox] block %r has no bbox (provenance must never be fabricated)" % b["kind"], file=sys.stderr) + sys.exit(1) + y0 = bb["y"] / h + if not (0.0 <= bb["x"] / w <= 1.0 and 0.0 <= y0 <= 1.0 and 0.0 <= (bb["y"] + bb["height"]) / h <= 1.0): + print("SMOKE FAIL [bbox] block %r bbox outside page after normalization" % b["kind"], file=sys.stderr) + sys.exit(1) + if min_y is None or y0 < min_y: + min_y = y0 + if first_block_top is None: + first_block_top = y0 + +# Reading-order evidence: the first block emitted must be the topmost +# block on the page. NOTE this is deliberately NOT a monotonic-Y check -- +# observed vendor output is not strictly monotonic (multi-column), and +# inventing that property would assert a guarantee the parser never made. +if first_block_top is not None and min_y is not None and first_block_top > min_y + 1e-9: + reading_order_ok = False + +print(" ok: pages=%d blocks=%d tables=%d cells=%d kinds=%s" % ( + len(pages), total_blocks, total_tables, total_cells, sorted(kinds))) +print(" ok: bounding boxes normalized to [0,1] on every block and table cell") + +if not total_tables: + print("SMOKE FAIL [tables] no table reconstructed (CASE-001 is a hospital bill)", file=sys.stderr) + sys.exit(1) +if not total_cells: + print("SMOKE FAIL [tables] tables present but no cells", file=sys.stderr) + sys.exit(1) +print(" ok: table structure intact (%d cells)" % total_cells) + +if "heading" not in kinds: + print("SMOKE FAIL [block-types] no heading blocks (block-type mapping lost)", file=sys.stderr) + sys.exit(1) +print(" ok: block types preserved: %s" % sorted(kinds)) + +if not reading_order_ok: + print("SMOKE FAIL [reading-order] first block is not the topmost block", file=sys.stderr) + sys.exit(1) +print(" ok: reading-order evidence intact (stream starts at page top)") +PYEOF + +echo "SMOKE PASS: worker image can parse a real corpus document"