Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ The root [README](../README.md) introduces Secure Tools. This directory owns det
| [Privacy model](./privacy-model.md) | Local-processing and network boundaries, storage, security controls, and bounded privacy claims |
| [Dependencies](./dependencies.md) | Production runtime inventory, versions, vendoring, licenses, and integrity ownership |
| [Local OCR foundation](./ocr-foundation.md) | Self-hosted Tesseract assets, languages, lifecycle, cancellation, caching, and privacy guarantees |
| [PDF OCR foundation](./pdf-ocr-foundation.md) | Typed local page rendering, OCR sequencing, progress, cancellation, and resource boundaries |
| [Sprint 16B Image → Text QA](./sprint-16b-qa.md) | Automated and Chromium browser evidence for the v2.1.0 Image → Text workflow |
| [Sprint 16C v2.1.0 release hardening](./sprint-16c-v2.1-release-hardening.md) | Release-candidate regression, OCR, privacy, browser, performance, and readiness evidence |
| [v2.1.0 release notes draft](./v2.1.0-release-notes-draft.md) | Unpublished release-note copy for the later promotion and release task |
Expand Down
1 change: 1 addition & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ The current production inventory and tool-specific behavior live in [tool status
- `js/config.js` centralizes repository links.
- `tools/shared/` owns common file admission, signature validation, image/PDF helpers, queue conventions, local save behavior, and shared tool presentation.
- `tools/shared/ocr.js` owns language selection, same-origin OCR paths, normalized progress, orientation-aware image preparation, worker reuse, cancellation, and disposal for the public Image → Text workflow.
- `tools/shared/pdf-ocr.ts` composes the existing PDF.js renderer and OCR service into a sequential, cancellable per-page text pipeline for future PDF OCR interfaces. It does not generate searchable PDFs.
- The File System Access API is used when available; a revoking Blob-download fallback serves other browsers.

Tool implementations retain specialized models when their workflows differ. Organizer uses a page grid and PDF rendering lifecycle; Metadata tools use bounded inspection models and fail-closed output verification. Shared UI does not erase these tool-specific guarantees.
Expand Down
21 changes: 21 additions & 0 deletions docs/pdf-ocr-foundation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# PDF OCR foundation

The PDF OCR foundation is a shared, strict TypeScript pipeline for future PDF-to-text and Searchable PDF workflows. It is not currently exposed as a public tool route. Searchable PDF generation remains a separate Sprint.

## Pipeline

`tools/shared/pdf-ocr.ts` loads a local PDF with the existing same-origin PDF.js renderer, renders each selected page to a white-background PNG canvas, passes that page image to the existing Tesseract OCR service, and returns an ordered array of `{ pageNumber, text }` results. The pipeline preserves English, Korean, and English + Korean recognition only.

Pages are processed sequentially with one OCR service. A render scale of `2` corresponds to approximately 144 pixels per PDF inch and balances OCR detail with browser memory. The existing 16,384-pixel dimension and 50-megapixel canvas safeguards remain in force. Documents with 25 or more pages surface a `largeDocument` signal rather than receiving an arbitrary hard rejection.

## Lifecycle and progress

Progress reports document loading, page rendering, real OCR stages, completed pages, and overall progress derived from actual OCR values. Unknown OCR progress remains indeterminate. Callers may select every page or an explicit ordered list of one-based page numbers.

Cancellation destroys the active PDF renderer and forwards an abort signal to OCR. Generation checks prevent late progress and page results after cancellation or source replacement. Every page proxy is cleaned, every canvas is reset to 1 × 1, the renderer is destroyed after each job, and OCR disposal terminates its worker. No page image, text, or PDF is uploaded, and no object URL is required.

The core exposes review-ready page boundaries but no searchable PDF, positioned text layer, geometry mapping, or PDF output. A future public UI can add range parsing, review, copy, and TXT export around these typed contracts without changing the local processing boundary.

## Verification

Unit tests cover page selection, page numbering, all three OCR language contracts, progress, invalid input, retry, cancellation during rendering and recognition, source replacement, stale callbacks, large-document signaling, and resource cleanup. The internal browser harness performs real one-page PDF rendering and English OCR, real multi-page and larger-page rendering, cancellation timing, and same-origin request checks. Existing OCR smoke continues to execute real English, Korean, and combined recognition.
2 changes: 1 addition & 1 deletion scripts/stage-typescript-test-dependencies.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,6 @@ const root = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..");
const stagingDirectory = path.join(root, ".ts-build", "tools", "shared");

fs.mkdirSync(stagingDirectory, { recursive: true });
for (const file of ["image.js", "ocr.js", "save.js"]) {
for (const file of ["image.js", "ocr.js", "pdf.js", "save.js"]) {
fs.copyFileSync(path.join(root, "tools", "shared", file), path.join(stagingDirectory, file));
}
5 changes: 5 additions & 0 deletions scripts/typescript-modules.mjs
Original file line number Diff line number Diff line change
@@ -1,4 +1,9 @@
export const compiledBrowserModules = Object.freeze([
Object.freeze({
source: "tools/shared/pdf-ocr.ts",
compiled: "tools/shared/pdf-ocr.js",
public: "shared/pdf-ocr.js",
}),
Object.freeze({
source: "tools/image/to-text/controller.ts",
compiled: "tools/image/to-text/controller.js",
Expand Down
18 changes: 18 additions & 0 deletions tests/browser/pdf-ocr-smoke.html
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta http-equiv="Content-Security-Policy" content="default-src 'self'; script-src 'self'; style-src 'self'; img-src 'self' blob: data:; connect-src 'none'; worker-src 'self'; object-src 'none'; frame-src 'none'; base-uri 'self'; form-action 'self'">
<title>Secure Tools local PDF OCR smoke test</title>
</head>
<body>
<main>
<h1>Local PDF OCR smoke test</h1>
<output id="result">Running…</output>
</main>
<script src="../../assets/vendor/pdf-lib/pdf-lib.min.js"></script>
<script src="../../assets/vendor/tesseract/engine/tesseract.min.js"></script>
<script type="module" src="./pdf-ocr-smoke.js"></script>
</body>
</html>
93 changes: 93 additions & 0 deletions tests/browser/pdf-ocr-smoke.js
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
import { createPdfOcrService } from "/shared/pdf-ocr.js";

const output = document.querySelector("#result");
const requestsBefore = performance.getEntriesByType("resource").map((entry) => entry.name);

async function createPdf(labels, dimensions = [612, 792]) {
const document = await window.PDFLib.PDFDocument.create();
const font = await document.embedFont(window.PDFLib.StandardFonts.HelveticaBold);
for (const label of labels) {
const page = document.addPage(dimensions);
page.drawText(label, { x: 90, y: dimensions[1] / 2, size: 88, font, color: window.PDFLib.rgb(0, 0, 0) });
}
const bytes = await document.save();
return bytes.buffer.slice(bytes.byteOffset, bytes.byteOffset + bytes.byteLength);
}

async function runRealOcr() {
const service = createPdfOcrService();
const started = performance.now();
try {
const result = await service.recognizeDocument({ sourceBytes: await createPdf(["HELLO"]), language: "eng" });
if (result.pages.length !== 1 || result.pages[0].pageNumber !== 1 || !/HELLO/i.test(result.pages[0].text)) {
throw new Error(`Unexpected PDF OCR result: ${JSON.stringify(result.pages)}`);
}
return { milliseconds: Math.round(performance.now() - started), text: result.pages[0].text.trim() };
} finally {
await service.dispose();
}
}

async function runRenderingSanity(labels, dimensions) {
let calls = 0;
const service = createPdfOcrService({
ocrService: {
async recognizeImage(image, options) {
calls += 1;
if (image.type !== "image/png") throw new Error(`Unexpected rendered type: ${image.type}`);
options.onProgress?.({ stage: "recognizing", progress: 1 });
return { text: `page ${calls}` };
},
async dispose() {},
},
});
const started = performance.now();
try {
const result = await service.recognizeDocument({ sourceBytes: await createPdf(labels, dimensions), language: "eng" });
if (result.pages.length !== labels.length) throw new Error(`Expected ${labels.length} pages, received ${result.pages.length}`);
return { milliseconds: Math.round(performance.now() - started), pages: result.pages.length };
} finally {
await service.dispose();
}
}

async function runCancellationSanity() {
let recognizing;
const reachedRecognition = new Promise((resolve) => { recognizing = resolve; });
const service = createPdfOcrService({
ocrService: {
recognizeImage(_image, options) {
recognizing();
return new Promise((_resolve, reject) => {
options.signal.addEventListener("abort", () => reject(Object.assign(new Error("OCR_CANCELLED"), { code: "OCR_CANCELLED" })), { once: true });
});
},
async dispose() {},
},
});
const task = service.recognizeDocument({ sourceBytes: await createPdf(["CANCEL"]), language: "eng" });
await reachedRecognition;
const started = performance.now();
await service.cancel();
let code = null;
try { await task; } catch (error) { code = error.code; }
await service.dispose();
if (code !== "PDF_OCR_CANCELLED") throw new Error(`Unexpected cancellation result: ${code}`);
return { milliseconds: Math.round(performance.now() - started) };
}

try {
const onePage = await runRealOcr();
const multiPage = await runRenderingSanity(["ONE", "TWO", "THREE"], [612, 792]);
const largerPage = await runRenderingSanity(["LARGE"], [900, 1200]);
const cancellation = await runCancellationSanity();
const requests = performance.getEntriesByType("resource").map((entry) => entry.name).slice(requestsBefore.length);
const externalRequests = requests.filter((value) => new URL(value).origin !== location.origin);
if (externalRequests.length) throw new Error(`External requests: ${externalRequests.join(", ")}`);
window.__pdfOcrSmokeResult = { ok: true, onePage, multiPage, largerPage, cancellation, requests, externalRequests };
output.textContent = `PASS: PDF OCR — ${onePage.text} | 1 page ${onePage.milliseconds} ms | 3 pages ${multiPage.milliseconds} ms | large page ${largerPage.milliseconds} ms | cancel ${cancellation.milliseconds} ms`;
} catch (error) {
window.__pdfOcrSmokeResult = { ok: false, error: error.message, cause: error.cause?.message || null };
output.textContent = `FAIL: ${error.message}`;
throw error;
}
Loading
Loading