Build a local-only, TypeScript-first PDF OCR foundation that renders PDF pages sequentially, reuses the existing OCR service, exposes structured per-page results and real progress, and preserves cancellation, stale-result protection, retry, and resource cleanup.
Scope:
- audit and reuse existing PDF loading/rendering and OCR infrastructure
- bounded page selection and deterministic OCR rendering
- typed page/result/progress contracts
- cancellation and source-replacement safety
- sequential resource-conscious processing
- regression, privacy, CSP, and browser integration coverage
- documentation for the future Searchable PDF Sprint
Explicitly excluded:
- searchable text-layer PDF generation
- cloud OCR or external OCR APIs
- AI correction
- handwriting guarantees
- additional OCR recognition languages
Build a local-only, TypeScript-first PDF OCR foundation that renders PDF pages sequentially, reuses the existing OCR service, exposes structured per-page results and real progress, and preserves cancellation, stale-result protection, retry, and resource cleanup.
Scope:
Explicitly excluded: