Skip to content

feat(core): hyphenate every language whose hyph-utf8 patterns we may carry - #619

Merged
yuroyami merged 1 commit into
mainfrom
claude/hopeful-dirac-4b5wda
Oct 6, 2026
Merged

yuroyami merged 1 commit into
mainfrom
claude/hopeful-dirac-4b5wda

Conversation

@yuroyami

@yuroyami yuroyami commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Resolves #207, and four defects found on the way (#615, #616, #617, #618).

What changes

Data (#207). tools/generate_hyphenation.py writes one Kotlin object for each hyph-utf8 set from a checkout of hyphenation/tex-hyphen. It carries the patterns and, for the first time, the upstream exception lists, both unmodified, under a header that quotes the copyright and the licence text and names the upstream commit. That gives 71 sets for more than 60 languages, up from 7. The seven sets bundled before regenerate with identical patterns.

Left out, with the reason recorded in the script: GPL-only or LGPL-only sets (Czech, Macedonian, Indonesian, upstream Serbian, Armenian, Latvian), sets with no stated licence (Romanian, mn-cyrl-x-lmc), empty sets (Arabic, Persian, Hebrew, Vietnamese), and Thai and Ethiopic, which break lines with no visible hyphen. A set that upstream adds later fails the script until someone decides on it. Serbian is still covered by the LPPL sh-cyrl and sh-latn sets.

Selection. English uses the full hyph-en-us set instead of about 60 hand-picked patterns. forLanguage now reads the rest of a BCP 47 tag:

  • British-spelling regions get hyph-en-gb.
  • 1901 picks traditional German, with the Swiss set for CH and LI.
  • polyton picks polytonic Greek.
  • The script picks Latin or Cyrillic Serbian, and Serbian with no script gets both sets, which share no letters.
  • x-classic, x-liturgic and x-school pick the private-use sets.

Only the sets a book asks for are loaded.

Memory. The trie now lives in flat arrays with sorted edges. Measured retained heap on the JVM: Hungarian (63k patterns) goes from about 15 MB with a map per node to 1.7 MB, German from 9.3 MB to 1.1 MB. A test keeps Hungarian under 4 MB.

Fixes

Tests

  • HyphenationGolden holds words for every set, real ones plus pseudo-words built from pattern letters taken from across each file. The expected breaks come from the script's independent implementation, which uses a dictionary lookup over the upstream files, not the trie or the Kotlin strings.
  • Mutation checks: dropping one Hungarian chunk fails the golden test. Reverting to lowercase(), removing the mark filter, or restoring the old EPUB word check and English fallback each fail their tests.
  • The EPUB tests that compare German against English now sweep a range of page widths. At a single width the English side passed without proving anything.

Gate

  • All 11 JVM suites of the CONTRIBUTING gate pass: core 450, pdf 853, epub 917, native-renderer 221, skia 31, difftest 6, kitepdf 32, media 48, net 18, compose-viewer 469, javascript 205. No failures.
  • The EPUB sweep passed with every book's page count unchanged.
  • This Linux host has no mutool, so the PDF differential ran as a KitePDF-only smoke pass: 58 pages, 0 render failures, no mean-error number. Nothing here touches PDF rendering.
  • checkKotlinAbi could not run here (no Android SDK). The public JVM surface of Hyphenator matches the committed dump member for member. The new constructor and trieBytes are internal.

Notes for review

  • Source size: the pattern data grows from about 0.5 MB to 3.1 MB of string constants in kitepdf-core. The largest sets are Hungarian (528 KB), the three German sets (about 270 KB each) and Classical Latin (216 KB). If that is too much for the JS or iOS bundles, the rarely selected sets (de-1901, de-ch-1901, la-x-classic, cu) are the first candidates to move into an optional module.
  • CONTRIBUTING says to work on main; this session was told to push to a branch, so this arrives as a PR.

Generated by Claude Code

…carry

tools/generate_hyphenation.py writes one Kotlin object for each set of the
hyph-utf8 project from a checkout of it: the patterns and, new, the exception
lists, carried unmodified under a header that quotes the copyright and the
licence and names the upstream commit. 71 sets cover more than 60 languages.
The script lists the sets it leaves out and why: the GPL-only and LGPL-only
ones (Czech, Macedonian, Indonesian, upstream Serbian, Armenian, Latvian),
the ones with no licence (Romanian), the empty ones, and Thai and Ethiopic,
which break a line with no visible hyphen. A set that upstream adds fails the
script until it is decided. The seven sets bundled before come out with the
same patterns.

English uses the full hyph-en-us set instead of about sixty patterns, and
forLanguage reads the rest of a BCP 47 tag: a British-spelling region picks
hyph-en-gb, the variant 1901 traditional German, polyton polytonic Greek,
the script Latin or Cyrillic Serbian, and Serbian with no script gets both.

The trie now lives in flat arrays with sorted edges. A map for each node held
Hungarian's 63k patterns in about 15 MB of heap and German's in 9 MB; the
arrays hold them in 1.7 MB and 1.1 MB.

HyphenationGolden holds words of every set with the breaks that the script's
own implementation finds, which looks patterns up in a dictionary read from
the upstream files, so neither the trie nor the packing can confirm itself.
Dropping one chunk of the Hungarian set fails it.

EPUB layout gives English breaks only to a chapter with no language or an
English one; a language with no bundled set is not hyphenated. A word's
leading and trailing punctuation no longer stops it hyphenating, and an
apostrophe, a joiner or a combining mark may stand inside it. The hyphenator
lower-cases one character for one, so a Turkish İ no longer shifts the
breaks after it, and puts no break before a combining mark.

Fixes #207
Fixes #615
Fixes #616
Fixes #617
Fixes #618
@yuroyami
yuroyami marked this pull request as ready for review October 6, 2026 21:59
@yuroyami
yuroyami merged commit 07e73d9 into main Oct 6, 2026
6 checks passed
@yuroyami
yuroyami deleted the claude/hopeful-dirac-4b5wda branch October 7, 2026 04:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Hyphenation covers a limited set of languages

2 participants