Repository navigation
feat(core): hyphenate every language whose hyph-utf8 patterns we may carry - #619
Merged
Merged
Conversation
…carry tools/generate_hyphenation.py writes one Kotlin object for each set of the hyph-utf8 project from a checkout of it: the patterns and, new, the exception lists, carried unmodified under a header that quotes the copyright and the licence and names the upstream commit. 71 sets cover more than 60 languages. The script lists the sets it leaves out and why: the GPL-only and LGPL-only ones (Czech, Macedonian, Indonesian, upstream Serbian, Armenian, Latvian), the ones with no licence (Romanian), the empty ones, and Thai and Ethiopic, which break a line with no visible hyphen. A set that upstream adds fails the script until it is decided. The seven sets bundled before come out with the same patterns. English uses the full hyph-en-us set instead of about sixty patterns, and forLanguage reads the rest of a BCP 47 tag: a British-spelling region picks hyph-en-gb, the variant 1901 traditional German, polyton polytonic Greek, the script Latin or Cyrillic Serbian, and Serbian with no script gets both. The trie now lives in flat arrays with sorted edges. A map for each node held Hungarian's 63k patterns in about 15 MB of heap and German's in 9 MB; the arrays hold them in 1.7 MB and 1.1 MB. HyphenationGolden holds words of every set with the breaks that the script's own implementation finds, which looks patterns up in a dictionary read from the upstream files, so neither the trie nor the packing can confirm itself. Dropping one chunk of the Hungarian set fails it. EPUB layout gives English breaks only to a chapter with no language or an English one; a language with no bundled set is not hyphenated. A word's leading and trailing punctuation no longer stops it hyphenating, and an apostrophe, a joiner or a combining mark may stand inside it. The hyphenator lower-cases one character for one, so a Turkish İ no longer shifts the breaks after it, and puts no break before a combining mark. Fixes #207 Fixes #615 Fixes #616 Fixes #617 Fixes #618
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Resolves #207, and four defects found on the way (#615, #616, #617, #618).
What changes
Data (#207).
tools/generate_hyphenation.pywrites one Kotlin object for each hyph-utf8 set from a checkout ofhyphenation/tex-hyphen. It carries the patterns and, for the first time, the upstream exception lists, both unmodified, under a header that quotes the copyright and the licence text and names the upstream commit. That gives 71 sets for more than 60 languages, up from 7. The seven sets bundled before regenerate with identical patterns.Left out, with the reason recorded in the script: GPL-only or LGPL-only sets (Czech, Macedonian, Indonesian, upstream Serbian, Armenian, Latvian), sets with no stated licence (Romanian,
mn-cyrl-x-lmc), empty sets (Arabic, Persian, Hebrew, Vietnamese), and Thai and Ethiopic, which break lines with no visible hyphen. A set that upstream adds later fails the script until someone decides on it. Serbian is still covered by the LPPLsh-cyrlandsh-latnsets.Selection. English uses the full
hyph-en-usset instead of about 60 hand-picked patterns.forLanguagenow reads the rest of a BCP 47 tag:hyph-en-gb.1901picks traditional German, with the Swiss set for CH and LI.polytonpicks polytonic Greek.x-classic,x-liturgicandx-schoolpick the private-use sets.Only the sets a book asks for are loaded.
Memory. The trie now lives in flat arrays with sorted edges. Measured retained heap on the JVM: Hungarian (63k patterns) goes from about 15 MB with a map per node to 1.7 MB, German from 9.3 MB to 1.1 MB. A test keeps Hungarian under 4 MB.
Fixes
İno longer shifts every later break.Krankenhaus,,«Bonjour») no longer stops it hyphenating, and an apostrophe or joiner may stand inside a word (l’université).Tests
HyphenationGoldenholds words for every set, real ones plus pseudo-words built from pattern letters taken from across each file. The expected breaks come from the script's independent implementation, which uses a dictionary lookup over the upstream files, not the trie or the Kotlin strings.lowercase(), removing the mark filter, or restoring the old EPUB word check and English fallback each fail their tests.Gate
mutool, so the PDF differential ran as a KitePDF-only smoke pass: 58 pages, 0 render failures, no mean-error number. Nothing here touches PDF rendering.checkKotlinAbicould not run here (no Android SDK). The public JVM surface ofHyphenatormatches the committed dump member for member. The new constructor andtrieBytesareinternal.Notes for review
kitepdf-core. The largest sets are Hungarian (528 KB), the three German sets (about 270 KB each) and Classical Latin (216 KB). If that is too much for the JS or iOS bundles, the rarely selected sets (de-1901,de-ch-1901,la-x-classic,cu) are the first candidates to move into an optional module.main; this session was told to push to a branch, so this arrives as a PR.Generated by Claude Code