rio-unicode: one crate, one Unicode version — width and graphemes - #1853
Closed
raphamorim wants to merge 1 commit into
Closed
raphamorim wants to merge 1 commit into
raphamorim wants to merge 1 commit into
Conversation
Width tables vendored from the alacritty unicode-width fork the tree already ships (same 0.1.x semantics, same generator lineage), regenerated for Unicode 17: the audit test in tests/ walks every codepoint against unicode-width-16 and the 196 differences are all Unicode 16-to-17 data changes — new scripts' combining marks going zero-width, Tangut/Khitan additions and new emoji going wide. Grapheme cluster segmentation joins it in the same crate, generated from the same UCD release, closing the version skew where widths came from Unicode 16 while the preedit's unicode-segmentation dependency shipped Unicode 17. The break decision is a pairwise state machine — is_break(prev, next, &mut state) with emoji-ZWJ, RI-parity and Indic conjunct (GB9c) state — the shape grid code needs to cluster codepoints cell by cell without materializing strings, with Graphemes/GraphemeIndices iterators on top for string callers. All 766 UCD GraphemeBreakTest vectors pass. The workspace unicode-width alias now points here; every consumer (rio-vt width table, sugarloaf, rioterm) moves to Unicode 17 in one step. Generators are checked in under scripts/ (upstream unicode.py v0.1.11 lineage with Unicode 17 fetch paths, plus grapheme.py).
Owner
Author
|
Folded into the consolidated unicode-foundation PR (all four workstream-0 pieces in one branch for testability, commits preserved). |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Workstream 0a of the mode-2027 scoping — the last of the four foundation pieces.
The problem it closes
The tree had a Unicode version skew: widths came from
unicode-width-16(the alacritty fork, frozen at Unicode 16), whileunicode-segmentation(used by the preedit work in #1849) ships Unicode 17. Width and segmentation disagreeing about what Unicode says is exactly the class of bug ghostty's unifiedsrc/unicode/module exists to prevent — this is rio's version of that module.Width half — same semantics, current data
Provenance mattered here: unicode-width 0.1.13+/0.2.x quietly changed per-char semantics (NUL→None, soft hyphen→0, conjoining jamo→0, even a width-3), which is why the tree pinned the fork in the first place. I audited three generator generations empirically — a test walks every codepoint against
unicode-width-16:The audit test stays in
tests/as a permanent record of the drift.lib.rsis vendored from the fork itself; the generator (upstreamunicode.pyv0.1.11 lineage, patched for Unicode 17 fetch paths) is checked in underscripts/.Grapheme half — built for the grid
is_break(prev: GraphemeClass, next: GraphemeClass, &mut BreakState) -> bool— a pairwise state machine carrying the three things single-pair rules can't see: emoji ZWJ sequences (GB11), regional-indicator parity (GB12/13), and Indic conjunct linkage (GB9c, via InCB-refined classes). This is the shape mode-2027 grid code needs — clustering codepoints cell by cell against the previous cell, ghostty-style, without materializing strings.Graphemes/GraphemeIndicesiterators wrap it for string callers (the preedit in #1849 can dropunicode-segmentationonce rebased).All 766 UCD
GraphemeBreakTest.txtconformance vectors pass, generated into a test by the checked-inscripts/grapheme.py. A lockstep test asserts the grapheme data and width tables always report the same Unicode version.Wiring
The workspace
unicode-widthalias now points atrio-unicode(path dep), so every consumer — rio-vt's 128 KiB runtime width table, sugarloaf, rioterm — moves to Unicode 17 in one step with no source changes. Full downstream suites green: 441 rio-vt, 172 rioterm, 46 librio, 133 sugarloaf. Clippy/fmt clean.Follow-ups unblocked: #1849 switches its segmentation here; the mode-2027 core (workstream 2 of the scoping doc) now has its tables and state machine waiting.