Skip to content

rio-unicode: one crate, one Unicode version — width and graphemes - #1853

Closed
raphamorim wants to merge 1 commit into
mainfrom
rio-unicode
Closed

raphamorim wants to merge 1 commit into
mainfrom
rio-unicode

Conversation

@raphamorim

Copy link
Copy Markdown
Owner

Workstream 0a of the mode-2027 scoping — the last of the four foundation pieces.

The problem it closes

The tree had a Unicode version skew: widths came from unicode-width-16 (the alacritty fork, frozen at Unicode 16), while unicode-segmentation (used by the preedit work in #1849) ships Unicode 17. Width and segmentation disagreeing about what Unicode says is exactly the class of bug ghostty's unified src/unicode/ module exists to prevent — this is rio's version of that module.

Width half — same semantics, current data

Provenance mattered here: unicode-width 0.1.13+/0.2.x quietly changed per-char semantics (NUL→None, soft hyphen→0, conjoining jamo→0, even a width-3), which is why the tree pinned the fork in the first place. I audited three generator generations empirically — a test walks every codepoint against unicode-width-16:

  • 0.2.2 base: 4,149 diffs (semantic overhaul — rejected)
  • 0.1.14 base: 4,149 diffs (same modernized code — rejected)
  • 0.1.11 base (this PR): 196 diffs — every one a genuine Unicode 16→17 data change (new scripts' combining marks → 0, Tangut/Khitan/CJK additions and new emoji → 2)

The audit test stays in tests/ as a permanent record of the drift. lib.rs is vendored from the fork itself; the generator (upstream unicode.py v0.1.11 lineage, patched for Unicode 17 fetch paths) is checked in under scripts/.

Grapheme half — built for the grid

is_break(prev: GraphemeClass, next: GraphemeClass, &mut BreakState) -> bool — a pairwise state machine carrying the three things single-pair rules can't see: emoji ZWJ sequences (GB11), regional-indicator parity (GB12/13), and Indic conjunct linkage (GB9c, via InCB-refined classes). This is the shape mode-2027 grid code needs — clustering codepoints cell by cell against the previous cell, ghostty-style, without materializing strings. Graphemes/GraphemeIndices iterators wrap it for string callers (the preedit in #1849 can drop unicode-segmentation once rebased).

All 766 UCD GraphemeBreakTest.txt conformance vectors pass, generated into a test by the checked-in scripts/grapheme.py. A lockstep test asserts the grapheme data and width tables always report the same Unicode version.

Wiring

The workspace unicode-width alias now points at rio-unicode (path dep), so every consumer — rio-vt's 128 KiB runtime width table, sugarloaf, rioterm — moves to Unicode 17 in one step with no source changes. Full downstream suites green: 441 rio-vt, 172 rioterm, 46 librio, 133 sugarloaf. Clippy/fmt clean.

Follow-ups unblocked: #1849 switches its segmentation here; the mode-2027 core (workstream 2 of the scoping doc) now has its tables and state machine waiting.

Width tables vendored from the alacritty unicode-width fork the tree
already ships (same 0.1.x semantics, same generator lineage),
regenerated for Unicode 17: the audit test in tests/ walks every
codepoint against unicode-width-16 and the 196 differences are all
Unicode 16-to-17 data changes — new scripts' combining marks going
zero-width, Tangut/Khitan additions and new emoji going wide.

Grapheme cluster segmentation joins it in the same crate, generated
from the same UCD release, closing the version skew where widths came
from Unicode 16 while the preedit's unicode-segmentation dependency
shipped Unicode 17. The break decision is a pairwise state machine —
is_break(prev, next, &mut state) with emoji-ZWJ, RI-parity and Indic
conjunct (GB9c) state — the shape grid code needs to cluster
codepoints cell by cell without materializing strings, with
Graphemes/GraphemeIndices iterators on top for string callers. All
766 UCD GraphemeBreakTest vectors pass.

The workspace unicode-width alias now points here; every consumer
(rio-vt width table, sugarloaf, rioterm) moves to Unicode 17 in one
step. Generators are checked in under scripts/ (upstream unicode.py
v0.1.11 lineage with Unicode 17 fetch paths, plus grapheme.py).
@raphamorim

Copy link
Copy Markdown
Owner Author

Folded into the consolidated unicode-foundation PR (all four workstream-0 pieces in one branch for testability, commits preserved).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant