Unicode foundation: shaped marks, seam repair, extras interning, rio-unicode - #1854
Merged
Merged
Conversation
The emit loop hashed a cell's zerowidth codepoints into the run-cache key but only ever pushed the base char into the shaping buffer, so a combining sequence rendered as its bare base: e + U+0301 drew as e. Marks now shape with their base on both platforms. The glyph-to-cell walk moves to explicit per-cell starts everywhere (UTF-16 units on macOS, byte offsets on swash) since the swash path's per-char cursor assumed one char per cell. VS15/VS16 stay out of the buffer, matching the hash: presentation is already resolved into font_id, and a selector fed to a text font only produces a notdef. Kitty placeholder cells are excluded; their zerowidth carries image-slice coordinates, not text.
ECH, DCH, ICH and EL blanket-blanked or shifted cells with no regard for wide-pair boundaries, leaving stranded halves behind: a Wide lead whose spacer was erased, a Spacer whose lead was deleted out from under it by a shift, and a LeadingSpacer closing the previous row whose wide char at column 0 was cleared. write_cell has always repaired these for single-cell overwrites; range mutations now run the same cleanup as a seam pass over the mutated row.
Review fixes on the seam-repair pass: a repaired cell keeps its WRAPLINE bit (the flag marks the logical line for reflow and lives on the row's last cell — often a pair's spacer — which the op never targeted); ED's partial-row clears run the same repair (they split pairs at the cursor boundary exactly like EL); a LeadingSpacer that a DCH shift dragged into the row interior reverts to narrow; and the previous-row damage call skips history rows, matching write_cell.
Replace the whole-row sweep with split_wide_seam: each range mutation names the exact boundaries it cuts (ghostty Screen.splitCellBoundary), O(1) per seam instead of O(columns) per op. ECH and EL split at both range edges, DCH at start, source end and row end, ICH at the insert point, row end and the last shifted cell (a lead whose spacer will not survive the shift), ED partial rows at their cursor-side edges. Observable semantics are unchanged — all eight seam regression tests pass as written. Worst-case EL on a 320-column grid drops from 299ns to 61ns per op (42ns with no repair at all).
Extras slots are now interned: identical (marks, hyperlink) content shares one slot, so the u16 id space stops being a per-occurrence budget whose overflow silently drops extras — repeated emoji and combining sequences cost one slot total. Shared slots make in-place mutation illegal, which exposes and fixes a live bug: every cell written under one OSC 8 template references the same slot, so pushing a combining mark into it attached the mark to the whole hyperlink span. The attach path is copy-on-write now: clone, extend, re-intern. Hyperlink span detection (hints, hover) compares the hyperlink itself instead of raw extras_id equality: a marked cell inside a link holds a different slot than its neighbors while belonging to the same span — under interning, and already before it, id equality was the wrong key. The sweep and free paths purge the interning lookup so stale content can't resolve to a dead slot.
Width tables vendored from the alacritty unicode-width fork the tree already ships (same 0.1.x semantics, same generator lineage), regenerated for Unicode 17: the audit test in tests/ walks every codepoint against unicode-width-16 and the 196 differences are all Unicode 16-to-17 data changes — new scripts' combining marks going zero-width, Tangut/Khitan additions and new emoji going wide. Grapheme cluster segmentation joins it in the same crate, generated from the same UCD release, closing the version skew where widths came from Unicode 16 while the preedit's unicode-segmentation dependency shipped Unicode 17. The break decision is a pairwise state machine — is_break(prev, next, &mut state) with emoji-ZWJ, RI-parity and Indic conjunct (GB9c) state — the shape grid code needs to cluster codepoints cell by cell without materializing strings, with Graphemes/GraphemeIndices iterators on top for string callers. All 766 UCD GraphemeBreakTest vectors pass. The workspace unicode-width alias now points here; every consumer (rio-vt width table, sugarloaf, rioterm) moves to Unicode 17 in one step. Generators are checked in under scripts/ (upstream unicode.py v0.1.11 lineage with Unicode 17 fetch paths, plus grapheme.py).
CI runs clippy with --all-features, which enabled the bench feature vendored from unicode-width and its nightly-only feature(test) attr — E0554 on stable. The benches were never runnable here; remove them, the feature, and the no_std-gated test imports. Also rewrite the extras reclaim-cadence test for interning semantics: repeat content reuses its slot without counting as an allocation, and per-slot free no longer exists.
The width-drift audit asserts its 196-codepoint envelope and pins known Unicode 17 changes in both directions instead of narrating to a log nobody reads; a regenerated table that drifts now fails CI. The dead cjk feature knob goes away (width_cjk is unconditional in this vintage). grapheme.py fetches its own UCD inputs — urllib raises on HTTP errors, so a 404 can never be saved as table data — and the cold-start regeneration was verified byte-identical to the committed tables. The vendored doc header now describes this crate's actual provenance instead of unicode-width 0.1.5's.
…tern Review fixes on the interning half. The combining-mark attach read extras_id() from whatever cell it targeted — on a bg-only cell those bits are the color channel, so a garbage id could clone an unrelated slot (someone's hyperlink) onto the cell and writing back corrupted the color; non-Codepoint targets now drop the mark. On slot exhaustion the attach keeps the cell's existing extras instead of erasing them. A hyperlink left open across an alt-screen switch is re-interned into the alt grid's table (ids are table-local; the seeded cursor template stamped a foreign id onto every alt cell). The VS16 wrap move sets the row's has_extras hint itself instead of relying on the caller's later attach. get_mut on the table goes the way of free(): interned slots are shared, mutation is copy-on-write or nothing; clear() had no callers and free-list semantics that trap any future reset wiring. Stale comments updated.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Consolidates #1850 + #1851 + #1852 + #1853 (all closed with pointers here) into one testable branch, the complete workstream 0 of the mode-2027 scoping. Original commits preserved via merges. Two additional fixes from the consolidated review are the last commits.