Skip to content

Unicode foundation: shaped marks, seam repair, extras interning, rio-unicode - #1854

Merged
raphamorim merged 14 commits into
mainfrom
unicode-foundation
Aug 10, 2026
Merged

raphamorim merged 14 commits into
mainfrom
unicode-foundation

Conversation

@raphamorim

@raphamorim raphamorim commented Aug 10, 2026 •

Copy link
Copy Markdown
Owner

Consolidates #1850 + #1851 + #1852 + #1853 (all closed with pointers here) into one testable branch, the complete workstream 0 of the mode-2027 scoping. Original commits preserved via merges. Two additional fixes from the consolidated review are the last commits.

The emit loop hashed a cell's zerowidth codepoints into the run-cache
key but only ever pushed the base char into the shaping buffer, so a
combining sequence rendered as its bare base: e + U+0301 drew as e.
Marks now shape with their base on both platforms. The glyph-to-cell
walk moves to explicit per-cell starts everywhere (UTF-16 units on
macOS, byte offsets on swash) since the swash path's per-char cursor
assumed one char per cell. VS15/VS16 stay out of the buffer, matching
the hash: presentation is already resolved into font_id, and a
selector fed to a text font only produces a notdef. Kitty placeholder
cells are excluded; their zerowidth carries image-slice coordinates,
not text.
ECH, DCH, ICH and EL blanket-blanked or shifted cells with no regard
for wide-pair boundaries, leaving stranded halves behind: a Wide lead
whose spacer was erased, a Spacer whose lead was deleted out from
under it by a shift, and a LeadingSpacer closing the previous row
whose wide char at column 0 was cleared. write_cell has always
repaired these for single-cell overwrites; range mutations now run
the same cleanup as a seam pass over the mutated row.
Review fixes on the seam-repair pass: a repaired cell keeps its
WRAPLINE bit (the flag marks the logical line for reflow and lives on
the row's last cell — often a pair's spacer — which the op never
targeted); ED's partial-row clears run the same repair (they split
pairs at the cursor boundary exactly like EL); a LeadingSpacer that a
DCH shift dragged into the row interior reverts to narrow; and the
previous-row damage call skips history rows, matching write_cell.
Replace the whole-row sweep with split_wide_seam: each range mutation
names the exact boundaries it cuts (ghostty Screen.splitCellBoundary),
O(1) per seam instead of O(columns) per op. ECH and EL split at both
range edges, DCH at start, source end and row end, ICH at the insert
point, row end and the last shifted cell (a lead whose spacer will not
survive the shift), ED partial rows at their cursor-side edges.
Observable semantics are unchanged — all eight seam regression tests
pass as written. Worst-case EL on a 320-column grid drops from 299ns
to 61ns per op (42ns with no repair at all).
Extras slots are now interned: identical (marks, hyperlink) content
shares one slot, so the u16 id space stops being a per-occurrence
budget whose overflow silently drops extras — repeated emoji and
combining sequences cost one slot total. Shared slots make in-place
mutation illegal, which exposes and fixes a live bug: every cell
written under one OSC 8 template references the same slot, so pushing
a combining mark into it attached the mark to the whole hyperlink
span. The attach path is copy-on-write now: clone, extend, re-intern.

Hyperlink span detection (hints, hover) compares the hyperlink itself
instead of raw extras_id equality: a marked cell inside a link holds a
different slot than its neighbors while belonging to the same span —
under interning, and already before it, id equality was the wrong key.
The sweep and free paths purge the interning lookup so stale content
can't resolve to a dead slot.
Width tables vendored from the alacritty unicode-width fork the tree
already ships (same 0.1.x semantics, same generator lineage),
regenerated for Unicode 17: the audit test in tests/ walks every
codepoint against unicode-width-16 and the 196 differences are all
Unicode 16-to-17 data changes — new scripts' combining marks going
zero-width, Tangut/Khitan additions and new emoji going wide.

Grapheme cluster segmentation joins it in the same crate, generated
from the same UCD release, closing the version skew where widths came
from Unicode 16 while the preedit's unicode-segmentation dependency
shipped Unicode 17. The break decision is a pairwise state machine —
is_break(prev, next, &mut state) with emoji-ZWJ, RI-parity and Indic
conjunct (GB9c) state — the shape grid code needs to cluster
codepoints cell by cell without materializing strings, with
Graphemes/GraphemeIndices iterators on top for string callers. All
766 UCD GraphemeBreakTest vectors pass.

The workspace unicode-width alias now points here; every consumer
(rio-vt width table, sugarloaf, rioterm) moves to Unicode 17 in one
step. Generators are checked in under scripts/ (upstream unicode.py
v0.1.11 lineage with Unicode 17 fetch paths, plus grapheme.py).
CI runs clippy with --all-features, which enabled the bench feature
vendored from unicode-width and its nightly-only feature(test) attr —
E0554 on stable. The benches were never runnable here; remove them,
the feature, and the no_std-gated test imports. Also rewrite the
extras reclaim-cadence test for interning semantics: repeat content
reuses its slot without counting as an allocation, and per-slot free
no longer exists.
The width-drift audit asserts its 196-codepoint envelope and pins
known Unicode 17 changes in both directions instead of narrating to a
log nobody reads; a regenerated table that drifts now fails CI. The
dead cjk feature knob goes away (width_cjk is unconditional in this
vintage). grapheme.py fetches its own UCD inputs — urllib raises on
HTTP errors, so a 404 can never be saved as table data — and the
cold-start regeneration was verified byte-identical to the committed
tables. The vendored doc header now describes this crate's actual
provenance instead of unicode-width 0.1.5's.
…tern

Review fixes on the interning half. The combining-mark attach read
extras_id() from whatever cell it targeted — on a bg-only cell those
bits are the color channel, so a garbage id could clone an unrelated
slot (someone's hyperlink) onto the cell and writing back corrupted
the color; non-Codepoint targets now drop the mark. On slot
exhaustion the attach keeps the cell's existing extras instead of
erasing them. A hyperlink left open across an alt-screen switch is
re-interned into the alt grid's table (ids are table-local; the
seeded cursor template stamped a foreign id onto every alt cell).
The VS16 wrap move sets the row's has_extras hint itself instead of
relying on the caller's later attach. get_mut on the table goes the
way of free(): interned slots are shared, mutation is copy-on-write
or nothing; clear() had no callers and free-list semantics that trap
any future reset wiring. Stale comments updated.
@raphamorim
raphamorim merged commit 4798bc9 into main Aug 10, 2026
14 checks passed
@raphamorim
raphamorim deleted the unicode-foundation branch August 10, 2026 15:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant