Conversation
Nikhahit + Sara Aa (ํา) and Sara E + Sara E (เเ) render the same as
Sara Am (ำ) and Sara Ae (แ), but tcc_p left the Nikhahit or the first
Sara E as a cluster of its own. newmm breaks words only at TCC
boundaries, so it could split them off as separate tokens, e.g.
word_tokenize("เเข็ง") returned ["เ", "เข็ง"].
Accept both spellings in the TCC+ rules. The input text is not
modified, so joined tokens still equal the input. Clusters for text in
standard spelling do not change.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MLAScAmHFQoCy5oDHoGoo1
|
Contributor
Author
|
The CI failures on this PR are pre-existing on
I confirmed each of these already reproduces on |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



What do these changes do
pythainlp.tokenize.tcc_p(TCC+, used by the defaultnewmmwordtokenizer through
tcc_pos_array()) now treats two common alternativespellings of a single Thai vowel as one character cluster:
ํา(U+0E4D U+0E32),with an optional tone mark before or after the Nikhahit
เเ(U+0E40 U+0E40)Both render the same as
ำ(U+0E33) andแ(U+0E41).They are the first and third entries of
_REORDER_PAIRSinpythainlp/util/normalize.py.The input text is not changed, so
"".join(word_tokenize(text)) == textstill holds. This is the difference from calling
pythainlp.util.normalize()before tokenizing:normalize()rewritesthe text, so the joined tokens no longer match the original input.
What was wrong
tcc_phad no rule for these sequences:Because
newmmonly breaks words at TCC boundaries, these wrongboundaries showed up in word tokens:
The bundled word list itself has entries with this spelling, and
tcc_psplit them the same way:
ซํ้า,พลํ้า,รํ่า.How this fixes it
In the
_RE_TCCtemplate:ct[ะาำ]?kbecomesctA?k, whereAexpands to(?:[ะาำ]|ํ[่-๋]?า).แin the Sara Ae rules expands to(?:แ|เเ), and[เ-ไ]ctbecomes(?:เเ|[เ-ไ])ct.The rules match these sequences only when the text contains a
Nikhahit or two Sara E in a row. Neither occurs in standard spelling, so
clusters for standard text do not change.
Results after the fix:
Checks against the unchanged
upstream/devversion, using all 66,409words in
thai_words()+thai_syllables():ซํ้า,พลํ้า,รํ่า: already in the alternative spelling, now clustered correctly)ํา(3,531 words, tone before Nikhahit): clusters equal theำspellingํา(1,090 words, tone after Nikhahit): clusters equal theำspellingเเ(5,710 words): clusters equal theแspellingํา/ํา(tone after) /เเnewmmtokens starting with a mark or a lone lead vowel, same three sets"".join(tokens) != texttcc_p.segmenttime, 151,908 charactersThe clusters that still start with a mark or a lone lead vowel appear
with the standard spelling too, so this change does not add them. For
example, the Thanthakhat in
แวมไพร์andแท็งก์น้ำ, and thedictionary entry
คาเแร็กเตอร์.New test:
TokenizeTestCase.test_tcc_p_alternative_spellingsintests/core/test_tokenize.py.Other findings, not changed in this PR:
ก่ินinstead of
กิ่น) still leaves the vowel as its own cluster. This is acharacter-order problem that needs changes to many rules.
Your checklist for this pull request
Ruff passes and mypy reports no errors in
tcc_p.py.tests.core.test_tokenize,tests.core.test_tokenize_thread_safety, andtests.compact.testc_tokenizepass. In the fulltests.corerun,6 tests in
test_generate,test_tag, andtest_corpusfail. They failthe same way on unmodified
devin the same environment, because corpusfiles are missing there.
🤖 Generated with Claude Code
https://claude.ai/code/session_01MLAScAmHFQoCy5oDHoGoo1