Skip to content

Cluster alternative Sara Am and Sara Ae spellings in tcc_p - #1523

Open
kamthorn wants to merge 1 commit into
PyThaiNLP:devfrom
kamthorn:claude/cool-gauss-ivm4sq
Open

kamthorn wants to merge 1 commit into
PyThaiNLP:devfrom
kamthorn:claude/cool-gauss-ivm4sq

Conversation

@kamthorn

Copy link
Copy Markdown
Contributor

What do these changes do

pythainlp.tokenize.tcc_p (TCC+, used by the default newmm word
tokenizer through tcc_pos_array()) now treats two common alternative
spellings of a single Thai vowel as one character cluster:

  • Sara Am typed as Nikhahit + Sara Aa: ํา (U+0E4D U+0E32),
    with an optional tone mark before or after the Nikhahit
  • Sara Ae typed as Sara E + Sara E: เเ (U+0E40 U+0E40)

Both render the same as ำ (U+0E33) and แ (U+0E41).
They are the first and third entries of _REORDER_PAIRS in
pythainlp/util/normalize.py.

The input text is not changed, so "".join(word_tokenize(text)) == text
still holds. This is the difference from calling
pythainlp.util.normalize() before tokenizing: normalize() rewrites
the text, so the joined tokens no longer match the original input.

What was wrong

tcc_p had no rule for these sequences:

>>> tcc_p.segment("ทํานา")   # Nikhahit + Sara Aa
['ท', 'ํ', 'า', 'นา']        # Nikhahit alone as a cluster
>>> tcc_p.segment("เเปลก")   # Sara E + Sara E
['เ', 'เป', 'ล', 'ก']        # first Sara E alone as a cluster

Because newmm only breaks words at TCC boundaries, these wrong
boundaries showed up in word tokens:

>>> word_tokenize("เเข็ง")
['เ', 'เข็ง']
>>> word_tokenize("เเปลก")
['เ', 'เปล', 'ก']

The bundled word list itself has entries with this spelling, and tcc_p
split them the same way: ซํ้า, พลํ้า, รํ่า.

How this fixes it

In the _RE_TCC template:

  • The rule ct[ะาำ]?k becomes ctA?k, where A expands to
    (?:[ะาำ]|ํ[่-๋]?า).
  • แ in the Sara Ae rules expands to (?:แ|เเ), and [เ-ไ]ct becomes
    (?:เเ|[เ-ไ])ct.

The rules match these sequences only when the text contains a
Nikhahit or two Sara E in a row. Neither occurs in standard spelling, so
clusters for standard text do not change.

Results after the fix:

>>> tcc_p.segment("ทํานา")
['ทํา', 'นา']
>>> tcc_p.segment("เเปลก")
['เเป', 'ล', 'ก']              # same as 'แปลก' -> ['แป', 'ล', 'ก']
>>> word_tokenize("เเข็ง")
['เเข็ง']

Checks against the unchanged upstream/dev version, using all 66,409
words in thai_words() + thai_syllables():

Check Result
Words whose clusters changed 3 (ซํ้า, พลํ้า, รํ่า: already in the alternative spelling, now clustered correctly)
Words rewritten to ํา (3,531 words, tone before Nikhahit): clusters equal the ำ spelling 3,531 / 3,531
Words rewritten to ํา (1,090 words, tone after Nikhahit): clusters equal the ำ spelling 1,090 / 1,090
Words rewritten to เเ (5,710 words): clusters equal the แ spelling 5,710 / 5,710
Clusters starting with a mark or a lone lead vowel, ํา / ํา (tone after) / เเ 3,694 → 1, 2,280 → 1, 6,048 → 39
newmm tokens starting with a mark or a lone lead vowel, same three sets 281 → 0, 150 → 0, 3,870 → 1
"".join(tokens) != text 0
tcc_p.segment time, 151,908 characters 0.890 s → 0.837 s (no regression)

The clusters that still start with a mark or a lone lead vowel appear
with the standard spelling too, so this change does not add them. For
example, the Thanthakhat in แวมไพร์ and แท็งก์น้ำ, and the
dictionary entry คาเแร็กเตอร์.

New test: TokenizeTestCase.test_tcc_p_alternative_spellings in
tests/core/test_tokenize.py.

Other findings, not changed in this PR:

  • A tone mark typed before an above or below vowel (for example ก่ิน
    instead of กิ่น) still leaves the vowel as its own cluster. This is a
    character-order problem that needs changes to many rules.

Your checklist for this pull request

  • Passed code styles and structures
  • Passed code linting checks and unit test

Ruff passes and mypy reports no errors in tcc_p.py.
tests.core.test_tokenize, tests.core.test_tokenize_thread_safety, and
tests.compact.testc_tokenize pass. In the full tests.core run,
6 tests in test_generate, test_tag, and test_corpus fail. They fail
the same way on unmodified dev in the same environment, because corpus
files are missing there.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MLAScAmHFQoCy5oDHoGoo1

Nikhahit + Sara Aa (ํา) and Sara E + Sara E (เเ) render the same as
Sara Am (ำ) and Sara Ae (แ), but tcc_p left the Nikhahit or the first
Sara E as a cluster of its own. newmm breaks words only at TCC
boundaries, so it could split them off as separate tokens, e.g.
word_tokenize("เเข็ง") returned ["เ", "เข็ง"].

Accept both spellings in the TCC+ rules. The input text is not
modified, so joined tokens still equal the input. Clusters for text in
standard spelling do not change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MLAScAmHFQoCy5oDHoGoo1
@sonarqubecloud

Copy link
Copy Markdown

@kamthorn

Copy link
Copy Markdown
Contributor Author

The CI failures on this PR are pre-existing on dev and unrelated to this change. None of them touch pythainlp/tokenize/tcc_p.py, tests/core/test_tokenize.py, or the CHANGELOG lines this PR adds.

  • Markdown lint: CHANGELOG.md:27,28,32,34,35 and pythainlp/corpus/corpus_license.md:164 — these lines are unchanged by this PR and already fail lint on origin/dev.
  • bandit: hardcoded-password false positives in pythainlp/transliterate/fastthaig2p.py:473,478 (existing code from an earlier merge).
  • mypy: errors in pythainlp/transliterate/wiktionary.py, pythainlp/benchmarks/word_tokenization.py, pythainlp/util/wordtonum.py, pythainlp/ulmfit/core.py, and tests/core/test_transliterate.py.
  • ruff: import-order/__all__ issues in pythainlp/benchmarks/word_tokenization.py, pythainlp/chat/core.py, pythainlp/generate/thai2fit.py, pythainlp/generate/wangchanglm.py, tests/core/test_fastthaig2p.py.
  • unittest (several Python versions/OSes): test_transliterate_thaig2p_v4_dispatch fails with AttributeError: 'function' object has no attribute 'thaig2p_v4' (mock-patch target issue in the dispatch mechanism).
  • unittest (ubuntu-latest, 3.13): fails during dependency install — Cython.Compiler.Errors.CompileError: sklearn/linear_model/_cd_fast.pyx — an environment/build issue, not a test failure.

I confirmed each of these already reproduces on origin/dev before this PR's commit. Happy to help fix any of them separately if useful.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants