Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,9 @@ and this project adheres to

### Changed

- Move `pythainlp.phayathaibert`, `pythainlp.wangchanberta`, and
`pythainlp.ulmfit` to `pythainlp.lm`; deprecate the old import paths
([#1527])
- `pythainlp.tokenize.deepcut`: built-in ONNX engine replaces the
TensorFlow-based `deepcut`; `custom_dict` is no longer applied ([#1372])
- Improve guardrails in `check_sara()` and `nighit()` ([#1453])
Expand Down Expand Up @@ -85,6 +88,7 @@ and this project adheres to
[#1511]: https://github.com/PyThaiNLP/pythainlp/pull/1511
[#1512]: https://github.com/PyThaiNLP/pythainlp/pull/1512
[#1526]: https://github.com/PyThaiNLP/pythainlp/pull/1526
[#1527]: https://github.com/PyThaiNLP/pythainlp/pull/1527
[#1529]: https://github.com/PyThaiNLP/pythainlp/pull/1529
[#1541]: https://github.com/PyThaiNLP/pythainlp/pull/1541
[#1542]: https://github.com/PyThaiNLP/pythainlp/pull/1542
Expand Down
13 changes: 12 additions & 1 deletion docs/api/lm.rst
Original file line number Diff line number Diff line change
Expand Up @@ -3,8 +3,19 @@
pythainlp.lm
============

The `pythainlp.lm` package provides language models and language modeling utilities.

Modules
-------

.. autofunction:: calculate_ngram_counts
.. autofunction:: remove_repeated_ngrams
.. autofunction:: remove_repeated_ngrams
.. autoclass:: Qwen3
:members:

Submodules
----------

* :mod:`pythainlp.lm.phayathaibert`
* :mod:`pythainlp.lm.wangchanberta`
* :mod:`pythainlp.lm.ulmfit`
14 changes: 10 additions & 4 deletions docs/api/phayathaibert.rst
Original file line number Diff line number Diff line change
@@ -1,8 +1,14 @@
.. currentmodule:: pythainlp.phayathaibert
.. currentmodule:: pythainlp.lm.phayathaibert

pythainlp.phayathaibert
=======================
The `pythainlp.phayathaibert` module is built upon the phayathaibert base model.
pythainlp.lm.phayathaibert
==========================

.. note::
:mod:`pythainlp.phayathaibert` has moved to :mod:`pythainlp.lm.phayathaibert`.
Importing from :mod:`pythainlp.phayathaibert` still works but emits a
:class:`DeprecationWarning` and will be removed in 6.0.

The `pythainlp.lm.phayathaibert` module is built upon the phayathaibert base model.

Modules
-------
Expand Down
16 changes: 10 additions & 6 deletions docs/api/ulmfit.rst
Original file line number Diff line number Diff line change
@@ -1,8 +1,14 @@
.. currentmodule:: pythainlp.ulmfit
.. currentmodule:: pythainlp.lm.ulmfit

pythainlp.ulmfit
====================================
Welcome to the `pythainlp.ulmfit` module, where you'll find powerful tools for Universal Language Model Fine-tuning for Text Classification (ULMFiT). ULMFiT is a cutting-edge technique for training deep learning models on large text corpora and then fine-tuning them for specific text classification tasks.
pythainlp.lm.ulmfit
===================

.. note::
:mod:`pythainlp.ulmfit` has moved to :mod:`pythainlp.lm.ulmfit`.
Importing from :mod:`pythainlp.ulmfit` still works but emits a
:class:`DeprecationWarning` and will be removed in 6.0.

Welcome to the `pythainlp.lm.ulmfit` module, where you'll find powerful tools for Universal Language Model Fine-tuning for Text Classification (ULMFiT). ULMFiT is a cutting-edge technique for training deep learning models on large text corpora and then fine-tuning them for specific text classification tasks.

Modules
-------
Expand Down Expand Up @@ -86,5 +92,3 @@ Modules
:noindex:

The `ungroup_emoji` function is designed for ungrouping emojis in text data, which can be crucial for emoji recognition and classification tasks.

.. The `pythainlp.ulmfit` module provides a comprehensive set of tools for ULMFiT-based text classification. Whether you need to preprocess Thai text, tokenize it, compute document vectors, or perform various text cleaning tasks, this module has the utilities you need. ULMFiT is a state-of-the-art technique in NLP, and these tools empower you to use it effectively for text classification.
14 changes: 10 additions & 4 deletions docs/api/wangchanberta.rst
Original file line number Diff line number Diff line change
@@ -1,8 +1,14 @@
.. currentmodule:: pythainlp.wangchanberta
.. currentmodule:: pythainlp.lm.wangchanberta

pythainlp.wangchanberta
=======================
The `pythainlp.wangchanberta` module is built upon the WangchanBERTa base model, specifically the `wangchanberta-base-att-spm-uncased` model, as detailed in the paper by Lowphansirikul et al. [#Lowphansirikul_2021]_.
pythainlp.lm.wangchanberta
==========================

.. note::
:mod:`pythainlp.wangchanberta` has moved to :mod:`pythainlp.lm.wangchanberta`.
Importing from :mod:`pythainlp.wangchanberta` still works but emits a
:class:`DeprecationWarning` and will be removed in 6.0.

The `pythainlp.lm.wangchanberta` module is built upon the WangchanBERTa base model, specifically the `wangchanberta-base-att-spm-uncased` model, as detailed in the paper by Lowphansirikul et al. [#Lowphansirikul_2021]_.

This base model is utilized for various natural language processing tasks in the Thai language, including named entity recognition, part-of-speech tagging, and subword tokenization.

Expand Down
2 changes: 1 addition & 1 deletion pythainlp/augment/lm/phayathaibert.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
if TYPE_CHECKING:
from transformers import AutoModelForMaskedLM, AutoTokenizer, Pipeline

from pythainlp.phayathaibert.core import ThaiTextProcessor
from pythainlp.lm.phayathaibert.core import ThaiTextProcessor

_MODEL_NAME: str = "clicknext/phayathaibert"

Expand Down
20 changes: 20 additions & 0 deletions pythainlp/lm/phayathaibert/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# SPDX-FileCopyrightText: 2016-2026 PyThaiNLP Project
# SPDX-FileType: SOURCE
# SPDX-License-Identifier: Apache-2.0
"""PhayaThaiBERT language model."""

__all__: list[str] = [
"NamedEntityTagger",
"PartOfSpeechTagger",
"ThaiTextAugmenter",
"ThaiTextProcessor",
"segment",
]

from pythainlp.lm.phayathaibert.core import (
NamedEntityTagger,
PartOfSpeechTagger,
ThaiTextAugmenter,
ThaiTextProcessor,
segment,
)
Loading
Loading