Add optional Hugging Face Rust JSON tokenizer backend - #212
Open
JacobSzwejbka wants to merge 5 commits into
Open
JacobSzwejbka wants to merge 5 commits into
JacobSzwejbka wants to merge 5 commits into
Conversation
JacobSzwejbka
force-pushed
the
hf-rust-tokenizer-tok
branch
4 times, most recently
from
September 24, 2026 19:05
281ac35 to
228a1da
Compare
Preserve independent BOS/EOS counts, avoid guessing a missing token ID, and reject .tok files whose decoder cannot yet be represented safely. Add C++ regression coverage and run the Rust checks in CI. AI-assisted-by: Codex
JacobSzwejbka
force-pushed
the
hf-rust-tokenizer-tok
branch
from
September 24, 2026 19:13
228a1da to
1669237
Compare
JacobSzwejbka
force-pushed
the
hf-rust-tokenizer-tok
branch
2 times, most recently
from
September 24, 2026 20:30
b3c4f65 to
790e279
Compare
AI-assisted-by: Codex
JacobSzwejbka
force-pushed
the
hf-rust-tokenizer-tok
branch
from
September 24, 2026 20:31
790e279 to
5c68507
Compare
JacobSzwejbka
marked this pull request as ready for review
September 24, 2026 20:32
JacobSzwejbka
requested review from
mergennachin
and removed request for
mergennachin
September 24, 2026 20:32
Generated with assistance from Codex.
Generated with assistance from Codex.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This adds an opt-in Rust-backed Hugging Face tokenizer for standard
tokenizer.jsonfiles. It uses only the published Hugging Face v1.0.0-rc.2 crates from crates.io; the Cargo lockfile contains no git dependencies.The backend canonicalizes normal Hub JSON in memory, uses the upstream v1 inference and decoder pipelines, reads BOS/EOS roles from canonical metadata or sibling tokenizer configuration files, and falls back cleanly when the RC cannot load a tokenizer. It preserves ExecuTorch's independent and repeatable BOS/EOS-count contract.
The option remains fully off by default: without
TOKENIZERS_BUILD_HF_RUST_TOKENIZER, CMake does not require Cargo or add the Rust archive.Experimental
.toksupport has been split into draft meta-pytorch/tokenizers#218. That draft is blocked on the format landing in a protected Hugging Face branch or release.Companion ExecuTorch integration: pytorch/executorch#22548.
Review order
Start with
rust_tokenizer/Cargo.tomlandrust_tokenizer/CMakeLists.txt, thenrust_tokenizer/src/lib.rsfor JSON loading and the narrow C ABI, followed bysrc/rust_hf_tokenizer.cppfor the existing C++ interface adaptation.Upstream alignment
Hugging Face's v1 RC announcement publishes the Rust crates and identifies inference-only C/C++ bindings for ExecuTorch and llama.cpp as a v1 goal. Their experimental C bindings PR #2375 is not yet merged and does not expose all metadata required by the ExecuTorch
Tokenizerinterface. This PR keeps its local ABI intentionally small so it can migrate once the upstream ABI stabilizes.Validation
Rust formatting, Clippy with warnings denied, and unit tests pass. The focused C++ adapter suite passes. A clean CMake build and installed-package consumer successfully round-trip a real Hub GPT-2
tokenizer.json; a real BERT tokenizer was also exercised. The default-off build does not invoke Cargo.The current Apple arm64
minsizesmoke executable measures 1,921,872 bytes stripped and 901,688 bytes with gzip-9.AI-assisted by Codex.