Files
Toki NasinandSigbjørn Skjæret abeada335e vocab : implement PLaMo-3 tokenizer pre-segmentation (#30045)
* vocab : implement PLaMo-3 tokenizer pre-segmentation

The PLaMo-3 tokenizer inserts hard boundaries before running the Unigram
DP, around <|plamo:...|>-looking text, and around runs of at least 4
identical characters or 2 spaces. Without them llama.cpp tokenizes code
indentation and repeated punctuation differently from the reference.

Reproduce the two re.sub() passes in llm_tokenizer_plamo2 by encoding
each segment independently.

* add vocab type "plamo3"

* Update src/llama-vocab.cpp

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

* misc change

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
2026-10-06 19:07:54 +02:00
..