mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-07 13:30:37 +02:00
* vocab : implement PLaMo-3 tokenizer pre-segmentation The PLaMo-3 tokenizer inserts hard boundaries before running the Unigram DP, around <|plamo:...|>-looking text, and around runs of at least 4 identical characters or 2 spaces. Without them llama.cpp tokenizes code indentation and repeated punctuation differently from the reference. Reproduce the two re.sub() passes in llm_tokenizer_plamo2 by encoding each segment independently. * add vocab type "plamo3" * Update src/llama-vocab.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * misc change --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>