mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-07 13:30:37 +02:00
* model: K2 Horizon gguf conversion code
* model: loading hparams and tensors in k2-horizon.cpp
* model: K2 Horizon compute graph
* model: K2 Horizon compute graph adjustment and registering tokenizers
* model: K2 Horizon chat template and accomodate safetensors naming
* unicode : add the K2-Horizon pre-tokenizer splitter
The K2-Horizon regex had no arm in unicode_regex_split_custom and fell through to the
general std::regex fallback, which fails two ways.
On MSVC std::regex rejects \p{...}, so no K2-Horizon GGUF loads on Windows at all:
llama-quantize, llama-imatrix and llama-perplexity all abort with
regex_error(error_escape) before a token is produced.
Where the fallback does compile it is still wrong. unicode_regex_split collapses each
codepoint to a single byte naming its Unicode category before matching, and U+200C/U+200D
are category Control, which has no entry in k_ucat_cpt, so both become the 0xD0 fallback
byte. The literal and alternatives in K2's regex can then never match and
every ZWNJ or ZWJ ends a letter run.
The splitter is the existing llama3 one with a single rule widened, since K2's regex
differs from llama3's only in that a letter run also takes marks, ZWNJ and ZWJ.
tests/test-unicode.cpp gains a case for this: it fails before the change with
[Amy] [ZWNJ khaham] and passes after with the run intact.
* tests: expand K2 Horizon unicode splitter coverage
* unicode: handle K2 Horizon case folding and empty input
Assisted-by: Codex
* jinja : support sequence indices in selectattr and rejectattr
Assisted-by: Codex
* model : add K2 Horizon dense and MoVA support
Includes the K2 Horizon implementation from ifm-ai/llama.cpp with converter, tensor-parallel and model save/reload fixes.
Assisted-by: Codex
* chat : support K2 Horizon reasoning and tool calls
Assisted-by: Codex
* conversion: remove obsolete K2 Aurora alias
Assisted-by: Codex
* k2-horizon: enforce response schemas and load YaRN betas
Constrain final JSON after reasoning, accept flexible JSON tool envelopes,
enforce XML dialects, and handle repeated or alternate thinking markers.
Load YaRN beta metadata instead of retaining the default values.
Add schema, streaming, continuation, and model reload regressions. Validate
CUDA and CPU builds and 0.9B, 4B, and MoVA conversation/tool round trips.
Assisted-by: Codex
* renaming template fixture
* adressing cisc follows ups
* desloppify the parser / adress aldehir comments
* clean test-chat
* remove fallback : model trained mostly on high anyway
* fix k2 attn_v_exp tn splitting and metal fusion baseline
* k2-horizon : forward expand views before sums
* k2-horizon: copy embds before group norm to fix TP
* disable tesnor parallelism
---------
Co-authored-by: Ryandito Diandaru <ryandito.diandaru@mbzuai.ac.ae>
Co-authored-by: WestWaters <mario.papaleo2013@gmail.com>
Co-authored-by: Natani L. Mayday <71436458+TaskPuppyNatani@users.noreply.github.com>
Co-authored-by: West <100190545+WestWaters@users.noreply.github.com>
Co-authored-by: aaryamonvikram <aaryamonvikram@gmail.com>
Co-authored-by: aaryamonvikram <96529820+aaryamonvikram@users.noreply.github.com>