1. Per row FP8 wrongly fell to NVFP4 conversion, as NVFP4 checked on
dims alone. Also check on dtype for FP8
2. Need to reshape FP8 QKV projections scales in the same way that
weights are reshaped
* common : read a GGUF's trained context from its metadata
common_get_gguf_n_ctx_train opens only the file's metadata (no_alloc,
like common_get_decision_type) and reads <arch>.context_length, so a
caller can learn the trained context without loading the model. It
accepts both u32 and u64 values and returns 0 when the file is missing,
unreadable, invalid, or reports no context length.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* server : report the trained context in the models listing
update_caps already resolves the model file offline to read its
modalities, so it now reads the trained context from the same GGUF
metadata, and GET /models reports it as context_length when it is
known. A router listing then carries the context without any Hub
request, which lets the UI sort and filter by it offline.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : take the trained context from the models listing
The router now reports context_length per model, so the option mapping
fills contextLength from it and the manager reads it before the Hub
record. The Context column, the context sort and the context filter
then work with the Hugging Face Hub API turned off. A browser suite
guards the sort and the search, the Hub-cache driven context filter and
the re-sort when details arrive after the sort was clicked.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : mark favorite models with a heart
A favorited model shows a rose heart in the selector even before its
row is hovered, and the crossed heart takes its place on hover, so
unfavoriting stays one hover away. The manager table marks its
favorited rows with the same heart after the badges and capabilities.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* fix: UI text nit
* fix: UI nits
* fix: Favorite models grouping in models table
* feat: Remove sorting from Status column in Models Table
* server: read the GGUF metadata once per model
Read the decision type and the trained context in a single GGUF open,
accept only a UINT32 context length like the model loader, and reset
n_ctx_train with the other caps so a failed refresh drops it.
* fix: Post-review fixes
---------
Co-authored-by: Pascal <admin@serveurperso.com>
* CUDA: fuse copy of updated state snapshots into recurrent cache with ssm_scan
* CUDA: remove redundant cuda copies with K==1 (non spec-dec) scenario as well
* opencl: skip kernel_cpy_f32_f32_pack on A6X to avoid shader compiler crash
* The A6x compiler backend found in iot device with a623 (E031.50.31.01)
cannot handle kernels with a large number of arguments. Skip this
kernel for A6x to avoid compiler crash
* opencl: A6X constant-fold workaround for get_local_size in GEMV kernels
* opencl: add Adreno 623 to A6X GPU detection list
* model : use exact GELU for ModernBERT encoders
Assisted-by: Codex
* model : keep tanh GELU aliases on ggml_geglu
Assisted-by: Claude Opus 5.5
* model : map gelu_python to ggml_geglu_erf
Assisted-by: Claude Opus 5.5
The reserve builds n_outputs_max_per_seq sampling chains per sampler,
while a decode built one per output row, so the graph changed its
topology after the reserve and GGML_SCHED_NO_REALLOC builds aborted on
the next same sized graph. Every sampler now builds
n_outputs_max_per_seq chains, the ones without a row of the ubatch on
the padding row and not selected, and graph_max_nodes counts them.
ggml_acc_impl narrowed a size_t offset to int32_t without checking that it
fits, so a large offset could truncate to a negative int32_t. The forward
then sign-extended it to a huge size_t and the bounds assertion wrapped,
allowing an OOB write below the dst buffer. Check the offset before the
narrowing, matching the existing check in ggml_set_impl.
* meta : handle views of tensors allocated on the host
A view shares the memory of its view_src, so ggml-alloc never allocates a view in
the buffer of the split it lands in - the scheduler copies the source into the
split and the ops that use the view read that copy. The view node itself is a noop
and does not need a split of its own, but the meta backend asserted when one was
left inside a meta split:
- ggml_backend_meta_get_split_state() dereferenced tensor->buffer->context
- the graph rebuild mapped every node with ggml_backend_meta_buffer_simple_tensor()
Accept such nodes when they are views of host tensors, which also generalizes the
previous s_copy_main workaround. This fixes the assert hit by KV cache views when
using --split-mode tensor with partial offload.
Assisted-by: pi:llama.cpp/Qwen3.8-Flash-Next
* archs : re-enable sm tensor for K2 Horizon
* cont : add TODO and reference
The mixed batch path of PR-29622 writes the token rows with set_rows
into a dup of the embeddings. WebGPU did not support DUP, so the dup
ran on the CPU while the set_rows writing into it was scheduled on
WebGPU, which then bound a CPU buffer and crashed. DUP is the same copy
as CPY and CONT and now goes through the same path.
* ui : add the drawer, sheet and grouped list primitives
Add the drawer and sheet overlay components, the toggle and toggle
group, the shared searchable input, the collapsible sections and the
grouped list, and the near viewport helper they position with.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : add the shared model components and data layer
Add the model row pieces shared by the selector and the manager, the
models store with its download status feed, the huggingface and
migration services, and the model utils, enums and constants they
read through.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : rework the models selector around its providers
Group the selector components under their own folder, rework the list
around the provider grouping, add the mobile trigger and the download
item, and derive the selection state in one hook.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : add the models manager
Add the manager table with its repo, quant and status rows, the
filters and the toolbar, the row actions in a drawer, the downloads
section, the manage models dialog and the stories and tests.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : rework the chat form actions and the mobile experience
Move the add actions into a drawer and a sheet, float the model
actions in a bar on a phone, open the context panel in a drawer, and
derive the attachment, reasoning and tools menus in hooks.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : move the mcp servers to the sidebar rail and tidy the dialogs
Move the mcp servers dialog to the sidebar rail with its menu
entries, even out the dialogs on a phone, and group the mcp and
settings stores behind their own modules.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : polish the shell
The secondary button gets its own look back and the chat add button its
own light surface. Pressable elements get a pointer cursor again, the
font rendering smooths on the app shell, and the agent skills stay out
of prettier's way.
The root layout props probe rework that came with this polish reads the
providers api url and stays with the providers change instead.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : wrap the model id classes
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : load the model the new chat CTA picks
Start a new chat selected the model but left it unloaded, so the chat
opened against a server with nothing in memory. Both CTAs share a load
helper now.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* refactor: Post-review fixes