The reserve builds n_outputs_max_per_seq sampling chains per sampler,
while a decode built one per output row, so the graph changed its
topology after the reserve and GGML_SCHED_NO_REALLOC builds aborted on
the next same sized graph. Every sampler now builds
n_outputs_max_per_seq chains, the ones without a row of the ubatch on
the padding row and not selected, and graph_max_nodes counts them.
* ui : add the drawer, sheet and grouped list primitives
Add the drawer and sheet overlay components, the toggle and toggle
group, the shared searchable input, the collapsible sections and the
grouped list, and the near viewport helper they position with.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : add the shared model components and data layer
Add the model row pieces shared by the selector and the manager, the
models store with its download status feed, the huggingface and
migration services, and the model utils, enums and constants they
read through.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : rework the models selector around its providers
Group the selector components under their own folder, rework the list
around the provider grouping, add the mobile trigger and the download
item, and derive the selection state in one hook.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : add the models manager
Add the manager table with its repo, quant and status rows, the
filters and the toolbar, the row actions in a drawer, the downloads
section, the manage models dialog and the stories and tests.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : rework the chat form actions and the mobile experience
Move the add actions into a drawer and a sheet, float the model
actions in a bar on a phone, open the context panel in a drawer, and
derive the attachment, reasoning and tools menus in hooks.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : move the mcp servers to the sidebar rail and tidy the dialogs
Move the mcp servers dialog to the sidebar rail with its menu
entries, even out the dialogs on a phone, and group the mcp and
settings stores behind their own modules.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : polish the shell
The secondary button gets its own look back and the chat add button its
own light surface. Pressable elements get a pointer cursor again, the
font rendering smooths on the app shell, and the agent skills stay out
of prettier's way.
The root layout props probe rework that came with this polish reads the
providers api url and stays with the providers change instead.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : wrap the model id classes
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : load the model the new chat CTA picks
Start a new chat selected the model but left it unloaded, so the chat
opened against a server with nothing in memory. Both CTAs share a load
helper now.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* refactor: Post-review fixes
* hexagon: fix IM2COL patch-embed DMA ring overflow
The exact-tiling (stride == kernel, no pad/dilation) IM2COL DMA kernel
issues IC*KH DDR->VTCM descriptors per output row without checking the
return value of dma_queue_push(), and then pops IC*KH times. The per-thread
DMA ring holds 256 entries and a push into a full ring returns false
and drops the transfer, so for IC*KH > 255 the remaining rows of the
VTCM staging buffer were never written and stale data (often NaN/inf)
leaked into the output.
Solution is to retire the oldest descriptor when the ring is full, just as the blocked
kernel in the same file already does, and wait with dma_queue_flush().
For testing, added exact-tiling test cases with IC*KH > 256 (2D 1x1, 2D 2x2 patch
embed and 1D, F16 and F32 dst), which fail on HTP without this fix.
* Apply suggestion from @max-krasnyansky
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* CUDA: radix top-k for large row counts
Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select,
gated on GGML_CUDA_TOPK_RADIX_MIN_ROWS. On qwen4exp at 34,816 tokens this cuts
top-k from 1,671,253 launches / 5,761.8 ms to 2,329 / 941.8 ms.
* CUDA: select the TOP_K implementation by shape
Replace the nrows/ncols special case with the decision boundary from #28547
(as implemented in #29278): bitonic for short rows, radix select for several
long rows, and DeviceTopK or CUB argsort for a single long row. The
thresholds stay overridable at build time.
Two refinements on top of that boundary:
- bitonic stays in use for rows up to a padded 1024 while the rows fit in one
wave of blocks (nrows <= number of SMs); radix select pays a fixed cost of
about a dozen launches that only amortizes over more rows
- with DeviceTopK available, it handles up to two rows
Radix select now processes rows in chunks so its scratch memory stays bounded,
and the bitonic path keeps its chunking. HIP and MUSA keep their previous
thresholds.
Add perf cases around the bitonic/radix crossover to test-backend-ops.
* CUDA: make top-k comments less verbose
* CUDA: remove the TOP_K width limit from supports_op
* CUDA: use DeviceTopK for single-row TOP_K if available
* CUDA: avoid ncols overflow in the TOP_K bitonic check
* CUDA: share the row chunking helper between argsort and top-k
* CUDA: do the TOP_K radix blocks_per_row math in int64_t
* CUDA: rename GGML_CUDA_TOP_K_NROWS_THRESHOLD_DEVICETOPK to GGML_CUDA_TOP_K_NROWS_THRESHOLD
* CUDA: share one sort helper between the bitonic and CUB TOP_K paths
* CUDA: update the TOP_K TODO, threshold and chunking comments
* tests: add TOP_K cases that span several row chunks
* CUDA: use int64_t col in the TOP_K radix loops, fix threshold comment
* CUDA: limit TOP_K and ARGSORT support to ne[0] <= INT_MAX
---------
Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com>
Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
* CUDA: fix CCCL version guard breaking on major version rollover
The guard compared the major and minor components independently:
CCCL_MAJOR_VERSION >= 3 && CCCL_MINOR_VERSION >= 1
Minor resets to 0 whenever a new major series is cut, so on CCCL 4.x
this evaluates as 4 >= 3 && 0 >= 1, i.e. false. STRIDED_ITERATOR_AVAILABLE
stops being defined and argsort silently falls back to the
init_offsets path. Nothing warns and the build still succeeds, so the
regression is a quiet performance loss rather than a compile error.
CCCL already exposes the version as a single packed integer in
MMMmmmpp form, which is what its own version header uses:
CCCL_VERSION = MAJOR * 1000000 + MINOR * 1000 + PATCH
so 3.4.3 is 3004003 and ">= 3.1" is a plain ">= 3001000". One
comparison, with no component arithmetic left to get wrong.
Checked against a hand-written "version >= 3.1" reference over 2.9.9,
3.0.0, 3.1.0, 3.1.99, 3.2.0, 3.4.3, 3.9.9, 3.99.99, 4.0.0, 4.2.7 and
5.0.0: no divergences. The old guard disagreed at 4.0.0 and 5.0.0.
Verified on RTX 4070 (sm_89), CUDA 13.4, CCCL 3.4.3:
- cmake --build build --config Release: exit 0
- test-backend-ops test -o ARGSORT -b CUDA0: 98/98 passed, CUDA0 OK
Note that a passing regression test does not on its own prove the guard
is still taken, since the fallback path passes too. Preprocessing the
real translation unit confirms the strided-iterator branch is the one
compiled in: counting_iterator is present, init_offsets is not.
Signed-off-by: Heitor <heitorgm@outlook.com>
* Update ggml/src/ggml-cuda/argsort.cu
* Apply suggestion from @ORippler
---------
Signed-off-by: Heitor <heitorgm@outlook.com>
Co-authored-by: Oliver Simons <osimons@nvidia.com>
* server : preserve context checkpoints across slot save/restore
Append the checkpoints after the packed server_tokens payload added in #26640
and count them in n_written / n_read, so a restored slot can still roll back to
a checkpoint instead of re-processing the whole prompt.
* server : drop draft checkpoint data that does not match the draft context
Restoring a slot saved with a different draft KV cache type aborted in
load_dft(). Test-load one draft checkpoint on restore and drop the draft
data if it does not fit, instead of crashing. Adds a regression test.
Co-authored-by: Igor Okulist <okigan@gmail.com>
* server : harden the checkpoint appendix of slot save files
Bound each blob size by the bytes left in the file before allocating, open the
file with UTF-8 paths on Windows like the llama state payload, fall back to full
prompt re-processing when a checkpoint restored from a slot file fails to load,
and replace the 1024 count cap by keeping the last n_ctx_checkpoints while reading.
* server : report an incomplete checkpoint appendix as a failed slot save
Return an error to the client when the appendix cannot be written, like a
failed payload write, and make the oversized-blob test declare a size that
cannot be allocated, so an unbounded allocation fails the test.
* server : reject an empty target state in the checkpoint appendix
A saved checkpoint always holds a target state, an empty blob would roll back
without restoring anything. Also log with the slot id, and load the draft test
model from the HF cache instead of a second download.
* common : return bool from checkpoint load_tgt / load_dft
A checkpoint restored from a slot file falls back to full prompt re-processing
when it fails to load, a checkpoint created in memory still aborts.
---------
Co-authored-by: Igor Okulist <okigan@gmail.com>
The bucket search in topk_nary_search.comp started from the range
[0, 0xFF800000), which ends just below the ordered-uint mapping of +inf,
so +inf and NaN were never counted. A workgroup block with fewer than k
countable values left the ballot empty and the shader read uninitialized
shared state (hang/device lost on NVIDIA, wrong indices on AMD), and a few
+inf in a block were selected without being counted, dropping real top
values.
Map NaN to -inf on input, start from [0, 0xFFFFFFFF) so every value is
counted, and clamp the top bucket's end (2^32) instead of wrapping to 0.
The k = 1 path compared float bits as signed integers, which orders
negative values backwards; compare floats instead.
Add test_top_k_inf to test-backend-ops: negative values, fewer than k
+inf and many -inf, for k = 1, 10, 40.
Assisted-by: Claude Opus 5.5
* cuda : support arbitrary striding for unary ops on f16, f32, and bf16
* Remove added newline
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* sycl: fuse the delta-net alpha gate (add + unary + mul)
* tests: cover the fused add + unary + mul chain
* sycl: give the fused alpha gate a flat path and pin the node skip
* model : support classifier_activation for rerankers
Assisted-by: Claude Opus 5.5
* model : map classifier gelu to gelu_erf and accept tanh
Assisted-by: Claude Opus 5.5
* model : default act_cls to tanh, ModernBERT falls back to gelu_erf
Assisted-by: Claude Opus 5.5
* ggml-cuda: assign two GDN state columns per warp
* ggml-cuda: use 4 GDN state columns per warp at S_v=128
* ggml-cuda: default cols_per_warp=4
* ggml-cuda: address GDN review nits
* hex-bufs: add support for alloc_buffer_n
* hex-bufs: add support for splitting large tensors into separate buffers
* hex-bufs: update GGML_HEXAGON_MBUF to accept three values dyn,static,total
* hex-bufs: bump dyn. default to 512MB since 128MB causes perf regressions with big MOEs
* hex-run: add --no-embd-offload option to simplify command lines on devices that need it
* Update scripts/snapdragon/run.py
Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com>
---------
Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com>
* model : add LiquidAI/d1-omni-600M decision model
Assisted-by: Claude Opus 5.5
* mtmd : keep conformer GLU sigmoid on CUDA
Assisted-by: Claude Opus 5.5
* server : take d1omni audio through images and input_audio, scope memory-less lfm2 to non-causal
Assisted-by: Claude Opus 5.5
* common : rename decision type d1omni to lfm2-d1-omni, server : make images an alias of files
Assisted-by: Claude Opus 5.5
* bugfix: infinite recursion caused by a tool named 'call'(#29967)
* chat : use index for schema and argument rules
* tests : remove tests
* tests : add expect_rules to peg test parser
---------
Co-authored-by: Alde Rojas <hello@alde.dev>
* model : add LiquidAI/d1-3b decision model
mtmd : read LFM2 image resize algo from GGUF
Assisted-by: Claude Opus 5.5
* common : rename decision type d1 to lfm2-d1
Assisted-by: Claude Opus 5.5
The CUDA FWHT covers widths 64 to 512. It runs one row per warp and keeps N/32
values per lane, so wider blocks need more registers per lane than that layout
allows.
fwht_cuda_block runs one row per thread block with 256 threads, so each thread
keeps N/256 values. Stages below the warp width still shuffle, those up to the
block width go through shared memory, and the rest stay in registers. Same
butterfly and sign convention as the warp kernel.
Widths 64 to 512 keep the warp kernel. 1024 through 8192 use the new one, for
both F32 and F16 sources. ggml_cuda_op_mul_mat_use_fwht (the shared
supports_op/dispatch predicate added in #29096) does not check width, so it
needed no change here: any width it admits that ggml_cuda_op_fwht can't serve
already falls through correctly to the cuBLAS path.
Rebased onto current master with #29096's F16 commit underneath it, since this
depends on the same F16 template infrastructure; that commit applied cleanly,
the only conflict was in test-backend-ops.cpp where an unrelated intervening
commit's own test additions landed near this block.
test-backend-ops on an A10 (lambdalabs): MUL_MAT 1297/1297, including all
FWHT/Hadamard cases (18 existing, 4 new F32 wide, 4 new F16 wide, 2 new
many-rows, 1 too-big boundary moved to 16384).
* llama: share the nextn tensor flags between models
Follow-up of the TODO in glm5-next: move the trunk-only and MTP-only
detection that each model copied into a nextn_flags helper of
llama_model_base. It probes the first trunk layer and the first NextN
layer, and adds TENSOR_SKIP when MTP is not loaded. qwen4exp probes
hc_attn_norm since it has no attn_norm.
deepseek4, nemotron-h, qwen35, qwen35moe, qwen3next and qwen4exp now
also accept a trunk-only file, like the other models.
* llama: avoid capturing structured bindings in the nextn flags
Lambdas that capture structured bindings need C++20, and GCC 15
rejects them under -Werror, so the models read the trunk and MTP
flags into plain variables.
The generic few-row MMA kernel works for any type with a 16-weight
dequantizer, so it now also takes BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K,
TQ2_0 and the IQ types. Each type starts at the row count where it beats
the current kernels on an M3 Ultra: 5 rows for TQ2_0, 4 for BF16, 3
for MXFP4, Q2_0, Q2_K and IQ4_NL, and 2 for the others.
test-backend-ops perf -o MUL_MAT, m=4096, k=14336, M3 Ultra, time of this
change over master (mean of two interleaved runs each): 0.23 to 0.98 from
the threshold to 8 rows, 0.24 to 0.33 at 9 to 16 rows, and 0.99 to 1.01 at
1 and 512 rows.
* metal : fix MUL_MAT+ADD fusion when the residual is itself a MUL_MAT
ggml_metal_op_mul_mat_mma picks the residual of a fused MUL_MAT+ADD as
"the ADD operand whose op is not MUL_MAT". When both operands of the ADD
are mat-mul outputs (x = W1 @ u + W2 @ v), that test is true for both, so
the residual resolves to the fused mat-mul's own, never-written output and
the kernel adds whatever that buffer holds.
The fusion check (ggml_metal_mul_mat_add_operand) already selects the
operand by identity; make the encoder do the same.
Clef decision models hit this in their head (proj_option_context @ ctx +
proj_option_lexical @ lex, 9 option rows): on Metal, /v1/systemone
probabilities collapse toward uniform (billing 0.28 where the CPU backend
gives 0.977, Cloudflare_clef-flash Q8_0), deterministic per memory layout,
correct with GGML_METAL_FUSION_DISABLE=1. Not a quantization issue: the
same file is right on CPU.
Add a MUL_MAT_ADD mode to test-backend-ops where the residual is a second
mat-mul; on Metal it fails 27 of 28 cases before this change (the one pass
is f16 n=2, under the MMA row threshold, so nothing fuses).
* Update tests/test-backend-ops.cpp
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* llama : add GLM5-Next NextN (MTP) graph
Build the GLM5-Next multi-token-prediction head as graph_mtp: the NextN block
embeds enorm(tok)+hnorm(h) through eh_proj, runs one plain DSA layer and the
shared lm_head, reusing the trunk's builders through the no_build tag ctor.
llama_memory_recurrent also tolerates a partial seq_rm when the context holds
no recurrent layers, which is what the MTP draft context needs.
Assisted-by: Claude
* llama : glm5-next: skip dead compute in headless NextN forwards
A NextN forward with no output rows (the MTP catch-up and the draft-context
prefill) persists only through its cache writes, so the headless graph keeps
the MLA latent, indexer key|gate and pooled-key writes and drops the query
path, the indexer selection, the attention body, the FFN and the LM head. The
4-token catch-up falls from 6.9 ms to 0.33 ms of kernels; the greedy output
hashes and the draft acceptance are unchanged.
Assisted-by: Claude
* llama : glm5-next: fix NextN extraction contracts and shared-tail rollback
Three fixes from the architectural review. The headless graph prune now also
requires that no unmasked nextn extraction is live, because that mode reads
n_tokens hidden rows regardless of the logits flags. Masked extraction
publishes the hidden rows gathered by the output ids, so a batch whose output
flags are not a prefix exports the right rows. A partial recurrent rollback
whose tail cell is shared with another sequence is now rejected instead of
silently moving that sequence's tail.
Assisted-by: Claude
* llama : glm5-next: tidy comments in the MTP changes
Assisted-by: Claude
* llama : glm5-next: crop the MTP graph to the output rows instead of pruning it
Replace the headless NextN prune with the crop pattern the other MTP
graphs use: gather the attention output and the block input at the
output ids before the position-wise FFN and the shared head. A NextN
forward with no output rows (the MTP catch-up and the draft-context
prefill) then runs the FFN and the head over zero rows. The 4-token
catch-up falls from 6.9 ms to 2.9 ms of kernels; greedy output hashes
are unchanged.
Assisted-by: Claude
* glm5-next: use the nextn crop helpers in the MTP graph
Replace the local crop condition and the masked select of t_h_nextn
with crop_before_nextn and crop_after_nextn, so the MTP graph narrows
its rows the same way as the main graph and the other models.
Describe the shared cell and empty filter branches of the recurrent
partial rollback.
* glm5-next: load MTP-only and trunk-only GGUF files
Make the trunk tensors optional when the file only holds the NextN
layer, and the NextN tensors optional when the file only holds the
trunk, so the split MTP GGUF loads as a draft model.
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
---------
Co-authored-by: Pascal <admin@serveurperso.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>