* cuda: stage the lightning indexer queries in head passes for MUSA
MUSA archs 21 and 22 cap static shared memory at 28 KB, and the tile
kernel staged the queries of all four heads next to the key tile for
33 KB. The queries are now staged in passes of
LIGHTNING_INDEXER_TILE_HEADS_PER_PASS heads: two on MUSA for 25 KB,
four elsewhere where the single pass folds to the previous kernel.
* cuda: use the vector lightning indexer kernel on MUSA
Address review from am17an: the tile kernel stays off MUSA, whose archs
21 and 22 cap static shared memory at 28 KB, below the 33 KB the tile
needs, so MUSA keeps the vector kernel it ran before. This replaces the
head passes, CUDA and ROCm run the merged kernel unchanged.
* server: support vision input for Clef
* move input_attn_causal to private
* extend old server_batch::embd
* server_batch::token::pos to multi dim
* nits
* fix abort
* fix img tokens cap
* fix yield_to_queue mutate data
* server: reject partial media truncation
* server: keep only the keep_first fix
Drop the mtmd test helper change, which no longer builds since
clip_image_f32_batch stores its entries by value, and drop the
vision test: no test fixture reaches a cut between two adjacent
media chunks with a reused cache (tinygemma3 uses SWA and wraps
images in text tokens, tinyopenjev and small-test are recurrent),
so the test passed or failed independently of the fix.
---------
Co-authored-by: Pascal <admin@serveurperso.com>
* vulkan: sparse flash attention for quantized K/V
Assisted-by: Claude
* vulkan: single-scan sparse FA index compaction
The compaction ran one workgroup per mask row and walked the row in
BLOCK_SIZE chunks, with a workgroup scan per chunk. For decode that is
one workgroup doing KV/1024 barrier-bound iterations, so at 128k cells
it cost more than the sparse attention it feeds.
Split the row into contiguous segments instead: one per subgroup with
ballot counting over coalesced loads, or one per thread without
subgroups. A single scan over the segment counts then gives each
segment its output offset. The index list stays ascending.
* llama : fix unexpected graph reallocation in the k-pool models
Both k-pool models built a graph shape that depends on state the
full-context reserve cannot know:
- qwen4exp branched on inp->cache_safe, which turns false as soon as
llama_memory_seq_cp shares cells (e.g. batched-bench -pps): the QSA
layers swapped scatter+gather for fill+concat and dropped the
new_pool_rep leaf, so the decode graph had 12 fewer nodes than the
reserved one
- glm5-next branched on gather = n_tokens <= 16 && n_kv > n_sel, so the
TG decode built the gather shape (7564 nodes) while the last reserve,
the PP one, had the dense shape (7762 nodes)
Either mismatch forces a decode-time re-reserve that drops the
worst-case sizing and bakes in the current state, so the next state
growth (n_pool, n_kv, n_new) needs more room at an unchanged graph size
and aborts under GGML_SCHED_DEBUG_REALLOC=1. Reproduce with, e.g.:
GGML_SCHED_DEBUG_REALLOC=1 ./bin/llama-batched-bench \
-hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K -npp 2500 -ntg 32 -npl 1,2 \
-c 32768 -pps -kvu
Always scatter+gather the pooled keys, and pick gather from context
constants only: n_ubatch bounds every ubatch, top_k + kpool - 1 bounds
n_sel. Every graph of a context then shares one shape, which the
reserve covers, and the dense path measured faster than the gather path
at 2.5k and 16k context.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
* llama : drop the unused k-pool cache_safe graph API
The k-pool graphs no longer branch on cache_safe, so nothing reads
get_kpool_cache_safe() or the conditional new_pool_rep any more: both
models always pass the scatter target, which set_input_kpool now
requires instead of merely preferring.
Also drop the cache_safe copy in kpool_build_sizes(), a sizes-only
helper. The layout and state flag itself stays, it still decides which
pools a layout with shared cells must re-pool.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
* tests : add a shared-seq graph reserve regression test
Decode a prompt into seq 0, share its cells with seq 1 via
llama_memory_seq_cp (what llama-batched-bench does for -pps), then keep
decoding both sequences. For the k-pool models sharing clears
cache_safe, which changes the graph topology while the pools keep
growing, so a scheduler that re-reserves with the current state
instead of the worst-case one aborts under GGML_SCHED_DEBUG_REALLOC=1.
The test registration sets that flag, and the test aborts on both
k-pool models before 2220411ec1.
kimi-linear and minimax-01 are skipped: they reserve the final pp graph
with n_seqs = 1 (see [TAG_RESERVE_DIAG_DECAY] in llama-context.cpp), so
every multi-seq graph has a different layout and re-reserves by design.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
* cont : add TODOs
* cont : fix comment
* cuda: match the moe weighted reduction on empty ubatches
ggml_cuda_match_moe_weighted_reduction rejected tensors with zero
rows. A ubatch without outputs shrinks the last layer to zero rows
through inp_out_ids, so graph_optimize dropped its alloc dep there and
the scheduler graph lost one node compared to the reserved one. The
scheduler then re-reserved at the size of that ubatch, and the next
ubatch with the same node count but larger tensors aborted under
GGML_SCHED_DEBUG_REALLOC=1.
The compute loop already skips empty nodes before trying any fusion,
so the guard only made the alloc deps depend on the row count.
* tests: build the rollback test only where internal symbols link
The shared-seq case calls llm_arch_from_string, which libllama does
not export through LLAMA_API, so linking test-recurrent-state-rollback
fails on Windows with shared libraries. Its build now sits in the
NOT WIN32 OR NOT BUILD_SHARED_LIBS block, next to test-llama-archs and
the test registration it already lives under.
* tests: skip archs by name in the shared-seq reserve test
The skip of kimi-linear and minimax-01 went through llm_arch_from_string,
which libllama does not export through LLAMA_API, so the test could not
link on Windows with shared libraries. It now compares the
general.architecture string directly, and the test builds on every
platform again.
---------
Co-authored-by: Pascal <admin@serveurperso.com>
* kv-cache: save exact KV rotation metadata, reject restoring mismatched rotation
* tests : move the state rotation test to test-save-load-state
the test is now part of the save/load test matrix and runs against
every model under test, like the rest of the suite
it probes the KV cache type combinations supported by the model and
treats models that do not use attention rotation as passing vacuously
Assisted-by: pi:llama.cpp/Qwen3.8-27B
* cont : skip unsupported KV caches
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60
On the SpacemiT X60, IME matrix acceleration only covered Q4_0/Q4_1/Q4_K.
Q8_0 had no IME1 kernel, and since the SpacemiT build sets
GGML_CPU_REPACK=OFF there was no repack path compiled in either, so Q8_0
had no accelerated path at all and ran roughly ten times slower than
Q4_0 for prefill on the same board.
- add make_block_q8_0x16 and the Q8_0 repack entry: interleave the
weights into the 16-column layout the IME1 vmadot sequence expects
- add ime1::gemm_kernel_i8i8, an int8 x int8 IME1 kernel with a
single-row and a 4-row A path; the 4-row path loads each B panel once
and reuses it across 4 rows of A
- add quantize_a_4row_i8 for the 4-row activation quantization
- wire both into forward_mul_mat and the repack factory for Q8_0
- docs: mark Q8_0 as supported on X60
Correctness was checked against a quant-exact integer reference for
K = 32 up to 4096, with a max relative error of about 1e-6, and by
checking that generation stays coherent across several prompts.
Tested on Milk-V Jupiter (SpacemiT X60), Bianbu 2.1.1, gcc 14.2, with
Qwen2.5-0.5B-Instruct Q8_0. llama-bench -t 4 under taskset -c 0-3, 5
repetitions on an idle board: pp128 goes from 10.70 to 93.87 t/s. Q4_0
is unchanged at 106.40 -> 107.51 t/s, as expected since this does not
touch that path.
* ggml-cpu : move q8_0_16x32 decl to IME1 section
* ggml-cpu : align q8_0 IME1 kernel assignments
* cuda: tile the lightning indexer over keys and tokens for 4 heads
With too few heads for a wmma tile, a block scores 64 keys against 8
tokens: the keys are staged once in half precision, the queries one
head at a time, and each thread owns one key for two tokens, so no dot
product needs a cross thread reduction. Batches smaller than a token
tile keep the vector kernel. test-backend-ops measures 4 heads.
* cuda: multiply the lightning indexer tile in float
Address review from am17an: the half2 products overflow once a single
q * k exceeds the f16 range. The queries stay in float in shared memory
and each half2 of keys is widened once for both tokens, so every
product and sum is computed in float.
* cuda: widen each lightning indexer key once for all heads
The tile kernel stages the queries and weights of every head at once,
so each key element is widened from half once and feeds all heads,
with a single barrier. F16 keys are copied into the tile without a
float round trip. Keeping the keys in float in shared memory measures
slower, the occupancy drops.
* cuda: stop the lightning indexer tile from spilling registers on ROCm
Each thread of the tile kernel now scores two keys for a single token,
so a warp shares its token and the query reads are broadcasts: six
shared reads per element pair instead of nine for the same products.
The inner loop is unrolled by 8, which keeps gfx908 at 63 VGPRs with no
spill where the fully unrolled loop needed over a thousand, and makes
the kernel 36x faster on an R9700 and slightly faster on CUDA.
* metal : few-row MMA mat-mul and batched copies for speculative decoding
Speculative decoding verifies a few draft tokens per step. Without the tensor API, Metal ran these mat-muls with the mat-vec kernels, whose time grows with every src1 row, so DFlash2 decoding on an M3 Ultra was slower than serial decoding.
- add mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices: each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4_0, Q8_0 and Q5_K have their own kernels, F32, F16, Q4_1, Q5_0, Q5_1, Q4_K and Q6_K use a generic path over the 16-weight dequantizers, and Q4_0 at 2 rows uses a 2-row variant of the mat-vec kernel
- use them only on MTLGPUFamilyApple7+ without the tensor API, from the row count at which they beat the mat-vec kernels on an M3 Ultra (F32: 6, F16, Q4_K, Q5_0, Q5_1: 3, other types: 2)
- fusion table: MUL_MAT + ADD adds a same-shape residual in the MMA store, and up to 16 adjacent same-layout f32 copies between the same two tensors run as one dispatch
- the fusion checks and ggml_graph_optimize take the device props, so the reorder packs MUL_MAT + ADD only on devices that can fuse it, at every src1 row count
- views do not count toward GGML_METAL_FUSION_MAX when the reorder packs a group, so 16 recurrent state snapshot copies with views between them stay one group
- the encoder checks the inner nodes of a fused group for concurrency, tracks written views by their extent, and does not count the destination of a CPY as a read
- CONCAT splits long rows across threadgroups when there are few rows
- tests: few-row MUL_MAT, MUL_MAT_ADD, CPY_BATCH and CONCAT cases in test-backend-ops (with a prepare_graph hook for the copy order), test-metal-graph-optimize, test-metal-cpy-batch-alias
* metal : remove the CPY_BATCH fusion and the memory range changes
Remove the batched copy fusion with its kernel and tests, and revert the
memory range changes, as suggested in review. The memory ranges, the
graph reorder and the CPY encoder are again the same as on master.
* cont : clean-up
* cont : drop has_tensor gate
* cont : clean-up operand/residual logic
* cont : drop Q4_0 ne11=2 special-case
* cont : add kernels/mul_mv_mma.metal
* cont : consolidate mma pipeline selection logic
* cont : decouple fusion logic from device props
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* log, server: make router child lines carry their own colors
The logger writes the color reset after the trailing newline, so the
reset opens the next line. On the shared pipe of a router child it lands
in front of the next state command, which the router then misses, and
the line break that works around it shows up as an empty log line on
every progress update.
The reset now goes before the trailing newlines, so every line is self
contained and the command goes back to its plain framing. The router
passes its effective color setting to its children, whose output ends
up in its terminal, and leaves that option out when comparing presets
on reload.
* log: enable ANSI colors on the Windows console
A Windows console renders ANSI sequences only in virtual terminal mode,
which nothing turns on for the logger, so llama-server prints raw escape
codes on the Windows 10 console while llama-cli, whose console code
enables it, shows colors. The logger now enables virtual terminal mode
on stdout and stderr when it turns colors on, and keeps colors off when
a console cannot render them. Pipes and files take the sequences as is.
* server: separate the router child commands from its logs
The child sent its state commands on the same pipe as its logs, so the
router had to pick them out of the log stream by a line prefix, and any
unterminated write in front of a command made the router miss it. This
resolves the TODO at the spawn that called for splitting stdout and
stderr.
The child now keeps stdout for the commands and points everything else
written to stdout at stderr, before anything is written. The router
reads both pipes, handles the commands from stdout and forwards stderr
as the log, and warns about any other line on the command pipe.
* server: address review from ngxson
The single server_child is now created first in the entry point and its
constructor keeps stdout for the commands, so the stream is a member of
the instance instead of a static, and init() is gone. The instance is
passed down to the server, while the CLI entry point creates its own.
* Update tools/server/server.cpp
---------
Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
* llama: support both embd + raw tokens in batch
* add to test-llama-archs
* also check case llm_arch_supports_mixed_batch = false
* constant graph topology
* have dedicated input for mixed case
* rm set_tensor_backend
* is_embd --> type
* consolidate m-rope pos handling into one place
* nits
* ggml-cpu: vectorize BF16 K tails in tinyBLAS
* tests: Skip tinyBLAS when use_ref is enabled so CPU tests compare against the vec_dot path.
* ggml-cpu: vectorize tinyBLAS F16/F32 tails
A TOOL_ID node that arrives after TOOL_CLOSE wrote through `current_tool`,
which still pointed into the just-destroyed `pending_tool_call` optional
(use-after-free, then a second free of the id buffer). Clear the pointer on
reset.
This is part of the fs::path modernization series.
That was also the opportunity to remove fs_list().
Signed-off-by: Adrien Gallouët <angt@huggingface.co>