Commit Graph
11416 Commits
Author SHA1 Message Date
Aman Gupta 9f80c69e3b CUDA: make the alloc_deps check batch independent
Fixes #29980
2026-10-05 17:59:15 +08:00
Yufeng HeandPascal 8e1642198d server: reject partial media truncation (#24076)
* server: reject partial media truncation

* server: keep only the keep_first fix

Drop the mtmd test helper change, which no longer builds since
clip_image_f32_batch stores its entries by value, and drop the
vision test: no test fixture reaches a cut between two adjacent
media chunks with a reused cache (tinygemma3 uses SWA and wraps
images in text tokens, tinyopenjev and small-test are recurrent),
so the test passed or failed independently of the fix.

---------

Co-authored-by: Pascal <admin@serveurperso.com>
b11415
2026-10-05 11:22:23 +02:00
François-Xavier Gsell 806eee9841 vulkan: fix stale prealloc_y reuse across flash attention and soft_max (#29591)
Assisted-by: Claude
b11414
2026-10-05 10:38:26 +02:00
François-Xavier Gsell b3daa077a5 vulkan: sparse flash attention for quantized K/V (#29639)
* vulkan: sparse flash attention for quantized K/V

Assisted-by: Claude

* vulkan: single-scan sparse FA index compaction

The compaction ran one workgroup per mask row and walked the row in
BLOCK_SIZE chunks, with a workgroup scan per chunk. For decode that is
one workgroup doing KV/1024 barrier-bound iterations, so at 128k cells
it cost more than the sparse attention it feeds.

Split the row into contiguous segments instead: one per subgroup with
ballot counting over coalesced loads, or one per thread without
subgroups. A single scan over the segment counts then gives each
segment its output offset. The index list stays ascending.
b11413
2026-10-05 10:37:54 +02:00
Georgi GerganovandPascal c173a53bdf llama : fix unexpected graph reallocation in the k-pool models (#29958)
* llama : fix unexpected graph reallocation in the k-pool models

Both k-pool models built a graph shape that depends on state the
full-context reserve cannot know:

- qwen4exp branched on inp->cache_safe, which turns false as soon as
  llama_memory_seq_cp shares cells (e.g. batched-bench -pps): the QSA
  layers swapped scatter+gather for fill+concat and dropped the
  new_pool_rep leaf, so the decode graph had 12 fewer nodes than the
  reserved one
- glm5-next branched on gather = n_tokens <= 16 && n_kv > n_sel, so the
  TG decode built the gather shape (7564 nodes) while the last reserve,
  the PP one, had the dense shape (7762 nodes)

Either mismatch forces a decode-time re-reserve that drops the
worst-case sizing and bakes in the current state, so the next state
growth (n_pool, n_kv, n_new) needs more room at an unchanged graph size
and aborts under GGML_SCHED_DEBUG_REALLOC=1. Reproduce with, e.g.:

  GGML_SCHED_DEBUG_REALLOC=1 ./bin/llama-batched-bench \
    -hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K -npp 2500 -ntg 32 -npl 1,2 \
    -c 32768 -pps -kvu

Always scatter+gather the pooled keys, and pick gather from context
constants only: n_ubatch bounds every ubatch, top_k + kpool - 1 bounds
n_sel. Every graph of a context then shares one shape, which the
reserve covers, and the dense path measured faster than the gather path
at 2.5k and 16k context.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

* llama : drop the unused k-pool cache_safe graph API

The k-pool graphs no longer branch on cache_safe, so nothing reads
get_kpool_cache_safe() or the conditional new_pool_rep any more: both
models always pass the scatter target, which set_input_kpool now
requires instead of merely preferring.

Also drop the cache_safe copy in kpool_build_sizes(), a sizes-only
helper. The layout and state flag itself stays, it still decides which
pools a layout with shared cells must re-pool.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

* tests : add a shared-seq graph reserve regression test

Decode a prompt into seq 0, share its cells with seq 1 via
llama_memory_seq_cp (what llama-batched-bench does for -pps), then keep
decoding both sequences. For the k-pool models sharing clears
cache_safe, which changes the graph topology while the pools keep
growing, so a scheduler that re-reserves with the current state
instead of the worst-case one aborts under GGML_SCHED_DEBUG_REALLOC=1.
The test registration sets that flag, and the test aborts on both
k-pool models before 2220411ec1.

kimi-linear and minimax-01 are skipped: they reserve the final pp graph
with n_seqs = 1 (see [TAG_RESERVE_DIAG_DECAY] in llama-context.cpp), so
every multi-seq graph has a different layout and re-reserves by design.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

* cont : add TODOs

* cont : fix comment

* cuda: match the moe weighted reduction on empty ubatches

ggml_cuda_match_moe_weighted_reduction rejected tensors with zero
rows. A ubatch without outputs shrinks the last layer to zero rows
through inp_out_ids, so graph_optimize dropped its alloc dep there and
the scheduler graph lost one node compared to the reserved one. The
scheduler then re-reserved at the size of that ubatch, and the next
ubatch with the same node count but larger tensors aborted under
GGML_SCHED_DEBUG_REALLOC=1.

The compute loop already skips empty nodes before trying any fusion,
so the guard only made the alloc deps depend on the row count.

* tests: build the rollback test only where internal symbols link

The shared-seq case calls llm_arch_from_string, which libllama does
not export through LLAMA_API, so linking test-recurrent-state-rollback
fails on Windows with shared libraries. Its build now sits in the
NOT WIN32 OR NOT BUILD_SHARED_LIBS block, next to test-llama-archs and
the test registration it already lives under.

* tests: skip archs by name in the shared-seq reserve test

The skip of kimi-linear and minimax-01 went through llm_arch_from_string,
which libllama does not export through LLAMA_API, so the test could not
link on Windows with shared libraries. It now compares the
general.architecture string directly, and the test builds on every
platform again.

---------

Co-authored-by: Pascal <admin@serveurperso.com>
b11412
2026-10-05 11:36:25 +03:00
Evan HuusandGeorgi Gerganov 210791069b kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata (#28498)
* kv-cache: save exact KV rotation metadata, reject restoring mismatched rotation

* tests : move the state rotation test to test-save-load-state

the test is now part of the save/load test matrix and runs against
every model under test, like the rest of the suite

it probes the KV cache type combinations supported by the model and
treats models that do not use attention rotation as passing vacuously

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* cont : skip unsupported KV caches

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b11411
2026-10-05 11:27:30 +03:00
Sigbjørn Skjæret e5983d6704 ci : winget urls must be separate strings (#29978) b11410 2026-10-05 09:36:57 +02:00
Sigbjørn Skjæret 4ca6b76f0b ci : fix docker workflow permissions (#29979) 2026-10-05 09:36:03 +02:00
Alan Tseng 9f12cd4a4c ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479)
* ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60

On the SpacemiT X60, IME matrix acceleration only covered Q4_0/Q4_1/Q4_K.
Q8_0 had no IME1 kernel, and since the SpacemiT build sets
GGML_CPU_REPACK=OFF there was no repack path compiled in either, so Q8_0
had no accelerated path at all and ran roughly ten times slower than
Q4_0 for prefill on the same board.

- add make_block_q8_0x16 and the Q8_0 repack entry: interleave the
  weights into the 16-column layout the IME1 vmadot sequence expects
- add ime1::gemm_kernel_i8i8, an int8 x int8 IME1 kernel with a
  single-row and a 4-row A path; the 4-row path loads each B panel once
  and reuses it across 4 rows of A
- add quantize_a_4row_i8 for the 4-row activation quantization
- wire both into forward_mul_mat and the repack factory for Q8_0
- docs: mark Q8_0 as supported on X60

Correctness was checked against a quant-exact integer reference for
K = 32 up to 4096, with a max relative error of about 1e-6, and by
checking that generation stays coherent across several prompts.

Tested on Milk-V Jupiter (SpacemiT X60), Bianbu 2.1.1, gcc 14.2, with
Qwen2.5-0.5B-Instruct Q8_0. llama-bench -t 4 under taskset -c 0-3, 5
repetitions on an idle board: pp128 goes from 10.70 to 93.87 t/s. Q4_0
is unchanged at 106.40 -> 107.51 t/s, as expected since this does not
touch that path.

* ggml-cpu : move q8_0_16x32 decl to IME1 section

* ggml-cpu : align q8_0 IME1 kernel assignments
b11408
2026-10-05 10:34:15 +03:00
Ed Addario ebe18bee5a vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (#29912) b11407 2026-10-05 10:18:55 +03:00
Masashi Yoshimura 8216c84623 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (#29483)
* add supports for q1/q5/q3_k/q5_k/q6_k/mxfp4 of mmvq path

* Add K_QUANTS_HANDLING macro to q1_0 of mmvq path
b11406
2026-10-05 08:56:45 +02:00
Pascal 1b43d31169 cuda: tile the lightning indexer over keys and tokens for 4 heads (#29901)
* cuda: tile the lightning indexer over keys and tokens for 4 heads

With too few heads for a wmma tile, a block scores 64 keys against 8
tokens: the keys are staged once in half precision, the queries one
head at a time, and each thread owns one key for two tokens, so no dot
product needs a cross thread reduction. Batches smaller than a token
tile keep the vector kernel. test-backend-ops measures 4 heads.

* cuda: multiply the lightning indexer tile in float

Address review from am17an: the half2 products overflow once a single
q * k exceeds the f16 range. The queries stay in float in shared memory
and each half2 of keys is widened once for both tokens, so every
product and sum is computed in float.

* cuda: widen each lightning indexer key once for all heads

The tile kernel stages the queries and weights of every head at once,
so each key element is widened from half once and feeds all heads,
with a single barrier. F16 keys are copied into the tile without a
float round trip. Keeping the keys in float in shared memory measures
slower, the occupancy drops.

* cuda: stop the lightning indexer tile from spilling registers on ROCm

Each thread of the tile kernel now scores two keys for a single token,
so a warp shares its token and the query reads are broadcasts: six
shared reads per element pair instead of nine for the same products.
The inner loop is unrolled by 8, which keeps gfx908 at 63 VGPRs with no
spill where the fully unrolled loop needed over a thousand, and makes
the kernel 36x faster on an R9700 and slightly faster on CUDA.
b11405
2026-10-05 09:39:47 +03:00
pratiknarola-tandGeorgi Gerganov a3a1c4747f metal : few-row MMA mat-mul (#29869)
* metal : few-row MMA mat-mul and batched copies for speculative decoding

Speculative decoding verifies a few draft tokens per step. Without the tensor API, Metal ran these mat-muls with the mat-vec kernels, whose time grows with every src1 row, so DFlash2 decoding on an M3 Ultra was slower than serial decoding.

- add mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices: each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4_0, Q8_0 and Q5_K have their own kernels, F32, F16, Q4_1, Q5_0, Q5_1, Q4_K and Q6_K use a generic path over the 16-weight dequantizers, and Q4_0 at 2 rows uses a 2-row variant of the mat-vec kernel
- use them only on MTLGPUFamilyApple7+ without the tensor API, from the row count at which they beat the mat-vec kernels on an M3 Ultra (F32: 6, F16, Q4_K, Q5_0, Q5_1: 3, other types: 2)
- fusion table: MUL_MAT + ADD adds a same-shape residual in the MMA store, and up to 16 adjacent same-layout f32 copies between the same two tensors run as one dispatch
- the fusion checks and ggml_graph_optimize take the device props, so the reorder packs MUL_MAT + ADD only on devices that can fuse it, at every src1 row count
- views do not count toward GGML_METAL_FUSION_MAX when the reorder packs a group, so 16 recurrent state snapshot copies with views between them stay one group
- the encoder checks the inner nodes of a fused group for concurrency, tracks written views by their extent, and does not count the destination of a CPY as a read
- CONCAT splits long rows across threadgroups when there are few rows
- tests: few-row MUL_MAT, MUL_MAT_ADD, CPY_BATCH and CONCAT cases in test-backend-ops (with a prepare_graph hook for the copy order), test-metal-graph-optimize, test-metal-cpy-batch-alias

* metal : remove the CPY_BATCH fusion and the memory range changes

Remove the batched copy fusion with its kernel and tests, and revert the
memory range changes, as suggested in review. The memory ranges, the
graph reorder and the CPY encoder are again the same as on master.

* cont : clean-up

* cont : drop has_tensor gate

* cont : clean-up operand/residual logic

* cont : drop Q4_0 ne11=2 special-case

* cont : add kernels/mul_mv_mma.metal

* cont : consolidate mma pipeline selection logic

* cont : decouple fusion logic from device props

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b11404
2026-10-05 08:29:13 +03:00
ynankaniandJohannes Gäßler 9d3aba6b5e CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (#29633)
* CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size

Signed-off-by: ynankani <ynankani@nvidia.com>

* adjust kernel selection logic

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
b11403
2026-10-05 10:26:50 +05:30
anujj d89651a7b2 CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (#29435) b11402 2026-10-05 08:32:07 +05:30
PascalandXuan-Son Nguyen a7fb71fab8 log, server: self contained colors, split child commands from logs in router mode (#29895)
* log, server: make router child lines carry their own colors

The logger writes the color reset after the trailing newline, so the
reset opens the next line. On the shared pipe of a router child it lands
in front of the next state command, which the router then misses, and
the line break that works around it shows up as an empty log line on
every progress update.

The reset now goes before the trailing newlines, so every line is self
contained and the command goes back to its plain framing. The router
passes its effective color setting to its children, whose output ends
up in its terminal, and leaves that option out when comparing presets
on reload.

* log: enable ANSI colors on the Windows console

A Windows console renders ANSI sequences only in virtual terminal mode,
which nothing turns on for the logger, so llama-server prints raw escape
codes on the Windows 10 console while llama-cli, whose console code
enables it, shows colors. The logger now enables virtual terminal mode
on stdout and stderr when it turns colors on, and keeps colors off when
a console cannot render them. Pipes and files take the sequences as is.

* server: separate the router child commands from its logs

The child sent its state commands on the same pipe as its logs, so the
router had to pick them out of the log stream by a line prefix, and any
unterminated write in front of a command made the router miss it. This
resolves the TODO at the spawn that called for splitting stdout and
stderr.

The child now keeps stdout for the commands and points everything else
written to stdout at stderr, before anything is written. The router
reads both pipes, handles the commands from stdout and forwards stderr
as the log, and warns about any other line on the command pipe.

* server: address review from ngxson

The single server_child is now created first in the entry point and its
constructor keeps stdout for the commands, so the stream is a member of
the instance instead of a static, and init() is gone. The instance is
passed down to the server, while the CLI entry point creates its own.

* Update tools/server/server.cpp

---------

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
b11401
2026-10-05 01:59:37 +02:00
Xuan-Son Nguyen 0bb496dbd3 llama: support both embd + raw tokens in batch (#29622)
* llama: support both embd + raw tokens in batch

* add to test-llama-archs

* also check case llm_arch_supports_mixed_batch = false

* constant graph topology

* have dedicated input for mixed case

* rm set_tensor_backend

* is_embd --> type

* consolidate m-rope pos handling into one place

* nits
b11400
2026-10-05 01:35:49 +02:00
Johannes Gäßler 2ca15f5404 CUDA: refactor swizzling code (#29612)
* CUDA: refactor swizzling code

* fix templates/loop bounds
b11399
2026-10-04 22:48:30 +02:00
SXX a7b94df2c6 ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (#29806)
* ggml-cpu: vectorize BF16 K tails in tinyBLAS

* tests: Skip tinyBLAS when use_ref is enabled so CPU tests compare against the vec_dot path.

* ggml-cpu: vectorize tinyBLAS F16/F32 tails
b11398
2026-10-04 22:22:17 +03:00
Adrien Gallouët 0eb6d9a813 cuda : move neu_padded to where it is used (#29940)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11397
2026-10-04 22:21:31 +03:00
Sigbjørn Skjæret 2e7c58c547 ci : windows llvm build requires ninja multi-config (#29959) b11396 2026-10-04 19:28:24 +02:00
Sigbjørn Skjæret 7f2dd88b0a ci : add windows arm64 vulkan release (#29954)
* add windows vulkan arm64 release

* add link
b11395
2026-10-04 18:46:15 +02:00
Aman Gupta bf79dbbcd0 AGENTS.md : revamp (#29656)
* agents: add note about skipping forks

* rm critical line
2026-10-04 21:38:19 +05:30
Anas dbe4c3ed42 chat-peg-parser : clear current_tool when pending_tool_call is reset (#29942)
A TOOL_ID node that arrives after TOOL_CLOSE wrote through `current_tool`,
which still pointed into the just-destroyed `pending_tool_call` optional
(use-after-free, then a second free of the id buffer). Clear the pointer on
reset.
b11393
2026-10-04 17:59:21 +02:00
Sigbjørn Skjæret 46847e6158 ci : set default permissions (#29945) b11392 2026-10-04 16:13:26 +02:00
Adrien Gallouët 2bc5635734 cuda : move blocks_per_col to where it is used (#29939)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11391
2026-10-04 15:47:21 +02:00
Johannes Gäßler dd266785c2 CUDA: fix MMQ memory fault if n_expert >> n_ubatch (#29941) b11390 2026-10-04 14:10:43 +02:00
Ruben Ortlam 16c163d561 vulkan: fix rdna4 mat_vec tuning (#29934) b11389 2026-10-04 14:05:21 +02:00
0504396140 imatrix: calculate activation-based statistics for new format (GGUF) imatrices (#14891)
* Use activations to calculate the stats
* Determine calculation mode
* Compute entropy for activations
* Compute cosine similarity based on activations
* Compute l2 norm
* Add compute_layer_statistics() function
* Update aggregated statistic report layout
* Fix printing l2 norm when calc_mode = 1
* Refactor variable name
* Compute aggregated (per layer) l2 norm
* Update aggregated sum of squared activations per layer
* Make ZD Score two-tailed
* Update report layout
* Reverse conditional logic to match convention
* Rename report heading
* Add --activation-statistics parameter
* Add Euclidean–Cosine Score (ECS)
* Add --activation-statistics logic to avoid doubling the imatrix size by default
* Update stats output sort based on imatrix type
* Process external NextN draft files (-md / --model-draft)
* Refactor to use new llama_batch_ext

Co-authored-by: compilade <git@compilade.net>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
b11388
2026-10-04 11:24:53 +02:00
Pranesh GonegandlaandPranesh Gonegandla 8330e96967 spec : fix n-gram drafts rejected at temp > 0 after truncation (#29924)
Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
b11387
2026-10-04 11:45:22 +03:00
Adrien Gallouët 6716df694b common : prepare load_from_models_dir() for path conversion (#29674)
This is part of the fs::path modernization series.
That was also the opportunity to remove fs_list().

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11386
2026-10-04 11:35:33 +03:00
Adrien Gallouët bf9a0ccce7 server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (#29938)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11385
2026-10-04 11:34:43 +03:00
Sigbjørn Skjæret 0faee50042 ci : pushing tag needs deploy key (#29937) b11384 2026-10-04 09:45:49 +02:00
Sigbjørn Skjæret f98b31c67e ci : improve release flow (#29913)
* improve release flow

* fix copied typo

* fix permissions
2026-10-04 09:31:24 +02:00
Masashi Yoshimura 11fe02151f webgpu: add f16 support to fill/set_rows (#29897) b11382 2026-10-04 09:07:47 +09:00
Adrien Gallouët 836d57176d mtmd : fix deprecated strdup warning on Windows (#29863)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11381
2026-10-03 19:47:54 +02:00
Alessandro de Oliveira Faria (A.K.A.CABELO) eec18f5d32 vendor : update cpp-httplib to 0.59.0 (#29886) b11380 2026-10-03 19:08:47 +02:00
Nik Bogatyrev 1537a0a8b2 server : fix laya abort by limiting n_batch to n_ubatch (#29903)
* server : fix laya abort by limiting n_batch to n_ubatch

Fixes #29902

Assisted-by: Claude

* fix(review) : rm tests, embeddings cond
b11379
2026-10-03 17:19:13 +02:00
Adrien Gallouët edd6e2bbda common : add common_is_tty() helper and fix deprecated warnings on Windows (#29860)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11378
2026-10-03 15:48:24 +02:00
Yash Raj Pandey 9bf55f4a36 chat : honor json_schema in Ling 3.0 parser (#29813)
* chat : honor json_schema in Ling 3.0 parser

Ling 3.0 only built a grammar for tool calls and did not handle inputs.json_schema, so response_format requests were left unconstrained.

Add an eager response-format grammar path with precedence over tools, following the existing parser patterns. Require </think> before JSON when thinking is enabled and do not allow trailing prose after the JSON response.

Fixes #29652.

Assisted-by: Claude Opus 5.5

* chat : require Ling 3.0 think block for response formats
b11377
2026-10-03 15:40:50 +02:00
Pascal a55e952b85 ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (#29904) b11376 2026-10-03 14:56:07 +02:00
Pascal 436f6f89e1 graph: gather the recurrent states once so the reserve covers every split (#29856)
build_rs gathered the extra states (n_rs - n_seqs rows) with their own
get_rows. The worst-case reserve has n_rs == n_seqs, so that node was
sized at zero rows, and any ubatch whose cells are not contiguous forced
a graph reallocation at an unchanged node count, which aborts under
GGML_SCHED_NO_REALLOC.

A single get_rows now gathers the n_rs states: the ubatch states and the
extra states are views of it, and its size only depends on n_rs, which
the reserve already sets to the maximum. A custom getter (mamba ssm_scan)
gathers from the second state, so a single sequence ubatch copies no
state. The views are built once per graph in the input to keep the host
overhead of the graph unchanged.
b11375
2026-10-03 14:02:05 +02:00
b92761a515 ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (#29852)
* ggml-openvino : Qwen3.5 MoE perf (#312)

Squash of ravi9/llama.cpp#312:

- ggml-openvino: add detailed inference profiling (Yu, Zijun)
- ggml-openvino: use remote output tensors by default (Yu, Zijun)
- ggml-openvino: optimize single-sequence recurrent state (Yu, Zijun)
- opt1: remove recurrent reset for single sequence, opt2: direct gdn outputs (break parallel sequence) (Yu, Zijun)
- fix parallel sequences (Yu, Zijun)
- ggml-openvino: simplify graph cache key (ynimmaga)
- enable stateful for qwen35 single sequence (Yu, Zijun)
- Fix after rebasing (Yu, Zijun)
- Add k-requant option q4_asym64 (Yu, Zijun)
- Fix qwen35 llama-bench -p 0 (Yu, Zijun)
- Simplify RESHAPE translation (Yu, Zijun)
- openvino: fuse MoE routing (Yu, Zijun)
- openvino: fuse GDN qk normalization (Yu, Zijun)
- openvino: enable GPU MoE fusion by default (Yu, Zijun)
- ggml-openvino: add cache_only mode to import cached compiled model on disk directly (Yu, Zijun)
- openvino : report the device allocation limit to ggml (Łukasz Ślusarczyk)
- Fix windows build (Yu, Zijun)

Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com>
Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com>

* ggml-openvino: Update doc of compiled model cache

* openvino: implement PRD-compliant device enumeration and memory reporting

* openvino: fix multi-device listing issues from review

- Only the device selected by GGML_OPENVINO_DEVICE reports as GPU; the
  other OpenVINO devices report as IGPU so llama.cpp does not offload to
  them. Initializing a non-selected device logs a warning.
- Name devices OPENVINO<i> again and show the OpenVINO id in the
  description. Raw "CPU" names shadowed the ggml CPU backend.
- Support GPU.N: create the OpenCL queue on OpenVINO's own context for
  the selected device, and replace "GPU"/"NPU" string comparisons with
  ggml_openvino_is_gpu()/ggml_openvino_is_npu().
- An unavailable GGML_OPENVINO_DEVICE is now an error that lists the
  available devices, instead of silently falling back to CPU.
- Memory: cap iGPU/NPU free memory at system available memory, fall back
  to system memory instead of 0/0 when the plugin lacks memory
  properties, and ignore host USM allocations in GPU usage.
- Initialize the device config once under a lock, even if OpenCL setup
  fails.
- Fix supports_op return type for non-selected devices (build error).

* openvino : take USM entry points from the selected device platform

clGetExtensionFunctionAddressForPlatform was called on the first platform
returned by clGetPlatformIDs. The address it returns is only valid for the
platform it was queried on, and the first platform is not always the one that
holds the device OpenVINO selected.

On a host whose first platform comes from another vendor the lookup returns
null, and then every read, write and memset on a GPU buffer fails with
"clEnqueueMemcpyINTEL not available".

Look both entry points up in init(), on the platform of the device OpenVINO
picked, and keep them in the device config next to the command queue.

Assisted-by: Claude Opus 5

* openvino: fuse MoE experts for models with a fused gate_up weight

FuseMoeCompressed only matches models whose gate and up projections are
separate GatherMatmul ops. gemma-4 packs both into one expert weight and
splits the result after the GEMM, so its MoE block stayed unfused and ran
the expert GEMMs as per-token GEMVs.

Add FuseMoeCompressedFusedGateUp, which matches that shape
(one GatherMatmul -> Slice/Slice -> Gelu(ERF) -> Multiply) and folds it into
the same MOECompressed op, using GEMM3_SWIGLU with GEGLU_ERF. The fused
weight, scale and zero point are split into gate/up halves by copying raw
bytes, since a graph Slice would be rewritten to StridedSlice and constant
folded, whose reference evaluator crashes on sub-byte types.

gemma-4 also applies a per-expert output scale to the down projection before
the router weights. MOECompressed takes only one per-expert weight, so that
scale is folded into the routing weights, which is exact.

The op reads the zero point straight off a weight port and needs an integer
Constant there, so the matcher requires one and leaves natively quantized
experts (exact f16 zp) to the unfused path.

gemma-4-26B-A4B on Arc B390, GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all,
llama-bench -p 512 -n 128 -r 2, against a GGML_OPENVINO_MOE_OP=0 baseline:
pp512 66.16 -> 1608.73 t/s, tg128 25.94 -> 26.46 t/s. Perplexity over 12
chunks is unchanged (1451.3 +/- 177.9 unfused vs 1427.6 +/- 175.1 fused).

No effect without that requant option, on models with separate gate/up
weights, or on CPU. test-backend-ops -b OPENVINO0 is unchanged by this
commit: two MUL_MAT_ID m_v cases fail, the same two on the unmodified base.

* openvino: fix rank-3 axis handling so MoE works under stateful execution

Stateful execution drops the leading size-1 batch dim, so OV tensors are rank
3 while GgmlOvDecoder::get_shape/get_stride still report GGML_MAX_DIMS=4
reversed entries. Several MoE ops derive OV axis indices straight from that
metadata, so they picked the wrong axis. A MoE model with
GGML_OPENVINO_STATEFUL_EXECUTION=1 aborts while building the graph:

  Check 'is_axis_valid(axis, r)' failed at src/core/src/validation_util.cpp:336
  While validating node 'opset11::TopK ... _ffn_moe_probs ...'
  Axis 3 out of the tensor rank range [-3, 2].

Fix idiom throughout: take the axis from the real OV rank, or shift a
metadata-derived axis down by metadata_rank - actual_rank.

  argsort.cpp    the router top-k axis is 2 on rank 3, not 3. This is the
                 abort quoted above.
  add.cpp        the MoE expert-sum bypass collapses the 8-ADD chain into one
                 ReduceSum on hardcoded axis 2, which on rank 3 reduces n_embd
                 instead of the expert axis. Now rank-2, with the following
                 Unsqueeze at rank-3.
  get_rows.cpp   squeezing a hardcoded {0,1} also strips the batch dim
                 whenever it is 1, which is every decode step. Squeeze down to
                 the trailing two dims instead.
  mul_mat_id.cpp pick the reshape dims by actual rank, and skip the trailing
                 Unsqueeze that re-adds the batch dim.
  view.cpp       the expert-plane slice had the Slice axis, dst_ov_axis, the
                 ShapeOf+Gather index and the Reshape target all rank-4.
  utils.cpp      process_view_input_new's "translate_view already resolved
                 this VIEW, skip re-slicing" shortcut required equal ranks. 4
                 vs 3 never matched, so every resolved expert plane got
                 re-sliced. Now compares the common trailing dims. Same axis
                 shift for the Slice in the view-chain walker.

Stateless is unchanged by construction: every edit is gated on the actual
rank, so axis_shift == 0 reproduces the previous code exactly. Checked on
OV-CPU by diffing greedy output against the unmodified base for dense
gemma-4-E2B, granite-1b-a400m and gemma-4-26B-A4B; all identical.

granite-1b-a400m on OV-CPU aborts with the error above before this change;
after it, it generates and is byte-identical to stateless. Dense gemma-4-E2B
is identical stateless vs stateful both before and after. test-backend-ops
-b OPENVINO0 is unchanged: two pre-existing MUL_MAT_ID m_v cases fail, the
same two on the unmodified base.

gemma-4-26B-A4B is a poor correctness vehicle here. On OV it already drifts
into degenerate repetition a few tokens in, in stateless as much as stateful,
and the two modes diverge somewhere inside that degenerate region instead of
matching token for token. Each mode is self-reproducible across runs.

Known limitation: FuseMoeCompressedFusedGateUp does not match the rank-3
graph, so a MoE model run with GGML_OPENVINO_STATEFUL_EXECUTION=1 loses the
prefill fusion while gaining decode. gemma-4-26B-A4B on Arc B390,
GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all, llama-bench -p 512 -n 128 -r 2:

  unfused (GGML_OPENVINO_MOE_OP=0)  pp512   66.16   tg128  25.94
  fused, stateless (default)        pp512 1608.73   tg128  26.46
  fused, stateful                   pp512   66.18   tg128  29.91

Stateful is opt-in and off by default, and MoE did not run there at all
before this, so nothing that previously worked regresses. Making the pass
match rank 3 is the follow-up.

* OpenVINO Backend: Upgrade graph cache to use node_idx, src_idx, node type

* ggml-openvino : enable more comprehensive conv fusion

* enable conv ops

* Reject kernel size 0 and support IM2COL_3D

* openvino : abort when the GPU remote context cannot be created

init() logged the error and returned, which left the device name a GPU but
remote_context empty. The remote buffer and tensor paths assert only on the
device being a GPU and then dereference that empty optional.

Those paths have no host fallback, and a device that OpenVINO listed should
have a working OpenCL context, so stop instead of continuing. An OpenCL stack
that is broken as a whole is still caught earlier by the device availability
check, which falls back to CPU.

Assisted-by: Claude Opus 5

* openvino : fix build warnings

The single-argument form of the OpenVINO RTTI macros is the intended one, but
their selector macro leaves __VA_ARGS__ empty, which -Wpedantic reports on
every pass and op header. Turn that warning off for this backend only, the
way ggml-cuda and ggml-sycl already do for their own third-party warnings.

Also drop a break and a dead assignment around a GGML_ABORT, which is noreturn.

Assisted-by: Claude Opus 5

* OpenVINO Backend: Support common MTMD ops

* ggml-openvino: give a reshaping view its own ov::Tensor

* ggml-openvino : compute HARDSIGMOID and EXPM1 in f32

HARDSIGMOID used a 1/6 constant in the input type, which is not exact
in bf16, and EXPM1 lost precision for small inputs in f16. Both now
compute in f32 and convert back, except on NPU where the f32 path
gives wrong results.

Fixes the HARDSIGMOID/EXPM1 test-backend-ops failures on GPU.

* ggml-openvino : update device selection and --list-devices

Show the selecting GGML_OPENVINO_DEVICE value and active device in
--list-devices, startup logs, and backend tests.

Clarify OpenVINO selection uses GGML_OPENVINO_DEVICE, not -dev.

* openvino : remove unreachable OpenCL queue checks

A remote buffer exists only on a GPU device, and init() aborts there if the
queue cannot be created, so the queue is never null at these call sites.

Assisted-by: Claude Opus 5

* openvino : update OpenVINO to 2026.4.1 and GPU drivers to 26.35.39758.10

* docs : update OpenVINO validated models and GPU driver version

* ggml-openvino : skip empty views when giving a reshaping view its own tensor

A zero-size view can sit at the end of a GPU USM buffer (Qwen3.5 recurrent cache). Wrapping it as a remote tensor throws "shared USM buffer has smaller size (0)".

Assisted-by: Claude

* ggml-openvino : rebind the cached decoder when llama passes a different graph

llama keeps separate graphs for batches with and without outputs. llama-server splits the prompt into chunks for context checkpoints, so a cached decoder could be reused with a graph built in other memory and bind the previous chunk's input tensors. SWA and recurrent models then lost most of the prompt in llama-cli and llama-server.

Assisted-by: Claude

* docs : update OpenVINO validated models

Smoke test on Lunar Lake (32 GB) with the two fixes above. Re-add the Qwen3.5 and gemma models.

Assisted-by: Claude

---------

Co-authored-by: Yu, Zijun <zijun.yu@intel.com>
Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com>
Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com>
Co-authored-by: haarika-madaka <haarika.madaka@intel.com>
Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com>
Co-authored-by: Mostafa Faheem <mostafaaafaheem@gmail.com>
b11374
2026-10-03 11:59:25 +03:00
Tarek Dakhran cb7934c52c model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M (#29862)
Register `Lfm2BidirectionalForMaskedLM` architecture for LFM2.5-Encoder
models.
2026-10-03 08:44:45 +02:00
PascalandRuben Ortlam 889edf43dd qwen4exp : halve the indexer score memory (#29825)
* qwen4exp : halve the indexer score memory

The indexer scored all heads in one product and rectified a copy of it,
so two [n_pool, n_idx_h, n_tokens] f32 tensors were live at once, the
largest buffers of the graph at long context. Each head now gets its
own product, rectified and summed in place into one [n_pool, n_tokens]
score.

* qwen4exp: let the allocator reuse the indexer score buffers

Address review from CISC: use plain ggml_add and ggml_relu in the
indexer head loop. The graph allocator already runs them in place when
their source has no other consumer, so the _inplace variants are not
needed. The compute buffer and the speed are unchanged.

* cuda: support 4 heads in the lightning indexer

Dispatch 4 heads to the vector kernel, too few for a wmma tile, and
accept them in supports_op. test-backend-ops covers 4 heads.

* metal: take the lightning indexer head count as a function constant

The kernel reads the head count from a function constant and zero fills
the last head tile, so any head count runs and 64 heads is unchanged.

* qwen4exp: compute the indexer score with the lightning indexer

Address review from am17an: the unweighted sum of the rectified head
scores scaled by 1/sqrt(head_dim) is the lightning indexer with every
head weight set to that scale, so the indexer calls
ggml_lightning_indexer on the pooled keys with an f16 pool mask. The
keys are read once for all heads and no per head score is
materialized.

* vulkan: tile the lightning indexer over keys and tokens

A workgroup scores 64 keys against 8 tokens: the keys are staged once
in shared memory, the queries one head at a time, and each invocation
owns one key for two tokens, so no dot product needs a cross invocation
reduction. The subgroup variant and the flat dispatch are gone, the grid
is keys x tokens x streams.

* vectorize vulkan loads and use fp16 dot product

---------

Co-authored-by: Ruben Ortlam <rortlam@redhat.com>
b11372
2026-10-03 07:19:00 +02:00
Xuan-Son NguyenandSigbjørn Skjæret 99b95488ca model: add support for clef decision model (text-only) (#29831)
* init support for clef (text only)

* more static graph

* clean up

* nits

* nits 2

* Update gguf-py/gguf/constants.py

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
b11371
2026-10-03 02:50:48 +02:00
Aman Gupta bed0a85660 CUDA: fuse shared experts into MMVQ (#29184)
* CUDA: fuse shared experts into MMVQ

* check if buffer is null

* move stride_col_dst to fusion args
b11370
2026-10-02 21:26:27 +03:00
Sigbjørn Skjæret 4ebdf2c74a ci : use t4-medium for cuda jobs (#29842)
[no ci]
2026-10-02 17:31:19 +02:00
1fb7ef3e33 spec : add probabilistic sampling for simple draft and MTP (#27694)
* Make the drafter probabilistic and the target verify by rejection sampling

* Drop stale spec_draft_q before drafting

* Fallback to argmax sampling for grammar-constrained requests and adding flag for enabling probabilistic draft sampling. Default flag value is greedy.

* Support grammar-constrained requests in rejection sampling

* Fix - renormalize distribution after masking

* copy rng on sampler copy and re-accept drafted tokens on replay

* Fix draft sampler sharing the target's rng stream

* Simplify the rejection sampler's inputs and move replay to the server

* Truncate the draft candidates along with the draft

---------

Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com>
Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
b11368
2026-10-02 17:52:54 +03:00
Yash Raj Pandey 134b2bb756 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (#27663) 2026-10-02 22:47:30 +08:00