Probe matrix proved K is the broken side of turbo KV on Qwen3-4B (QK-norm gamma outliers vs PolarQuant's no-per-channel-range format) while V tolerates 2-bit for free. Applied to ornith duo main: K q8_0 + V turbo2 = 128K ctx in ~880MB RAM at full speed (8.7-10.1 t/s concurrent, vs 3x decode loss with turbo2 K). Qwen duo stays q4/q4 24K — only clean mix is bigger than q4/q4. InnerQ equalizer confirmed broken as shipped (compensation error grows with scale strength); documented in MOE-FINDINGS with the full matrix and fix ladder. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
8.0 KiB
MoE Findings — xps9700, 2026-07-10
Session: big-model-runner scope shifted to this laptop (15 GiB DDR4-2933, GTX 1650 Ti
3.7 GB VRAM, i7-10750H 6c/12t, SN730 PCIe3 NVMe). All numbers from server .timings
at temp 0, structure-gated (JSON validity + bracket balance), disk quiet.
Method + priors from the x570 flash-162b record (gitea: mozempk/big-model-runner).
Scoreboard
| model / config | tg t/s | pp t/s | gates |
|---|---|---|---|
| ornith-35b Q2_K_L (13.1 GB, 3 exp layers VRAM) | 29.1 | 66 | pass |
| qwen36-35b UD-Q2_K_XL (12.3 GB, 3 exp layers VRAM) | 23.0 | 36 | pass |
| qwen36-35b Q2 on upstream master | 22.8 | 28–32 | pass |
qwen36-35b Q2 on ik_llama (+-ser 6,1) |
18.0 | 42 | pass |
gpt-oss-20b MXFP4 (--n-cpu-moe 21) |
16.7 | 29.5 | pass |
gpt-oss-20b (--cpu-moe, all experts CPU) |
14.4 | 22.3 | pass |
| qwen36-35b UD-IQ4_XS 17.7 GB (thrash) | 2.8 | 3.7 | pass |
| ornith-9b / qwen3.5-9b dense Q8 (old ceiling) | 4.4 | ~45 | — |
Laws of this machine
- The page-cache cliff: MoE offload is fast iff the GGUF fits page cache (~12–13 GB with services running). 12.3 GB → 23 t/s; 17.7 GB → 2.8 t/s (measured 1.8 GB/s sustained NVMe page-in, ~650 MB faulted/token — eviction churn, warm == cold).
- ds4/x570 asymmetric quant recipe transfers: routed experts tolerate 2-bit; dense/attention/embeddings must stay high-bit. unsloth UD-Q2_K_XL and bartowski Q2_K_L are pre-made versions of this mix. Structure gates pass; 35B-A3B @ Q2-experts beats 20B @ 4-bit on both speed and (by benchmarks) quality.
- VRAM expert placement:
--n-cpu-moe N(first N layers' experts → CPU, rest → GPU; direction verified empirically). ~455 MB/layer (gpt-oss MXFP4), ~230 MB/layer (qwen36 Q2). 3 layers in spare VRAM = +16% tg, +32% pp on gpt-oss. Which layers doesn't matter when file is cache-resident (middle-hot vs last-3: 16.79 vs 16.73) — only how many. 6 layers OOMs (compute buffers need ~500 MB). - Builds are a wash for K-quant decode: turboquant (May) == upstream (Jul)
== within noise. ik_llama: −23% decode / +13% prefill here;
-ser 6,1marginal. Keep turboquant as default binary (turbo2 KV for the dense 9Bs). - MTP / speculative decode: skip (x570 measured −20% net at 80% acceptance — expert-union tax; worse when disk-bound; our GGUFs lack MTP tensors anyway).
- Codacus commits: skip (−5% on x570; its cudaHostRegister mmap-pinning would try to pin >RAM here).
- Benching discipline: never bench while downloads/builds run — page-cache
flushing fakes an 80% regression (measured 16.7 → 3.2 on identical config).
pkill -fpatterns self-match the invoking shell — SIGSTOP'd our own bench once.
Duo config (resident main + subagent, 2026-07-10)
llama-swap group duo (swap: false, exclusive: true): ornith-35b-duo +
qwen3-4b-duo stay loaded together for pi (main coding model + fast subagent).
Requesting any NON-duo model unloads the whole group — pi must use the -duo ids.
VRAM goes to the small model, not the big one. User observation confirmed:
ornith's routed experts never load into VRAM, and its dense-on-GPU split
(2354 MiB) starved qwen down to 176 MiB via --fit on. Flipped: qwen3-4b
ngl 99 owns the GPU (3240 MiB incl. KV+compute), ornith runs pure CPU.
| duo member | config | solo t/s | concurrent t/s |
|---|---|---|---|
| qwen3-4b-duo (before) | --fit on, 176 MiB VRAM |
10.3 | 5.8 |
| ornith-35b-duo (before) | dense GPU, experts CPU | 16.2 | 10.6 |
| qwen3-4b-duo (after) | ngl 99, full GPU | 43.4 | 43.0 |
| ornith-35b-duo (after) | pure CPU, CUDA_VISIBLE_DEVICES= |
9.2 | 8.0 |
Net: subagent 5.8 → 43 t/s (7.4×) under concurrent load; main pays −25% (10.6 → 8.0). Subagent is now contention-immune (GPU decode, 3 CPU threads).
Context sizes (2026-07-10, verified loaded + benched): ornith 128K (q8 K + turbo2 V), qwen 24K. Concurrent: ornith 8.7-10.1 / qwen 43-44.5 t/s.
- Ornith is GDN-hybrid: only 10/40 layers carry KV. Mixed KV types: K q8_0 + V turbo2 = ~880 MB @ 128K in RAM. Full-turbo2 gated clean on this arch (PPL ladder @8K: f16 6.813 / q8 6.838 / turbo2 7.009, Δ0.196 < 0.5) but costs 3× CPU decode on the K side (9.2 → 2.2-3.3 t/s) — K stays q8. V-side turbo2 is speed-free on CPU (probe-verified on qwen, confirmed here). Trained ctx 262K; 128K KV no longer steals page cache.
- Qwen3-4B is full-GQA: 40 KB/token even at q4_0 → KV 648 MB @ 16K, 1296 MB @ 32K = OOM (weights 2.3 GB + compute leave no room). 24K fits at 3564 MiB / 4096. turbo KV is broken on this model — see §Turbo-KV below.
Turbo-KV on Qwen3-4B: root cause + fix status (2026-07-10)
Full investigation: 3 agents + empirical matrix. Fork source:
github.com/TheTom/llama-cpp-turboquant, local clone ~/Sources/llama-cpp-turboquant
(branch fix/innerq-clamp), patched images *-innerq built.
Root cause (confirmed): turbo2/3/4 = PolarQuant per-128 head vector (one fp16 norm + fixed WHT + fixed Lloyd-Max centroids, no per-channel range). Qwen3's QK-norm gamma has extreme per-channel outliers (blk.0 ch51 γ=44 vs mean 1.7 = 73% of K energy) → K direction info lands below the 2-bit centroid gap. Outlier channels are the lowest-freq RoPE dims → error is common-mode at 4K (looks fine), phase-spreads by 8-32K → blow-up. SmolLM3 (no QK-norm) and Gemma4 (constant γ, mostly SWA) pass the same gate.
Empirical matrix (CPU, ctx 8K, chunk-1 PPL, ref q8/q8 = 18.03):
| ctk | ctv | PPL | verdict |
|---|---|---|---|
| turbo2 | turbo2 | 442.2 | broken |
| turbo2 | q8_0 | 404.2 | broken → K is the culprit |
| q8_0 | turbo2 | 18.12 (Δ+0.09) | clean → V tolerates 2-bit |
| q4_0 | turbo2 | 19.03 (Δ+1.00) | fails 0.5 gate — K needs ≥8-bit |
No zero-code ctx win for qwen: the only clean mix (q8K+turbo2V, 49.5 KB/tok) is bigger than q4/q4 (40.5). Also: turbo2-K costs 2.6× CPU decode; turbo2-V free.
InnerQ (fork's per-channel K equalizer): broken as shipped. Widened its [0.5,2.0] clamp to 64× (commit on fix/innerq-clamp) — but GPU PPL gate shows the compensation path is quantitatively wrong: strength 0.001 → 431 (= no-op control), 0.5 → 2482 (author defaults, WORSE than off), 1.0 → 15M. Error grows superlinearly with scale strength; not the calibration-window mismatch (chunk 2 equally broken). The 2× clamp was hiding a real compensation bug — feature is off-by-default for a reason. Real fix = load-time static per-LAYER scales derived from attn_k_norm γ (InnerQ state is global-128ch, γ outliers are per-layer) + verified Q/V compensation — parked, see fix ladder in the agents' reports (session scratchpad) if resumed.
Gotchas hit:
--n-gpu-layers 0is NOT CPU-only on a CUDA build: it still cudaMallocs a ~1 GB prompt-processing compute buffer → OOM + segfault when qwen holds the GPU. Must hide the device entirely (env: CUDA_VISIBLE_DEVICES=).- Ornith pure-CPU costs vs its solo config (29.1 → 9.2): dense backbone every token moves to DDR4, plus qwen's 2.4 GB GGUF competes for page cache (13.1 + 2.4 GB vs ~13 GB usable cache). Still fine as a thinking main model.
Tools added
scripts/moe_bench.py— temp-0 gates +.timingsthroughput via swap-stackscripts/expert_heatmap.py— mincore() page-residency per expert slice of a GGUF (expert index = slowest dim → contiguous slices). Confirms hot layers, verifies offload direction. Only discriminating in thrash regime.swap-stack/— llama-swap v236 tri-binary image (turboquant + upstream + ik), single endpoint :8080, per-model binary via macros +env: LD_LIBRARY_PATH.
Ideas parked
--parallel 2+aggregate throughput (x570 avenue #3: expert-read amortization)- ornith-9b/qwen35-9b could be re-run as MoE-style ngl=99 sweeps with upstream
--fit on(auto VRAM placement) - vmtouch/mlock hot-expert pinning + heatmap: only pays off in thrash regime — moot while everything daily-driver fits cache; revisit for IQ4-class quality runs
- RAM upgrade to 64 GB (2 SODIMM) → IQ4-class 35Bs cache-resident → this whole table shifts up a tier