Files
llama-cpp/docs/MOE-FINDINGS.md
T
mozempkandClaude Fable 5 2a2305490a swap-stack: ornith duo 128K via mixed KV (q8 K + turbo2 V); turbo-KV root cause
Probe matrix proved K is the broken side of turbo KV on Qwen3-4B
(QK-norm gamma outliers vs PolarQuant's no-per-channel-range format)
while V tolerates 2-bit for free. Applied to ornith duo main: K q8_0 +
V turbo2 = 128K ctx in ~880MB RAM at full speed (8.7-10.1 t/s concurrent,
vs 3x decode loss with turbo2 K). Qwen duo stays q4/q4 24K — only clean
mix is bigger than q4/q4. InnerQ equalizer confirmed broken as shipped
(compensation error grows with scale strength); documented in
MOE-FINDINGS with the full matrix and fix ladder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 15:44:19 +02:00

143 lines
8.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# MoE Findings — xps9700, 2026-07-10
Session: big-model-runner scope shifted to this laptop (15 GiB DDR4-2933, GTX 1650 Ti
3.7 GB VRAM, i7-10750H 6c/12t, SN730 PCIe3 NVMe). All numbers from server `.timings`
at temp 0, structure-gated (JSON validity + bracket balance), disk quiet.
Method + priors from the x570 flash-162b record (gitea: mozempk/big-model-runner).
## Scoreboard
| model / config | tg t/s | pp t/s | gates |
|---|---|---|---|
| **ornith-35b Q2_K_L** (13.1 GB, 3 exp layers VRAM) | **29.1** | **66** | pass |
| **qwen36-35b UD-Q2_K_XL** (12.3 GB, 3 exp layers VRAM) | **23.0** | 36 | pass |
| qwen36-35b Q2 on upstream master | 22.8 | 28–32 | pass |
| qwen36-35b Q2 on ik_llama (+`-ser 6,1`) | 18.0 | 42 | pass |
| gpt-oss-20b MXFP4 (`--n-cpu-moe 21`) | 16.7 | 29.5 | pass |
| gpt-oss-20b (`--cpu-moe`, all experts CPU) | 14.4 | 22.3 | pass |
| qwen36-35b UD-IQ4_XS 17.7 GB (thrash) | 2.8 | 3.7 | pass |
| ornith-9b / qwen3.5-9b dense Q8 (old ceiling) | 4.4 | ~45 | — |
## Laws of this machine
1. **The page-cache cliff**: MoE offload is fast iff the GGUF fits page cache
(~12–13 GB with services running). 12.3 GB → 23 t/s; 17.7 GB → 2.8 t/s
(measured 1.8 GB/s sustained NVMe page-in, ~650 MB faulted/token — eviction
churn, warm == cold).
2. **ds4/x570 asymmetric quant recipe transfers**: routed experts tolerate 2-bit;
dense/attention/embeddings must stay high-bit. unsloth UD-Q2_K_XL and
bartowski Q2_K_L are pre-made versions of this mix. Structure gates pass;
35B-A3B @ Q2-experts beats 20B @ 4-bit on both speed and (by benchmarks) quality.
3. **VRAM expert placement**: `--n-cpu-moe N` (first N layers' experts → CPU,
rest → GPU; direction verified empirically). ~455 MB/layer (gpt-oss MXFP4),
~230 MB/layer (qwen36 Q2). 3 layers in spare VRAM = +16% tg, +32% pp on gpt-oss.
**Which** layers doesn't matter when file is cache-resident (middle-hot vs last-3:
16.79 vs 16.73) — only how many. 6 layers OOMs (compute buffers need ~500 MB).
4. **Builds are a wash for K-quant decode**: turboquant (May) == upstream (Jul)
== within noise. ik_llama: −23% decode / +13% prefill here; `-ser 6,1` marginal.
Keep turboquant as default binary (turbo2 KV for the dense 9Bs).
5. **MTP / speculative decode: skip** (x570 measured −20% net at 80% acceptance —
expert-union tax; worse when disk-bound; our GGUFs lack MTP tensors anyway).
6. **Codacus commits: skip** (−5% on x570; its cudaHostRegister mmap-pinning would
try to pin >RAM here).
7. **Benching discipline**: never bench while downloads/builds run — page-cache
flushing fakes an 80% regression (measured 16.7 → 3.2 on identical config).
`pkill -f` patterns self-match the invoking shell — SIGSTOP'd our own bench once.
## Duo config (resident main + subagent, 2026-07-10)
llama-swap group `duo` (`swap: false, exclusive: true`): `ornith-35b-duo` +
`qwen3-4b-duo` stay loaded together for pi (main coding model + fast subagent).
Requesting any NON-duo model unloads the whole group — pi must use the `-duo` ids.
**VRAM goes to the small model, not the big one.** User observation confirmed:
ornith's routed experts never load into VRAM, and its dense-on-GPU split
(2354 MiB) starved qwen down to 176 MiB via `--fit on`. Flipped: qwen3-4b
ngl 99 owns the GPU (3240 MiB incl. KV+compute), ornith runs pure CPU.
| duo member | config | solo t/s | concurrent t/s |
|---|---|---|---|
| qwen3-4b-duo (before) | `--fit on`, 176 MiB VRAM | 10.3 | 5.8 |
| ornith-35b-duo (before) | dense GPU, experts CPU | 16.2 | 10.6 |
| **qwen3-4b-duo (after)** | ngl 99, full GPU | **43.4** | **43.0** |
| **ornith-35b-duo (after)** | pure CPU, `CUDA_VISIBLE_DEVICES=` | 9.2 | 8.0 |
Net: subagent 5.8 → 43 t/s (7.4×) under concurrent load; main pays −25%
(10.6 → 8.0). Subagent is now contention-immune (GPU decode, 3 CPU threads).
**Context sizes (2026-07-10, verified loaded + benched):** ornith **128K** (q8 K +
turbo2 V), qwen 24K. Concurrent: ornith 8.7-10.1 / qwen 43-44.5 t/s.
- Ornith is GDN-hybrid: only 10/40 layers carry KV. Mixed KV types: K q8_0 +
V turbo2 = ~880 MB @ 128K in RAM. Full-turbo2 gated clean on this arch
(PPL ladder @8K: f16 6.813 / q8 6.838 / turbo2 7.009, Δ0.196 < 0.5) but
costs 3× CPU decode on the K side (9.2 → 2.2-3.3 t/s) — K stays q8.
V-side turbo2 is speed-free on CPU (probe-verified on qwen, confirmed here).
Trained ctx 262K; 128K KV no longer steals page cache.
- Qwen3-4B is full-GQA: 40 KB/token even at q4_0 → KV 648 MB @ 16K, 1296 MB
@ 32K = **OOM** (weights 2.3 GB + compute leave no room). 24K fits at
3564 MiB / 4096. turbo KV is broken on this model — see §Turbo-KV below.
## Turbo-KV on Qwen3-4B: root cause + fix status (2026-07-10)
Full investigation: 3 agents + empirical matrix. Fork source:
github.com/TheTom/llama-cpp-turboquant, local clone ~/Sources/llama-cpp-turboquant
(branch fix/innerq-clamp), patched images `*-innerq` built.
**Root cause (confirmed):** turbo2/3/4 = PolarQuant per-128 head vector (one fp16
norm + fixed WHT + fixed Lloyd-Max centroids, no per-channel range). Qwen3's
QK-norm gamma has extreme per-channel outliers (blk.0 ch51 γ=44 vs mean 1.7 =
73% of K energy) → K direction info lands below the 2-bit centroid gap.
Outlier channels are the lowest-freq RoPE dims → error is common-mode at 4K
(looks fine), phase-spreads by 8-32K → blow-up. SmolLM3 (no QK-norm) and
Gemma4 (constant γ, mostly SWA) pass the same gate.
**Empirical matrix (CPU, ctx 8K, chunk-1 PPL, ref q8/q8 = 18.03):**
| ctk | ctv | PPL | verdict |
|---|---|---|---|
| turbo2 | turbo2 | 442.2 | broken |
| turbo2 | q8_0 | 404.2 | broken → **K is the culprit** |
| q8_0 | turbo2 | 18.12 (Δ+0.09) | **clean → V tolerates 2-bit** |
| q4_0 | turbo2 | 19.03 (Δ+1.00) | fails 0.5 gate — K needs ≥8-bit |
No zero-code ctx win for qwen: the only clean mix (q8K+turbo2V, 49.5 KB/tok)
is bigger than q4/q4 (40.5). Also: turbo2-K costs 2.6× CPU decode; turbo2-V free.
**InnerQ (fork's per-channel K equalizer): broken as shipped.** Widened its
[0.5,2.0] clamp to 64× (commit on fix/innerq-clamp) — but GPU PPL gate shows
the compensation path is quantitatively wrong: strength 0.001 → 431 (= no-op
control), 0.5 → 2482 (author defaults, WORSE than off), 1.0 → 15M. Error grows
superlinearly with scale strength; not the calibration-window mismatch (chunk 2
equally broken). The 2× clamp was hiding a real compensation bug — feature is
off-by-default for a reason. Real fix = load-time static per-LAYER scales
derived from attn_k_norm γ (InnerQ state is global-128ch, γ outliers are
per-layer) + verified Q/V compensation — parked, see fix ladder in the
agents' reports (session scratchpad) if resumed.
Gotchas hit:
- `--n-gpu-layers 0` is NOT CPU-only on a CUDA build: it still cudaMallocs a
~1 GB prompt-processing compute buffer → OOM + segfault when qwen holds the
GPU. Must hide the device entirely (`env: CUDA_VISIBLE_DEVICES=`).
- Ornith pure-CPU costs vs its solo config (29.1 → 9.2): dense backbone every
token moves to DDR4, plus qwen's 2.4 GB GGUF competes for page cache
(13.1 + 2.4 GB vs ~13 GB usable cache). Still fine as a thinking main model.
## Tools added
- `scripts/moe_bench.py` — temp-0 gates + `.timings` throughput via swap-stack
- `scripts/expert_heatmap.py` — mincore() page-residency per expert slice of a GGUF
(expert index = slowest dim → contiguous slices). Confirms hot layers, verifies
offload direction. Only discriminating in thrash regime.
- `swap-stack/` — llama-swap v236 tri-binary image (turboquant + upstream + ik),
single endpoint :8080, per-model binary via macros + `env: LD_LIBRARY_PATH`.
## Ideas parked
- `--parallel 2+` aggregate throughput (x570 avenue #3: expert-read amortization)
- ornith-9b/qwen35-9b could be re-run as MoE-style ngl=99 sweeps with upstream
`--fit on` (auto VRAM placement)
- vmtouch/mlock hot-expert pinning + heatmap: only pays off in thrash regime —
moot while everything daily-driver fits cache; revisit for IQ4-class quality runs
- RAM upgrade to 64 GB (2 SODIMM) → IQ4-class 35Bs cache-resident → this whole
table shifts up a tier