swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model. Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to 176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s. CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB pp compute buffer on CUDA builds (OOM+segfault). Duo section in MOE-FINDINGS.md; also snapshots prior swap-stack migration state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,93 @@
|
||||
# MoE Findings — xps9700, 2026-07-10
|
||||
|
||||
Session: big-model-runner scope shifted to this laptop (15 GiB DDR4-2933, GTX 1650 Ti
|
||||
3.7 GB VRAM, i7-10750H 6c/12t, SN730 PCIe3 NVMe). All numbers from server `.timings`
|
||||
at temp 0, structure-gated (JSON validity + bracket balance), disk quiet.
|
||||
Method + priors from the x570 flash-162b record (gitea: mozempk/big-model-runner).
|
||||
|
||||
## Scoreboard
|
||||
|
||||
| model / config | tg t/s | pp t/s | gates |
|
||||
|---|---|---|---|
|
||||
| **ornith-35b Q2_K_L** (13.1 GB, 3 exp layers VRAM) | **29.1** | **66** | pass |
|
||||
| **qwen36-35b UD-Q2_K_XL** (12.3 GB, 3 exp layers VRAM) | **23.0** | 36 | pass |
|
||||
| qwen36-35b Q2 on upstream master | 22.8 | 28–32 | pass |
|
||||
| qwen36-35b Q2 on ik_llama (+`-ser 6,1`) | 18.0 | 42 | pass |
|
||||
| gpt-oss-20b MXFP4 (`--n-cpu-moe 21`) | 16.7 | 29.5 | pass |
|
||||
| gpt-oss-20b (`--cpu-moe`, all experts CPU) | 14.4 | 22.3 | pass |
|
||||
| qwen36-35b UD-IQ4_XS 17.7 GB (thrash) | 2.8 | 3.7 | pass |
|
||||
| ornith-9b / qwen3.5-9b dense Q8 (old ceiling) | 4.4 | ~45 | — |
|
||||
|
||||
## Laws of this machine
|
||||
|
||||
1. **The page-cache cliff**: MoE offload is fast iff the GGUF fits page cache
|
||||
(~12–13 GB with services running). 12.3 GB → 23 t/s; 17.7 GB → 2.8 t/s
|
||||
(measured 1.8 GB/s sustained NVMe page-in, ~650 MB faulted/token — eviction
|
||||
churn, warm == cold).
|
||||
2. **ds4/x570 asymmetric quant recipe transfers**: routed experts tolerate 2-bit;
|
||||
dense/attention/embeddings must stay high-bit. unsloth UD-Q2_K_XL and
|
||||
bartowski Q2_K_L are pre-made versions of this mix. Structure gates pass;
|
||||
35B-A3B @ Q2-experts beats 20B @ 4-bit on both speed and (by benchmarks) quality.
|
||||
3. **VRAM expert placement**: `--n-cpu-moe N` (first N layers' experts → CPU,
|
||||
rest → GPU; direction verified empirically). ~455 MB/layer (gpt-oss MXFP4),
|
||||
~230 MB/layer (qwen36 Q2). 3 layers in spare VRAM = +16% tg, +32% pp on gpt-oss.
|
||||
**Which** layers doesn't matter when file is cache-resident (middle-hot vs last-3:
|
||||
16.79 vs 16.73) — only how many. 6 layers OOMs (compute buffers need ~500 MB).
|
||||
4. **Builds are a wash for K-quant decode**: turboquant (May) == upstream (Jul)
|
||||
== within noise. ik_llama: −23% decode / +13% prefill here; `-ser 6,1` marginal.
|
||||
Keep turboquant as default binary (turbo2 KV for the dense 9Bs).
|
||||
5. **MTP / speculative decode: skip** (x570 measured −20% net at 80% acceptance —
|
||||
expert-union tax; worse when disk-bound; our GGUFs lack MTP tensors anyway).
|
||||
6. **Codacus commits: skip** (−5% on x570; its cudaHostRegister mmap-pinning would
|
||||
try to pin >RAM here).
|
||||
7. **Benching discipline**: never bench while downloads/builds run — page-cache
|
||||
flushing fakes an 80% regression (measured 16.7 → 3.2 on identical config).
|
||||
`pkill -f` patterns self-match the invoking shell — SIGSTOP'd our own bench once.
|
||||
|
||||
## Duo config (resident main + subagent, 2026-07-10)
|
||||
|
||||
llama-swap group `duo` (`swap: false, exclusive: true`): `ornith-35b-duo` +
|
||||
`qwen3-4b-duo` stay loaded together for pi (main coding model + fast subagent).
|
||||
Requesting any NON-duo model unloads the whole group — pi must use the `-duo` ids.
|
||||
|
||||
**VRAM goes to the small model, not the big one.** User observation confirmed:
|
||||
ornith's routed experts never load into VRAM, and its dense-on-GPU split
|
||||
(2354 MiB) starved qwen down to 176 MiB via `--fit on`. Flipped: qwen3-4b
|
||||
ngl 99 owns the GPU (3240 MiB incl. KV+compute), ornith runs pure CPU.
|
||||
|
||||
| duo member | config | solo t/s | concurrent t/s |
|
||||
|---|---|---|---|
|
||||
| qwen3-4b-duo (before) | `--fit on`, 176 MiB VRAM | 10.3 | 5.8 |
|
||||
| ornith-35b-duo (before) | dense GPU, experts CPU | 16.2 | 10.6 |
|
||||
| **qwen3-4b-duo (after)** | ngl 99, full GPU | **43.4** | **43.0** |
|
||||
| **ornith-35b-duo (after)** | pure CPU, `CUDA_VISIBLE_DEVICES=` | 9.2 | 8.0 |
|
||||
|
||||
Net: subagent 5.8 → 43 t/s (7.4×) under concurrent load; main pays −25%
|
||||
(10.6 → 8.0). Subagent is now contention-immune (GPU decode, 3 CPU threads).
|
||||
|
||||
Gotchas hit:
|
||||
- `--n-gpu-layers 0` is NOT CPU-only on a CUDA build: it still cudaMallocs a
|
||||
~1 GB prompt-processing compute buffer → OOM + segfault when qwen holds the
|
||||
GPU. Must hide the device entirely (`env: CUDA_VISIBLE_DEVICES=`).
|
||||
- Ornith pure-CPU costs vs its solo config (29.1 → 9.2): dense backbone every
|
||||
token moves to DDR4, plus qwen's 2.4 GB GGUF competes for page cache
|
||||
(13.1 + 2.4 GB vs ~13 GB usable cache). Still fine as a thinking main model.
|
||||
|
||||
## Tools added
|
||||
|
||||
- `scripts/moe_bench.py` — temp-0 gates + `.timings` throughput via swap-stack
|
||||
- `scripts/expert_heatmap.py` — mincore() page-residency per expert slice of a GGUF
|
||||
(expert index = slowest dim → contiguous slices). Confirms hot layers, verifies
|
||||
offload direction. Only discriminating in thrash regime.
|
||||
- `swap-stack/` — llama-swap v236 tri-binary image (turboquant + upstream + ik),
|
||||
single endpoint :8080, per-model binary via macros + `env: LD_LIBRARY_PATH`.
|
||||
|
||||
## Ideas parked
|
||||
|
||||
- `--parallel 2+` aggregate throughput (x570 avenue #3: expert-read amortization)
|
||||
- ornith-9b/qwen35-9b could be re-run as MoE-style ngl=99 sweeps with upstream
|
||||
`--fit on` (auto VRAM placement)
|
||||
- vmtouch/mlock hot-expert pinning + heatmap: only pays off in thrash regime —
|
||||
moot while everything daily-driver fits cache; revisit for IQ4-class quality runs
|
||||
- RAM upgrade to 64 GB (2 SODIMM) → IQ4-class 35Bs cache-resident → this whole
|
||||
table shifts up a tier
|
||||
Reference in New Issue
Block a user