swap-stack: duo bigger contexts — ornith 64K (CPU q8 KV), qwen 24K (GPU q4 KV ceiling)
Ornith GDN-hybrid KV is cheap (10/40 layers, 680MB @ 64K in RAM). Qwen3-4B full-GQA KV is 40KB/tok at q4_0: 32K OOMs on 4GB VRAM, 24K fits (3564 MiB). turbo KV rejected — broken on Qwen3-4B (PPL 438, FINDINGS §2). Concurrent throughput unchanged: ornith 9.2 / qwen 43.2 t/s. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -65,6 +65,16 @@ ngl 99 owns the GPU (3240 MiB incl. KV+compute), ornith runs pure CPU.
|
||||
Net: subagent 5.8 → 43 t/s (7.4×) under concurrent load; main pays −25%
|
||||
(10.6 → 8.0). Subagent is now contention-immune (GPU decode, 3 CPU threads).
|
||||
|
||||
**Context sizes (2026-07-10, verified loaded + benched):** ornith 64K, qwen 24K.
|
||||
Throughput unchanged (ornith 9.2 / qwen 43.2 concurrent).
|
||||
- Ornith is GDN-hybrid: only 10/40 layers carry KV → 64K q8 = just 680 MB, in
|
||||
RAM (CPU model). 128K would work RAM-wise (+680 MB) but steals page cache
|
||||
from the 13.1 GB GGUF — thrash risk, not attempted. Trained ctx 262K.
|
||||
- Qwen3-4B is full-GQA: 40 KB/token even at q4_0 → KV 648 MB @ 16K, 1296 MB
|
||||
@ 32K = **OOM** (weights 2.3 GB + compute leave no room). 24K fits at
|
||||
3564 MiB / 4096. turbo2/3/4 KV would halve it but is BROKEN on this model
|
||||
(PPL 438 @ 32K — FINDINGS §2). 24K is the hard ceiling on 4 GB VRAM.
|
||||
|
||||
Gotchas hit:
|
||||
- `--n-gpu-layers 0` is NOT CPU-only on a CUDA build: it still cudaMallocs a
|
||||
~1 GB prompt-processing compute buffer → OOM + segfault when qwen holds the
|
||||
|
||||
Reference in New Issue
Block a user