swap-stack: ornith duo 128K via mixed KV (q8 K + turbo2 V); turbo-KV root cause

Probe matrix proved K is the broken side of turbo KV on Qwen3-4B
(QK-norm gamma outliers vs PolarQuant's no-per-channel-range format)
while V tolerates 2-bit for free. Applied to ornith duo main: K q8_0 +
V turbo2 = 128K ctx in ~880MB RAM at full speed (8.7-10.1 t/s concurrent,
vs 3x decode loss with turbo2 K). Qwen duo stays q4/q4 24K — only clean
mix is bigger than q4/q4. InnerQ equalizer confirmed broken as shipped
(compensation error grows with scale strength); documented in
MOE-FINDINGS with the full matrix and fix ladder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-10 15:44:19 +02:00
co-authored by Claude Fable 5
parent da655e074d
commit 2a2305490a
2 changed files with 50 additions and 10 deletions
+4 -3
View File
@@ -60,14 +60,15 @@ models:
"ornith-35b-duo":
name: "Ornith 35B (duo main)"
description: "Duo main: fully CPU, 64K ctx (GDN hybrid — KV only 680MB q8 in RAM). CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer"
description: "Duo main: fully CPU, 128K ctx, q8 K + turbo2 V (K-side turbo costs 3x CPU decode, V-side is free — probe-verified; full-turbo2 quality gate Δ0.196 upper-bounds this mix). ~880MB KV. CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer"
env:
- "CUDA_VISIBLE_DEVICES="
cmd: |
${server-base} ${q8-kv}
${server-base}
--cache-type-k q8_0 --cache-type-v turbo2
--model /models/deepreinforce-ai_Ornith-1.0-35B-Q2_K_L.gguf
--n-gpu-layers 0 --jinja
--ctx-size 65536
--ctx-size 131072
--batch-size 1024 --ubatch-size 512
--cont-batching --parallel 1
ttl: 0