swap-stack: ornith duo 128K via mixed KV (q8 K + turbo2 V); turbo-KV root cause
Probe matrix proved K is the broken side of turbo KV on Qwen3-4B (QK-norm gamma outliers vs PolarQuant's no-per-channel-range format) while V tolerates 2-bit for free. Applied to ornith duo main: K q8_0 + V turbo2 = 128K ctx in ~880MB RAM at full speed (8.7-10.1 t/s concurrent, vs 3x decode loss with turbo2 K). Qwen duo stays q4/q4 24K — only clean mix is bigger than q4/q4. InnerQ equalizer confirmed broken as shipped (compensation error grows with scale strength); documented in MOE-FINDINGS with the full matrix and fix ladder. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -60,14 +60,15 @@ models:
|
||||
|
||||
"ornith-35b-duo":
|
||||
name: "Ornith 35B (duo main)"
|
||||
description: "Duo main: fully CPU, 64K ctx (GDN hybrid — KV only 680MB q8 in RAM). CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer"
|
||||
description: "Duo main: fully CPU, 128K ctx, q8 K + turbo2 V (K-side turbo costs 3x CPU decode, V-side is free — probe-verified; full-turbo2 quality gate Δ0.196 upper-bounds this mix). ~880MB KV. CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer"
|
||||
env:
|
||||
- "CUDA_VISIBLE_DEVICES="
|
||||
cmd: |
|
||||
${server-base} ${q8-kv}
|
||||
${server-base}
|
||||
--cache-type-k q8_0 --cache-type-v turbo2
|
||||
--model /models/deepreinforce-ai_Ornith-1.0-35B-Q2_K_L.gguf
|
||||
--n-gpu-layers 0 --jinja
|
||||
--ctx-size 65536
|
||||
--ctx-size 131072
|
||||
--batch-size 1024 --ubatch-size 512
|
||||
--cont-batching --parallel 1
|
||||
ttl: 0
|
||||
|
||||
Reference in New Issue
Block a user