swap-stack: duo bigger contexts — ornith 64K (CPU q8 KV), qwen 24K (GPU q4 KV ceiling)

Ornith GDN-hybrid KV is cheap (10/40 layers, 680MB @ 64K in RAM).
Qwen3-4B full-GQA KV is 40KB/tok at q4_0: 32K OOMs on 4GB VRAM, 24K fits
(3564 MiB). turbo KV rejected — broken on Qwen3-4B (PPL 438, FINDINGS §2).
Concurrent throughput unchanged: ornith 9.2 / qwen 43.2 t/s.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-10 08:35:07 +02:00
co-authored by Claude Fable 5
parent 5e68d30d31
commit da655e074d
2 changed files with 14 additions and 4 deletions
+4 -4
View File
@@ -60,21 +60,21 @@ models:
"ornith-35b-duo":
name: "Ornith 35B (duo main)"
description: "Duo variant: fully CPU (GPU reserved for qwen3-4b-duo). CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer"
description: "Duo main: fully CPU, 64K ctx (GDN hybrid — KV only 680MB q8 in RAM). CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer"
env:
- "CUDA_VISIBLE_DEVICES="
cmd: |
${server-base} ${q8-kv}
--model /models/deepreinforce-ai_Ornith-1.0-35B-Q2_K_L.gguf
--n-gpu-layers 0 --jinja
--ctx-size 32768
--ctx-size 65536
--batch-size 1024 --ubatch-size 512
--cont-batching --parallel 1
ttl: 0
"qwen3-4b-duo":
name: "Qwen3 4B (duo subagent)"
description: "Duo subagent: full GPU (ngl 99), 3 threads, gate-clean JSON"
description: "Duo subagent: full GPU (ngl 99), 24K ctx (VRAM ceiling: q4 KV = 40KB/tok, 32K OOMs; turbo KV broken on this model), 3 threads"
cmd: |
/app/llama-server
--host 127.0.0.1 --port ${PORT}
@@ -82,7 +82,7 @@ models:
--flash-attn on ${q4-kv}
--model /models/Qwen3-4B-Q4_K_M.gguf
--n-gpu-layers 99
--ctx-size 16384
--ctx-size 24576
--batch-size 512 --ubatch-size 256
--cont-batching --parallel 1
ttl: 0