swap-stack: duo bigger contexts — ornith 64K (CPU q8 KV), qwen 24K (GPU q4 KV ceiling)
Ornith GDN-hybrid KV is cheap (10/40 layers, 680MB @ 64K in RAM). Qwen3-4B full-GQA KV is 40KB/tok at q4_0: 32K OOMs on 4GB VRAM, 24K fits (3564 MiB). turbo KV rejected — broken on Qwen3-4B (PPL 438, FINDINGS §2). Concurrent throughput unchanged: ornith 9.2 / qwen 43.2 t/s. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -60,21 +60,21 @@ models:
|
||||
|
||||
"ornith-35b-duo":
|
||||
name: "Ornith 35B (duo main)"
|
||||
description: "Duo variant: fully CPU (GPU reserved for qwen3-4b-duo). CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer"
|
||||
description: "Duo main: fully CPU, 64K ctx (GDN hybrid — KV only 680MB q8 in RAM). CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer"
|
||||
env:
|
||||
- "CUDA_VISIBLE_DEVICES="
|
||||
cmd: |
|
||||
${server-base} ${q8-kv}
|
||||
--model /models/deepreinforce-ai_Ornith-1.0-35B-Q2_K_L.gguf
|
||||
--n-gpu-layers 0 --jinja
|
||||
--ctx-size 32768
|
||||
--ctx-size 65536
|
||||
--batch-size 1024 --ubatch-size 512
|
||||
--cont-batching --parallel 1
|
||||
ttl: 0
|
||||
|
||||
"qwen3-4b-duo":
|
||||
name: "Qwen3 4B (duo subagent)"
|
||||
description: "Duo subagent: full GPU (ngl 99), 3 threads, gate-clean JSON"
|
||||
description: "Duo subagent: full GPU (ngl 99), 24K ctx (VRAM ceiling: q4 KV = 40KB/tok, 32K OOMs; turbo KV broken on this model), 3 threads"
|
||||
cmd: |
|
||||
/app/llama-server
|
||||
--host 127.0.0.1 --port ${PORT}
|
||||
@@ -82,7 +82,7 @@ models:
|
||||
--flash-attn on ${q4-kv}
|
||||
--model /models/Qwen3-4B-Q4_K_M.gguf
|
||||
--n-gpu-layers 99
|
||||
--ctx-size 16384
|
||||
--ctx-size 24576
|
||||
--batch-size 512 --ubatch-size 256
|
||||
--cont-batching --parallel 1
|
||||
ttl: 0
|
||||
|
||||
Reference in New Issue
Block a user