4 Commits
Author SHA1 Message Date
mozempkandClaude Fable 5 2a2305490a swap-stack: ornith duo 128K via mixed KV (q8 K + turbo2 V); turbo-KV root cause
Probe matrix proved K is the broken side of turbo KV on Qwen3-4B
(QK-norm gamma outliers vs PolarQuant's no-per-channel-range format)
while V tolerates 2-bit for free. Applied to ornith duo main: K q8_0 +
V turbo2 = 128K ctx in ~880MB RAM at full speed (8.7-10.1 t/s concurrent,
vs 3x decode loss with turbo2 K). Qwen duo stays q4/q4 24K — only clean
mix is bigger than q4/q4. InnerQ equalizer confirmed broken as shipped
(compensation error grows with scale strength); documented in
MOE-FINDINGS with the full matrix and fix ladder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 15:44:19 +02:00
mozempkandClaude Fable 5 da655e074d swap-stack: duo bigger contexts — ornith 64K (CPU q8 KV), qwen 24K (GPU q4 KV ceiling)
Ornith GDN-hybrid KV is cheap (10/40 layers, 680MB @ 64K in RAM).
Qwen3-4B full-GQA KV is 40KB/tok at q4_0: 32K OOMs on 4GB VRAM, 24K fits
(3564 MiB). turbo KV rejected — broken on Qwen3-4B (PPL 438, FINDINGS §2).
Concurrent throughput unchanged: ornith 9.2 / qwen 43.2 t/s.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:35:07 +02:00
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00
mozempk 4ad296608b Initial commit: tuned multi-model llama.cpp stack
- 5 models: SmolLM3-3B, Gemma4-E2B/E4B, Qwen3-4B, Qwen3.5-9B
- TurboQuant image (FORCE_MMQ): +6-11% free speed on Turing GPUs
- Bigctx profiles (-nkvo KV in RAM): 2-16x context gain
- turbo2 KV: 2x smaller, benchmarked against PPL quality gate
- Per-model env files with justified parameters
- kv_quant_test.sh + cpu_ctx_test.sh benchmark scripts
- docs/FINDINGS.md: surprises, pitfalls, recommendations
- docs/ARCHITECTURE.md: compose + test script design
2026-05-06 15:56:40 +02:00