Commit Graph
2 Commits
Author SHA1 Message Date
mozempkandClaude Opus 4.8 c8ab0b93ce feat(bonsai): add Bonsai-27B-Q1_0 fully-on-GPU (1650 Ti 4GB)
Add a 4th binary to the swap-stack: PrismML llama-server-ternary (Q1_0_g128
ternary kernels, built sm75/CUDA-12.8 static) for Bonsai-27B-Q1_0.

Config: 3.8GB weights ALL on the 4GB GPU (ngl 99) + KV in CPU RAM
(--no-kv-offload, since VRAM cannot hold both) + q4_0 KV + --flash-attn
(GPU attention; iq4_nl forces CPU attention) + tiny batch (-b 64) to fit
the ~100MB VRAM headroom. Fits at 3611/3718 MiB. ~7 t/s decode, ~27 pp
prefill (1650 Ti has no tensor cores + KV over PCIe = the ceiling).

Measured alt configs (all worse): ngl48+KV-VRAM = 8.5 t/s decode but only
10 pp prefill (CPU layers tank prefill); ngl99+KV-VRAM OOMs even at ctx 2048.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 05:17:13 +02:00
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00