Files
llama-cpp/docs
mozempkandClaude Fable 5 da655e074d swap-stack: duo bigger contexts — ornith 64K (CPU q8 KV), qwen 24K (GPU q4 KV ceiling)
Ornith GDN-hybrid KV is cheap (10/40 layers, 680MB @ 64K in RAM).
Qwen3-4B full-GQA KV is 40KB/tok at q4_0: 32K OOMs on 4GB VRAM, 24K fits
(3564 MiB). turbo KV rejected — broken on Qwen3-4B (PPL 438, FINDINGS §2).
Concurrent throughput unchanged: ornith 9.2 / qwen 43.2 t/s.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:35:07 +02:00
..