Files
llama-cpp/envs/.env.qwen3-4b
T
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00

43 lines
1.4 KiB
Bash

# ==============================================================================
# Qwen3-4B-Instruct Q4_K_M — Alibaba (May 2025)
# Architecture: 4B params, 32 layers, full GQA (32 KV heads)
# Model size: ~2.4 GB Q4_K_M | Full GPU fit (ngl=99)
# Features: thinking mode, tool calling, 119 languages, Apache 2.0
#
# Benchmark (TurboQuant SM75, 2026-05-05):
# pp=191 t/s tg=44.3 t/s @ ctx=16K, fa=1
# KV: 39.6 KB/token (full GQA = double SmolLM3's KV)
#
# ⚠️ CRITICAL: turbo2/3/4 KV is BROKEN for Qwen3-4B (PPL catastrophic @ ctx≥8K)
# Always use q4_0 KV! See FINDINGS.md
#
# Optimization: q4_0 KV only, ctx=16K max (full-attn VRAM wall)
# parallel=1 (limited headroom with large KV)
# ==============================================================================
MODEL_FILE=Qwen3-4B-Q4_K_M.gguf
# All layers fit — ~2.4 GB leaves ~1.3 GB free for KV + compute
N_GPU_LAYERS=99
# 16K practical max (full GQA eats VRAM fast, 24K OOM)
CTX_SIZE=16384
# i7-10750H: t=6 physical cores optimal
THREADS=6
THREADS_BATCH=6
# Standard batches, ubatch=batch/2 for pure-GPU (ngl=99, all 32 layers on GPU)
BATCH_SIZE=512
UBATCH_SIZE=256
# q4_0 KV ONLY — turbo2/3/4 catastrophically broken for Qwen3-4B!
CACHE_TYPE_K=q4_0
CACHE_TYPE_V=q4_0
# 1 parallel slot — limited VRAM with large KV @ 16K ctx
PARALLEL=1
# fa=1 gives +6% boost on full-attention Qwen3
EXTRA_ARGS="--flash-attn on --cont-batching"