Duo group (resident main+subagent for pi): flip VRAM to the small model. Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to 176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s. CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB pp compute buffer on CUDA builds (OOM+segfault). Duo section in MOE-FINDINGS.md; also snapshots prior swap-stack migration state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
43 lines
1.4 KiB
Bash
43 lines
1.4 KiB
Bash
# ==============================================================================
|
|
# Qwen3-4B-Instruct Q4_K_M — Alibaba (May 2025)
|
|
# Architecture: 4B params, 32 layers, full GQA (32 KV heads)
|
|
# Model size: ~2.4 GB Q4_K_M | Full GPU fit (ngl=99)
|
|
# Features: thinking mode, tool calling, 119 languages, Apache 2.0
|
|
#
|
|
# Benchmark (TurboQuant SM75, 2026-05-05):
|
|
# pp=191 t/s tg=44.3 t/s @ ctx=16K, fa=1
|
|
# KV: 39.6 KB/token (full GQA = double SmolLM3's KV)
|
|
#
|
|
# ⚠️ CRITICAL: turbo2/3/4 KV is BROKEN for Qwen3-4B (PPL catastrophic @ ctx≥8K)
|
|
# Always use q4_0 KV! See FINDINGS.md
|
|
#
|
|
# Optimization: q4_0 KV only, ctx=16K max (full-attn VRAM wall)
|
|
# parallel=1 (limited headroom with large KV)
|
|
# ==============================================================================
|
|
|
|
MODEL_FILE=Qwen3-4B-Q4_K_M.gguf
|
|
|
|
# All layers fit — ~2.4 GB leaves ~1.3 GB free for KV + compute
|
|
N_GPU_LAYERS=99
|
|
|
|
# 16K practical max (full GQA eats VRAM fast, 24K OOM)
|
|
CTX_SIZE=16384
|
|
|
|
# i7-10750H: t=6 physical cores optimal
|
|
THREADS=6
|
|
THREADS_BATCH=6
|
|
|
|
# Standard batches, ubatch=batch/2 for pure-GPU (ngl=99, all 32 layers on GPU)
|
|
BATCH_SIZE=512
|
|
UBATCH_SIZE=256
|
|
|
|
# q4_0 KV ONLY — turbo2/3/4 catastrophically broken for Qwen3-4B!
|
|
CACHE_TYPE_K=q4_0
|
|
CACHE_TYPE_V=q4_0
|
|
|
|
# 1 parallel slot — limited VRAM with large KV @ 16K ctx
|
|
PARALLEL=1
|
|
|
|
# fa=1 gives +6% boost on full-attention Qwen3
|
|
EXTRA_ARGS="--flash-attn on --cont-batching"
|