Duo group (resident main+subagent for pi): flip VRAM to the small model. Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to 176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s. CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB pp compute buffer on CUDA builds (OOM+segfault). Duo section in MOE-FINDINGS.md; also snapshots prior swap-stack migration state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
40 lines
1.4 KiB
Bash
40 lines
1.4 KiB
Bash
# ==============================================================================
|
||
# Gemma 4 E4B-it Q4_K_M — Google DeepMind (April 2025)
|
||
# Architecture: 4.5B effective (8B total with PLE), 42 layers, hybrid attention
|
||
# Model size: ~4.7 GB Q4_K_M | All layers fit on GPU! (ngl=42)
|
||
# Modalities: text + image + audio + video
|
||
#
|
||
# Benchmark (TurboQuant SM75, 2026-05-05):
|
||
# ngl=42: pp=133 t/s tg=32.0 t/s @ ctx=24K, fa=1
|
||
# Surprise: ALL 42 layers fit despite file > VRAM (paged weight loading)
|
||
#
|
||
# Optimization: ngl=42 (all layers), q4_0 KV, parallel=1 (VRAM tight at 24K)
|
||
# Batch 1024/256 for better throughput on CPU-split layers
|
||
# ==============================================================================
|
||
|
||
MODEL_FILE=google_gemma-4-E4B-it-Q4_K_M.gguf
|
||
|
||
# ALL 42 layers fit on GPU when no other containers hold VRAM
|
||
# ngl sweep confirmed: ngl=42 → 133 pp / 32.0 tg t/s (vs ngl=28 → 59/16.5)
|
||
N_GPU_LAYERS=42
|
||
|
||
# 24K max — hybrid sliding-window keeps most KV tiny, 32K OOM
|
||
CTX_SIZE=24576
|
||
|
||
# t=6 still optimal even for pure-GPU (hyperthreading hurts)
|
||
THREADS=6
|
||
THREADS_BATCH=6
|
||
|
||
# Larger batches for multimodal. ubatch=512 for pure-GPU (ngl=42, all layers)
|
||
BATCH_SIZE=1024
|
||
UBATCH_SIZE=512
|
||
|
||
# turbo2 KV for 6.4× compression (hybrid attention benefits from tiny KV)
|
||
CACHE_TYPE_K=turbo2
|
||
CACHE_TYPE_V=turbo2
|
||
|
||
PARALLEL=1
|
||
|
||
# fa=1 confirmed +3% boost on hybrid Gemma4 attention
|
||
EXTRA_ARGS="--flash-attn on --cont-batching"
|