Files
llama-cpp/envs/.env.gemma4-e4b
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00

40 lines
1.4 KiB
Bash
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ==============================================================================
# Gemma 4 E4B-it Q4_K_M — Google DeepMind (April 2025)
# Architecture: 4.5B effective (8B total with PLE), 42 layers, hybrid attention
# Model size: ~4.7 GB Q4_K_M | All layers fit on GPU! (ngl=42)
# Modalities: text + image + audio + video
#
# Benchmark (TurboQuant SM75, 2026-05-05):
# ngl=42: pp=133 t/s tg=32.0 t/s @ ctx=24K, fa=1
# Surprise: ALL 42 layers fit despite file > VRAM (paged weight loading)
#
# Optimization: ngl=42 (all layers), q4_0 KV, parallel=1 (VRAM tight at 24K)
# Batch 1024/256 for better throughput on CPU-split layers
# ==============================================================================
MODEL_FILE=google_gemma-4-E4B-it-Q4_K_M.gguf
# ALL 42 layers fit on GPU when no other containers hold VRAM
# ngl sweep confirmed: ngl=42 → 133 pp / 32.0 tg t/s (vs ngl=28 → 59/16.5)
N_GPU_LAYERS=42
# 24K max — hybrid sliding-window keeps most KV tiny, 32K OOM
CTX_SIZE=24576
# t=6 still optimal even for pure-GPU (hyperthreading hurts)
THREADS=6
THREADS_BATCH=6
# Larger batches for multimodal. ubatch=512 for pure-GPU (ngl=42, all layers)
BATCH_SIZE=1024
UBATCH_SIZE=512
# turbo2 KV for 6.4× compression (hybrid attention benefits from tiny KV)
CACHE_TYPE_K=turbo2
CACHE_TYPE_V=turbo2
PARALLEL=1
# fa=1 confirmed +3% boost on hybrid Gemma4 attention
EXTRA_ARGS="--flash-attn on --cont-batching"