Duo group (resident main+subagent for pi): flip VRAM to the small model. Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to 176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s. CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB pp compute buffer on CUDA builds (OOM+segfault). Duo section in MOE-FINDINGS.md; also snapshots prior swap-stack migration state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
40 lines
1.6 KiB
Bash
40 lines
1.6 KiB
Bash
# ==============================================================================
|
|
# gpt-oss-20b MXFP4 — OpenAI (Aug 2025), Apache 2.0
|
|
# Architecture: 20.9B total / 3.6B active MoE (24 layers, 32 experts, top-4)
|
|
# Model size: 12.1 GB native MXFP4 (quant-aware trained — no further quant loss)
|
|
# Strategy: MoE offload — attention/KV on GPU, expert FFNs in RAM (--cpu-moe)
|
|
# mmap (NO mlock): 12.1 GB file pages in on demand, hot experts stay
|
|
# in page cache. Reads ~2 GB/token → RAM-BW bound, est. 8-15 t/s
|
|
#
|
|
# NOTE: quality-sensitive (JSON/code) — KV at q8_0 until turbo2 passes the
|
|
# per-model PPL gate (see FINDINGS.md §2: turbo2 broke Qwen3-4B).
|
|
# ==============================================================================
|
|
|
|
MODEL_FILE=gpt-oss-20b-mxfp4.gguf
|
|
|
|
# All attention layers on GPU; --cpu-moe in EXTRA_ARGS keeps expert weights in RAM.
|
|
# Dense (non-expert) part is well under 3.7 GB VRAM.
|
|
N_GPU_LAYERS=99
|
|
|
|
# 131K native; start at 32K. KV is small (GQA, 24 layers).
|
|
CTX_SIZE=32768
|
|
|
|
# t=6 optimal for i7-10750H (6 physical cores). HT hurts (FINDINGS.md §5).
|
|
THREADS=6
|
|
THREADS_BATCH=6
|
|
|
|
# Larger batches amortize per-expert work during prefill (batch-union effect).
|
|
BATCH_SIZE=1024
|
|
UBATCH_SIZE=512
|
|
|
|
# q8_0 KV: quality-first until turbo2 is PPL-gated for this arch.
|
|
CACHE_TYPE_K=q8_0
|
|
CACHE_TYPE_V=q8_0
|
|
|
|
PARALLEL=1
|
|
|
|
# --cpu-moe: expert FFN tensors stay in CPU RAM (the MoE offload trick)
|
|
# mmap default (no --no-mmap/--mlock): lets page cache manage the 12.1 GB file
|
|
# --jinja: required for gpt-oss harmony chat template
|
|
EXTRA_ARGS="--flash-attn on --cpu-moe --jinja"
|