Files
llama-cpp/envs/.env.gpt-oss-20b
T
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00

40 lines
1.6 KiB
Bash

# ==============================================================================
# gpt-oss-20b MXFP4 — OpenAI (Aug 2025), Apache 2.0
# Architecture: 20.9B total / 3.6B active MoE (24 layers, 32 experts, top-4)
# Model size: 12.1 GB native MXFP4 (quant-aware trained — no further quant loss)
# Strategy: MoE offload — attention/KV on GPU, expert FFNs in RAM (--cpu-moe)
# mmap (NO mlock): 12.1 GB file pages in on demand, hot experts stay
# in page cache. Reads ~2 GB/token → RAM-BW bound, est. 8-15 t/s
#
# NOTE: quality-sensitive (JSON/code) — KV at q8_0 until turbo2 passes the
# per-model PPL gate (see FINDINGS.md §2: turbo2 broke Qwen3-4B).
# ==============================================================================
MODEL_FILE=gpt-oss-20b-mxfp4.gguf
# All attention layers on GPU; --cpu-moe in EXTRA_ARGS keeps expert weights in RAM.
# Dense (non-expert) part is well under 3.7 GB VRAM.
N_GPU_LAYERS=99
# 131K native; start at 32K. KV is small (GQA, 24 layers).
CTX_SIZE=32768
# t=6 optimal for i7-10750H (6 physical cores). HT hurts (FINDINGS.md §5).
THREADS=6
THREADS_BATCH=6
# Larger batches amortize per-expert work during prefill (batch-union effect).
BATCH_SIZE=1024
UBATCH_SIZE=512
# q8_0 KV: quality-first until turbo2 is PPL-gated for this arch.
CACHE_TYPE_K=q8_0
CACHE_TYPE_V=q8_0
PARALLEL=1
# --cpu-moe: expert FFN tensors stay in CPU RAM (the MoE offload trick)
# mmap default (no --no-mmap/--mlock): lets page cache manage the 12.1 GB file
# --jinja: required for gpt-oss harmony chat template
EXTRA_ARGS="--flash-attn on --cpu-moe --jinja"