Duo group (resident main+subagent for pi): flip VRAM to the small model. Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to 176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s. CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB pp compute buffer on CUDA builds (OOM+segfault). Duo section in MOE-FINDINGS.md; also snapshots prior swap-stack migration state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
39 lines
1.3 KiB
Bash
39 lines
1.3 KiB
Bash
# ==============================================================================
|
||
# Ornith-1.0-9B Q8_0 — Coding-focused, MIT license
|
||
# Architecture: ~9B parameters, transformer-based
|
||
# Model size: 8.9 GB | VRAM usage: ~3.4 GB (11 layers on GPU)
|
||
# RAM usage: ~5.5 GB (remaining layers pinned via mlock)
|
||
#
|
||
# Optimization: Similar architecture to Qwen3.5-9B, apply same tuning
|
||
# Target: ~4.4 t/s with turbo2 KV, 32K context
|
||
#
|
||
# Benchmark: Thread sweep shows t=6 optimal (physical cores only, HT hurts)
|
||
# ==============================================================================
|
||
|
||
MODEL_FILE=ornith-9b-Q8_0.gguf
|
||
|
||
# GPU: 11 layers fit in 3.7 GB VRAM. ngl=12 causes OOM at ctx>2048.
|
||
N_GPU_LAYERS=11
|
||
|
||
# 32K context with turbo2 KV (~104 MiB vs ~3.3 GB for f16)
|
||
CTX_SIZE=32768
|
||
|
||
# t=6 optimal for i7-10750H (6 physical cores). t>6 uses HT which hurts.
|
||
THREADS=6
|
||
THREADS_BATCH=6
|
||
|
||
# Larger batch for better throughput. ubatch=256 for CPU-split (ngl=11/36 layers)
|
||
BATCH_SIZE=512
|
||
UBATCH_SIZE=256
|
||
|
||
# turbo2: 2-bit KV cache, 6.4× smaller than f16. Requires TurboQuant image.
|
||
CACHE_TYPE_K=turbo2
|
||
CACHE_TYPE_V=turbo2
|
||
|
||
PARALLEL=1
|
||
|
||
# --no-mmap --mlock: pins entire model in RAM (prevents paging, avoids cold reads)
|
||
# --flash-attn on: +2-3% prefill boost, required for bigctx
|
||
# --jinja: required for Ornith chat formatting
|
||
EXTRA_ARGS="--flash-attn on --no-mmap --mlock --jinja"
|