Files
llama-cpp/envs/.env.ornith-9b
T
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00

39 lines
1.3 KiB
Bash
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ==============================================================================
# Ornith-1.0-9B Q8_0 — Coding-focused, MIT license
# Architecture: ~9B parameters, transformer-based
# Model size: 8.9 GB | VRAM usage: ~3.4 GB (11 layers on GPU)
# RAM usage: ~5.5 GB (remaining layers pinned via mlock)
#
# Optimization: Similar architecture to Qwen3.5-9B, apply same tuning
# Target: ~4.4 t/s with turbo2 KV, 32K context
#
# Benchmark: Thread sweep shows t=6 optimal (physical cores only, HT hurts)
# ==============================================================================
MODEL_FILE=ornith-9b-Q8_0.gguf
# GPU: 11 layers fit in 3.7 GB VRAM. ngl=12 causes OOM at ctx>2048.
N_GPU_LAYERS=11
# 32K context with turbo2 KV (~104 MiB vs ~3.3 GB for f16)
CTX_SIZE=32768
# t=6 optimal for i7-10750H (6 physical cores). t>6 uses HT which hurts.
THREADS=6
THREADS_BATCH=6
# Larger batch for better throughput. ubatch=256 for CPU-split (ngl=11/36 layers)
BATCH_SIZE=512
UBATCH_SIZE=256
# turbo2: 2-bit KV cache, 6.4× smaller than f16. Requires TurboQuant image.
CACHE_TYPE_K=turbo2
CACHE_TYPE_V=turbo2
PARALLEL=1
# --no-mmap --mlock: pins entire model in RAM (prevents paging, avoids cold reads)
# --flash-attn on: +2-3% prefill boost, required for bigctx
# --jinja: required for Ornith chat formatting
EXTRA_ARGS="--flash-attn on --no-mmap --mlock --jinja"