Duo group (resident main+subagent for pi): flip VRAM to the small model. Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to 176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s. CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB pp compute buffer on CUDA builds (OOM+segfault). Duo section in MOE-FINDINGS.md; also snapshots prior swap-stack migration state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
34 lines
1.1 KiB
Bash
34 lines
1.1 KiB
Bash
# ==============================================================================
|
|
# Ornith-1.0-35B IQ4_XS — DeepReinforce (Jun 2026), MIT
|
|
# Architecture: Qwen3_5MoeForConditionalGeneration — RL-finetune on the
|
|
# Qwen3.5-35B-A3B skeleton → 35B total / ~3B active MoE
|
|
# Model size: 18.8 GB IQ4_XS (bartowski). Coding-focused, self-improving RL.
|
|
# Strategy: identical to qwen36-35b — MoE offload (--cpu-moe), mmap > RAM.
|
|
#
|
|
# ⚠ RAM: 18.8 GB file on 15 GiB machine — stop heavy containers first.
|
|
# ==============================================================================
|
|
|
|
MODEL_FILE=deepreinforce-ai_Ornith-1.0-35B-IQ4_XS.gguf
|
|
|
|
# Dense backbone on GPU, routed experts → CPU via --cpu-moe.
|
|
N_GPU_LAYERS=99
|
|
|
|
CTX_SIZE=32768
|
|
|
|
# t=6 optimal for i7-10750H (6 physical cores). HT hurts.
|
|
THREADS=6
|
|
THREADS_BATCH=6
|
|
|
|
BATCH_SIZE=1024
|
|
UBATCH_SIZE=512
|
|
|
|
# q8_0 KV: quality-first until turbo2 is PPL-gated for this arch.
|
|
CACHE_TYPE_K=q8_0
|
|
CACHE_TYPE_V=q8_0
|
|
|
|
PARALLEL=1
|
|
|
|
# --cpu-moe: routed experts in CPU RAM/page cache; mmap default (file > RAM)
|
|
# --jinja: Ornith chat template (same requirement as ornith-9b)
|
|
EXTRA_ARGS="--flash-attn on --cpu-moe --jinja"
|