Files
llama-cpp/envs/.env.ornith-35b
T
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00

34 lines
1.1 KiB
Bash

# ==============================================================================
# Ornith-1.0-35B IQ4_XS — DeepReinforce (Jun 2026), MIT
# Architecture: Qwen3_5MoeForConditionalGeneration — RL-finetune on the
# Qwen3.5-35B-A3B skeleton → 35B total / ~3B active MoE
# Model size: 18.8 GB IQ4_XS (bartowski). Coding-focused, self-improving RL.
# Strategy: identical to qwen36-35b — MoE offload (--cpu-moe), mmap > RAM.
#
# ⚠ RAM: 18.8 GB file on 15 GiB machine — stop heavy containers first.
# ==============================================================================
MODEL_FILE=deepreinforce-ai_Ornith-1.0-35B-IQ4_XS.gguf
# Dense backbone on GPU, routed experts → CPU via --cpu-moe.
N_GPU_LAYERS=99
CTX_SIZE=32768
# t=6 optimal for i7-10750H (6 physical cores). HT hurts.
THREADS=6
THREADS_BATCH=6
BATCH_SIZE=1024
UBATCH_SIZE=512
# q8_0 KV: quality-first until turbo2 is PPL-gated for this arch.
CACHE_TYPE_K=q8_0
CACHE_TYPE_V=q8_0
PARALLEL=1
# --cpu-moe: routed experts in CPU RAM/page cache; mmap default (file > RAM)
# --jinja: Ornith chat template (same requirement as ornith-9b)
EXTRA_ARGS="--flash-attn on --cpu-moe --jinja"