compose: replace open-webui with AnythingLLM on swap-stack endpoint; drop IQ4 GGUFs

AnythingLLM (port 3001, profile webui) points at llama-swap-stack:8080
via generic-openai, default model ornith-35b-duo (131K ctx). Deleted the
two >RAM IQ4 GGUFs (36.5GB — thrash-only until 64GB RAM upgrade) and
their swap-stack entries; re-download paths noted in config comment.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-10 16:02:33 +02:00
co-authored by Claude Fable 5
parent aba5dd0a45
commit 85e27a1a47
2 changed files with 19 additions and 30 deletions
+3 -21
View File
@@ -102,17 +102,6 @@ models:
--cont-batching --parallel 1
ttl: 300
"qwen36-35b":
name: "Qwen3.6-35B-A3B UD-IQ4_XS"
description: "35B/3B-active MoE, 17.7GB mmap > RAM. Stop heavy containers first"
cmd: |
${server-base} ${moe-offload} ${q8-kv}
--model /models/Qwen3.6-35B-A3B-UD-IQ4_XS.gguf
--ctx-size 32768
--batch-size 1024 --ubatch-size 512
--cont-batching --parallel 1
ttl: 300
"qwen36-35b-q2":
name: "Qwen3.6-35B-A3B UD-Q2_K_XL"
description: "DAILY DRIVER 35B: 23.0 tg / 36 pp. ds4 asymmetric recipe (dense high-bit, experts 2-bit), 12.3GB page-cache resident, last 3 expert layers in VRAM. Gates pass"
@@ -137,16 +126,9 @@ models:
--cont-batching --parallel 1
ttl: 300
"ornith-35b-iq4":
name: "Ornith-1.0-35B IQ4_XS"
description: "Quality-first variant, 18.8GB mmap > RAM = ~2.8 t/s thrash. Batch jobs only (or post-RAM-upgrade)"
cmd: |
${server-base} ${moe-offload} ${q8-kv}
--model /models/deepreinforce-ai_Ornith-1.0-35B-IQ4_XS.gguf
--ctx-size 32768
--batch-size 1024 --ubatch-size 512
--cont-batching --parallel 1
ttl: 300
# (IQ4 variants of ornith-35b/qwen36-35b deleted 2026-07-10 — 36.5GB disk,
# both were >RAM thrash-only until the 64GB upgrade. Re-download if needed:
# bartowski Ornith-1.0-35B IQ4_XS, unsloth Qwen3.6-35B-A3B UD-IQ4_XS.)
# ── Dense 9B (RAM-bandwidth-bound, ~4.4 t/s) ───────────────────────────────
# (ik_llama / upstream-master A/B variants removed 2026-07-10 — questions