Commit Graph
6 Commits
Author SHA1 Message Date
mozempkandClaude Fable 5 f078677c88 warmboot: heal fully-down duo when stack is idle
Stack restart after the sidecar's boot pass left duo_up=0 and the heal
branch (which required exactly one member) never fired. Heal now also
triggers when zero duo members run AND the running list is empty, so a
bounced stack repopulates automatically. A deliberate swap to a non-duo
model still isn't fought (running list non-empty).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:21:46 +02:00
mozempkandClaude Fable 5 67c1b6b827 warmboot: self-heal crashed duo members + boot-restore guard
Suspend/resume kills the GPU-side duo member (CUDA context lost; often
wedged nvidia_uvm on this machine). Daemon now ticks every 60s: if
exactly one duo member is running, the missing one is reloaded via its
/upstream health endpoint and its snapshot restored. Saves stay at
5-min cadence. Boot restore skips members that are already running.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:13:01 +02:00
mozempkandClaude Fable 5 b27c37a552 swap-stack: duo-warmboot sidecar — automatic prompt-KV restore/save
curlimages/curl sidecar (profile stack, depends_on healthy): restores
duo slot caches on boot (pre-loads both duo servers) and auto-saves
every 300s while loaded. No manual duo-warmboot.sh calls needed; the
script remains for ad-hoc use. Snapshot tracks the latest real session;
prefix matching (--cache-reuse) handles divergent new sessions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 16:16:24 +02:00
mozempkandClaude Fable 5 597b882743 swap-stack: slot KV save/restore for duo prompt caches (x570 S5 pattern)
--slot-save-path + --cache-reuse 256 on both duo members, ./kv-cache
mounted at /cache. Measured on ornith (3K-token system prompt): warm
prefix reuse 296ms vs 106.7s cold prefill (361x); slot save 84MB in
59ms, GDN recurrent state serializes. scripts/duo-warmboot.sh wraps
save/restore/status via llama-swap /upstream passthrough. Restore
after restart: run 'duo-warmboot.sh restore' post stack boot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 16:13:20 +02:00
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00
mozempk 4ad296608b Initial commit: tuned multi-model llama.cpp stack
- 5 models: SmolLM3-3B, Gemma4-E2B/E4B, Qwen3-4B, Qwen3.5-9B
- TurboQuant image (FORCE_MMQ): +6-11% free speed on Turing GPUs
- Bigctx profiles (-nkvo KV in RAM): 2-16x context gain
- turbo2 KV: 2x smaller, benchmarked against PPL quality gate
- Per-model env files with justified parameters
- kv_quant_test.sh + cpu_ctx_test.sh benchmark scripts
- docs/FINDINGS.md: surprises, pitfalls, recommendations
- docs/ARCHITECTURE.md: compose + test script design
2026-05-06 15:56:40 +02:00