Files
llama-cpp/swap-stack/Dockerfile
T
mozempkandClaude Fable 5 5e68d30d31 swap-stack: duo config — qwen3-4b full GPU, ornith-35b pure CPU
Duo group (resident main+subagent for pi): flip VRAM to the small model.
Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to
176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s.
CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB
pp compute buffer on CUDA builds (OOM+segfault). Duo section in
MOE-FINDINGS.md; also snapshots prior swap-stack migration state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 08:28:32 +02:00

20 lines
1.2 KiB
Docker

# Tri-binary swap-stack: llama-swap + three llama-server builds.
# /app/llama-server TurboQuant fork (May 2026) — turbo2/3/4 KV, needed by 9B configs
# /app-upstream/llama-server upstream master (Jul 2026) — newest MoE/arch work
# /app/llama-server-ik ik_llama.cpp — -ser / -fmoe / -rtr, fast IQ-quant CPU kernels
# config.yaml picks the binary per model (env: LD_LIBRARY_PATH for upstream).
FROM local/llama-cpp-upstream:server-cuda-sm75-mmq AS upstream
FROM local/ik-llama:server-cuda-sm75 AS ik
FROM local/llama-cpp-turboquant:server-cuda-sm75-mmq
COPY --from=upstream /app /app-upstream
COPY --from=ik /llama-server /app-ik/llama-server
COPY --from=ik /usr/local/lib/libllama.so /usr/local/lib/libggml.so /usr/local/lib/libmtmd.so /app-ik/
COPY --from=ik /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcudart.so.12 /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcublas.so.12 /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcublasLt.so.12 /app-ik/
COPY llama-swap /app/llama-swap
# config.yaml is bind-mounted at runtime (see compose.yaml) so edits
# only need a container restart, not a rebuild.
ENTRYPOINT ["/app/llama-swap", "-config", "/app/config.yaml", "-listen", ":8080"]