Duo group (resident main+subagent for pi): flip VRAM to the small model. Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to 176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s. CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB pp compute buffer on CUDA builds (OOM+segfault). Duo section in MOE-FINDINGS.md; also snapshots prior swap-stack migration state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
20 lines
1.2 KiB
Docker
20 lines
1.2 KiB
Docker
# Tri-binary swap-stack: llama-swap + three llama-server builds.
|
|
# /app/llama-server TurboQuant fork (May 2026) — turbo2/3/4 KV, needed by 9B configs
|
|
# /app-upstream/llama-server upstream master (Jul 2026) — newest MoE/arch work
|
|
# /app/llama-server-ik ik_llama.cpp — -ser / -fmoe / -rtr, fast IQ-quant CPU kernels
|
|
# config.yaml picks the binary per model (env: LD_LIBRARY_PATH for upstream).
|
|
FROM local/llama-cpp-upstream:server-cuda-sm75-mmq AS upstream
|
|
FROM local/ik-llama:server-cuda-sm75 AS ik
|
|
|
|
FROM local/llama-cpp-turboquant:server-cuda-sm75-mmq
|
|
|
|
COPY --from=upstream /app /app-upstream
|
|
COPY --from=ik /llama-server /app-ik/llama-server
|
|
COPY --from=ik /usr/local/lib/libllama.so /usr/local/lib/libggml.so /usr/local/lib/libmtmd.so /app-ik/
|
|
COPY --from=ik /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcudart.so.12 /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcublas.so.12 /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcublasLt.so.12 /app-ik/
|
|
COPY llama-swap /app/llama-swap
|
|
|
|
# config.yaml is bind-mounted at runtime (see compose.yaml) so edits
|
|
# only need a container restart, not a rebuild.
|
|
ENTRYPOINT ["/app/llama-swap", "-config", "/app/config.yaml", "-listen", ":8080"]
|