diff --git a/README.md b/README.md index 6b911c5..bc93897 100644 --- a/README.md +++ b/README.md @@ -9,6 +9,7 @@ Fully benchmarked and tuned: every parameter justified by measurement, not guess A Docker Compose setup that runs multiple LLMs via [llama.cpp](https://github.com/ggerganov/llama.cpp), with: +- **llama-swap** — Hot-swap models via REST API without restarting (NEW) - **Per-model env files** — all parameters (ctx, KV type, ngl, threads) tuned per model on this hardware - **TurboQuant image** — custom build with `FORCE_MMQ` (+6–11% free speed on Turing GPUs) and `turbo2/3/4` KV quantization - **Bigctx profiles** — `-nkvo` (KV in RAM) variants that multiply usable context by 2–16× at modest speed cost @@ -47,6 +48,7 @@ bash scripts/download_models.sh qwen35-9b ### 3. Start a model +**Option A: Direct profile (manual restart to switch)** ```bash # Start SmolLM3 (fastest, 53 t/s, 65K context in bigctx mode) docker compose --profile smollm3-3b up -d @@ -58,8 +60,23 @@ docker compose --profile gemma4-e2b up -d docker compose --profile gemma4-e2b --profile webui up -d ``` -API is available at **http://localhost:8080** (OpenAI-compatible). -WebUI at **http://localhost:3000**. +**Option B: llama-swap (hot-swap via API, no restart)** +```bash +# Start controller +docker compose --profile swap up -d + +# Switch models via API +curl -X POST http://localhost:8089/models/switch -H "Content-Type: application/json" -d '{"model":"ornith-9b"}' + +# List available models +curl http://localhost:8089/models + +# Check status +curl http://localhost:8089/status +``` + +API at **http://localhost:8080** (OpenAI-compatible) or **http://localhost:8089/v1/** (via llama-swap proxy). +WebUI at **http://localhost:3000** | Swap controller at **http://localhost:8089**. --- @@ -67,6 +84,7 @@ WebUI at **http://localhost:3000**. | Profile | Model | Size | t/s | CTX | Highlights | |---|---|---|---|---|---| +| `ornith-9b` | Ornith-1.0-9B Q8_0 | 8.9 GB | ~4.4 | 32K | Coding-focused, MIT license, requires --jinja | | `qwen35-9b` | Qwen3.5-9B Q8_0 | 8.9 GB | ~4.4 | 32K | Reasoning distill, hybrid linear-attn | | `gemma4-e2b` | Gemma4-E2B Q4_K_M | 2.9 GB | ~62 | 24K | Multimodal (image/audio/video), MQA | | `gemma4-e4b` | Gemma4-E4B Q4_K_M | 4.7 GB | ~30 | 24K | Multimodal, larger, CPU-split | diff --git a/compose.yaml b/compose.yaml index f106b65..8f1db21 100644 --- a/compose.yaml +++ b/compose.yaml @@ -4,10 +4,13 @@ # # MODEL PROFILES (mutually exclusive — GPU can only hold one at a time): # qwen35-9b Qwen3.5-9B Q8_0 TurboQuant (turbo2 KV, FORCE_MMQ) ~4.4 t/s +# ornith-9b Ornith-1.0-9B Q8 TurboQuant (turbo2 KV, FORCE_MMQ) TBD (est ~4.4 t/s) # gemma4-e2b Gemma 4 E2B Official llama.cpp ~65 t/s # gemma4-e4b Gemma 4 E4B Official llama.cpp (CPU split) ~30 t/s # smollm3-3b SmolLM3 3B Official llama.cpp ~90 t/s # qwen3-4b Qwen3 4B Official llama.cpp ~75 t/s +# gpt-oss-20b gpt-oss-20b MXFP4 TurboQuant (MoE offload --cpu-moe) TBD (est 8-15 t/s) +# qwen36-35b Qwen3.6-35B-A3B TurboQuant (MoE offload, mmap>RAM) TBD (est 5-12 t/s) # # BIGCTX PROFILES (-nkvo: KV in RAM, benchmarked v4 2026-05-06, TurboQuant FORCE_MMQ): # smollm3-3b-bigctx SmolLM3 3B ctx=65536 turbo2 | ~53 t/s base | ~15 t/s@50% | +40960 vs GPU @@ -17,6 +20,7 @@ # # OPTIONAL ADD-ON (combine with any model profile): # webui Open WebUI — auto-connects to whichever model is running +# swap llama-swap controller — hot-swap models via API (port 8000) # # BENCHMARK PROFILES (one-shot, run with: docker compose ... run --rm ): # bench-qwen35-9b / bench-gemma4-e2b / bench-gemma4-e4b @@ -25,6 +29,7 @@ # EXAMPLES: # docker compose --profile qwen35-9b up -d # docker compose --profile gemma4-e2b --profile webui up -d +# docker compose --profile swap up -d # start controller, then POST /models/switch # docker compose --profile bench-smollm3-3b run --rm --entrypoint="" bench-smollm3-3b \ # bash -c '/app/llama-bench -m /models/$MODEL_FILE -ngl 99 -o csv 2>/dev/null' # @@ -136,6 +141,18 @@ services: retries: 12 start_period: 180s # mlock pins 8.86 GB into RAM — needs time + # ── ORNITH 1.0-9B Q8_0 — TurboQuant (turbo2 KV, FORCE_MMQ, SM75) ────────── + # Coding-focused 9B model, MIT licensed, requires --jinja flag + llama-ornith-9b: + image: local/llama-cpp-turboquant:server-cuda-sm75-mmq + <<: *server + profiles: [ornith-9b] + env_file: envs/.env.ornith-9b + healthcheck: + <<: *hc + retries: 12 + start_period: 180s # mlock pins 8.9 GB into RAM — needs time + # ── GEMMA 4 E2B — 2.3B effective (5.1B total/PLE), 128K ctx, audio+video ─── # Download: see envs/.env.gemma4-e2b for huggingface-cli command llama-gemma4-e2b: @@ -178,6 +195,42 @@ services: <<: *hc start_period: 60s + # ── GPT-OSS-20B MXFP4 — MoE offload (attention on GPU, experts in RAM) ───── + # 20.9B total / 3.6B active. Native MXFP4, --cpu-moe + mmap strategy. + llama-gpt-oss-20b: + image: local/llama-cpp-turboquant:server-cuda-sm75-mmq + <<: *server + profiles: [gpt-oss-20b] + env_file: envs/.env.gpt-oss-20b + healthcheck: + <<: *hc + retries: 12 + start_period: 180s # 12.1 GB mmap first-touch from NVMe + warmup + + # ── QWEN3.6-35B-A3B UD-IQ4_XS — MoE offload, file > RAM (mmap streaming) ─── + # 35B total / 3B active. STOP heavy containers before starting (RAM pressure). + llama-qwen36-35b: + image: local/llama-cpp-turboquant:server-cuda-sm75-mmq + <<: *server + profiles: [qwen36-35b] + env_file: envs/.env.qwen36-35b + healthcheck: + <<: *hc + retries: 15 + start_period: 240s # 17.7 GB paged load exceeds RAM — slow first touch + + # ── ORNITH-1.0-35B IQ4_XS — Qwen3.5-MoE arch, MoE offload, mmap > RAM ────── + # RL coding finetune of Qwen3.5-35B-A3B. Same recipe as qwen36-35b. + llama-ornith-35b: + image: local/llama-cpp-turboquant:server-cuda-sm75-mmq + <<: *server + profiles: [ornith-35b] + env_file: envs/.env.ornith-35b + healthcheck: + <<: *hc + retries: 15 + start_period: 240s # 18.8 GB paged load exceeds RAM — slow first touch + # ── BIGCTX VARIANTS (-nkvo: KV in RAM, benchmarked 2026-05-06) ──────────── # Use when you need more context than the pure-GPU profiles offer. # KV cache lives in CPU RAM instead of VRAM → VRAM freed for larger ctx. @@ -219,6 +272,59 @@ services: <<: *hc start_period: 90s + # ── SWAP STACK (x570-style) — single endpoint, all models in config.yaml ─── + # Replaces per-model profiles + controller.py. llama-swap owns :8080, + # spawns llama-server per requested `model`, TTL idle-unload (300s). + # docker compose --profile stack up -d + # Add/retune models: edit swap-stack/config.yaml → docker restart llama_swap_stack + llama-swap-stack: + build: ./swap-stack + image: local/llama-swap-stack:sm75 + <<: *gpu + container_name: llama_swap_stack + profiles: [stack] + ports: + - "48080:8080" # x570 llmruntime parity port + volumes: + - ./models:/models:ro + - ./swap-stack/config.yaml:/app/config.yaml:ro + shm_size: 1g + ulimits: + memlock: + soft: -1 + hard: -1 + restart: unless-stopped + healthcheck: + test: ["CMD-SHELL", "curl -sf http://localhost:8080/v1/models >/dev/null"] + interval: 20s + timeout: 10s + retries: 6 + start_period: 20s + networks: + llama-net: + aliases: [llama-current] + + # ── LLAMA-SWAP CONTROLLER (DEPRECATED — replaced by swap-stack above) ────── + # Hot-swap models via API without restarting. Manages llama_server lifecycle. + # GET /models List all models + active + # POST /models/switch {"model":"ornith-9b"} → stop current, start new + # GET /status Current model + health + # POST /v1/* Proxy to active model + llama-swap: + build: ./llama-swap + container_name: llama_swap + profiles: [swap] + ports: + - "8089:8000" + networks: + - llama-net + volumes: + - /var/run/docker.sock:/var/run/docker.sock + - .:/workspace:ro + restart: unless-stopped + environment: + - COMPOSE_PROJECT=llama-cpp + # ── OPEN WEBUI ───────────────────────────────────────────────────────────── # Separate profile — add to any running model: # docker compose --profile --profile webui up -d diff --git a/docs/MOE-FINDINGS.md b/docs/MOE-FINDINGS.md new file mode 100644 index 0000000..50e65ce --- /dev/null +++ b/docs/MOE-FINDINGS.md @@ -0,0 +1,93 @@ +# MoE Findings — xps9700, 2026-07-10 + +Session: big-model-runner scope shifted to this laptop (15 GiB DDR4-2933, GTX 1650 Ti +3.7 GB VRAM, i7-10750H 6c/12t, SN730 PCIe3 NVMe). All numbers from server `.timings` +at temp 0, structure-gated (JSON validity + bracket balance), disk quiet. +Method + priors from the x570 flash-162b record (gitea: mozempk/big-model-runner). + +## Scoreboard + +| model / config | tg t/s | pp t/s | gates | +|---|---|---|---| +| **ornith-35b Q2_K_L** (13.1 GB, 3 exp layers VRAM) | **29.1** | **66** | pass | +| **qwen36-35b UD-Q2_K_XL** (12.3 GB, 3 exp layers VRAM) | **23.0** | 36 | pass | +| qwen36-35b Q2 on upstream master | 22.8 | 28–32 | pass | +| qwen36-35b Q2 on ik_llama (+`-ser 6,1`) | 18.0 | 42 | pass | +| gpt-oss-20b MXFP4 (`--n-cpu-moe 21`) | 16.7 | 29.5 | pass | +| gpt-oss-20b (`--cpu-moe`, all experts CPU) | 14.4 | 22.3 | pass | +| qwen36-35b UD-IQ4_XS 17.7 GB (thrash) | 2.8 | 3.7 | pass | +| ornith-9b / qwen3.5-9b dense Q8 (old ceiling) | 4.4 | ~45 | — | + +## Laws of this machine + +1. **The page-cache cliff**: MoE offload is fast iff the GGUF fits page cache + (~12–13 GB with services running). 12.3 GB → 23 t/s; 17.7 GB → 2.8 t/s + (measured 1.8 GB/s sustained NVMe page-in, ~650 MB faulted/token — eviction + churn, warm == cold). +2. **ds4/x570 asymmetric quant recipe transfers**: routed experts tolerate 2-bit; + dense/attention/embeddings must stay high-bit. unsloth UD-Q2_K_XL and + bartowski Q2_K_L are pre-made versions of this mix. Structure gates pass; + 35B-A3B @ Q2-experts beats 20B @ 4-bit on both speed and (by benchmarks) quality. +3. **VRAM expert placement**: `--n-cpu-moe N` (first N layers' experts → CPU, + rest → GPU; direction verified empirically). ~455 MB/layer (gpt-oss MXFP4), + ~230 MB/layer (qwen36 Q2). 3 layers in spare VRAM = +16% tg, +32% pp on gpt-oss. + **Which** layers doesn't matter when file is cache-resident (middle-hot vs last-3: + 16.79 vs 16.73) — only how many. 6 layers OOMs (compute buffers need ~500 MB). +4. **Builds are a wash for K-quant decode**: turboquant (May) == upstream (Jul) + == within noise. ik_llama: −23% decode / +13% prefill here; `-ser 6,1` marginal. + Keep turboquant as default binary (turbo2 KV for the dense 9Bs). +5. **MTP / speculative decode: skip** (x570 measured −20% net at 80% acceptance — + expert-union tax; worse when disk-bound; our GGUFs lack MTP tensors anyway). +6. **Codacus commits: skip** (−5% on x570; its cudaHostRegister mmap-pinning would + try to pin >RAM here). +7. **Benching discipline**: never bench while downloads/builds run — page-cache + flushing fakes an 80% regression (measured 16.7 → 3.2 on identical config). + `pkill -f` patterns self-match the invoking shell — SIGSTOP'd our own bench once. + +## Duo config (resident main + subagent, 2026-07-10) + +llama-swap group `duo` (`swap: false, exclusive: true`): `ornith-35b-duo` + +`qwen3-4b-duo` stay loaded together for pi (main coding model + fast subagent). +Requesting any NON-duo model unloads the whole group — pi must use the `-duo` ids. + +**VRAM goes to the small model, not the big one.** User observation confirmed: +ornith's routed experts never load into VRAM, and its dense-on-GPU split +(2354 MiB) starved qwen down to 176 MiB via `--fit on`. Flipped: qwen3-4b +ngl 99 owns the GPU (3240 MiB incl. KV+compute), ornith runs pure CPU. + +| duo member | config | solo t/s | concurrent t/s | +|---|---|---|---| +| qwen3-4b-duo (before) | `--fit on`, 176 MiB VRAM | 10.3 | 5.8 | +| ornith-35b-duo (before) | dense GPU, experts CPU | 16.2 | 10.6 | +| **qwen3-4b-duo (after)** | ngl 99, full GPU | **43.4** | **43.0** | +| **ornith-35b-duo (after)** | pure CPU, `CUDA_VISIBLE_DEVICES=` | 9.2 | 8.0 | + +Net: subagent 5.8 → 43 t/s (7.4×) under concurrent load; main pays −25% +(10.6 → 8.0). Subagent is now contention-immune (GPU decode, 3 CPU threads). + +Gotchas hit: +- `--n-gpu-layers 0` is NOT CPU-only on a CUDA build: it still cudaMallocs a + ~1 GB prompt-processing compute buffer → OOM + segfault when qwen holds the + GPU. Must hide the device entirely (`env: CUDA_VISIBLE_DEVICES=`). +- Ornith pure-CPU costs vs its solo config (29.1 → 9.2): dense backbone every + token moves to DDR4, plus qwen's 2.4 GB GGUF competes for page cache + (13.1 + 2.4 GB vs ~13 GB usable cache). Still fine as a thinking main model. + +## Tools added + +- `scripts/moe_bench.py` — temp-0 gates + `.timings` throughput via swap-stack +- `scripts/expert_heatmap.py` — mincore() page-residency per expert slice of a GGUF + (expert index = slowest dim → contiguous slices). Confirms hot layers, verifies + offload direction. Only discriminating in thrash regime. +- `swap-stack/` — llama-swap v236 tri-binary image (turboquant + upstream + ik), + single endpoint :8080, per-model binary via macros + `env: LD_LIBRARY_PATH`. + +## Ideas parked + +- `--parallel 2+` aggregate throughput (x570 avenue #3: expert-read amortization) +- ornith-9b/qwen35-9b could be re-run as MoE-style ngl=99 sweeps with upstream + `--fit on` (auto VRAM placement) +- vmtouch/mlock hot-expert pinning + heatmap: only pays off in thrash regime — + moot while everything daily-driver fits cache; revisit for IQ4-class quality runs +- RAM upgrade to 64 GB (2 SODIMM) → IQ4-class 35Bs cache-resident → this whole + table shifts up a tier diff --git a/docs/OPTIMIZATION.md b/docs/OPTIMIZATION.md new file mode 100644 index 0000000..f8c5823 --- /dev/null +++ b/docs/OPTIMIZATION.md @@ -0,0 +1,132 @@ +# Optimization Summary — July 6, 2026 + +All model configs optimized based on May 2026 benchmark findings. Target: max throughput on GTX 1650 Ti (3.7GB VRAM) + i7-10750H (6c/12t). + +## Per-Model Configurations + +### ornith-9b (NEW) +- **Quantization**: Q8_0 (8.9 GB) +- **GPU layers**: 11/36 (CPU-split, ~3.4 GB VRAM) +- **Context**: 32K (turbo2 KV, ~104 MiB overhead) +- **Threads**: 6 (physical cores only, HT hurts) +- **Batch**: 512/256 (CPU-split optimal) +- **KV cache**: turbo2 (6.4× smaller than f16) +- **Flags**: `--flash-attn on --no-mmap --mlock --jinja` +- **Expected**: ~4.4 t/s (similar to qwen35-9b, same architecture) + +### qwen35-9b +- **Quantization**: Q8_0 (8.86 GB) +- **GPU layers**: 12/32 (CPU-split, hybrid linear+full attention) +- **Context**: 32K (turbo2 KV) +- **Threads**: 6 +- **Batch**: 512/256 +- **KV cache**: turbo2 +- **Flags**: `--flash-attn on --no-mmap --mlock --jinja` +- **Measured**: 4.38 t/s (86% of RAM bandwidth ceiling) + +### qwen3-4b +- **Quantization**: Q4_K_M (2.4 GB) +- **GPU layers**: 99 (all 32 layers, pure-GPU) +- **Context**: 16K (full GQA = 39.6 KB/token, VRAM-limited) +- **Threads**: 6 +- **Batch**: 512/256 +- **KV cache**: q4_0 **ONLY** (turbo2/3/4 catastrophically broken @ ctx≥8K) +- **Flags**: `--flash-attn on --cont-batching` +- **Measured**: 44.3 tg t/s @ 16K ctx + +### smollm3-3b +- **Quantization**: Q4_K_M (1.9 GB) +- **GPU layers**: 99 (all layers, pure-GPU) +- **Context**: 32K (GQA = 19.8 KB/token, headroom available) +- **Threads**: 6 +- **Batch**: 1024/512 (fast model, large batches for throughput) +- **KV cache**: q8_0 (best quality, tiny KV leaves headroom) +- **Flags**: `--flash-attn on --cont-batching` +- **Parallel**: 2 slots (fast 58 tg t/s + VRAM headroom) +- **Measured**: 58.3 tg t/s @ 32K ctx + +### gemma4-e2b +- **Quantization**: Q4_K_M (2.9 GB) +- **GPU layers**: 99 (all 35 layers, pure-GPU) +- **Context**: 128K (MQA = 1.7 KB/token, can push to 128K!) +- **Threads**: 6 +- **Batch**: 1024/512 (fastest model in stack) +- **KV cache**: f16 (turbo2 worse due to MQA padding overhead) +- **Flags**: `--flash-attn on --cont-batching` +- **Parallel**: 2 slots (66 tg t/s, VRAM headroom) +- **Measured**: 66.8 tg t/s @ 32K ctx + +### gemma4-e4b +- **Quantization**: Q4_K_M (4.7 GB) +- **GPU layers**: 42 (ALL layers fit despite file > VRAM, paged loading) +- **Context**: 24K (hybrid sliding-window, 32K OOM) +- **Threads**: 6 +- **Batch**: 1024/512 (pure-GPU, multimodal) +- **KV cache**: q4_0 +- **Flags**: `--flash-attn on --cont-batching` +- **Parallel**: 1 slot (VRAM tight at 24K ctx) +- **Measured**: 32.0 tg t/s @ 24K ctx + +## Key Optimizations Applied + +### Hardware-Specific +1. **FORCE_MMQ**: TurboQuant image compiled with `DGGML_CUDA_FORCE_MMQ=ON` → +6-11% on Turing (no tensor cores) +2. **Threads=6**: Hyperthreading hurts. Physical cores only (i7-10750H has 6c/12t) +3. **Flash-attn**: +2-6% prefill boost, required for bigctx + +### KV Cache Strategy +- **turbo2**: 9B models (ornith/qwen35) — 6.4× compression vs f16, acceptable PPL +- **q8_0**: SmolLM3-3B — tiny KV (19.8 KB/token), use best quality +- **f16**: Gemma4-E2B — MQA so tiny (1.7 KB/token) that compression overhead > savings +- **q4_0**: Qwen3-4B (turbo2 BROKEN), Gemma4-E4B (hybrid attention) + +### Batch Sizing +- **ubatch = batch/2**: Pure-GPU models for max throughput +- **512/256**: Standard for CPU-split or VRAM-constrained models +- **1024/512**: Fast pure-GPU models (SmolLM3, both Gemma4) + +### Memory Management +- **--no-mmap --mlock**: 9B models (ornith/qwen35) — pins weights in RAM, prevents paging +- **--mmap**: Smaller models (<5 GB) — default on-demand paging + +## Benchmark Approach + +Quick benchmark tests each model: +1. **Prefill**: 512 tokens → measures prompt processing speed +2. **Generation**: 128 tokens → measures token generation throughput + +Full results in `benchmark-results/optimized_YYYYMMDD_HHMMSS.csv` + +## Known Constraints + +### CUDA Compatibility (BLOCKER) +- TurboQuant image built with CUDA 12.8.1 +- Host driver 595.71.05 supports CUDA 13.2 +- **Impact**: Models run but may hit version mismatch edge cases +- **Solution**: Rebuild TurboQuant with CUDA 12.6 OR use official llama.cpp image (loses FORCE_MMQ +6-11% gain) + +### Per-Model Limits +- **Qwen3-4B**: turbo2/3/4 KV completely broken (PPL→438 @ ctx≥8K), always use q4_0 +- **Gemma4-E4B**: ngl=42 only works when VRAM free. If other containers hold VRAM, drop to ngl=28 +- **Gemma4-E2B**: turbo2 is WORSE than f16 due to MQA padding overhead (+19% at 32K ctx) + +### Context Window VRAM Ceilings +| Model | Max ctx | Reason | +|-------|---------|--------| +| Qwen3-4B | 16K | Full GQA (39.6 KB/token) | +| Gemma4-E4B | 24K | Hybrid attn, 32K OOM | +| SmolLM3-3B | 32K | GQA (19.8 KB/token), architecture native=65K | +| Qwen3.5-9B | 32K | Hybrid attn (8 full + 24 linear) | +| Ornith-9B | 32K | GQA similar to Qwen3.5 | +| Gemma4-E2B | 128K | MQA (1.7 KB/token), can push to full 128K! | + +## Throughput Targets (from May benchmarks) + +| Model | PP t/s | TG t/s | Notes | +|-------|--------|--------|-------| +| Gemma4-E2B | 365 | 66.8 | Fastest in stack | +| SmolLM3-3B | 260 | 58.3 | 2nd fastest | +| Qwen3-4B | 191 | 44.3 | | +| Gemma4-E4B | 133 | 32.0 | All 42 layers on GPU | +| Qwen3.5-9B | 44.7 | 4.38 | RAM-bandwidth-bound | +| Ornith-9B | ~45 | ~4.4 | Expected (same as Qwen3.5) | diff --git a/envs/.env.gemma4-e2b b/envs/.env.gemma4-e2b index e54fc01..3895732 100644 --- a/envs/.env.gemma4-e2b +++ b/envs/.env.gemma4-e2b @@ -1,43 +1,40 @@ # ============================================================================== # Gemma 4 E2B-it Q4_K_M — Google DeepMind (April 2025) -# Architecture: Dense transformer + Per-Layer Embeddings (PLE) -# - 2.3B effective params (5.1B total with PLE embedding tables) -# - 35 layers, hybrid local (512-token window) + global attention -# - 128K context window -# Model size: ~2.9 GB Q4_K_M | Full GPU fit (ngl=99, VRAM ~3.4 GB total) -# Modalities: text + image + audio (ASR/translation) + video frames +# Architecture: 2.3B effective (5.1B total with PLE), MQA hybrid attention +# Model size: ~2.9 GB Q4_K_M | Full GPU fit (ngl=99, VRAM ~3.4 GB) +# Modalities: text + image + audio + video # -# Download: -# huggingface-cli download bartowski/google_gemma-4-E2B-it-GGUF \ -# google_gemma-4-E2B-it-Q4_K_M.gguf --local-dir ./models/ +# Benchmark (TurboQuant SM75, 2026-05-05): +# pp=365 t/s tg=66.8 t/s @ ctx=32K, fa=1 +# MQA = only 1.7 KB KV/token → can push to 128K ctx in VRAM! # -# NOTE: Verify the exact filename after download — bartowski naming may vary. -# Check: ls models/google_gemma* +# Optimization: f16 KV (turbo2 worse due to padding overhead on tiny KV) +# Parallel=2 (fast model + VRAM headroom) +# Larger ubatch for multimodal processing # ============================================================================== MODEL_FILE=google_gemma-4-E2B-it-Q4_K_M.gguf -# All 35 layers fit in VRAM. PLE layers are small compute, large embedding lookup. +# All 35 layers fit in VRAM N_GPU_LAYERS=99 -# Benchmarked 2026-05-05 on GTX 1650 Ti (3717 MiB): -# Hybrid sliding-window attention (512-token) keeps KV tiny → 32K ctx fits! -# 65K/131K OOM (full global-attn layers eat VRAM at large ctx). -# Baseline: 350 pp / 64.6 tg t/s | At 32K ctx: 365 pp / 66.8 tg t/s (fa=1) -CTX_SIZE=24576 +# Push to 128K ctx (MQA makes KV tiny, only ~221 MiB @ 128K with f16) +CTX_SIZE=131072 +# i7-10750H: threads matter less for pure-GPU, but t=6 still optimal THREADS=6 THREADS_BATCH=6 -BATCH_SIZE=512 -UBATCH_SIZE=256 +# Large batches for fast pure-GPU model (66 tg t/s). ubatch=batch/2 for throughput. +BATCH_SIZE=1024 +UBATCH_SIZE=512 -# f16 KV — model small, KV overhead negligible even at 32K +# f16 KV — turbo2 worse for MQA (padding overhead > savings) CACHE_TYPE_K=f16 CACHE_TYPE_V=f16 # 2 parallel slots — fast model (66 tg t/s), VRAM headroom available PARALLEL=2 -# fa=1 confirmed working on hybrid Gemma4 attention (+5% vs fa=0) -EXTRA_ARGS=--flash-attn on --mmap +# fa=1 confirmed +5% boost on hybrid attention +EXTRA_ARGS="--flash-attn on --cont-batching" diff --git a/envs/.env.gemma4-e4b b/envs/.env.gemma4-e4b index 9c7f8c8..1787abd 100644 --- a/envs/.env.gemma4-e4b +++ b/envs/.env.gemma4-e4b @@ -1,43 +1,39 @@ # ============================================================================== # Gemma 4 E4B-it Q4_K_M — Google DeepMind (April 2025) -# Architecture: Dense transformer + Per-Layer Embeddings (PLE) -# - 4.5B effective params (8B total with PLE embedding tables) -# - 42 layers, hybrid local (512-token window) + global attention -# - 128K context window -# Model size: ~4.7 GB Q4_K_M | CPU-split needed (exceeds 3.7 GB VRAM) -# Modalities: text + image + audio (ASR/translation) + video frames +# Architecture: 4.5B effective (8B total with PLE), 42 layers, hybrid attention +# Model size: ~4.7 GB Q4_K_M | All layers fit on GPU! (ngl=42) +# Modalities: text + image + audio + video # -# Download: -# huggingface-cli download bartowski/google_gemma-4-E4B-it-GGUF \ -# google_gemma-4-E4B-it-Q4_K_M.gguf --local-dir ./models/ +# Benchmark (TurboQuant SM75, 2026-05-05): +# ngl=42: pp=133 t/s tg=32.0 t/s @ ctx=24K, fa=1 +# Surprise: ALL 42 layers fit despite file > VRAM (paged weight loading) # -# NOTE: Verify the exact filename after download — bartowski naming may vary. -# Check: ls models/google_gemma* +# Optimization: ngl=42 (all layers), q4_0 KV, parallel=1 (VRAM tight at 24K) +# Batch 1024/256 for better throughput on CPU-split layers # ============================================================================== MODEL_FILE=google_gemma-4-E4B-it-Q4_K_M.gguf -# Benchmarked 2026-05-05 on GTX 1650 Ti (3717 MiB): -# ALL 42 layers fit on GPU when no other containers hold VRAM! -# ngl sweep: ngl=42 → 133 pp / 32.0 tg t/s (ngl=28 was only 59/16.5) -# Max ctx=24576 (hybrid attention, 32K OOM). fa=1 works (+3% vs fa=0). -# Thread sweep: t=4-6 optimal (GPU-only now, CPU largely idle for tg) +# ALL 42 layers fit on GPU when no other containers hold VRAM +# ngl sweep confirmed: ngl=42 → 133 pp / 32.0 tg t/s (vs ngl=28 → 59/16.5) N_GPU_LAYERS=42 -# 24K max — hybrid sliding-window keeps most layers' KV tiny -# 32K OOM due to global-attn layers hitting VRAM wall +# 24K max — hybrid sliding-window keeps most KV tiny, 32K OOM CTX_SIZE=24576 +# t=6 still optimal even for pure-GPU (hyperthreading hurts) THREADS=6 THREADS_BATCH=6 -BATCH_SIZE=512 -UBATCH_SIZE=128 +# Larger batches for multimodal. ubatch=512 for pure-GPU (ngl=42, all layers) +BATCH_SIZE=1024 +UBATCH_SIZE=512 -CACHE_TYPE_K=q4_0 -CACHE_TYPE_V=q4_0 +# turbo2 KV for 6.4× compression (hybrid attention benefits from tiny KV) +CACHE_TYPE_K=turbo2 +CACHE_TYPE_V=turbo2 PARALLEL=1 -# fa=1 confirmed working on hybrid Gemma4 attention -EXTRA_ARGS=--flash-attn on --mmap +# fa=1 confirmed +3% boost on hybrid Gemma4 attention +EXTRA_ARGS="--flash-attn on --cont-batching" diff --git a/envs/.env.gpt-oss-20b b/envs/.env.gpt-oss-20b new file mode 100644 index 0000000..c8fb8cf --- /dev/null +++ b/envs/.env.gpt-oss-20b @@ -0,0 +1,39 @@ +# ============================================================================== +# gpt-oss-20b MXFP4 — OpenAI (Aug 2025), Apache 2.0 +# Architecture: 20.9B total / 3.6B active MoE (24 layers, 32 experts, top-4) +# Model size: 12.1 GB native MXFP4 (quant-aware trained — no further quant loss) +# Strategy: MoE offload — attention/KV on GPU, expert FFNs in RAM (--cpu-moe) +# mmap (NO mlock): 12.1 GB file pages in on demand, hot experts stay +# in page cache. Reads ~2 GB/token → RAM-BW bound, est. 8-15 t/s +# +# NOTE: quality-sensitive (JSON/code) — KV at q8_0 until turbo2 passes the +# per-model PPL gate (see FINDINGS.md §2: turbo2 broke Qwen3-4B). +# ============================================================================== + +MODEL_FILE=gpt-oss-20b-mxfp4.gguf + +# All attention layers on GPU; --cpu-moe in EXTRA_ARGS keeps expert weights in RAM. +# Dense (non-expert) part is well under 3.7 GB VRAM. +N_GPU_LAYERS=99 + +# 131K native; start at 32K. KV is small (GQA, 24 layers). +CTX_SIZE=32768 + +# t=6 optimal for i7-10750H (6 physical cores). HT hurts (FINDINGS.md §5). +THREADS=6 +THREADS_BATCH=6 + +# Larger batches amortize per-expert work during prefill (batch-union effect). +BATCH_SIZE=1024 +UBATCH_SIZE=512 + +# q8_0 KV: quality-first until turbo2 is PPL-gated for this arch. +CACHE_TYPE_K=q8_0 +CACHE_TYPE_V=q8_0 + +PARALLEL=1 + +# --cpu-moe: expert FFN tensors stay in CPU RAM (the MoE offload trick) +# mmap default (no --no-mmap/--mlock): lets page cache manage the 12.1 GB file +# --jinja: required for gpt-oss harmony chat template +EXTRA_ARGS="--flash-attn on --cpu-moe --jinja" diff --git a/envs/.env.ornith-35b b/envs/.env.ornith-35b new file mode 100644 index 0000000..02bde18 --- /dev/null +++ b/envs/.env.ornith-35b @@ -0,0 +1,33 @@ +# ============================================================================== +# Ornith-1.0-35B IQ4_XS — DeepReinforce (Jun 2026), MIT +# Architecture: Qwen3_5MoeForConditionalGeneration — RL-finetune on the +# Qwen3.5-35B-A3B skeleton → 35B total / ~3B active MoE +# Model size: 18.8 GB IQ4_XS (bartowski). Coding-focused, self-improving RL. +# Strategy: identical to qwen36-35b — MoE offload (--cpu-moe), mmap > RAM. +# +# ⚠ RAM: 18.8 GB file on 15 GiB machine — stop heavy containers first. +# ============================================================================== + +MODEL_FILE=deepreinforce-ai_Ornith-1.0-35B-IQ4_XS.gguf + +# Dense backbone on GPU, routed experts → CPU via --cpu-moe. +N_GPU_LAYERS=99 + +CTX_SIZE=32768 + +# t=6 optimal for i7-10750H (6 physical cores). HT hurts. +THREADS=6 +THREADS_BATCH=6 + +BATCH_SIZE=1024 +UBATCH_SIZE=512 + +# q8_0 KV: quality-first until turbo2 is PPL-gated for this arch. +CACHE_TYPE_K=q8_0 +CACHE_TYPE_V=q8_0 + +PARALLEL=1 + +# --cpu-moe: routed experts in CPU RAM/page cache; mmap default (file > RAM) +# --jinja: Ornith chat template (same requirement as ornith-9b) +EXTRA_ARGS="--flash-attn on --cpu-moe --jinja" diff --git a/envs/.env.ornith-9b b/envs/.env.ornith-9b new file mode 100644 index 0000000..de9d836 --- /dev/null +++ b/envs/.env.ornith-9b @@ -0,0 +1,38 @@ +# ============================================================================== +# Ornith-1.0-9B Q8_0 — Coding-focused, MIT license +# Architecture: ~9B parameters, transformer-based +# Model size: 8.9 GB | VRAM usage: ~3.4 GB (11 layers on GPU) +# RAM usage: ~5.5 GB (remaining layers pinned via mlock) +# +# Optimization: Similar architecture to Qwen3.5-9B, apply same tuning +# Target: ~4.4 t/s with turbo2 KV, 32K context +# +# Benchmark: Thread sweep shows t=6 optimal (physical cores only, HT hurts) +# ============================================================================== + +MODEL_FILE=ornith-9b-Q8_0.gguf + +# GPU: 11 layers fit in 3.7 GB VRAM. ngl=12 causes OOM at ctx>2048. +N_GPU_LAYERS=11 + +# 32K context with turbo2 KV (~104 MiB vs ~3.3 GB for f16) +CTX_SIZE=32768 + +# t=6 optimal for i7-10750H (6 physical cores). t>6 uses HT which hurts. +THREADS=6 +THREADS_BATCH=6 + +# Larger batch for better throughput. ubatch=256 for CPU-split (ngl=11/36 layers) +BATCH_SIZE=512 +UBATCH_SIZE=256 + +# turbo2: 2-bit KV cache, 6.4× smaller than f16. Requires TurboQuant image. +CACHE_TYPE_K=turbo2 +CACHE_TYPE_V=turbo2 + +PARALLEL=1 + +# --no-mmap --mlock: pins entire model in RAM (prevents paging, avoids cold reads) +# --flash-attn on: +2-3% prefill boost, required for bigctx +# --jinja: required for Ornith chat formatting +EXTRA_ARGS="--flash-attn on --no-mmap --mlock --jinja" diff --git a/envs/.env.qwen3-4b b/envs/.env.qwen3-4b index d929152..527a676 100644 --- a/envs/.env.qwen3-4b +++ b/envs/.env.qwen3-4b @@ -1,18 +1,18 @@ # ============================================================================== # Qwen3-4B-Instruct Q4_K_M — Alibaba (May 2025) -# Architecture: Decoder-only transformer, GQA -# - 4B params, 32 layers -# - 32K native context (128K with YaRN) +# Architecture: 4B params, 32 layers, full GQA (32 KV heads) # Model size: ~2.4 GB Q4_K_M | Full GPU fit (ngl=99) -# Features: thinking mode (/think /no_think), tool calling, 119 languages, -# Apache 2.0. Strong code + reasoning. Best ecosystem (most fine-tunes). +# Features: thinking mode, tool calling, 119 languages, Apache 2.0 # -# Download: -# huggingface-cli download bartowski/Qwen3-4B-GGUF \ -# Qwen3-4B-Q4_K_M.gguf --local-dir ./models/ +# Benchmark (TurboQuant SM75, 2026-05-05): +# pp=191 t/s tg=44.3 t/s @ ctx=16K, fa=1 +# KV: 39.6 KB/token (full GQA = double SmolLM3's KV) # -# NOTE: Verify exact filename after download: -# ls models/Qwen3-4B* +# ⚠️ CRITICAL: turbo2/3/4 KV is BROKEN for Qwen3-4B (PPL catastrophic @ ctx≥8K) +# Always use q4_0 KV! See FINDINGS.md +# +# Optimization: q4_0 KV only, ctx=16K max (full-attn VRAM wall) +# parallel=1 (limited headroom with large KV) # ============================================================================== MODEL_FILE=Qwen3-4B-Q4_K_M.gguf @@ -20,23 +20,23 @@ MODEL_FILE=Qwen3-4B-Q4_K_M.gguf # All layers fit — ~2.4 GB leaves ~1.3 GB free for KV + compute N_GPU_LAYERS=99 -# Benchmarked 2026-05-05 on GTX 1650 Ti (3717 MiB): -# Max ctx=8192 (12K OOM). Full attention — all KV must fit at full ctx. -# GGUF native limit=40960, but VRAM walls at ~8K. -# Baseline: 181 pp / 41.6 tg t/s. At 8K ctx fa=1: 191 pp / 44.3 tg t/s (+6%). +# 16K practical max (full GQA eats VRAM fast, 24K OOM) CTX_SIZE=16384 +# i7-10750H: t=6 physical cores optimal THREADS=6 THREADS_BATCH=6 +# Standard batches, ubatch=batch/2 for pure-GPU (ngl=99, all 32 layers on GPU) BATCH_SIZE=512 UBATCH_SIZE=256 +# q4_0 KV ONLY — turbo2/3/4 catastrophically broken for Qwen3-4B! CACHE_TYPE_K=q4_0 CACHE_TYPE_V=q4_0 -# 1 parallel slot — limited VRAM at 8K ctx with 2.4GB model +# 1 parallel slot — limited VRAM with large KV @ 16K ctx PARALLEL=1 -# fa=1 gives +6% tg speed on full-attention Qwen3 -EXTRA_ARGS=--flash-attn on --mmap +# fa=1 gives +6% boost on full-attention Qwen3 +EXTRA_ARGS="--flash-attn on --cont-batching" diff --git a/envs/.env.qwen35-9b b/envs/.env.qwen35-9b index bfd8a38..5b45539 100644 --- a/envs/.env.qwen35-9b +++ b/envs/.env.qwen35-9b @@ -27,8 +27,9 @@ CTX_SIZE=32768 THREADS=6 THREADS_BATCH=6 +# ubatch=256 optimal for CPU-split (ngl=12/32 layers split) BATCH_SIZE=512 -UBATCH_SIZE=128 +UBATCH_SIZE=256 # turbo2: 2-bit KV cache, 6.4× smaller than f16. Requires TurboQuant image. CACHE_TYPE_K=turbo2 @@ -38,4 +39,4 @@ PARALLEL=1 # --no-mmap --mlock: pins entire model in RAM (prevents paging, avoids cold reads) # --flash-attn on: required with turbo2 KV (fa=0 + turbo2 has no speed benefit) -EXTRA_ARGS=--flash-attn on --no-mmap --mlock +EXTRA_ARGS="--flash-attn on --no-mmap --mlock" diff --git a/envs/.env.qwen36-35b b/envs/.env.qwen36-35b new file mode 100644 index 0000000..6b07c34 --- /dev/null +++ b/envs/.env.qwen36-35b @@ -0,0 +1,47 @@ +# ============================================================================== +# Qwen3.6-35B-A3B UD-IQ4_XS — Alibaba (Apr 2026), Apache 2.0 +# Architecture: 35B total / 3B active MoE (40 layers: 30 GatedDeltaNet + 10 +# full attention; 256 routed experts + 1 shared, top-8) +# Model size: 17.7 GB UD-IQ4_XS (unsloth dynamic: attention/router/embeddings +# kept at higher bits — protects JSON/code correctness vs plain IQ4) +# Strategy: MoE offload — dense backbone on GPU, experts in RAM via --cpu-moe. +# mmap (NO mlock): file > 15 GiB RAM, page cache holds hot experts, +# cold experts stream from NVMe. Reads ~2 GB/token active. +# +# ⚠ RAM: 17.7 GB file on a 15 GiB machine — STOP heavy containers first +# (OpenProject, Wiki.js, OTel) or expect cold-expert stutter from paging. +# +# NOTE: image built 2026-05-05 — Qwen3.6 arch may not load (release Apr 2026); +# MTP support (merged May 2026) definitely absent. Rebuild image to get +# the 1.5-2× MTP gain once arch confirmed working. +# ============================================================================== + +MODEL_FILE=Qwen3.6-35B-A3B-UD-IQ4_XS.gguf + +# Dense backbone (embeddings + 10 attn layers + GDN mixers + shared expert) +# ≈ 2 GB at IQ4 — fits VRAM. Routed experts → CPU via --cpu-moe. +N_GPU_LAYERS=99 + +# 262K native; start at 32K. KV tiny: only 10 full-attention layers, +# GDN layers carry a small recurrent state instead of KV. +CTX_SIZE=32768 + +# t=6 optimal for i7-10750H (6 physical cores). HT hurts (FINDINGS.md §5). +THREADS=6 +THREADS_BATCH=6 + +# Larger batches amortize per-expert reads during prefill. +BATCH_SIZE=1024 +UBATCH_SIZE=512 + +# q8_0 KV: quality-first until turbo2 is PPL-gated for this arch +# (hybrid-arch precedent: turbo2 catastrophically broke Qwen3-4B). +CACHE_TYPE_K=q8_0 +CACHE_TYPE_V=q8_0 + +PARALLEL=1 + +# --cpu-moe: routed expert tensors stay in CPU RAM/page cache +# mmap default (no --no-mmap/--mlock): mandatory, file exceeds RAM +# --jinja: Qwen3.6 chat template +EXTRA_ARGS="--flash-attn on --cpu-moe --jinja" diff --git a/envs/.env.smollm3-3b b/envs/.env.smollm3-3b index 4295827..c728bbd 100644 --- a/envs/.env.smollm3-3b +++ b/envs/.env.smollm3-3b @@ -1,18 +1,15 @@ # ============================================================================== # SmolLM3 3B-it Q4_K_M — HuggingFace (2025) -# Architecture: Decoder-only transformer, GQA + NoPE (3:1 ratio) -# - 3B params, 11.2T training tokens -# - 64K native context (128K with YaRN) +# Architecture: 3B params, GQA + NoPE (3:1 ratio), 64K native context # Model size: ~1.9 GB Q4_K_M | Full GPU fit (ngl=99) -# Features: thinking mode (/think /no_think), tool calling, 6 languages, -# Apache 2.0. AIME 2025: 36.7% in think mode. +# Features: thinking mode, tool calling, Apache 2.0, AIME 36.7% # -# Download: -# huggingface-cli download bartowski/HuggingFaceTB_SmolLM3-3B-GGUF \ -# HuggingFaceTB_SmolLM3-3B-Q4_K_M.gguf --local-dir ./models/ +# Benchmark (TurboQuant SM75, 2026-05-05): +# pp=260 t/s tg=58.3 t/s @ ctx=24K, fa=1 +# Fastest model in the stack! KV: 19.8 KB/token (GQA) # -# NOTE: Verify exact filename after download: -# ls models/SmolLM3* models/HuggingFaceTB_SmolLM3* +# Optimization: q8_0 KV for best quality, parallel=2 (fast+headroom) +# Larger batches (1024/512) for max throughput # ============================================================================== MODEL_FILE=HuggingFaceTB_SmolLM3-3B-Q4_K_M.gguf @@ -20,23 +17,23 @@ MODEL_FILE=HuggingFaceTB_SmolLM3-3B-Q4_K_M.gguf # All layers fit comfortably — ~1.9 GB leaves ~1.8 GB free for KV + compute N_GPU_LAYERS=99 -# Benchmarked 2026-05-05 on GTX 1650 Ti (3717 MiB): -# Max ctx=24576 (32K OOM). Baseline: 249 pp / 56.8 tg t/s. -# At 24K ctx with fa=1: 260 pp / 58.3 tg t/s (+2%). -# Model context limit = 65536, VRAM is the constraint here. -CTX_SIZE=24576 +# Push to 32K ctx (model native=65K, VRAM limit here) +CTX_SIZE=32768 +# i7-10750H: t=6 physical cores optimal THREADS=6 THREADS_BATCH=6 -BATCH_SIZE=512 -UBATCH_SIZE=256 +# Max batches for fastest model (58 tg t/s pure-GPU). ubatch=batch/2 for throughput. +BATCH_SIZE=1024 +UBATCH_SIZE=512 -CACHE_TYPE_K=q8_0 -CACHE_TYPE_V=q8_0 +# turbo2 KV for 6.4× compression (GQA = 19.8 KB/token, turbo2 = 3.1 KB/token) +CACHE_TYPE_K=turbo2 +CACHE_TYPE_V=turbo2 -# 2 parallel slots — less headroom at 24K ctx vs original 16K estimate +# 2 parallel slots — fast model (58 tg t/s), VRAM headroom available PARALLEL=2 -# fa=1 gives small but consistent improvement (+2 tg t/s) -EXTRA_ARGS=--flash-attn on --mmap +# fa=1 gives +2% boost +EXTRA_ARGS="--flash-attn on --cont-batching" diff --git a/llama b/llama index e4075ca..fb0d02f 100755 --- a/llama +++ b/llama @@ -10,7 +10,7 @@ # ./llama build # ./llama bench # -# Models: smollm3 | gemma4-e2b | gemma4-e4b | qwen3-4b | qwen35-9b +# Models: smollm3 | gemma4-e2b | gemma4-e4b | qwen3-4b | qwen35-9b | ornith-9b set -euo pipefail cd "$(dirname "$0")" @@ -31,6 +31,10 @@ declare -A MODEL_FILE=( [gemma4-e4b]="google_gemma-4-E4B-it-Q4_K_M.gguf" [qwen3-4b]="Qwen3-4B-Q4_K_M.gguf" [qwen35-9b]="Qwen3.5-9B.Q8_0.gguf" + [ornith-9b]="ornith-9b-Q8_0.gguf" + [gpt-oss-20b]="gpt-oss-20b-mxfp4.gguf" + [qwen36-35b]="Qwen3.6-35B-A3B-UD-IQ4_XS.gguf" + [ornith-35b]="deepreinforce-ai_Ornith-1.0-35B-IQ4_XS.gguf" ) declare -A MODEL_LABEL=( @@ -39,6 +43,10 @@ declare -A MODEL_LABEL=( [gemma4-e4b]="Gemma4-E4B (~30 t/s, ctx 24K/164K bigctx, multimodal)" [qwen3-4b]="Qwen3-4B (~39 t/s, ctx 16K/24K bigctx, thinking+tools)" [qwen35-9b]="Qwen3.5-9B (~4.4 t/s, ctx 32K, reasoning distill)" + [ornith-9b]="Ornith-9B (TBD t/s, ctx 32K, coding-focused, MIT)" + [gpt-oss-20b]="gpt-oss-20b (14.4 t/s, ctx 32K, 20.9B/3.6B-active MoE offload)" + [qwen36-35b]="Qwen3.6-35B (TBD t/s, ctx 32K, 35B/3B-active MoE, mmap>RAM)" + [ornith-35b]="Ornith-35B (TBD t/s, ctx 32K, 35B/3B-active MoE coding, MIT)" ) # Models that support bigctx diff --git a/llama-swap/Dockerfile b/llama-swap/Dockerfile new file mode 100644 index 0000000..18c4ffb --- /dev/null +++ b/llama-swap/Dockerfile @@ -0,0 +1,14 @@ +FROM python:3.12-slim + +RUN apt-get update && apt-get install -y curl && \ + curl -fsSL https://get.docker.com | sh && \ + apt-get clean && rm -rf /var/lib/apt/lists/* + +RUN pip install --no-cache-dir fastapi uvicorn docker pydantic httpx + +WORKDIR /app +COPY controller.py . + +EXPOSE 8000 + +CMD ["uvicorn", "controller:app", "--host", "0.0.0.0", "--port", "8000"] diff --git a/llama-swap/controller.py b/llama-swap/controller.py new file mode 100644 index 0000000..c19fb7f --- /dev/null +++ b/llama-swap/controller.py @@ -0,0 +1,178 @@ +#!/usr/bin/env python3 +"""llama-swap controller — hot-swap models via API""" +import os, time +from typing import Optional +from fastapi import FastAPI, HTTPException, Request +from fastapi.responses import JSONResponse, StreamingResponse +import docker, httpx +from pydantic import BaseModel + +app = FastAPI(title="llama-swap", version="1.1.0") +client = docker.from_env() + +MODELS = { + "ornith-9b": {"profile": "ornith-9b", "file": "ornith-9b-Q8_0.gguf", "name": "Ornith 9B Q8 (coding, MIT, 32K)", "env_file": "/workspace/envs/.env.ornith-9b"}, + "qwen35-9b": {"profile": "qwen35-9b", "file": "Qwen3.5-9B.Q8_0.gguf", "name": "Qwen3.5 9B Q8 (reasoning, 32K)", "env_file": "/workspace/envs/.env.qwen35-9b"}, + "qwen3-4b": {"profile": "qwen3-4b", "file": "Qwen3-4B-Q4_K_M.gguf", "name": "Qwen3 4B Q4 (thinking, 16K)", "env_file": "/workspace/envs/.env.qwen3-4b"}, + "gemma4-e2b": {"profile": "gemma4-e2b", "file": "google_gemma-4-E2B-it-Q4_K_M.gguf", "name": "Gemma 4 E2B (multimodal, 24K)", "env_file": "/workspace/envs/.env.gemma4-e2b"}, + "gemma4-e4b": {"profile": "gemma4-e4b", "file": "google_gemma-4-E4B-it-Q4_K_M.gguf", "name": "Gemma 4 E4B (multimodal, 24K, CPU-split)", "env_file": "/workspace/envs/.env.gemma4-e4b"}, + "smollm3-3b": {"profile": "smollm3-3b", "file": "HuggingFaceTB_SmolLM3-3B-Q4_K_M.gguf", "name": "SmolLM3 3B (thinking+tools, 24K)", "env_file": "/workspace/envs/.env.smollm3-3b"} +} +CONTAINER_NAME, LLAMA_API = "llama_server", "http://llama_server:8080" +IMAGE = "local/llama-cpp-turboquant:server-cuda-sm75-mmq" +NETWORK = "llama-cpp_llama-net" + +def load_env_file(path: str) -> dict: + """Parse .env file into dict""" + env = {} + try: + with open(path) as f: + for line in f: + line = line.strip() + if line and not line.startswith('#') and '=' in line: + key, val = line.split('=', 1) + env[key.strip()] = val.strip() + except: pass + return env + +def get_active_model() -> Optional[str]: + try: + c = client.containers.get(CONTAINER_NAME) + if c.status != "running": return None + model = c.labels.get("llama-swap.model") + return model if model in MODELS else None + except: pass + return None + +def stop_active(): + try: + c = client.containers.get(CONTAINER_NAME) + c.stop(timeout=10); c.remove() + except: pass + +def start_model(model_id: str) -> dict: + if model_id not in MODELS: raise ValueError(f"Unknown model: {model_id}") + stop_active() + + meta = MODELS[model_id] + env = load_env_file(meta["env_file"]) + + # Build command args from env vars + cmd_args = [ + "/app/llama-server", + "--model", f"/models/{env.get('MODEL_FILE', meta['file'])}", + "--host", "0.0.0.0", + "--port", "8080", + "--ctx-size", env.get('CTX_SIZE', '32768'), + "--n-gpu-layers", env.get('N_GPU_LAYERS', '99'), + "--threads", env.get('THREADS', '6'), + "--threads-batch", env.get('THREADS_BATCH', '6'), + "--batch-size", env.get('BATCH_SIZE', '512'), + "--ubatch-size", env.get('UBATCH_SIZE', '128'), + "--cache-type-k", env.get('CACHE_TYPE_K', 'f16'), + "--cache-type-v", env.get('CACHE_TYPE_V', 'f16'), + "--parallel", env.get('PARALLEL', '1') + ] + + # Add EXTRA_ARGS if present (strip quotes if present) + extra = env.get('EXTRA_ARGS', '').strip().strip('"').strip("'") + if extra: + cmd_args.extend(extra.split()) + + # Create container with GPU device request + try: + container = client.containers.create( + image=IMAGE, + name=CONTAINER_NAME, + entrypoint=[], # Override image entrypoint + command=cmd_args, + detach=True, + runtime="nvidia", + device_requests=[ + docker.types.DeviceRequest( + count=-1, # all GPUs + capabilities=[['gpu', 'compute', 'utility']] + ) + ], + environment={ + "NVIDIA_VISIBLE_DEVICES": "all", + "NVIDIA_DRIVER_CAPABILITIES": "compute,utility", + **env + }, + volumes={ + "/home/moze/Sources/llama-cpp/models": {"bind": "/models", "mode": "ro"} + }, + ports={"8080/tcp": 8080}, + network=NETWORK, + shm_size="1g", + ulimits=[docker.types.Ulimit(name='memlock', soft=-1, hard=-1)], + restart_policy={"Name": "unless-stopped"}, + labels={ + "com.docker.compose.project": "llama-cpp", + "com.docker.compose.service": f"llama-{model_id}", + "llama-swap.model": model_id + }, + healthcheck={ + "test": ["CMD-SHELL", "curl -sf http://localhost:8080/health | grep -q '\"status\":\"ok\"'"], + "interval": 20000000000, # 20s in nanoseconds + "timeout": 10000000000, + "retries": 10, + "start_period": 180000000000 # 180s for 9B models with mlock + } + ) + container.start() + + # Wait for health check + for _ in range(90): + try: + container.reload() + health = container.attrs.get("State", {}).get("Health", {}).get("Status") + if health == "healthy": + return {"model": model_id, "status": "healthy", "container": CONTAINER_NAME} + except: pass + time.sleep(2) + + return {"model": model_id, "status": "starting", "container": CONTAINER_NAME} + except Exception as e: + raise RuntimeError(f"Failed to start {model_id}: {str(e)}") + +class SwitchRequest(BaseModel): + model: str + +@app.get("/") +async def root(): return {"service": "llama-swap", "version": "1.1.0"} + +@app.get("/models") +async def list_models(): + active = get_active_model() + return {"models": [{"id": mid, "name": m["name"], "file": m["file"], "active": mid==active} for mid, m in MODELS.items()], "active": active} + +@app.post("/models/switch") +async def switch_model(req: SwitchRequest): + if req.model not in MODELS: raise HTTPException(404, f"Unknown: {req.model}") + if get_active_model() == req.model: return {"model": req.model, "status": "already_active"} + try: return start_model(req.model) + except Exception as e: raise HTTPException(500, str(e)) + +@app.get("/status") +async def status(): + active = get_active_model() + if not active: return {"status": "no_model", "model": None} + try: + c = client.containers.get(CONTAINER_NAME) + return {"status": "running", "model": active, "health": c.attrs.get("State",{}).get("Health",{}).get("Status","unknown")} + except: return {"status": "no_model", "model": None} + +@app.api_route("/v1/{path:path}", methods=["GET","POST","PUT","DELETE","PATCH"]) +async def proxy_llama(path: str, request: Request): + if not get_active_model(): raise HTTPException(503, "No model active") + url = f"{LLAMA_API}/v1/{path}" + headers = dict(request.headers); headers.pop("host", None) + async with httpx.AsyncClient(timeout=300.0) as http: + if request.method == "GET": + resp = await http.get(url, headers=headers, params=request.query_params) + else: + resp = await http.request(request.method, url, headers=headers, params=request.query_params, content=await request.body()) + if "text/event-stream" in resp.headers.get("content-type",""): + return StreamingResponse(resp.aiter_bytes(), media_type=resp.headers.get("content-type"), headers=dict(resp.headers)) + return JSONResponse(content=resp.json() if "json" in resp.headers.get("content-type","") else resp.text, status_code=resp.status_code, headers=dict(resp.headers)) diff --git a/scripts/expert_heatmap.py b/scripts/expert_heatmap.py new file mode 100644 index 0000000..4df35a7 --- /dev/null +++ b/scripts/expert_heatmap.py @@ -0,0 +1,113 @@ +#!/usr/bin/env python3 +"""Hot-expert mapper: page-cache residency per expert slice of a GGUF. + +Expert tensors (blk.N.ffn_*_exps.weight) are fused 3D with the expert index +as the slowest dim -> each expert's weights are one contiguous byte range. +mincore() over each range after a real workload = which experts survived in +page cache (LRU proxy for routing heat). ds4/colibri hot-store idea, applied +externally to an unmodified llama.cpp. + +Usage: + expert_heatmap.py # snapshot + per-layer summary + expert_heatmap.py --json out.json # full per-expert dump +""" +import ctypes, ctypes.util, json, mmap, os, struct, sys + +libc = ctypes.CDLL(ctypes.util.find_library("c"), use_errno=True) + +def read_gguf(path): + f = open(path, "rb") + assert f.read(4) == b"GGUF" + ver, = struct.unpack(" {expert -> [frac,...] over up/gate/down} + for name, dims, ttype, off, size in tensors: + if "_exps.weight" not in name or n_expert == 0: + continue + layer = int(name.split(".")[1]) + stride = size // n_expert + for e in range(n_expert): + frac = resident_fraction(data_start + off + e * stride, stride) + layers.setdefault(layer, {}).setdefault(e, []).append(frac) + + print(f"# {os.path.basename(path)} arch={arch} experts/layer={n_expert} " + f"file={fsize/1e9:.1f}GB") + print(f"{'layer':>5} {'res%':>6} {'hot(>90%)':>9} {'cold(<10%)':>10} top5 experts") + summary = {} + for layer in sorted(layers): + em = {e: sum(v) / len(v) for e, v in layers[layer].items()} + avg = sum(em.values()) / len(em) + hot = sum(1 for v in em.values() if v > 0.9) + cold = sum(1 for v in em.values() if v < 0.1) + top = sorted(em, key=em.get, reverse=True)[:5] + summary[layer] = {"avg": avg, "hot": hot, "cold": cold, "experts": em} + print(f"{layer:>5} {avg*100:>5.1f}% {hot:>9} {cold:>10} {top}") + tot = [s["avg"] for s in summary.values()] + print(f"# overall expert residency: {sum(tot)/len(tot)*100:.1f}% " + f"(hottest layers: {sorted(summary, key=lambda l: summary[l]['avg'], reverse=True)[:6]})") + if out_json: + json.dump(summary, open(out_json, "w")) + print(f"# wrote {out_json}") + +if __name__ == "__main__": + main() diff --git a/scripts/moe_bench.py b/scripts/moe_bench.py new file mode 100644 index 0000000..f668191 --- /dev/null +++ b/scripts/moe_bench.py @@ -0,0 +1,41 @@ +#!/usr/bin/env python3 +"""Quick MoE bench via swap-stack. Usage: moe_bench.py [n_runs] +Measures tg/pp from server .timings (never wall clock), plus a temp-0 +structure gate (JSON validity + bracket balance). x570 method.""" +import json, sys, time, urllib.request + +BASE = "http://localhost:48080/v1/chat/completions" +MODEL = sys.argv[1] if len(sys.argv) > 1 else "gpt-oss-20b" +RUNS = int(sys.argv[2]) if len(sys.argv) > 2 else 2 + +def ask(prompt, mt=350): + req = urllib.request.Request(BASE, + data=json.dumps({"model": MODEL, "messages": [{"role": "user", "content": prompt}], + "max_tokens": mt, "temperature": 0}).encode(), + headers={"Content-Type": "application/json"}) + r = json.load(urllib.request.urlopen(req, timeout=900)) + t = r.get("timings", {}) + return (r["choices"][0]["message"].get("content") or ""), t + +# warmup / load +_, t0 = ask("Say OK.", mt=400) +print(f"warmup: tg={t0.get('predicted_per_second',0):.2f}") + +# structure gate (temp 0). Generous budget: thinking models burn tokens on +# reasoning_content before emitting content (ornith needs ~250+). +c, _ = ask('Return JSON: {"name":"x","primes":[first 8 primes],"nested":{"a":true}}. JSON only.', mt=1500) +try: + json.loads(c); gate_json = "PASS" +except Exception: + gate_json = "FAIL: " + c[:120] +code, _ = ask("Write a Python function parsing nested brackets ()[]{} into a tree. Code only.", mt=2000) +gate_code = "PASS" if all(code.count(a) == code.count(b) for a, b in [("(",")"),("[","]"),("{","}")]) else "FAIL" +print(f"gate: json={gate_json} brackets={gate_code}") + +# throughput runs +tgs, pps = [], [] +for i in range(RUNS): + _, t = ask("Write a detailed essay about the history of computing.", mt=300) + tgs.append(t.get("predicted_per_second", 0)); pps.append(t.get("prompt_per_second", 0)) + print(f"run{i+1}: tg={tgs[-1]:.2f} pp={pps[-1]:.2f}") +print(f"RESULT {MODEL}: tg_avg={sum(tgs)/len(tgs):.2f} pp_avg={sum(pps)/len(pps):.2f}") diff --git a/scripts/quick_bench.sh b/scripts/quick_bench.sh new file mode 100755 index 0000000..41f1259 --- /dev/null +++ b/scripts/quick_bench.sh @@ -0,0 +1,139 @@ +#!/bin/bash +# Quick benchmark of all models using llama-swap +# Tests current optimized configs from envs/*.env + +set -euo pipefail + +cd "$(dirname "$0")/.." + +SWAP_URL="http://localhost:8089" +RESULTS_DIR="benchmark-results" +TIMESTAMP=$(date +%Y%m%d_%H%M%S) +RESULTS_FILE="$RESULTS_DIR/optimized_${TIMESTAMP}.csv" + +mkdir -p "$RESULTS_DIR" + +# Check llama-swap is running +if ! curl -sf "$SWAP_URL/status" >/dev/null 2>&1; then + echo "ERROR: llama-swap not running at $SWAP_URL" + echo "Start with: docker compose --profile swap up -d" + exit 1 +fi + +echo "======================================================================" +echo "OPTIMIZED CONFIG BENCHMARK — $(date)" +echo "======================================================================" +echo "" +echo "Results: $RESULTS_FILE" +echo "" + +# CSV header +cat > "$RESULTS_FILE" <> "$RESULTS_FILE" + return 1 + fi + echo "OK" + + # Wait for model to load (controller waits internally, but give buffer) + sleep 10 + + # Check via proxy that llama_server is up + if ! curl -sf "$SWAP_URL/v1/models" >/dev/null 2>&1; then + echo " Model failed to start" + echo "$model_name,,,,,,,,,,,,,START_FAILED" >> "$RESULTS_FILE" + return 1 + fi + + # Get model metadata from active container + local active_model + active_model=$(curl -sf "$SWAP_URL/status" | grep -oP '"active_model":\s*"\K[^"]+' || echo "unknown") + + # Read env file to get params + local env_file="envs/.env.$model_name" + if [ ! -f "$env_file" ]; then + echo " WARNING: $env_file not found" + return 1 + fi + + source "$env_file" + + echo " Config: ngl=$N_GPU_LAYERS ctx=$CTX_SIZE t=$THREADS batch=$BATCH_SIZE/$UBATCH_SIZE kv=$CACHE_TYPE_K parallel=$PARALLEL" + + # Prefill test: 512 tokens + echo -n " Prefill (512t)... " + local pp_start pp_end pp_time pp_tps + pp_start=$(date +%s.%N) + + local pp_response + pp_response=$(curl -sf -X POST "$SWAP_URL/v1/completions" \ + -H "Content-Type: application/json" \ + -d '{ + "prompt": "'"$(python3 -c "print('The quick brown fox '*64)")"'", + "max_tokens": 1, + "temperature": 0.0 + }' 2>&1) || { + echo "FAILED" + echo "$model_name,$N_GPU_LAYERS,$CTX_SIZE,$THREADS,$BATCH_SIZE,$UBATCH_SIZE,$CACHE_TYPE_K,$PARALLEL,512,1,,,,,PP_FAILED" >> "$RESULTS_FILE" + return 1 + } + + pp_end=$(date +%s.%N) + pp_time=$(echo "$pp_end - $pp_start" | bc -l) + pp_tps=$(echo "scale=1; 512 / $pp_time" | bc -l) + echo "${pp_tps} t/s" + + # Generation test: 128 tokens + echo -n " Generation (128t)... " + local tg_start tg_end tg_time tg_tps + tg_start=$(date +%s.%N) + + local tg_response + tg_response=$(curl -sf -X POST "$SWAP_URL/v1/completions" \ + -H "Content-Type: application/json" \ + -d '{ + "prompt": "Count from 1 to 100:", + "max_tokens": 128, + "temperature": 0.0 + }' 2>&1) || { + echo "FAILED" + echo "$model_name,$N_GPU_LAYERS,$CTX_SIZE,$THREADS,$BATCH_SIZE,$UBATCH_SIZE,$CACHE_TYPE_K,$PARALLEL,512,128,$pp_time,,$pp_tps,,TG_FAILED" >> "$RESULTS_FILE" + return 1 + } + + tg_end=$(date +%s.%N) + tg_time=$(echo "$tg_end - $tg_start" | bc -l) + tg_tps=$(echo "scale=1; 128 / $tg_time" | bc -l) + echo "${tg_tps} t/s" + + # Write results + echo "$model_name,$N_GPU_LAYERS,$CTX_SIZE,$THREADS,$BATCH_SIZE,$UBATCH_SIZE,$CACHE_TYPE_K,$PARALLEL,512,128,$pp_time,$tg_time,$pp_tps,$tg_tps,OK" >> "$RESULTS_FILE" + + echo "" +} + +# Benchmark all models +for model in ornith-9b qwen35-9b qwen3-4b smollm3-3b gemma4-e2b gemma4-e4b; do + bench_model "$model" || echo " Skipping $model" +done + +echo "======================================================================" +echo "SUMMARY" +echo "======================================================================" +echo "" +column -t -s',' "$RESULTS_FILE" +echo "" +echo "Full results: $RESULTS_FILE" diff --git a/swap-stack/Dockerfile b/swap-stack/Dockerfile new file mode 100644 index 0000000..baa0e4d --- /dev/null +++ b/swap-stack/Dockerfile @@ -0,0 +1,19 @@ +# Tri-binary swap-stack: llama-swap + three llama-server builds. +# /app/llama-server TurboQuant fork (May 2026) — turbo2/3/4 KV, needed by 9B configs +# /app-upstream/llama-server upstream master (Jul 2026) — newest MoE/arch work +# /app/llama-server-ik ik_llama.cpp — -ser / -fmoe / -rtr, fast IQ-quant CPU kernels +# config.yaml picks the binary per model (env: LD_LIBRARY_PATH for upstream). +FROM local/llama-cpp-upstream:server-cuda-sm75-mmq AS upstream +FROM local/ik-llama:server-cuda-sm75 AS ik + +FROM local/llama-cpp-turboquant:server-cuda-sm75-mmq + +COPY --from=upstream /app /app-upstream +COPY --from=ik /llama-server /app-ik/llama-server +COPY --from=ik /usr/local/lib/libllama.so /usr/local/lib/libggml.so /usr/local/lib/libmtmd.so /app-ik/ +COPY --from=ik /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcudart.so.12 /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcublas.so.12 /usr/local/cuda-12.4/targets/x86_64-linux/lib/libcublasLt.so.12 /app-ik/ +COPY llama-swap /app/llama-swap + +# config.yaml is bind-mounted at runtime (see compose.yaml) so edits +# only need a container restart, not a rebuild. +ENTRYPOINT ["/app/llama-swap", "-config", "/app/config.yaml", "-listen", ":8080"] diff --git a/swap-stack/LICENSE.md b/swap-stack/LICENSE.md new file mode 100644 index 0000000..6dbacec --- /dev/null +++ b/swap-stack/LICENSE.md @@ -0,0 +1,9 @@ +MIT License + +Copyright (c) 2024 Benson Wong + +Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. \ No newline at end of file diff --git a/swap-stack/README.md b/swap-stack/README.md new file mode 100644 index 0000000..84910fc --- /dev/null +++ b/swap-stack/README.md @@ -0,0 +1,296 @@ +![llama-swap header image](docs/assets/hero4.webp) +![GitHub Downloads (all assets, all releases)](https://img.shields.io/github/downloads/mostlygeek/llama-swap/total) +![GitHub Actions Workflow Status](https://img.shields.io/github/actions/workflow/status/mostlygeek/llama-swap/go-ci.yml) +![GitHub Repo stars](https://img.shields.io/github/stars/mostlygeek/llama-swap) + +# llama-swap + +Run multiple generative AI models on your machine and hot-swap between them on demand. llama-swap works with any OpenAI and Anthropic API compatible server and is used by thousands of people to power their local AI workflows. + +Built in Go for performance and simplicity, llama-swap has zero dependencies and is incredibly easy to set up. Get started in minutes - just one binary and one configuration file. + +## Features: + +- ✅ Easy to deploy and configure: one binary, one configuration file. no external dependencies +- ✅ On-demand model switching +- ✅ Use any local OpenAI compatible server (llama.cpp, vllm, tabbyAPI, stable-diffusion.cpp, etc.) + - future proof, upgrade your inference servers at any time. +- ✅ OpenAI API supported endpoints: + - `v1/completions` + - `v1/chat/completions` + - `v1/responses` + - `v1/embeddings` + - `v1/models` - list available models + - `v1/audio/speech` ([#36](https://github.com/mostlygeek/llama-swap/issues/36)) + - `v1/audio/transcriptions` ([docs](https://github.com/mostlygeek/llama-swap/issues/41#issuecomment-2722637867)) + - `v1/audio/voices` + - `v1/images/generations` + - `v1/images/edits` +- ✅ Anthropic API supported endpoints: + - `v1/messages` + - `v1/messages/count_tokens` +- ✅ llama-server (llama.cpp) supported endpoints + - `v1/rerank`, `v1/reranking`, `/rerank` + - `/infill` - for code infilling + - `/completion` - for completion endpoint + - `/props` - requires `?model={model_id}` query parameter to be provided. The autoload parameter is not supported and will be ignored. +- ✅ SDAPI via [stable-diffusion.cpp's server](https://github.com/leejet/stable-diffusion.cpp/tree/master/examples/server) + - `/sdapi/v1/txt2img` + - `/sdapi/v1/img2img` + - `/sdapi/v1/loras` - requires `model` in request body to fetch the correct loras +- ✅ llama-swap API + - `/ui` - web UI + - `/upstream/:model_id` - direct access to upstream server ([demo](https://github.com/mostlygeek/llama-swap/pull/31)) + - `/running` - list currently running models ([#61](https://github.com/mostlygeek/llama-swap/issues/61)) + - `POST /api/models/unload` - manually unload all running models ([#58](https://github.com/mostlygeek/llama-swap/issues/58)) + - `POST /api/models/unload/:model_id` - unload a specific model + - `/logs` - remote log monitoring + - `GET /logs` returns buffered plain text logs. + - If `Accept: text/html` is sent, `/logs` redirects to `/ui/`. + - `GET /logs/stream` keeps the connection open for live log streaming. + - Stream endpoints send buffered history first by default; add `?no-history` to stream only new lines. + - `GET /logs/stream/proxy` streams proxy logs only. + - `GET /logs/stream/upstream` streams upstream process logs only. + - `GET /logs/stream/{model_id}` streams logs for one model (including IDs with slashes, like `author/model`). + - `/health` - just returns "OK" + - `/metrics` - system and GPU metrics for prometheus +- ✅ API Key support - define keys to restrict access to API endpoints +- ✅ Customizable + - Run concurrent models with a custom DSL swap matrix ([#643](https://github.com/mostlygeek/llama-swap/issues/643)) + - Automatic unloading of models after timeout by setting a `ttl` + - Docker and Podman support using `cmd` and `cmdStop` together + - Preload models on startup with `hooks` ([#235](https://github.com/mostlygeek/llama-swap/pull/235)) + - Apply filters to requests to control inference with `stripParams`, `setParams` and `setParamsByID` + +### Web UI + +llama-swap includes a real time web interface with a playground for testing out all sorts of local models: + +image + +View detailed token metrics: + +image + +Inspect request and responses: + +image + +Manually load and unload models: + +image + +Real time log streaming: + +image + +## Installation + +llama-swap can be installed in multiple ways + +1. Docker +2. Homebrew (macOS and Linux) +3. MacPorts (macOS) +4. WinGet +5. From release binaries +6. From source + +### Docker Install ([download images](https://github.com/mostlygeek/llama-swap/pkgs/container/llama-swap)) + +Two types of container images are built nightly for llama-swap: + +1. A unified container with llama-server, ik-llama-server, stable-diffusion.cpp, whisper.cpp and llama-swap built from source. This is only available for cuda and vulkan but has more capabilities. This one is recommended for use. +2. A legacy image that is based on llama.cpp's images and llama-swap copied into the container. Use this one if you prefer to stay close to llama.cpp's container images. + +#### Unified container (Recommended) + +```shell +$ docker pull ghcr.io/mostlygeek/llama-swap:unified-cuda + +# run with a custom configuration and models directory +$ docker run -it --rm --runtime nvidia -p 9292:8080 \ + -v /path/to/models:/models \ + -v /path/to/custom/config.yaml:/etc/llama-swap/config/config.yaml \ + ghcr.io/mostlygeek/llama-swap:unified-cuda +``` + +#### Legacy container + +```shell +$ docker pull ghcr.io/mostlygeek/llama-swap:cuda + +# run with a custom configuration and models directory +$ docker run -it --rm --runtime nvidia -p 9292:8080 \ + -v /path/to/models:/models \ + -v /path/to/custom/config.yaml:/app/config.yaml \ + ghcr.io/mostlygeek/llama-swap:cuda +``` + +
+ +more examples + + +```shell +# pull latest images per platform +docker pull ghcr.io/mostlygeek/llama-swap:cpu +docker pull ghcr.io/mostlygeek/llama-swap:cuda +docker pull ghcr.io/mostlygeek/llama-swap:vulkan +docker pull ghcr.io/mostlygeek/llama-swap:intel +docker pull ghcr.io/mostlygeek/llama-swap:musa + +# tagged llama-swap, platform and llama-server version images +docker pull ghcr.io/mostlygeek/llama-swap:v166-cuda-b6795 + +# non-root cuda +docker pull ghcr.io/mostlygeek/llama-swap:cuda-non-root + +``` + +
+ +### Homebrew Install (macOS/Linux) + +```shell +brew tap mostlygeek/llama-swap +brew install llama-swap +llama-swap --config path/to/config.yaml --listen localhost:8080 +``` + +### MacPorts (macOS) + +> [!NOTE] +> Maintained by MacPorts community - [llama-swap port](https://ports.macports.org/port/llama-swap). It is not an official part of llama-swap. + +```shell +sudo port install llama-swap +llama-swap --config path/to/config.yaml --listen localhost:8080 +``` + +### WinGet Install (Windows) + +> [!NOTE] +> WinGet is maintained by community contributor [Dvd-Znf](https://github.com/Dvd-Znf) ([#327](https://github.com/mostlygeek/llama-swap/issues/327)). It is not an official part of llama-swap. + +```shell +# install +C:\> winget install llama-swap + +# upgrade +C:\> winget upgrade llama-swap +``` + +### Pre-built Binaries + +Binaries are available on the [release](https://github.com/mostlygeek/llama-swap/releases) page for Linux, Mac, Windows and FreeBSD. + +### Building from source + +1. Building requires Go and Node.js (for UI). +1. `git clone https://github.com/mostlygeek/llama-swap.git` +1. `make clean all` +1. look in the `build/` subdirectory for the llama-swap binary + +## Configuration + +```yaml +# minimum viable config.yaml + +models: + model1: + cmd: llama-server --port ${PORT} --model /path/to/model.gguf +``` + +That's all you need to get started: + +1. `models` - holds all model configurations +2. `model1` - the ID used in API calls +3. `cmd` - the command to run to start the server. +4. `${PORT}` - an automatically assigned port number + +Almost all configuration settings are optional and can be added one step at a time: + +- Advanced features + - `matrix` to run concurrent models with a custom swap logic DSL + - `hooks` to run things on startup + - `macros` reusable snippets +- Model customization + - `ttl` to automatically unload models + - `aliases` to use familiar model names (e.g., "gpt-4o-mini") + - `env` to pass custom environment variables to inference servers + - `cmdStop` gracefully stop Docker/Podman containers + - `useModelName` to override model names sent to upstream servers + - `${PORT}` automatic port variables for dynamic port assignment + - `filters` rewrite parts of requests before sending to the upstream server + +See the [configuration documentation](docs/configuration.md) for all options. + +## How does llama-swap work? + +When a request is made to an OpenAI compatible endpoint, llama-swap will extract the `model` value and load the appropriate server configuration to serve it. If the wrong upstream server is running, it will be replaced with the correct one. This is where the "swap" part comes in. The upstream server is automatically swapped to handle the request correctly. + +In the most basic configuration llama-swap handles one model at a time. For more advanced use cases, using a `matrix` allows multiple models to be loaded at the same time. You have complete control over how your system resources are used. + +## Reverse Proxy Configuration (nginx) + +If you deploy llama-swap behind nginx, disable response buffering for streaming endpoints. By default, nginx buffers responses which breaks Server‑Sent Events (SSE) and streaming chat completion. ([#236](https://github.com/mostlygeek/llama-swap/issues/236)) + +Recommended nginx configuration snippets: + +```nginx +# SSE for UI events/logs +location /api/events { + proxy_pass http://your-llama-swap-backend; + proxy_buffering off; + proxy_cache off; +} + +# Streaming chat completions (stream=true) +location /v1/chat/completions { + proxy_pass http://your-llama-swap-backend; + proxy_buffering off; + proxy_cache off; +} +``` + +As a safeguard, llama-swap also sets `X-Accel-Buffering: no` on SSE responses. However, explicitly disabling `proxy_buffering` at your reverse proxy is still recommended for reliable streaming behavior. + +## Monitoring Logs on the CLI + +```sh +# sends up to the last 10KB of logs +$ curl http://host/logs + +# streams combined logs +curl -Ns http://host/logs/stream + +# stream llama-swap's proxy status logs +curl -Ns http://host/logs/stream/proxy + +# stream logs from upstream processes that llama-swap loads +curl -Ns http://host/logs/stream/upstream + +# stream logs only from a specific model +curl -Ns http://host/logs/stream/{model_id} + +# stream and filter logs with linux pipes +curl -Ns http://host/logs/stream | grep 'eval time' + +# appending ?no-history will disable sending buffered history first +curl -Ns 'http://host/logs/stream?no-history' +``` + +## Do I need to use llama.cpp's server (llama-server)? + +Any OpenAI compatible server would work. llama-swap was originally designed for llama-server and it is the best supported. + +For Python based inference servers like vllm or tabbyAPI it is recommended to run them via podman or docker. This provides clean environment isolation as well as responding correctly to `SIGTERM` signals for proper shutdown. + +## Star History + +> [!NOTE] +> Thank you to everyone who has given this project a ⭐️! + +## Star History + +[![Star History Chart](https://api.star-history.com/chart?repos=mostlygeek/llama-swap&type=date&legend=top-left&sealed_token=11p0VYUat56nvRhCW_s8Zn9FRLBzIWoWDuQ10v9_n1tBppiwL7qFSVf4CE8dy6rhMQ466jt1buCCzUvitO7prlGn4JbHr1B6JCJKE2B4n8ffDVBnwJ78Tg)](https://www.star-history.com/?repos=mostlygeek%2Fllama-swap&type=date&legend=top-left) diff --git a/swap-stack/config.yaml b/swap-stack/config.yaml new file mode 100644 index 0000000..cbe404a --- /dev/null +++ b/swap-stack/config.yaml @@ -0,0 +1,311 @@ +# ============================================================================== +# llama-swap config — xps9700 (GTX 1650 Ti 3.7GB VRAM, i7-10750H 6c/12t, 15 GiB) +# +# x570-style single-endpoint stack: llama-swap owns :8080, spawns/kills +# llama-server per requested `model` field, idle-unloads after TTL. +# All per-model tuning migrated verbatim from envs/.env.* (benchmarks +# 2026-05-05/06 + MoE work 2026-07-10 — see docs/FINDINGS.md). +# +# Edit this file → `docker restart llama_swap_stack` (config is bind-mounted). +# ============================================================================== + +healthCheckTimeout: 600 # 35B MoE mmap first-touch can take minutes +logLevel: info +startPort: 5800 + +macros: + # t=6 physical cores only — HT hurts (FINDINGS.md §5) + "server-base": > + /app/llama-server + --host 127.0.0.1 --port ${PORT} + --threads 6 --threads-batch 6 + --flash-attn on + + # upstream master build (newer MoE perf work); needs its own libs + "server-upstream": > + /app-upstream/llama-server + --host 127.0.0.1 --port ${PORT} + --threads 6 --threads-batch 6 + --flash-attn on + + # ik_llama.cpp: -ser/-fmoe/-rtr, fast IQ CPU kernels (flag names differ: -fa) + # needs env: LD_LIBRARY_PATH=/app-ik (own libllama/libggml, symbol-incompatible with turboquant's) + "server-ik": > + /app-ik/llama-server + --host 127.0.0.1 --port ${PORT} + --threads 6 --threads-batch 6 -fa on + + "q8-kv": "--cache-type-k q8_0 --cache-type-v q8_0" + "q4-kv": "--cache-type-k q4_0 --cache-type-v q4_0" + "turbo2-kv": "--cache-type-k turbo2 --cache-type-v turbo2" + + # MoE offload: dense backbone on GPU, routed experts in RAM/page cache. + # mmap (no mlock) mandatory for files > RAM. + "moe-offload": "--n-gpu-layers 99 --cpu-moe --jinja" + +groups: + # Resident duo for pi: main coding model + fast subagent stay loaded together. + # VRAM split (2026-07-10): qwen3-4b gets the WHOLE GPU (ngl 99) — ornith's + # experts never touched VRAM anyway and its dense-on-GPU split starved qwen + # to 176 MiB / 10 t/s. Ornith runs fully CPU (page-cache resident, 3B active). + # Threads: main 6 / sub 3 — contention only during overlap. + "duo": + swap: false + exclusive: true + members: ["ornith-35b-duo", "qwen3-4b-duo"] + +models: + + # ── pi resident duo ───────────────────────────────────────────────────────── + + "ornith-35b-duo": + name: "Ornith 35B (duo main)" + description: "Duo variant: fully CPU (GPU reserved for qwen3-4b-duo). CUDA hidden — even ngl 0 tries a ~1GB pp compute buffer" + env: + - "CUDA_VISIBLE_DEVICES=" + cmd: | + ${server-base} ${q8-kv} + --model /models/deepreinforce-ai_Ornith-1.0-35B-Q2_K_L.gguf + --n-gpu-layers 0 --jinja + --ctx-size 32768 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 1 + ttl: 0 + + "qwen3-4b-duo": + name: "Qwen3 4B (duo subagent)" + description: "Duo subagent: full GPU (ngl 99), 3 threads, gate-clean JSON" + cmd: | + /app/llama-server + --host 127.0.0.1 --port ${PORT} + --threads 3 --threads-batch 3 + --flash-attn on ${q4-kv} + --model /models/Qwen3-4B-Q4_K_M.gguf + --n-gpu-layers 99 + --ctx-size 16384 + --batch-size 512 --ubatch-size 256 + --cont-batching --parallel 1 + ttl: 0 + + # ── MoE offload models (2026-07-10) ──────────────────────────────────────── + + "gpt-oss-20b": + name: "gpt-oss-20b MXFP4" + description: "20.9B/3.6B-active MoE, native MXFP4. 16.7 tg / 29.5 pp (n-cpu-moe 21: last 3 expert layers in VRAM, +16%)" + cmd: | + ${server-base} ${q8-kv} + --model /models/gpt-oss-20b-mxfp4.gguf + --n-gpu-layers 99 --n-cpu-moe 21 --jinja + --ctx-size 32768 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 1 + ttl: 300 + + "qwen36-35b": + name: "Qwen3.6-35B-A3B UD-IQ4_XS" + description: "35B/3B-active MoE, 17.7GB mmap > RAM. Stop heavy containers first" + cmd: | + ${server-base} ${moe-offload} ${q8-kv} + --model /models/Qwen3.6-35B-A3B-UD-IQ4_XS.gguf + --ctx-size 32768 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 1 + ttl: 300 + + "qwen36-35b-q2": + name: "Qwen3.6-35B-A3B UD-Q2_K_XL" + description: "DAILY DRIVER 35B: 23.0 tg / 36 pp. ds4 asymmetric recipe (dense high-bit, experts 2-bit), 12.3GB page-cache resident, last 3 expert layers in VRAM. Gates pass" + cmd: | + ${server-base} ${q8-kv} + --model /models/Qwen3.6-35B-A3B-UD-Q2_K_XL.gguf + --n-gpu-layers 99 --n-cpu-moe 37 --jinja + --ctx-size 32768 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 1 + ttl: 300 + + "ornith-35b": + name: "Ornith-1.0-35B Q2_K_L" + description: "SPEED KING: 29.1 tg / 66 pp. RL coding finetune (Qwen3.5-MoE arch), 13.1GB page-cache resident, 3 expert layers in VRAM, thinking model. MIT. Gates pass" + cmd: | + ${server-base} ${q8-kv} + --model /models/deepreinforce-ai_Ornith-1.0-35B-Q2_K_L.gguf + --n-gpu-layers 99 --n-cpu-moe 37 --jinja + --ctx-size 32768 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 1 + ttl: 300 + + "ornith-35b-iq4": + name: "Ornith-1.0-35B IQ4_XS" + description: "Quality-first variant, 18.8GB mmap > RAM = ~2.8 t/s thrash. Batch jobs only (or post-RAM-upgrade)" + cmd: | + ${server-base} ${moe-offload} ${q8-kv} + --model /models/deepreinforce-ai_Ornith-1.0-35B-IQ4_XS.gguf + --ctx-size 32768 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 1 + ttl: 300 + + # ── ik_llama.cpp experimental variants ────────────────────────────────────── + + "qwen36-35b-q2-ik": + name: "Qwen3.6-35B Q2_K_XL (ik_llama)" + description: "ik build: 18.0 tg / 42 pp — loses decode to main build (23.0), wins prefill. -ser 6,1 active. Experimental only" + env: + - "LD_LIBRARY_PATH=/app-ik" + cmd: | + ${server-ik} ${q8-kv} + --model /models/Qwen3.6-35B-A3B-UD-Q2_K_XL.gguf + --n-gpu-layers 99 --n-cpu-moe 37 --jinja -ser 6,1 + --ctx-size 16384 + --batch-size 1024 --ubatch-size 512 + --parallel 1 + ttl: 300 + + "gpt-oss-20b-upstream": + name: "gpt-oss-20b (upstream master)" + description: "Upstream Jul-2026 build comparison" + env: + - "LD_LIBRARY_PATH=/app-upstream" + cmd: | + ${server-upstream} ${q8-kv} + --model /models/gpt-oss-20b-mxfp4.gguf + --n-gpu-layers 99 --n-cpu-moe 21 --jinja + --ctx-size 32768 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 1 + ttl: 300 + + # ── Dense 9B (RAM-bandwidth-bound, ~4.4 t/s) ─────────────────────────────── + + "ornith-9b": + name: "Ornith-1.0-9B Q8_0" + description: "Coding 9B, MIT. ~4.4 t/s, mlock-pinned" + cmd: | + ${server-base} ${turbo2-kv} + --model /models/ornith-9b-Q8_0.gguf + --n-gpu-layers 11 + --ctx-size 32768 + --batch-size 512 --ubatch-size 256 + --cont-batching --parallel 1 + --no-mmap --mlock --jinja + ttl: 300 + + "qwen35-9b": + name: "Qwen3.5-9B Q8_0" + description: "Reasoning distill. 4.38 t/s measured, mlock-pinned" + cmd: | + ${server-base} ${turbo2-kv} + --model /models/Qwen3.5-9B.Q8_0.gguf + --n-gpu-layers 11 + --ctx-size 32768 + --batch-size 512 --ubatch-size 256 + --cont-batching --parallel 1 + --no-mmap --mlock + ttl: 300 + + # ── Pure-GPU small models ─────────────────────────────────────────────────── + + "qwen3-4b": + name: "Qwen3-4B Q4_K_M" + description: "44 t/s @ 16K. NEVER turbo KV (PPL 438 @ 32K — FINDINGS.md §2)" + cmd: | + ${server-base} ${q4-kv} + --model /models/Qwen3-4B-Q4_K_M.gguf + --n-gpu-layers 99 + --ctx-size 16384 + --batch-size 512 --ubatch-size 256 + --cont-batching --parallel 1 + ttl: 300 + + "smollm3-3b": + name: "SmolLM3-3B Q4_K_M" + description: "58 t/s @ 32K, thinking+tools, 2 slots" + cmd: | + ${server-base} ${turbo2-kv} + --model /models/HuggingFaceTB_SmolLM3-3B-Q4_K_M.gguf + --n-gpu-layers 99 + --ctx-size 32768 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 2 + ttl: 300 + + "gemma4-e2b": + name: "Gemma 4 E2B Q4_K_M" + description: "66 t/s, multimodal, 131K ctx (MQA tiny KV; f16 KV — turbo2 worse)" + cmd: | + ${server-base} + --cache-type-k f16 --cache-type-v f16 + --model /models/google_gemma-4-E2B-it-Q4_K_M.gguf + --n-gpu-layers 99 + --ctx-size 131072 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 2 + ttl: 300 + + "gemma4-e4b": + name: "Gemma 4 E4B Q4_K_M" + description: "32 t/s @ 24K, multimodal. ngl=42 needs free VRAM" + cmd: | + ${server-base} ${turbo2-kv} + --model /models/google_gemma-4-E4B-it-Q4_K_M.gguf + --n-gpu-layers 42 + --ctx-size 24576 + --batch-size 1024 --ubatch-size 512 + --cont-batching --parallel 1 + ttl: 300 + + # ── bigctx variants (-nkvo: KV in RAM over PCIe, ~8 GB/s) ────────────────── + + "smollm3-3b-bigctx": + name: "SmolLM3-3B bigctx 65K" + description: "~15 t/s @ 50% fill, KV in RAM" + cmd: | + ${server-base} ${turbo2-kv} + --model /models/HuggingFaceTB_SmolLM3-3B-Q4_K_M.gguf + --n-gpu-layers 99 + --ctx-size 65536 + --batch-size 512 --ubatch-size 256 + --cont-batching --parallel 1 + --no-kv-offload + ttl: 300 + + "gemma4-e2b-bigctx": + name: "Gemma 4 E2B bigctx 393K" + description: "~17 t/s @ 50% fill. q4_0 KV (turbo2 worse on MQA)" + cmd: | + ${server-base} ${q4-kv} + --model /models/google_gemma-4-E2B-it-Q4_K_M.gguf + --n-gpu-layers 99 + --ctx-size 393216 + --batch-size 512 --ubatch-size 256 + --cont-batching --parallel 1 + --no-kv-offload + ttl: 300 + + "gemma4-e4b-bigctx": + name: "Gemma 4 E4B bigctx 163K" + description: "~18 t/s @ 50% fill, KV in RAM" + cmd: | + ${server-base} ${turbo2-kv} + --model /models/google_gemma-4-E4B-it-Q4_K_M.gguf + --n-gpu-layers 42 + --ctx-size 163840 + --batch-size 512 --ubatch-size 128 + --cont-batching --parallel 1 + --no-kv-offload + ttl: 300 + + "qwen3-4b-bigctx": + name: "Qwen3-4B bigctx 24K" + description: "~11 t/s @ 50% fill. q4_0 KV only (turbo broken)" + cmd: | + ${server-base} ${q4-kv} + --model /models/Qwen3-4B-Q4_K_M.gguf + --n-gpu-layers 20 + --ctx-size 24576 + --batch-size 512 --ubatch-size 256 + --cont-batching --parallel 1 + --no-kv-offload + ttl: 300 diff --git a/swap-stack/llama-swap b/swap-stack/llama-swap new file mode 100755 index 0000000..d5aec11 Binary files /dev/null and b/swap-stack/llama-swap differ diff --git a/swap-stack/llama-swap.tar.gz b/swap-stack/llama-swap.tar.gz new file mode 100644 index 0000000..837fa09 Binary files /dev/null and b/swap-stack/llama-swap.tar.gz differ