Duo group (resident main+subagent for pi): flip VRAM to the small model. Ornith experts never touch VRAM; its dense-on-GPU split starved qwen to 176 MiB / 5.8 t/s concurrent. After flip: qwen 43 t/s, ornith 8 t/s. CUDA_VISIBLE_DEVICES= required for ornith — ngl 0 still allocates ~1GB pp compute buffer on CUDA builds (OOM+segfault). Duo section in MOE-FINDINGS.md; also snapshots prior swap-stack migration state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
15 lines
363 B
Docker
15 lines
363 B
Docker
FROM python:3.12-slim
|
|
|
|
RUN apt-get update && apt-get install -y curl && \
|
|
curl -fsSL https://get.docker.com | sh && \
|
|
apt-get clean && rm -rf /var/lib/apt/lists/*
|
|
|
|
RUN pip install --no-cache-dir fastapi uvicorn docker pydantic httpx
|
|
|
|
WORKDIR /app
|
|
COPY controller.py .
|
|
|
|
EXPOSE 8000
|
|
|
|
CMD ["uvicorn", "controller:app", "--host", "0.0.0.0", "--port", "8000"]
|