Full OpenAPI Compatible API including multimodal text generation, text to speech, speech to text, and image generation with cutting edge open source models and llama-swap on Kubernetes on a Strix Halo (AMD Ryzen AI Max+ 395) and Vulkan with llama.cpp, whisper.cpp, stablediffusion.cpp, Kokoro-FastAPI, and Qwen3 TTS OpenAI FastAPI
Published: 2026-10-04
In the following blog post I am describing my setup for self-hosting a almost fully compatible OpenAI server on a Strix Halo computer with 128GB of RAM. The computer has enough RAM to run everything together at the same time in an efficient manner. This setup is the basis upon which all my other AI projects are built.
Overall, there are quite a few components that need to be built into a single docker image - yielding an image size of around 3 - 4 GB. However, thanks to using the Vulkan backend, the size is still manageable - and compared to LLM model sizes the images are a drop in the ocean.
On the other hand, this setup offers all of the following:
- Chatting with an OpenAI compatible API based Chatbot, including multimodal capabilities (such as transcribing scanned images, or reasoning about embedded pdfs)
- Generating diffusion based AI images
- Editing images
- Converting speech to text
- Converting text to speech
The only major open feature I have not yet expored deeply is video and audio generation. I am sure I will be looking into these blank spots in the future - for now I am happy that I can run all the above features concurrently on my little AI box.
By any means the setup is not complete; I need to look into audio.cpp and gufo next, those are great projects that should offer clear performance / dependency size benefits for my strix halo setup. However, the current setup covers enough in order to be valuable and usable for chatting / vibing.
First I am building a base image for llama-swap, including llama.cpp, wbisper.cpp, stablediffusion.cpp, Kokoto-FastAPI and Qwen3 TTS OpenAI FastAPI. The repository / image are publicly available here. However, you probably want to customize the Dockerfile. I have copied the Dockerfile as a reference further below. The Dockerfile creates an image that is around 3GB to 4GB big, which is much smaller thanks to the Vulkan backend rather than using a ROCm backend for instance. The second advantage is that the image should be hardware agnostic - but I've only tested it on AMD hardware.
Afterwards, I am running the image on my Strix Halo node via my Kubernetes cluster. Further below I have also included the full kubernetes config of the llm-namespace which is controlling & running the image on my cluster. If you want to use it you'll have to adapt a few things (such as the node selection), and I am also using flux for keeping the deployed image updated with the latest published image in the repository. You'd have to setup your own docker registry / flux repository references and secrets. Still, if you're running a kubernetes cluster I am sure this example can be a good starting point.
For anyone else, the easiest setup is probably via docker-compose, a setup which I have previously described here. Most of that should still remain true with the newer llama-swap image. Basically you can apply the same docker-compose file and update the config.yaml and Dockerfile to match the current examples, download the necessary models, and you're good to go. In the end of the blog post, I have also included these config files as examples. They're not fully tested though, so there could be some issues.
I have also added the example files to the repository. To get started, you just need to clone the repository here, clone all the desired models as per the config, and run docker compose up.
Have fun! In a future blog post, I will show how I use this setup together with opencode, but for now, you could just point any tool (such as opencode) at the OpenAI-compatible REST endpoints. Or you can open the llama-swap UI under the local port, such as http://localhost:12345/, or https://llm.example.com.
Dockerfile
FROM debian:testing
ARG NODE_VERSION=24
ENV DEBIAN_FRONTEND=noninteractive \
PHONEMIZER_ESPEAK_PATH=/usr/bin \
PHONEMIZER_ESPEAK_DATA=/usr/share/espeak-ng-data \
ESPEAK_DATA_PATH=/usr/share/espeak-ng-data \
PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
NUMBA_CACHE_DIR=/tmp/numba_cache
## Container
RUN mkdir /models
RUN mkdir /conf
## Install dependencies
RUN apt update \
&& apt upgrade -y \
&& apt install -y \
build-essential \
git \
wget \
cmake \
npm \
curl \
glslc \
nodejs \
python3 \
xz-utils \
python3-pip \
python3-wheel \
spirv-headers \
libcurl4-openssl-dev \
libcpp-httplib-dev \
libminiaudio-dev \
libxcb-xinput0 \
libxcb-xinerama0 \
libxcb-cursor-dev \
libvulkan-dev \
libavcodec-dev \
libavformat-dev \
libavutil-dev \
espeak-ng \
espeak-ng-data \
ffmpeg \
zstd \
g++ \
libsndfile1 \
libgomp1 \
libvulkan1 \
mesa-vulkan-drivers \
libsox-dev \
sox \
vulkan-tools \
radeontop \
&& apt autoremove -y \
&& apt clean -y \
&& rm -rf /tmp/* /var/tmp/* \
&& find /var/cache/apt/archives /var/lib/apt/lists -not -name lock -type f -delete \
&& find /var/cache -type f -delete
## Install Go
RUN curl --output go-linux-amd64.tar.gz --location https://go.dev/dl/go1.26.1.linux-amd64.tar.gz \
&& rm -rf /usr/local/go && tar -C /usr/local -xzf go-linux-amd64.tar.gz \
&& rm go-linux-amd64.tar.gz
ENV PATH=$PATH:/usr/local/go/bin
## Install Rust
RUN curl https://sh.rustup.rs -sSf | sh -s -- -y
ENV PATH=$PATH:/root/.cargo/bin
## Install UV
RUN curl -LsSf https://astral.sh/uv/install.sh | sh \
&& mv /root/.local/bin/uv /usr/local/bin/ \
&& mv /root/.local/bin/uvx /usr/local/bin/
## Clone repositories
WORKDIR /app
RUN git clone https://github.com/ggml-org/llama.cpp.git
RUN git clone https://github.com/ggml-org/whisper.cpp.git
RUN git clone https://github.com/mostlygeek/llama-swap
RUN git clone https://github.com/remsky/Kokoro-FastAPI.git
RUN git clone https://github.com/leejet/stable-diffusion.cpp.git
RUN git clone https://github.com/groxaxo/Qwen3-TTS-Openai-Fastapi.git
## Build Qwen3 TTS
WORKDIR /app/Qwen3-TTS-Openai-Fastapi
RUN uv venv --python 3.12 \
&& uv pip install --no-cache-dir --upgrade pip setuptools wheel \
&& uv pip install --no-cache-dir \
torch>=2.0.0 \
torchaudio>=2.0.0 \
--index-url https://download.pytorch.org/whl/cpu \
&& uv pip install --no-cache-dir \
transformers>=4.40.0 \
accelerate>=1.0.0 \
librosa \
soundfile \
pydub \
numpy \
scipy \
einops \
onnxruntime \
fastapi>=0.109.0 \
uvicorn[standard]>=0.27.0 \
python-multipart \
pydantic>=2.0.0 \
inflect \
aiofiles \
&& uv pip install --no-cache-dir -e .
RUN mkdir -p /tmp/numba_cache
## Build stable-diffusion.cpp
WORKDIR /app/stable-diffusion.cpp
RUN git submodule sync --recursive && git submodule update --init --recursive
RUN cmake . -B ./build -DSD_VULKAN=ON && \
cmake --build ./build --config Release --parallel
ENV PATH=$PATH:/app/stable-diffusion.cpp/build/bin
## Build kokoro-fastapi
WORKDIR /app/Kokoro-FastAPI
RUN uv venv --python 3.12 && \
uv sync --extra cpu --no-cache
ENV PATH=$PATH:/app/Kokoro-FastAPI/.venv/bin
## Build Whisper.cpp
WORKDIR /app/whisper.cpp
RUN cmake -B build -DGGML_VULKAN=1 -D WHISPER_FFMPEG=yes && \
cmake --build build --config Release -j$(nproc)
ENV PATH=$PATH:/app/whisper.cpp/build/bin
## Build llama.cpp
WORKDIR /app/llama.cpp
RUN cmake -B build -DGGML_NATIVE=OFF -DGGML_VULKAN=ON -DLLAMA_BUILD_TESTS=OFF -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON && \
cmake --build build --config Release -j$(nproc)
RUN uv venv && uv pip install -r requirements.txt --index-strategy unsafe-best-match
ENV PATH=$PATH:/app/llama.cpp/build/bin
## Build llama-swap
WORKDIR /app/llama-swap
RUN make clean all
ENV PATH=$PATH:/app/llama-swap/build
WORKDIR /app
CMD ["/bin/bash"]
Kubernetes Config
apiVersion: v1
kind: Namespace
metadata:
name: llm
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: llm
namespace: llm
annotations:
cert-manager.io/cluster-issuer: letsencrypt
traefik.ingress.kubernetes.io/router.middlewares: default-redirect-https@kubernetescrd
spec:
ingressClassName: traefik
rules:
- host: llm.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: llm
port:
number: 15463
tls:
- hosts:
- llm.example.com
secretName: llm-example-com
---
kind: Service
apiVersion: v1
metadata:
name: llm
namespace: llm
spec:
selector:
app: llm
ports:
- protocol: TCP
port: 15463
targetPort: 15463
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: llm
namespace: llm
spec:
selector:
matchLabels:
app: llm
replicas: 1
template:
metadata:
labels:
app: llm
spec:
securityContext:
seccompProfile:
type: Unconfined
fsGroup: 0
runAsUser: 0
runAsGroup: 0
runAsNonRoot: false
supplementalGroups:
- 44
- 991
hostIPC: true
containers:
- name: llm
securityContext:
privileged: true
allowPrivilegeEscalation: true
capabilities:
add:
- SYS_PTRACE
image: registry.example.com/infra/llama-swap-llama-cpp-vulkan/llama-release:master-3cf447df-1790412791 # {"$imagepolicy": "llm:image-policy"}
command: ['/app/llama-swap/build/llama-swap-linux-amd64']
args: ['--config', '/app/config.yaml', '--listen', '0.0.0.0:15463']
ports:
- containerPort: 15463
volumeMounts:
- name: llama-swap-config
mountPath: /app/config.yaml
subPath: config.yaml
readOnly: true
- name: dev-kfd
mountPath: /dev/kfd
securityContext:
privileged: true
- name: dev-dri
mountPath: /dev/dri
securityContext:
privileged: true
- name: run-lactd
mountPath: /run/lactd.sock
securityContext:
privileged: true
- name: models
mountPath: /models
- name: huggingface
mountPath: /root/.cache/huggingface
volumes:
- name: llama-swap-config
configMap:
name: llama-swap-config-v73
items:
- key: config.yaml
path: config.yaml
- name: dev-kfd
hostPath:
path: /dev/kfd
- name: dev-dri
hostPath:
path: /dev/dri
- name: run-lactd
hostPath:
path: /run/lactd.sock
- name: models
hostPath:
path: /models
- name: huggingface
hostPath:
path: /models/hf
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/arch
operator: In
values:
- amd64
- key: kubernetes.io/hostname
operator: In
values:
- srv-7
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: module
operator: In
values:
- llm
topologyKey: 'kubernetes.io/hostname'
---
apiVersion: image.toolkit.fluxcd.io/v1
kind: ImageRepository
metadata:
name: image-repository
namespace: llm
spec:
image: registry.example.com/infra/llama-swap-llama-cpp-vulkan/llama-release
interval: 5m
---
apiVersion: image.toolkit.fluxcd.io/v1
kind: ImagePolicy
metadata:
name: image-policy
namespace: llm
spec:
imageRepositoryRef:
name: image-repository
filterTags:
pattern: '^master-[a-fA-F0-9]+-(?P<ts>[1-9][0-9]*)'
extract: '$ts'
policy:
numerical:
order: asc
---
apiVersion: image.toolkit.fluxcd.io/v1
kind: ImageUpdateAutomation
metadata:
name: image-update-automation
namespace: llm
spec:
interval: 5m
sourceRef:
kind: GitRepository
name: flux
git:
checkout:
ref:
branch: master
commit:
author:
email: mr.robot@example.com
name: mr.robot
messageTemplate: |
Automated image update
Automation name: {{ .AutomationObject }}
Files:
{{ range $filename, $_ := .Changed.FileChanges -}}
- {{ $filename }}
{{ end -}}
Objects:
{{ range $resource, $changes := .Changed.Objects -}}
- {{ $resource.Kind }} {{ $resource.Name }}
Changes:
{{- range $_, $change := $changes }}
- {{ $change.OldValue }} -> {{ $change.NewValue }}
{{ end -}}
{{ end -}}
push:
branch: master
update:
path: ./clusters/k8s-cluster-1
strategy: Setters
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: flux
namespace: llm
spec:
interval: 1m0s
ref:
branch: master
url: https://git.example.com/scope/repo.git
secretRef:
name: secretname
---
apiVersion: v1
kind: ConfigMap
metadata:
name: llama-swap-config-v73
namespace: llm
data:
config.yaml: |
healthCheckTimeout: 600
startPort: 14001
globalTTL: 1337
path: /models/metrics/llama-swap.sqlite
metricsMaxInMemory: 10000
captureBuffer: 33
models:
gemma-4-26b:
cmd: llama-server --port ${PORT} --host 0.0.0.0 --model /models/chat/gemma-4-26B-A4B-qat/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --mmproj /models/chat/gemma-4-26B-A4B-qat/mmproj-F32.gguf --chat-template-file /app/llama.cpp/models/templates/google-gemma-4-31B-it-interleaved.jinja --jinja --image-min-tokens 560 --image-max-tokens 2240 --batch-size 4096 --ubatch-size 4096 --temp 1.0 --top-p 0.95 --top-k 64 -c 255555
aliases:
- "gpt-4.1-mini"
- "default"
gpt-oss-120b:
cmd: llama-server --port ${PORT} --host 0.0.0.0 --model /models/chat/gpt-oss-120b/gpt-oss-120b-UD-Q8_K_XL-00001-of-00002.gguf --jinja --temp 1.0 --top-p 1.0 --top-k 0 -c 111111
qwen-3.8-27b:
cmd: llama-server --port ${PORT} --host 0.0.0.0 --model /models/chat/qwen-3.8-27b/Qwen3.8-27B-UD-Q8_K_XL.gguf --mmproj /models/chat/qwen-3.8-27b/mmproj-BF16.gguf --jinja --temp 1.0 --top-p 0.95 --min-p 0 --top-k 20 --repeat-penalty 1.0 --presence-penalty 0.0 -c 262144 --spec-type draft-mtp --batch-size 4096 --ubatch-size 4096
automatic-speech-recognition--whisper:
checkEndpoint: /v1/audio/transcriptions/
cmd: /app/whisper.cpp/build/bin/whisper-server --port ${PORT} --host 0.0.0.0 --convert --request-path /v1/audio/transcriptions --inference-path "" --model /models/audio/whisper/ggml-large-v3.bin
tts--kokoro:
useModelName: kokoro
env:
- "USE_GPU=false"
- "USE_ONNX=false"
- "UV_PYTHON=3.12"
- "UV_WORKING_DIR=/app/Kokoro-FastAPI"
- "PYTHONPATH=/app/Kokoro-FastAPI:/app/Kokoro-FastAPI/api"
- "MODEL_DIR=/models/audio/kokoro"
- "VOICES_DIR=/models/audio/kokoro/voices"
- "WEB_PLAYER_PATH=/app/Kokoro-FastAPI/web"
- "ESPEAK_DATA_PATH=/usr/lib/x86_64-linux-gnu/espeak-ng-data"
cmd: uv run uvicorn api.src.main:app --port ${PORT} --host 0.0.0.0
tts--qwen-3:
useModelName: tts-1
env:
- "UV_PYTHON=3.12"
- "UV_NO_PROJECT=1"
- "UV_WORKING_DIR=/app/Qwen3-TTS-Openai-Fastapi"
- "PYTHONPATH=/app/Qwen3-TTS-Openai-Fastapi"
- "TTS_BACKEND=official" # TTS_BACKEND=pytorch
- "TTS_MODEL_NAME=Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice"
cmd: uv run uvicorn api.main:app --port ${PORT} --host 0.0.0.0
stable-diffusion--flux2-klein:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/flux2-klein/flux-2-klein-4b-Q8_0.gguf --vae /models/image/flux2-klein/full_encoder_small_decoder.safetensors --llm /models/chat/qwen-3.0-4b-it/Qwen3-4B-Instruct-2507-UD-Q8_K_XL.gguf --cfg-scale 1.0 --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
stable-diffusion--z-image:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/z-image/z-image-Q8_0.gguf --vae /models/image/z-image/ae.safetensors --llm /models/chat/qwen-3.0-4b-it/Qwen3-4B-Instruct-2507-UD-Q8_K_XL.gguf --cfg-scale 1.0 --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
stable-diffusion--z-image-turbo:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/z-image-turbo/z_image_turbo-Q8_0.gguf --vae /models/image/z-image-turbo/ae.safetensors --llm /models/chat/qwen-3.0-4b-it/Qwen3-4B-Instruct-2507-UD-Q8_K_XL.gguf --cfg-scale 1.0 --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
stable-diffusion--qwen-image:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/qwen-image/Qwen_Image-Q8_0.gguf --vae /models/image/qwen-image/qwen_image_vae.safetensors --llm /models/chat/qwen-2.5-vl-7b/Qwen2.5-VL-7B-Instruct.Q8_0.gguf --llm_vision /models/chat/qwen-2.5-vl-7b/Qwen2.5-VL-7B-Instruct.mmproj-f16.gguf --cfg-scale 2.5 --sampling-method euler --flow-shift 3 --qwen-image-zero-cond-t --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
stable-diffusion--qwen-image-edit:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/qwen-image-edit/Qwen_Image_Edit-Q8_0.gguf --vae /models/image/qwen-image-edit/qwen_image_vae.safetensors --llm /models/chat/qwen-2.5-vl-7b/Qwen2.5-VL-7B-Instruct.Q8_0.gguf --llm_vision /models/chat/qwen-2.5-vl-7b/Qwen2.5-VL-7B-Instruct.mmproj-f16.gguf --cfg-scale 2.5 --sampling-method euler --flow-shift 3 --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
groups:
default:
swap: false
members:
- "gemma-4-26b"
- "qwen-3.8-27b"
- "stable-diffusion--flux2-klein"
- "stable-diffusion--qwen-image-edit"
- "stable-diffusion--qwen-image"
- "tts--kokoro"
- "tts--qwen-3"
- "automatic-speech-recognition--whisper"
hooks:
on_startup:
preload:
- "gemma-4-26b"
---
docker-compose
docker-compose.yml
services:
server:
build: ..
ports:
- '12345:12345'
volumes:
- /models:/models
- ./config.yaml:/conf/config.yaml
devices:
- '/dev/kfd:/dev/kfd'
- '/dev/dri:/dev/dri'
security_opt:
- seccomp:unconfined
group_add:
- video
cap_add:
- SYS_PTRACE
ipc: 'host'
command: /app/llama-swap/build/llama-swap-linux-amd64 --listen 0.0.0.0:12345 --config /conf/config.yaml
config.yaml
healthCheckTimeout: 600
startPort: 14001
globalTTL: 1337
path: /models/metrics/llama-swap.sqlite
metricsMaxInMemory: 10000
captureBuffer: 33
models:
gemma-4-26b:
cmd: llama-server --port ${PORT} --host 0.0.0.0 --model /models/chat/gemma-4-26B-A4B-qat/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --mmproj /models/chat/gemma-4-26B-A4B-qat/mmproj-F32.gguf --chat-template-file /app/llama.cpp/models/templates/google-gemma-4-31B-it-interleaved.jinja --jinja --image-min-tokens 560 --image-max-tokens 2240 --batch-size 4096 --ubatch-size 4096 --temp 1.0 --top-p 0.95 --top-k 64 -c 255555
aliases:
- "gpt-4.1-mini"
- "default"
gpt-oss-120b:
cmd: llama-server --port ${PORT} --host 0.0.0.0 --model /models/chat/gpt-oss-120b/gpt-oss-120b-UD-Q8_K_XL-00001-of-00002.gguf --jinja --temp 1.0 --top-p 1.0 --top-k 0 -c 111111
qwen-3.8-27b:
cmd: llama-server --port ${PORT} --host 0.0.0.0 --model /models/chat/qwen-3.8-27b/Qwen3.8-27B-UD-Q8_K_XL.gguf --mmproj /models/chat/qwen-3.8-27b/mmproj-BF16.gguf --jinja --temp 1.0 --top-p 0.95 --min-p 0 --top-k 20 --repeat-penalty 1.0 --presence-penalty 0.0 -c 262144 --spec-type draft-mtp --batch-size 4096 --ubatch-size 4096
automatic-speech-recognition--whisper:
checkEndpoint: /v1/audio/transcriptions/
cmd: /app/whisper.cpp/build/bin/whisper-server --port ${PORT} --host 0.0.0.0 --convert --request-path /v1/audio/transcriptions --inference-path "" --model /models/audio/whisper/ggml-large-v3.bin
tts--kokoro:
useModelName: kokoro
env:
- "USE_GPU=false"
- "USE_ONNX=false"
- "UV_PYTHON=3.12"
- "UV_WORKING_DIR=/app/Kokoro-FastAPI"
- "PYTHONPATH=/app/Kokoro-FastAPI:/app/Kokoro-FastAPI/api"
- "MODEL_DIR=/models/audio/kokoro"
- "VOICES_DIR=/models/audio/kokoro/voices"
- "WEB_PLAYER_PATH=/app/Kokoro-FastAPI/web"
- "ESPEAK_DATA_PATH=/usr/lib/x86_64-linux-gnu/espeak-ng-data"
cmd: uv run uvicorn api.src.main:app --port ${PORT} --host 0.0.0.0
tts--qwen-3:
useModelName: tts-1
env:
- "UV_PYTHON=3.12"
- "UV_NO_PROJECT=1"
- "UV_WORKING_DIR=/app/Qwen3-TTS-Openai-Fastapi"
- "PYTHONPATH=/app/Qwen3-TTS-Openai-Fastapi"
- "TTS_BACKEND=official" # TTS_BACKEND=pytorch
- "TTS_MODEL_NAME=Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice"
cmd: uv run uvicorn api.main:app --port ${PORT} --host 0.0.0.0
stable-diffusion--flux2-klein:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/flux2-klein/flux-2-klein-4b-Q8_0.gguf --vae /models/image/flux2-klein/full_encoder_small_decoder.safetensors --llm /models/chat/qwen-3.0-4b-it/Qwen3-4B-Instruct-2507-UD-Q8_K_XL.gguf --cfg-scale 1.0 --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
stable-diffusion--z-image:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/z-image/z-image-Q8_0.gguf --vae /models/image/z-image/ae.safetensors --llm /models/chat/qwen-3.0-4b-it/Qwen3-4B-Instruct-2507-UD-Q8_K_XL.gguf --cfg-scale 1.0 --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
stable-diffusion--z-image-turbo:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/z-image-turbo/z_image_turbo-Q8_0.gguf --vae /models/image/z-image-turbo/ae.safetensors --llm /models/chat/qwen-3.0-4b-it/Qwen3-4B-Instruct-2507-UD-Q8_K_XL.gguf --cfg-scale 1.0 --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
stable-diffusion--qwen-image:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/qwen-image/Qwen_Image-Q8_0.gguf --vae /models/image/qwen-image/qwen_image_vae.safetensors --llm /models/chat/qwen-2.5-vl-7b/Qwen2.5-VL-7B-Instruct.Q8_0.gguf --llm_vision /models/chat/qwen-2.5-vl-7b/Qwen2.5-VL-7B-Instruct.mmproj-f16.gguf --cfg-scale 2.5 --sampling-method euler --flow-shift 3 --qwen-image-zero-cond-t --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
stable-diffusion--qwen-image-edit:
checkEndpoint: /
cmd: sd-server --diffusion-model /models/image/qwen-image-edit/Qwen_Image_Edit-Q8_0.gguf --vae /models/image/qwen-image-edit/qwen_image_vae.safetensors --llm /models/chat/qwen-2.5-vl-7b/Qwen2.5-VL-7B-Instruct.Q8_0.gguf --llm_vision /models/chat/qwen-2.5-vl-7b/Qwen2.5-VL-7B-Instruct.mmproj-f16.gguf --cfg-scale 2.5 --sampling-method euler --flow-shift 3 --diffusion-fa -v --listen-ip 0.0.0.0 --listen-port ${PORT} --vae-tiling
groups:
default:
swap: false
members:
- "gemma-4-26b"
- "qwen-3.8-27b"
- "stable-diffusion--flux2-klein"
- "stable-diffusion--qwen-image-edit"
- "stable-diffusion--qwen-image"
- "tts--kokoro"
- "tts--qwen-3"
- "automatic-speech-recognition--whisper"
hooks:
on_startup:
preload:
- "gemma-4-26b"