Files
VoiceStudio/deploy/docker-compose.yml
T
99e01610bb feat(docker): publish ROCm/AMD GPU image variant (#1165) (#1166)
The Docker image was CUDA-only, so AMD GPUs (e.g. RX 7900 XTX under
Podman) silently ran on CPU. Every preview and release now also ships a
ROCm variant built from the same Dockerfile:

- deploy/Dockerfile: parameterize the runtime base with a BASE_IMAGE
  build-arg (default unchanged: pytorch/pytorch 2.8.0 CUDA). Add
  PIP/UV_BREAK_SYSTEM_PACKAGES for the ROCm base's PEP-668-marked
  Ubuntu 24.04 Python (no-op on the conda CUDA base), and a build-time
  GPU_FLAVOR guard asserting the dependency install did not clobber the
  base image's GPU torch/torchaudio — a future dep bump that forces a
  torch reinstall now fails the build instead of shipping a CPU-only
  "ROCm" image.
- .github/workflows/docker.yml: new build-and-push-rocm job (separate
  job for runner disk — the ROCm base is ~25 GB unpacked, so it frees
  the preinstalled toolchains first). Tags mirror the CUDA semantics
  with a -rocm suffix (:rocm rolling preview, :stable-rocm, :X.Y.Z-rocm,
  :X.Y-rocm, :sha-xxxx-rocm) on both GHCR and Docker Hub, same secret
  gating. flavor latest=false so release tags can't clobber :latest.
  No cache-to: the ROCm layers would blow the 10 GB GHA cache budget.
- deploy/docker-compose.yml: new opt-in 'rocm' profile passing the GPU
  through via /dev/kfd + /dev/dri, with HSA_OVERRIDE_GFX_VERSION=11.0.0
  documented (user-set, not baked in — backend auto-sets it for known
  consumer GFX IDs).
- Docs-sync: docker.md (ROCm quick start incl. Podman/Quadlet, tag
  table, troubleshooting), dockerhub-overview.md, README AMD note,
  linux.md ROCm section cross-link, CHANGELOG [Unreleased].

Base image: rocm/pytorch:rocm7.2.4_ubuntu24.04_py3.12_pytorch_release_2.8.0
— torch 2.8.0 exactly matches the CUDA image (identical resolution, so
uv keeps it), py3.12 satisfies requires-python >=3.11 (the ubuntu22.04
variants are py3.10 and do not).

Closes #1165

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 15:34:28 +05:30

153 lines
6.2 KiB
YAML

# ──────────────────────────────────────────────────────────────
# OmniVoice Studio — Docker Compose
#
# Quick start:
# docker compose -f deploy/docker-compose.yml --profile cpu up # CPU mode
# docker compose -f deploy/docker-compose.yml --profile gpu up # NVIDIA GPU
# docker compose -f deploy/docker-compose.yml --profile rocm up # AMD GPU (ROCm)
#
# All services bind to port 3900, so they MUST be opt-in via profiles —
# otherwise `compose up` would race them and one would fail to bind.
#
# First run downloads ~4 GB of models. Progress is shown in logs.
# Open http://localhost:3900 once the health check passes.
#
# SECURITY: The port is bound to 127.0.0.1 by default — only this
# machine can reach the API. To expose OmniVoice on your LAN (or
# through a reverse proxy / tunnel), change the port mapping to
# "0.0.0.0:3900:3900" or "3900:3900". OmniVoice itself ships no
# authentication — if you expose it, put it behind a reverse proxy
# with auth (Caddy basic_auth, nginx + htpasswd, Tailscale, etc.).
# ──────────────────────────────────────────────────────────────
services:
# ── CPU mode — activate with: docker compose --profile cpu up
omnivoice:
image: ghcr.io/debpalash/omnivoice-studio:latest
# To build from source instead of pulling, comment out `image:` and
# uncomment the two lines below:
build:
context: ..
dockerfile: deploy/Dockerfile
container_name: omnivoice-studio
profiles: ["cpu"]
ports:
- "127.0.0.1:3900:3900"
volumes:
- omnivoice-data:/app/omnivoice_data
environment:
- HF_HOME=/app/omnivoice_data/huggingface
- HF_TOKEN=${HF_TOKEN:-}
- OMNIVOICE_DATA_DIR=/app/omnivoice_data
- PYTHONPATH=/app/backend
- PYTHONUNBUFFERED=1
# Bind uvicorn to 0.0.0.0 *inside* the container so the host-side port
# mapping above can forward traffic in. The 127.0.0.1 prefix on the
# `ports:` mapping is what enforces loopback-only on the host —
# OMNIVOICE_BIND_HOST=0.0.0.0 here only opens the container's own
# interface. The backend default is 127.0.0.1 (see backend/main.py).
- OMNIVOICE_BIND_HOST=0.0.0.0
# Headless server: relax the desktop-only loopback origin gate so the
# web UI's /system/* and /api/settings/* routes work through Docker's
# NAT (issue #261). Already baked into the image; shown here so it's
# discoverable. If you front the container with your own auth proxy on
# loopback, set this to 0 to re-enable the strict gate.
- OMNIVOICE_SERVER_MODE=1
healthcheck:
test: ["CMD", "curl", "-sf", "http://localhost:3900/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 120s
restart: unless-stopped
# ── GPU mode — activate with: docker compose --profile gpu up
omnivoice-gpu:
image: ghcr.io/debpalash/omnivoice-studio:latest
build:
context: ..
dockerfile: deploy/Dockerfile
container_name: omnivoice-studio-gpu
profiles: ["gpu"]
ports:
- "127.0.0.1:3900:3900"
volumes:
- omnivoice-data:/app/omnivoice_data
environment:
- HF_HOME=/app/omnivoice_data/huggingface
- HF_TOKEN=${HF_TOKEN:-}
- OMNIVOICE_DATA_DIR=/app/omnivoice_data
- PYTHONPATH=/app/backend
- PYTHONUNBUFFERED=1
# Bind uvicorn to 0.0.0.0 *inside* the container — same as the CPU
# service above. The host-side `127.0.0.1:3900:3900` mapping keeps
# LAN reachability off by default.
- OMNIVOICE_BIND_HOST=0.0.0.0
# See the CPU service above — relaxes the loopback origin gate for the
# headless Docker deployment (issue #261). Set to 0 to re-enable it.
- OMNIVOICE_SERVER_MODE=1
healthcheck:
test: ["CMD", "curl", "-sf", "http://localhost:3900/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
restart: unless-stopped
# ── AMD GPU (ROCm) mode — activate with: docker compose --profile rocm up
# Uses the dedicated `:rocm` image variant (#1165). The GPU is passed
# through as plain device nodes — no toolkit needed, the host only needs
# the amdgpu kernel driver (the ROCm userspace ships inside the image).
# Podman works with the same two --device flags (Quadlet: AddDevice=).
omnivoice-rocm:
image: ghcr.io/debpalash/omnivoice-studio:rocm
# To build from source instead of pulling, comment out `image:` and
# uncomment the lines below. BASE_IMAGE/GPU_FLAVOR are required — the
# Dockerfile's defaults build the CUDA variant.
# build:
# context: ..
# dockerfile: deploy/Dockerfile
# args:
# BASE_IMAGE: rocm/pytorch:rocm7.2.4_ubuntu24.04_py3.12_pytorch_release_2.8.0
# GPU_FLAVOR: rocm
container_name: omnivoice-studio-rocm
profiles: ["rocm"]
ports:
- "127.0.0.1:3900:3900"
devices:
- /dev/kfd
- /dev/dri
volumes:
- omnivoice-data:/app/omnivoice_data
environment:
- HF_HOME=/app/omnivoice_data/huggingface
- HF_TOKEN=${HF_TOKEN:-}
- OMNIVOICE_DATA_DIR=/app/omnivoice_data
- PYTHONPATH=/app/backend
- PYTHONUNBUFFERED=1
# See the CPU service above — container-internal bind + relaxed
# loopback origin gate for the headless Docker deployment.
- OMNIVOICE_BIND_HOST=0.0.0.0
- OMNIVOICE_SERVER_MODE=1
# RDNA3 consumer cards (RX 7900 XTX/XT and friends, gfx1100): if the
# GPU is not detected, uncomment the override below. The backend
# auto-sets it for known consumer GFX IDs, so try without it first.
# - HSA_OVERRIDE_GFX_VERSION=11.0.0
healthcheck:
test: ["CMD", "curl", "-sf", "http://localhost:3900/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
restart: unless-stopped
volumes:
omnivoice-data: