Files
VoiceStudio/deploy/docker-compose.yml
T
Palash Debnath 93025d9a81 fix(dictation): stop the widget stranding an empty square, repair the swept data dirs (#1398)
The dictation hotkey could leave a blank dark square stuck on the desktop with no way to dismiss it. Three defects compounded: the tray listener's effect depended on [state], so it detached across an await on every state change and a press landing in that gap was lost; an idle pill renders null, so the window Rust had already shown was empty; and the opaque chrome background made that empty window a hard-edged square. Nothing could hide it — dismiss() is only reachable from the X button, Esc, or a post-session timer, none of which exist for a session that never started.

Fixed at the invariant rather than the call sites: the listener subscribes once for the component's lifetime, the widget window's chrome background is transparent, and an idle-but-visible window reconciles itself to hidden. The reconcile is polled (a dropped press changes no React state, so there is nothing to key an effect off) and aborts if its effect is torn down mid-check, so it can never hide a dictation that has just started.

Also in scope:

- The rename sweep had repointed three data-dir literals at a brand-named directory that does not exist, so smoke-test.sh verified a directory the backend never writes and desktop-prod.sh silently stopped clearing backend state on Windows. Both invisible on macOS, where they are usually run. A guard test now pins the assignments specifically.
- The dictation model picker's download sizes were wrong for all seven models, in both directions — Parakeet TDT v3 (the recommended default) understated 180 MB against an actual 670 MB, while the low-RAM fallbacks were overstated threefold, discouraging exactly the choice that would have helped. Measured from the published repos and pinned by a test.
- The 0.6B Parakeet models now decode on more threads, capped by host cores and still overridable.
- uninstall.ps1 gained a UTF-8 BOM (Windows PowerShell 5.1 mis-decodes its non-ASCII output without one), and sponsor.yml lost its last OmniVoice references.
2026-08-07 05:39:41 +05:30

153 lines
6.2 KiB
YAML

# ──────────────────────────────────────────────────────────────
# VoiceStudio — Docker Compose
#
# Quick start:
# docker compose -f deploy/docker-compose.yml --profile cpu up # CPU mode
# docker compose -f deploy/docker-compose.yml --profile gpu up # NVIDIA GPU
# docker compose -f deploy/docker-compose.yml --profile rocm up # AMD GPU (ROCm)
#
# All services bind to port 3900, so they MUST be opt-in via profiles —
# otherwise `compose up` would race them and one would fail to bind.
#
# First run downloads ~4 GB of models. Progress is shown in logs.
# Open http://localhost:3900 once the health check passes.
#
# SECURITY: The port is bound to 127.0.0.1 by default — only this
# machine can reach the API. To expose OmniVoice on your LAN (or
# through a reverse proxy / tunnel), change the port mapping to
# "0.0.0.0:3900:3900" or "3900:3900". OmniVoice itself ships no
# authentication — if you expose it, put it behind a reverse proxy
# with auth (Caddy basic_auth, nginx + htpasswd, Tailscale, etc.).
# ──────────────────────────────────────────────────────────────
services:
# ── CPU mode — activate with: docker compose --profile cpu up
omnivoice:
image: ghcr.io/debpalash/omnivoice-studio:latest
# To build from source instead of pulling, comment out `image:` and
# uncomment the two lines below:
build:
context: ..
dockerfile: deploy/Dockerfile
container_name: omnivoice-studio
profiles: ["cpu"]
ports:
- "127.0.0.1:3900:3900"
volumes:
- omnivoice-data:/app/omnivoice_data
environment:
- HF_HOME=/app/omnivoice_data/huggingface
- HF_TOKEN=${HF_TOKEN:-}
- OMNIVOICE_DATA_DIR=/app/omnivoice_data
- PYTHONPATH=/app/backend
- PYTHONUNBUFFERED=1
# Bind uvicorn to 0.0.0.0 *inside* the container so the host-side port
# mapping above can forward traffic in. The 127.0.0.1 prefix on the
# `ports:` mapping is what enforces loopback-only on the host —
# OMNIVOICE_BIND_HOST=0.0.0.0 here only opens the container's own
# interface. The backend default is 127.0.0.1 (see backend/main.py).
- OMNIVOICE_BIND_HOST=0.0.0.0
# Headless server: relax the desktop-only loopback origin gate so the
# web UI's /system/* and /api/settings/* routes work through Docker's
# NAT (issue #261). Already baked into the image; shown here so it's
# discoverable. If you front the container with your own auth proxy on
# loopback, set this to 0 to re-enable the strict gate.
- OMNIVOICE_SERVER_MODE=1
healthcheck:
test: ["CMD", "curl", "-sf", "http://localhost:3900/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 120s
restart: unless-stopped
# ── GPU mode — activate with: docker compose --profile gpu up
omnivoice-gpu:
image: ghcr.io/debpalash/omnivoice-studio:latest
build:
context: ..
dockerfile: deploy/Dockerfile
container_name: omnivoice-studio-gpu
profiles: ["gpu"]
ports:
- "127.0.0.1:3900:3900"
volumes:
- omnivoice-data:/app/omnivoice_data
environment:
- HF_HOME=/app/omnivoice_data/huggingface
- HF_TOKEN=${HF_TOKEN:-}
- OMNIVOICE_DATA_DIR=/app/omnivoice_data
- PYTHONPATH=/app/backend
- PYTHONUNBUFFERED=1
# Bind uvicorn to 0.0.0.0 *inside* the container — same as the CPU
# service above. The host-side `127.0.0.1:3900:3900` mapping keeps
# LAN reachability off by default.
- OMNIVOICE_BIND_HOST=0.0.0.0
# See the CPU service above — relaxes the loopback origin gate for the
# headless Docker deployment (issue #261). Set to 0 to re-enable it.
- OMNIVOICE_SERVER_MODE=1
healthcheck:
test: ["CMD", "curl", "-sf", "http://localhost:3900/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
restart: unless-stopped
# ── AMD GPU (ROCm) mode — activate with: docker compose --profile rocm up
# Uses the dedicated `:rocm` image variant (#1165). The GPU is passed
# through as plain device nodes — no toolkit needed, the host only needs
# the amdgpu kernel driver (the ROCm userspace ships inside the image).
# Podman works with the same two --device flags (Quadlet: AddDevice=).
omnivoice-rocm:
image: ghcr.io/debpalash/omnivoice-studio:rocm
# To build from source instead of pulling, comment out `image:` and
# uncomment the lines below. BASE_IMAGE/GPU_FLAVOR are required — the
# Dockerfile's defaults build the CUDA variant.
# build:
# context: ..
# dockerfile: deploy/Dockerfile
# args:
# BASE_IMAGE: rocm/pytorch:rocm7.2.4_ubuntu24.04_py3.12_pytorch_release_2.8.0
# GPU_FLAVOR: rocm
container_name: omnivoice-studio-rocm
profiles: ["rocm"]
ports:
- "127.0.0.1:3900:3900"
devices:
- /dev/kfd
- /dev/dri
volumes:
- omnivoice-data:/app/omnivoice_data
environment:
- HF_HOME=/app/omnivoice_data/huggingface
- HF_TOKEN=${HF_TOKEN:-}
- OMNIVOICE_DATA_DIR=/app/omnivoice_data
- PYTHONPATH=/app/backend
- PYTHONUNBUFFERED=1
# See the CPU service above — container-internal bind + relaxed
# loopback origin gate for the headless Docker deployment.
- OMNIVOICE_BIND_HOST=0.0.0.0
- OMNIVOICE_SERVER_MODE=1
# RDNA3 consumer cards (RX 7900 XTX/XT and friends, gfx1100): if the
# GPU is not detected, uncomment the override below. The backend
# auto-sets it for known consumer GFX IDs, so try without it first.
# - HSA_OVERRIDE_GFX_VERSION=11.0.0
healthcheck:
test: ["CMD", "curl", "-sf", "http://localhost:3900/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
restart: unless-stopped
volumes:
omnivoice-data: