The repository was renamed. 724 references across 59 files now point at the new URL — README badges, docs, install guides, the updater's releases API call, CONTRIBUTING, the Colab link and the probe harness. GitHub redirects the old URLs, so nothing was broken in the meantime. Deliberately NOT renamed, because each breaks something on a user's machine: the Tauri bundle identifier (the path to every existing user's data), /usr/lib/omnivoice-studio and the compose container names, and the published Docker image paths. The image path needed a code change to STAY still: docker.yml derived it from github.repository, so the next build would have published to ghcr.io/debpalash/voicestudio while Docker Hub, a hardcoded literal, stayed put — everyone pulling the documented GHCR path would have kept receiving the last pre-rename image forever. It is now pinned, with a test that fails if it ever derives from the repo name again. Also makes the probe's repo-name assertion shape-based: it hardcoded the old name and failed on every PR after the rename while the code it tests worked perfectly.
9.8 KiB
OmniVoice Studio — Install with Docker
For headless servers, dedicated GPUs, or "I want one command" deployments. The docker image bundles the backend; the UI is served over HTTP and you open it in a normal browser.
Official images: ghcr.io/debpalash/omnivoice-studio
and palashdeb/omnivoice-studio on Docker Hub — same images, same tags.
Image ↔ version mapping
Tag What you get :latestRolling preview — latest commit on main, at or ahead of the last release. This is the preview channel; pin:stablefor production.:stableMost recent versioned release (updated on every v*git tag):0.4.1Exact release version :0.4Latest patch within the 0.4 minor :mainAlias of the same rolling mainbuild as:latest:sha-xxxxxxxSpecific commit (produced by manual workflow dispatch) :rocmAMD GPU (ROCm) build of the rolling preview — the ROCm analogue of :latest:stable-rocm,:0.4.1-rocm,:0.4-rocm,:sha-xxxxxxx-rocmROCm builds of the corresponding CUDA tags above Versioning rule: preview builds always come from
mainand never version-sort below:stable— upgrades flow naturally.Note on the update-channel toggle: The update-channel UI (Settings → About → Update channel) is part of the Tauri desktop app's built-in auto-updater. It does not apply to the Docker image — the Docker image is the headless web-server build. To update your Docker deployment, pull the new image tag and recreate the container (
docker compose pull && docker compose up -d).
Pull and run (CPU)
docker pull ghcr.io/debpalash/omnivoice-studio:latest
docker run -d --name omnivoice \
-p 127.0.0.1:3900:3900 \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:latest
Docker Hub mirror: the same images are published to
palashdeb/omnivoice-studioon Docker Hub with identical tags — swap the image forpalashdeb/omnivoice-studio:latestif you prefer Docker Hub. Tag semantics (:latest= rolling main preview,:stable/:X.Y.Z= releases) are the same on both registries.
Open http://localhost:3900. The first run downloads
~2.4 GB of model weights — follow docker logs -f omnivoice to watch.
Pull and run (NVIDIA GPU)
docker run -d --name omnivoice --gpus all \
-p 127.0.0.1:3900:3900 \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:latest
GPU mode requires the NVIDIA Container Toolkit on the host.
Pull and run (AMD GPU / ROCm)
AMD GPUs use the dedicated :rocm image variant — the default (CUDA)
image runs CPU-only on AMD hardware. The ROCm userspace ships inside the
image; the host only needs the amdgpu kernel driver. Pass the GPU through
as plain device nodes (no container toolkit needed):
docker run -d --name omnivoice \
--device /dev/kfd --device /dev/dri \
-p 127.0.0.1:3900:3900 \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:rocm
The same flags work with Podman (podman run --device /dev/kfd --device /dev/dri …); in a Quadlet unit that's two AddDevice= lines:
# ~/.config/containers/systemd/omnivoice.container
[Container]
Image=ghcr.io/debpalash/omnivoice-studio:rocm
AddDevice=/dev/kfd
AddDevice=/dev/dri
PublishPort=127.0.0.1:3900:3900
Volume=omnivoice-data:/app/omnivoice_data
Release pins exist too: :stable-rocm, :0.4.1-rocm, :0.4-rocm mirror
the CUDA tags exactly.
Consumer cards and APUs (RX 6000/7000, Strix Point/Halo): the backend auto-sets
HSA_OVERRIDE_GFX_VERSIONwhen — and only when — your card's GFX ID is missing from the shipped ROCm build's architecture list, so try without any override first. Overriding a natively-supported GPU (gfx1151 on ROCm 7.x, for example) only forces it onto foreign kernels. If the GPU still isn't used, force it explicitly with-e HSA_OVERRIDE_GFX_VERSION=11.0.0(user-set on the container — it is deliberately not baked into the image, because the right value depends on your card); a value you set is always respected as-is.Rootless / non-root hosts: if
/dev/kfdis group-owned, the container user needs those groups too — add--group-addfor your host'srenderandvideoGIDs (getent group render video).
Verify the container sees the GPU:
docker exec omnivoice python3 -c \
"import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"
(ROCm-built PyTorch reports through torch.cuda.* — True plus your card's
name means torch can see the GPU.) That check alone isn't proof the app is
using it: Settings → System shows the device OmniVoice actually resolved.
If it reads cpu while the command above prints True, the backend log line
starting Falling back to CPU: names the architecture mismatch it hit.
Docker Compose (recommended)
# CPU
docker compose -f deploy/docker-compose.yml --profile cpu up -d
# NVIDIA GPU
docker compose -f deploy/docker-compose.yml --profile gpu up -d
# AMD GPU (ROCm)
docker compose -f deploy/docker-compose.yml --profile rocm up -d
The docker-compose.yml shipped in deploy/ defaults to 127.0.0.1:3900
on the host. The backend inside the container binds to 0.0.0.0 so the
host port mapping can forward — the host-side 127.0.0.1 binding is what
enforces loopback-only.
LAN access
To expose OmniVoice on your LAN (e.g. you're running it on a homelab box and opening the UI from a laptop), change the host port mapping:
# deploy/docker-compose.yml
services:
omnivoice:
ports:
- "0.0.0.0:3900:3900" # ← was 127.0.0.1:3900:3900
The OmniVoice frontend defaults to the same origin the page was served
from, so opening the UI from http://<lan-ip>:3900 Just Works for both the
page load and the API/media requests it makes afterwards.
If you front the app with a reverse proxy and the API and UI land on
different origins, pin the API base explicitly. Use OMNIVOICE_PUBLIC_API_BASE
— a runtime env var the backend injects into the page, so it works with the
prebuilt image via docker run -e (the older VITE_OMNIVOICE_API is inlined at
build time and cannot be set on a prebuilt image):
docker run -e OMNIVOICE_PUBLIC_API_BASE=https://api.your-host.example \
-p 0.0.0.0:3900:3900 \
ghcr.io/debpalash/omnivoice-studio:latest
OMNIVOICE_PUBLIC_API_BASEmust be a plainhttp(s)://…URL; anything else is ignored and the app falls back to same-origin. If you build from source you may instead bakeVITE_OMNIVOICE_APIat build time, but the runtime var above is simpler and image-agnostic.
Security: OmniVoice ships no authentication. Anything on your LAN with the URL can use the app. Put it behind a reverse proxy with
basic_auth(Caddy / nginx + htpasswd) or a private network overlay (Tailscale, ZeroTier) before exposing publicly.
Volume mounts
Two paths are worth persisting across container restarts:
| Mount | Purpose | Why |
|---|---|---|
omnivoice_data:/app/omnivoice_data |
Project DB, user voices, settings | Survives upgrade; encrypted HF token lives here |
~/.cache/huggingface:/root/.cache/huggingface |
HF model cache | Re-using your host's cache saves ~2.4 GB of re-downloads |
Troubleshooting
- Container reports 0.2.7 but image is tagged 0.3.x: This was a workflow bug
(fixes #249, #251) — the
:latesttag was not being updated on release tag pushes. Pull the image again after the fix is merged:docker pull ghcr.io/debpalash/omnivoice-studio:latest. The running version is now shown in Settings → About → Version (read live from the backend), so the web UI no longer displays a dash in Docker. - Checking which version is running:
docker exec omnivoice python -c "import importlib.metadata; print(importlib.metadata.version('omnivoice'))", or hit the/healthendpoint — it returns{"status": "ok", "device": ..., "version": "0.3.x"}. - "Loopback origin required" errors (and a blank version): the desktop
build restricts the
/system/*and/api/settings/*routes to a loopback origin, but Docker's NAT makes every request look non-loopback, so the gate used to 403 the whole admin UI (issue #261). The image now ships withOMNIVOICE_SERVER_MODE=1, which relaxes that gate for the headless deployment — exposure is instead governed by your-pport mapping (keep the127.0.0.1:prefix to stay local) plus the optional share PIN. If you front the container with your own auth proxy on loopback, setOMNIVOICE_SERVER_MODE=0to re-enable the strict gate. - Media-preview 404 in LAN mode: see the LAN access section
above — the
window.location.hostfix shipped in v0.3. - GPU not detected (NVIDIA): verify
docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu22.04 nvidia-smisucceeds first. - GPU not detected (AMD): make sure you pulled the
:rocmtag (the default image is CUDA-only) and passed--device /dev/kfd --device /dev/dri. Check the container sees the card withdocker exec omnivoice rocminfo | grep -i gfx; on RDNA3 consumer cards try-e HSA_OVERRIDE_GFX_VERSION=11.0.0— see Pull and run (AMD GPU / ROCm) above. - More entries: docs/install/troubleshooting.md.