The About page reported 'Compute device: cpu / GPU active: no / VRAM 0.00 GB' on a machine with a working GPU. Every line was true and none was usable — it is also exactly what a machine with no GPU at all reports, so the report could not distinguish a driver that isn't loaded from a container that cannot open the device from a ROCm older than the card. The probe already knew all of it; torch.cuda.is_available() returning False simply produced no note. Each cause now reads differently: a missing device node names the --device flags, a permissions failure names --group-add and how to find the host's real render/video GIDs (copied numbers are the most common way this ends up on CPU in Docker), a card newer than the shipped ROCm points at rocminfo, an HSA_OVERRIDE_GFX_VERSION that is doing more harm than good is named first because it is both likelier and cheaper to test, and an unreachable NVIDIA driver gets its own advice. The probe never diagnoses from a measurement it did not complete: when torch.cuda.is_available() itself raises, the exception is reported and no device findings are asserted beside it. Metadata access that raises is contained too — this runs on the path whose whole job is to explain a failure, so it cannot become one.
11 KiB
VoiceStudio — Install with Docker
For headless servers, dedicated GPUs, or "I want one command" deployments. The docker image bundles the backend; the UI is served over HTTP and you open it in a normal browser.
Official images: ghcr.io/debpalash/omnivoice-studio
and palashdeb/omnivoice-studio on Docker Hub — same images, same tags.
Image ↔ version mapping
Tag What you get :latestRolling preview — latest commit on main, at or ahead of the last release. This is the preview channel; pin:stablefor production.:stableMost recent versioned release (updated on every v*git tag):0.4.1Exact release version :0.4Latest patch within the 0.4 minor :mainAlias of the same rolling mainbuild as:latest:sha-xxxxxxxSpecific commit (produced by manual workflow dispatch) :rocmAMD GPU (ROCm) build of the rolling preview — the ROCm analogue of :latest:stable-rocm,:0.4.1-rocm,:0.4-rocm,:sha-xxxxxxx-rocmROCm builds of the corresponding CUDA tags above Versioning rule: preview builds always come from
mainand never version-sort below:stable— upgrades flow naturally.Note on the update-channel toggle: The update-channel UI (Settings → About → Update channel) is part of the Tauri desktop app's built-in auto-updater. It does not apply to the Docker image — the Docker image is the headless web-server build. To update your Docker deployment, pull the new image tag and recreate the container (
docker compose pull && docker compose up -d).
Pull and run (CPU)
docker pull ghcr.io/debpalash/omnivoice-studio:latest
docker run -d --name omnivoice \
-p 127.0.0.1:3900:3900 \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:latest
Docker Hub mirror: the same images are published to
palashdeb/omnivoice-studioon Docker Hub with identical tags — swap the image forpalashdeb/omnivoice-studio:latestif you prefer Docker Hub. Tag semantics (:latest= rolling main preview,:stable/:X.Y.Z= releases) are the same on both registries.
Open http://localhost:3900. The first run downloads
~2.4 GB of model weights — follow docker logs -f omnivoice to watch.
Pull and run (NVIDIA GPU)
docker run -d --name omnivoice --gpus all \
-p 127.0.0.1:3900:3900 \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:latest
GPU mode requires the NVIDIA Container Toolkit on the host.
Pull and run (AMD GPU / ROCm)
AMD GPUs use the dedicated :rocm image variant — the default (CUDA)
image runs CPU-only on AMD hardware. The ROCm userspace ships inside the
image; the host only needs the amdgpu kernel driver. Pass the GPU through
as plain device nodes (no container toolkit needed):
docker run -d --name omnivoice \
--device /dev/kfd --device /dev/dri \
-p 127.0.0.1:3900:3900 \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:rocm
The same flags work with Podman (podman run --device /dev/kfd --device /dev/dri …); in a Quadlet unit that's two AddDevice= lines:
# ~/.config/containers/systemd/omnivoice.container
[Container]
Image=ghcr.io/debpalash/omnivoice-studio:rocm
AddDevice=/dev/kfd
AddDevice=/dev/dri
PublishPort=127.0.0.1:3900:3900
Volume=omnivoice-data:/app/omnivoice_data
Release pins exist too: :stable-rocm, :0.4.1-rocm, :0.4-rocm mirror
the CUDA tags exactly.
Consumer cards and APUs (RX 6000/7000, Strix Point/Halo): the backend auto-sets
HSA_OVERRIDE_GFX_VERSIONwhen — and only when — your card's GFX ID is missing from the shipped ROCm build's architecture list, so try without any override first. Overriding a natively-supported GPU (gfx1151 on ROCm 7.x, for example) only forces it onto foreign kernels. If the GPU still isn't used, force it explicitly with-e HSA_OVERRIDE_GFX_VERSION=11.0.0(user-set on the container — it is deliberately not baked into the image, because the right value depends on your card); a value you set is always respected as-is.Rootless / non-root hosts: if
/dev/kfdis group-owned, the container user needs those groups too — add--group-addfor your host'srenderandvideoGIDs (getent group render video).
Verify the container sees the GPU:
docker exec omnivoice python3 -c \
"import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"
(ROCm-built PyTorch reports through torch.cuda.* — True plus your card's
name means torch can see the GPU.) That check alone isn't proof the app is
using it: Settings → System shows the device VoiceStudio actually resolved.
If it reads cpu while the command above prints True, the backend log line
starting Falling back to CPU: names the architecture mismatch it hit.
If the command prints False, Settings → System now says why, and the
three answers need different fixes:
| What it says | What to do |
|---|---|
/dev/kfd is not present |
The container was started without --device /dev/kfd --device /dev/dri, or the host's amdgpu driver isn't loaded. |
this process cannot open it |
A group problem. Run ls -l /dev/kfd /dev/dri/render* on the host, and pass those GIDs with --group-add. The numbers differ between machines — a --group-add 39 copied from someone else's command grants nothing. |
no GPU was enumerated |
The device nodes are fine and the runtime still found nothing — usually a card newer than the image's ROCm. Check rocminfo on the host, and see the HSA_OVERRIDE_GFX_VERSION note above. |
Docker Compose (recommended)
# CPU
docker compose -f deploy/docker-compose.yml --profile cpu up -d
# NVIDIA GPU
docker compose -f deploy/docker-compose.yml --profile gpu up -d
# AMD GPU (ROCm)
docker compose -f deploy/docker-compose.yml --profile rocm up -d
The docker-compose.yml shipped in deploy/ defaults to 127.0.0.1:3900
on the host. The backend inside the container binds to 0.0.0.0 so the
host port mapping can forward — the host-side 127.0.0.1 binding is what
enforces loopback-only.
LAN access
To expose VoiceStudio on your LAN (e.g. you're running it on a homelab box and opening the UI from a laptop), change the host port mapping:
# deploy/docker-compose.yml
services:
omnivoice:
ports:
- "0.0.0.0:3900:3900" # ← was 127.0.0.1:3900:3900
The VoiceStudio frontend defaults to the same origin the page was served
from, so opening the UI from http://<lan-ip>:3900 Just Works for both the
page load and the API/media requests it makes afterwards.
If you front the app with a reverse proxy and the API and UI land on
different origins, pin the API base explicitly. Use OMNIVOICE_PUBLIC_API_BASE
— a runtime env var the backend injects into the page, so it works with the
prebuilt image via docker run -e (the older VITE_OMNIVOICE_API is inlined at
build time and cannot be set on a prebuilt image):
docker run -e OMNIVOICE_PUBLIC_API_BASE=https://api.your-host.example \
-p 0.0.0.0:3900:3900 \
ghcr.io/debpalash/omnivoice-studio:latest
OMNIVOICE_PUBLIC_API_BASEmust be a plainhttp(s)://…URL; anything else is ignored and the app falls back to same-origin. If you build from source you may instead bakeVITE_OMNIVOICE_APIat build time, but the runtime var above is simpler and image-agnostic.
Security: VoiceStudio ships no authentication. Anything on your LAN with the URL can use the app. Put it behind a reverse proxy with
basic_auth(Caddy / nginx + htpasswd) or a private network overlay (Tailscale, ZeroTier) before exposing publicly.
Volume mounts
Two paths are worth persisting across container restarts:
| Mount | Purpose | Why |
|---|---|---|
omnivoice_data:/app/omnivoice_data |
Project DB, user voices, settings | Survives upgrade; encrypted HF token lives here |
~/.cache/huggingface:/root/.cache/huggingface |
HF model cache | Re-using your host's cache saves ~2.4 GB of re-downloads |
Troubleshooting
- Container reports 0.2.7 but image is tagged 0.3.x: This was a workflow bug
(fixes #249, #251) — the
:latesttag was not being updated on release tag pushes. Pull the image again after the fix is merged:docker pull ghcr.io/debpalash/omnivoice-studio:latest. The running version is now shown in Settings → About → Version (read live from the backend), so the web UI no longer displays a dash in Docker. - Checking which version is running:
docker exec omnivoice python -c "import importlib.metadata; print(importlib.metadata.version('omnivoice'))", or hit the/healthendpoint — it returns{"status": "ok", "device": ..., "version": "0.3.x"}. - "Loopback origin required" errors (and a blank version): the desktop
build restricts the
/system/*and/api/settings/*routes to a loopback origin, but Docker's NAT makes every request look non-loopback, so the gate used to 403 the whole admin UI (issue #261). The image now ships withOMNIVOICE_SERVER_MODE=1, which relaxes that gate for the headless deployment — exposure is instead governed by your-pport mapping (keep the127.0.0.1:prefix to stay local) plus the optional share PIN. If you front the container with your own auth proxy on loopback, setOMNIVOICE_SERVER_MODE=0to re-enable the strict gate. - Media-preview 404 in LAN mode: see the LAN access section
above — the
window.location.hostfix shipped in v0.3. - GPU not detected (NVIDIA): verify
docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu22.04 nvidia-smisucceeds first. - GPU not detected (AMD): make sure you pulled the
:rocmtag (the default image is CUDA-only) and passed--device /dev/kfd --device /dev/dri. Check the container sees the card withdocker exec omnivoice rocminfo | grep -i gfx. On consumer cards, run without anyHSA_OVERRIDE_GFX_VERSIONfirst — the backend sets it itself when your card needs it, and overriding a natively-supported GPU only forces it onto foreign kernels. See Pull and run (AMD GPU / ROCm) above for when to set one by hand. - More entries: docs/install/troubleshooting.md.