diff --git a/docs/downloading-models.md b/docs/downloading-models.md index 8aaac6fb..7c13f6a0 100644 --- a/docs/downloading-models.md +++ b/docs/downloading-models.md @@ -60,7 +60,7 @@ State is reported at **Settings → About** / `GET /system/info`: - `fast_download.xet_installed` — `hf_xet` present (true) - `fast_download.xet_active` — whether Xet actually drives downloads (false by default, because of `HF_HUB_DISABLE_XET`) -- the **⚡ fast download** badge in **Model Catalogue → Downloaded weights** appears only when Xet +- the **⚡ fast download** badge in **Model Catalogue → Downloaded weights** (under the TTS or ASR tab) appears only when Xet is *active*. The backend logs one line at startup, e.g. @@ -148,7 +148,7 @@ failed download at once. Caveats: ## Cancelling a download -**Model Catalogue → Downloaded weights** lets you cancel an in-flight install. Cancellation stops +**Model Catalogue → Downloaded weights** (under the engine's family tab) lets you cancel an in-flight install. Cancellation stops further retries and clears the failure cooldown so you can restart immediately. A file that's already streaming finishes first — cancellation takes effect at the next retry boundary. @@ -162,6 +162,6 @@ takes effect at the next retry boundary. High-performance mode only helps if RAM and bandwidth are plentiful. - **"download finished but no model weights were found"** — the download was interrupted and left a partial snapshot. Delete the model in - **Model Catalogue → Downloaded weights** and install it again. + **Model Catalogue → Downloaded weights** (under the engine's family tab) and install it again. - **Out of disk** — model sizes are shown in the catalog; free space or change the cache location with `HF_HOME` / `HF_HUB_CACHE`. diff --git a/docs/engines/audio-cpp.md b/docs/engines/audio-cpp.md index 32934937..008c407b 100644 --- a/docs/engines/audio-cpp.md +++ b/docs/engines/audio-cpp.md @@ -80,11 +80,11 @@ instead of that dedicated-VRAM floor. ``` Alternatively set `OMNIVOICE_AUDIOCPP_DIR` to the directory containing it. -3. Restart VoiceStudio, open **Model Catalogue → Downloaded weights**, find +3. Restart VoiceStudio, open **Model Catalogue (TTS tab) → Downloaded weights**, find **Breeze-TTS-2 Q8_0 for audio.cpp**, review its research/non-commercial license note, and click **Install**. Generation never starts this ~4.73 GiB download automatically. -4. Pick `audiocpp` in **Model Catalogue**. The server starts +4. Pick `audiocpp` in **Model Catalogue** (TTS tab → **Use**). The server starts lazily on first generate (`server.json` + `server.log` live under the app data `audiocpp/` directory). @@ -131,7 +131,7 @@ The managed loopback port may be taken. Check `server.log` next to ### `Breeze-TTS-2 ... not installed` or `package ... not completely installed` -Install the model from **Model Catalogue → Downloaded weights**. If an interrupted install +Install the model from **Model Catalogue (TTS tab) → Downloaded weights**. If an interrupted install left it incomplete, use **Reinstall** there. If the error persists after a complete reinstall, file an issue with the package listing. diff --git a/docs/engines/confucius4-tts.md b/docs/engines/confucius4-tts.md index 9b20875f..2df6ef0f 100644 --- a/docs/engines/confucius4-tts.md +++ b/docs/engines/confucius4-tts.md @@ -56,7 +56,7 @@ Then point VoiceStudio at the clone and restart: - **macOS/Linux:** `export OMNIVOICE_CONFUCIUS4_TTS_DIR=/path/to/Confucius4-TTS` - **Windows (PowerShell):** `[Environment]::SetEnvironmentVariable("OMNIVOICE_CONFUCIUS4_TTS_DIR","C:\path\to\Confucius4-TTS","User")` -Select **Confucius4-TTS** in Model Catalogue. The first synthesize triggers +Select **Confucius4-TTS** in Model Catalogue (TTS tab → **Use**). The first synthesize triggers the weight downloads above, then generates. ### Optional overrides diff --git a/docs/engines/gpt-sovits.md b/docs/engines/gpt-sovits.md index 72df3139..886f4d56 100644 --- a/docs/engines/gpt-sovits.md +++ b/docs/engines/gpt-sovits.md @@ -24,7 +24,7 @@ HTTP. python api_v2.py -a 127.0.0.1 -p 9880 -c GPT_SoVITS/configs/tts_infer.yaml ``` -2. Select the engine via **Model Catalogue** or +2. Select the engine via **Model Catalogue** (TTS tab → **Use**) or `OMNIVOICE_TTS_BACKEND=gpt-sovits`. VoiceStudio marks the engine available only when the server responds diff --git a/docs/engines/kittentts.md b/docs/engines/kittentts.md index 1e5631a9..f7fa284c 100644 --- a/docs/engines/kittentts.md +++ b/docs/engines/kittentts.md @@ -20,7 +20,7 @@ only — but a much faster and much smaller install. pip install kittentts ``` -Then select the engine via **Model Catalogue** or +Then select the engine via **Model Catalogue** (TTS tab → **Use**) or `OMNIVOICE_TTS_BACKEND=kittentts`. ## Voices diff --git a/docs/engines/mlx-audio.md b/docs/engines/mlx-audio.md index a5b4a455..da85bb56 100644 --- a/docs/engines/mlx-audio.md +++ b/docs/engines/mlx-audio.md @@ -19,7 +19,7 @@ platforms never reports as available pip install mlx-audio ``` -Then select the engine via **Model Catalogue** or +Then select the engine via **Model Catalogue** (TTS tab → **Use**) or `OMNIVOICE_TTS_BACKEND=mlx-audio`. ## Model selection diff --git a/docs/engines/moss-tts-nano.md b/docs/engines/moss-tts-nano.md index 3c3f884b..b32ddf9a 100644 --- a/docs/engines/moss-tts-nano.md +++ b/docs/engines/moss-tts-nano.md @@ -25,7 +25,7 @@ cd MOSS-TTS-Nano uv pip install -e . ``` -Then select the engine via **Model Catalogue** or +Then select the engine via **Model Catalogue** (TTS tab → **Use**) or `OMNIVOICE_TTS_BACKEND=moss-tts-nano`. ## Model selection diff --git a/docs/engines/omnivoice-gguf.md b/docs/engines/omnivoice-gguf.md index bf5ce7b8..6192a68b 100644 --- a/docs/engines/omnivoice-gguf.md +++ b/docs/engines/omnivoice-gguf.md @@ -32,10 +32,10 @@ against the same table (an F32 reference quant, ~3.2 GB, is override-only). ## Setup Nothing to install: installer and CI builds bundle the binary for your -platform. Select the engine via **Model Catalogue** or +platform. Select the engine via **Model Catalogue** (TTS tab → **Use**) or `OMNIVOICE_TTS_BACKEND=omnivoice-gguf`. The quant weights download on first use (see [downloading-models.md](../downloading-models.md)) — install them -ahead of time from **Model Catalogue → Downloaded weights** if you want the first +ahead of time from **Model Catalogue (TTS tab) → Downloaded weights** if you want the first generation to be quick; a long first render is the download, not a hang. **Source checkouts:** the repo ships zero-byte placeholders in `bin/` — real diff --git a/docs/engines/omnivoice.md b/docs/engines/omnivoice.md index 8e5bee58..3057dcc1 100644 --- a/docs/engines/omnivoice.md +++ b/docs/engines/omnivoice.md @@ -103,7 +103,7 @@ The env var overrides the persisted UI choice. above — switch to OmniVoice GGUF or close other GPU apps. - First generation is slow: the first call downloads multi-GB weights. To keep the first render quick, install the model ahead of time from - **Model Catalogue → Downloaded weights** — a long first generate is almost always the + **Model Catalogue (TTS tab) → Downloaded weights** — a long first generate is almost always the download, not a hang. - General install issues: [install/troubleshooting.md](../install/troubleshooting.md). diff --git a/docs/engines/parakeet-mlx.md b/docs/engines/parakeet-mlx.md index fca24e92..a17369ad 100644 --- a/docs/engines/parakeet-mlx.md +++ b/docs/engines/parakeet-mlx.md @@ -15,7 +15,7 @@ Apple Silicon source installs since 0.3.22**. - **Model Catalogue**, ASR tab → **Use** on the Parakeet TDT v3 (MLX) row, or `OMNIVOICE_ASR_BACKEND=parakeet-mlx`. - **Dictation prefers it automatically**: once the model weights are - installed (Model Catalogue → Downloaded weights — the auto-pick never triggers a + installed (Model Catalogue (ASR tab) → Downloaded weights — the auto-pick never triggers a download), live dictation/capture uses it whenever your system language is one of the 25 covered European languages. Other languages keep the multilingual Whisper engine, so dictation coverage never regresses. diff --git a/docs/engines/pockettts.md b/docs/engines/pockettts.md index 681b87e0..6bf401e3 100644 --- a/docs/engines/pockettts.md +++ b/docs/engines/pockettts.md @@ -35,7 +35,7 @@ for this model. You also need HuggingFace access to the gated repo (see [downloading-models.md](../downloading-models.md) for token setup). -3. Select the engine via **Model Catalogue** or +3. Select the engine via **Model Catalogue** (TTS tab → **Use**) or `OMNIVOICE_TTS_BACKEND=pockettts`. ## Platform notes diff --git a/docs/engines/sherpa-onnx-asr.md b/docs/engines/sherpa-onnx-asr.md index fbc74dab..80ad3802 100644 --- a/docs/engines/sherpa-onnx-asr.md +++ b/docs/engines/sherpa-onnx-asr.md @@ -10,7 +10,7 @@ partials either way. ## Selecting it - Ensure `sherpa-onnx` is installed (`uv add sherpa-onnx` on source installs). -- Pick a dictation model in the app (Model Catalogue → Downloaded weights lists the +- Pick a dictation model in the app (Model Catalogue (ASR tab) → Downloaded weights lists the selectable set below), or **Model Catalogue**, ASR tab → **Use**, or pin `OMNIVOICE_ASR_BACKEND=sherpa-onnx-asr`. - `OMNIVOICE_SHERPA_ASR_MODEL` selects the model — default diff --git a/docs/engines/sherpa-onnx.md b/docs/engines/sherpa-onnx.md index 5313b1d4..c2d4a843 100644 --- a/docs/engines/sherpa-onnx.md +++ b/docs/engines/sherpa-onnx.md @@ -30,7 +30,7 @@ TTS model directory. export OMNIVOICE_SHERPA_MODEL=/path/to/model-dir ``` -4. Select the engine via **Model Catalogue** or +4. Select the engine via **Model Catalogue** (TTS tab → **Use**) or `OMNIVOICE_TTS_BACKEND=sherpa-onnx`. The directory must contain `model.onnx` and `tokens.txt`. Sherpa-ONNX ships diff --git a/docs/engines/supertonic3.md b/docs/engines/supertonic3.md index 27f0a3c9..58bb1445 100644 --- a/docs/engines/supertonic3.md +++ b/docs/engines/supertonic3.md @@ -28,7 +28,7 @@ crashes and cold init never block the rest of VoiceStudio. unavailable until you review and accept in **Model Catalogue → Supertonic-3**. -3. Select the engine via **Model Catalogue** or +3. Select the engine via **Model Catalogue** (TTS tab → **Use**) or `OMNIVOICE_TTS_BACKEND=supertonic3`. The first synthesis cold-downloads ~400 MB of model weights, pinned to an diff --git a/docs/install/macos.md b/docs/install/macos.md index fc00f950..e68d70ea 100644 --- a/docs/install/macos.md +++ b/docs/install/macos.md @@ -157,7 +157,7 @@ without the quarantine step. - **Apple Silicon (M-series):** VoiceStudio automatically picks the `mlx-whisper` and `mlx-audio` backends where available — these use the Apple Neural Engine and Metal Performance Shaders for ~2× the throughput of the CPU path. - Installing the **Parakeet TDT v3 (MLX)** model from **Model Catalogue → Downloaded weights** + Installing the **Parakeet TDT v3 (MLX)** model from **Model Catalogue** (ASR tab → **Downloaded weights**) additionally makes dictation/capture prefer the `parakeet-mlx` engine (25 European languages, word timestamps, ~2 GB unified memory) — it is never downloaded without that explicit install, and it is only auto-preferred when diff --git a/docs/install/troubleshooting.md b/docs/install/troubleshooting.md index 29dab25d..2720b125 100644 --- a/docs/install/troubleshooting.md +++ b/docs/install/troubleshooting.md @@ -89,7 +89,7 @@ uv pip install --reinstall transformers ``` Or, as a quick workaround, switch ASR to **faster-whisper** in -**Model Catalogue → Downloaded weights**. If it recurs, add the backend **`.venv`** to your +**Model Catalogue** (ASR tab → **Use**). If it recurs, add the backend **`.venv`** to your antivirus exclusions (see §1). Newer builds classify this error and show the reinstall hint directly instead of a bare path + "try restarting". @@ -408,7 +408,7 @@ Intel-Mac wheels, so this entry only applies to historical installs (see ## 10. Windows: `Could not locate cudnn_ops_infer64_8.dll` during transcription **Symptom:** on Windows + NVIDIA, transcription/dubbing fails and the backend -log shows `Could not locate cudnn_ops_infer64_8.dll`. Model Catalogue → Downloaded weights shows +log shows `Could not locate cudnn_ops_infer64_8.dll`. Model Catalogue (ASR tab) shows WhisperX or faster-whisper selected. On builds before this was fixed, the failure looked much worse than a failed @@ -444,7 +444,7 @@ uv pip install --no-deps --python .venv\Scripts\python.exe --target .venv\Lib\si (On Linux the target is `.venv/lib/pythonX.Y/site-packages/cudnn8_compat`.) Or sidestep cuDNN 8 entirely: switch the ASR backend to **PyTorch Whisper** in -**Model Catalogue → Downloaded weights**. It runs on PyTorch's own stack (cuDNN 9, bundled with +**Model Catalogue** (ASR tab → **Use**). It runs on PyTorch's own stack (cuDNN 9, bundled with torch) and needs no cuDNN-8 DLL — it loads its Whisper pipeline on demand (no extra env var). @@ -600,12 +600,12 @@ did was `generate:start (audio)`, a dub, or a dictation. **Fix — reduce ASR load (any one of these):** -1. **Pick a smaller ASR model / engine** in **Model Catalogue → Downloaded weights** — e.g. +1. **Pick a smaller ASR model / engine** in **Model Catalogue** (ASR tab: **Use** an engine, then a smaller model under **Downloaded weights**) — e.g. faster-whisper **medium** or **small**, instead of large-v3. Biggest win on low-VRAM GPUs. 2. **Free VRAM**: **Flush the TTS model** before dubbing so ASR isn't competing for memory (top toolbar → Flush → "Unload all + flush", or per-model from - Model Catalogue → Downloaded weights — see [Flush caches / Unload resident model](../performance.md#flush-caches--unload-resident-model) + Model Catalogue → Downloaded weights under the engine's family tab — see [Flush caches / Unload resident model](../performance.md#flush-caches--unload-resident-model) for exactly what it frees and the API equivalents for scripts), or 3. **Run ASR on CPU** (slower but reliable) if your GPU is small. 4. **Test with a 10-second clip** first — if that returns quickly, it confirms a diff --git a/docs/performance.md b/docs/performance.md index 2c41e8aa..478ea84e 100644 --- a/docs/performance.md +++ b/docs/performance.md @@ -183,7 +183,7 @@ drain, or restart the backend, and then Flush. - **Unload all + flush** — the above **plus** fully unloads the resident TTS model. Frees the most memory; the next generation pays the ~8 s reload. -- **Model Catalogue → Downloaded weights** — rows whose weights are resident right now show an +- **Model Catalogue → Downloaded weights** (under the engine's family tab) — rows whose weights are resident right now show an "In memory" badge with the same per-model **Unload** button. **From a script** (the local API on port 3900), the same operations: diff --git a/frontend/src/components/catalogue/SetupSummary.jsx b/frontend/src/components/catalogue/SetupSummary.jsx index dabce062..97e6717e 100644 --- a/frontend/src/components/catalogue/SetupSummary.jsx +++ b/frontend/src/components/catalogue/SetupSummary.jsx @@ -39,6 +39,8 @@ export function summarizeFamily(family, familyData) { const name = entry?.display_name || active || null; if (!entry || entry.available === false) return { name, device: null, status: 'setup' }; const routing = entry.routing_status; + // Installed but with no usable device path (select would be refused too). + if (routing === 'unavailable') return { name, device: null, status: 'setup' }; const device = routing === 'n/a' ? 'remote' : DEVICE_LABEL[entry.effective_device] || entry.effective_device; return { @@ -128,13 +130,24 @@ export default function SetupSummary({ onChange }) { const installRest = async () => { if (missing.length === 0) return; setInstalling(true); - try { - await Promise.all(missing.map((m) => installMutation.mutateAsync(m.repo_id))); - toast.success(t('models.started_downloading', { count: missing.length })); - } catch (e) { - toast.error(t('models.install_failed', { message: e?.message || e })); - } finally { - setInstalling(false); + // Every request settles before the button re-enables, so a single early + // rejection cannot re-arm the action while sibling installs are still in + // flight (which would let the same repo be requested twice). + const results = await Promise.allSettled( + missing.map((m) => installMutation.mutateAsync(m.repo_id)), + ); + setInstalling(false); + const failed = results + .map((r, i) => (r.status === 'rejected' ? { repo: missing[i].repo_id, err: r.reason } : null)) + .filter(Boolean); + const started = results.length - failed.length; + if (started > 0) toast.success(t('models.started_downloading', { count: started })); + if (failed.length > 0) { + toast.error( + t('models.install_failed', { + message: failed.map((f) => `${f.repo}: ${f.err?.message || f.err}`).join(' · '), + }), + ); } }; @@ -158,7 +171,29 @@ export default function SetupSummary({ onChange }) { {row.label}