Compare commits

...
Author SHA1 Message Date
Palash Debnath 135ccd09ec Merge pull request #2158 from debpalash/fix/release-review-hardening
Harden Electron release authorization and playback continuity
2026-09-16 14:51:56 -07:00
Palash Debnath 7114315e04 Check both signing credentials stay scoped to packaging 2026-09-17 03:03:04 +05:30
Palash Debnath 3d7c7710a1 Hide unavailable remote GPU metrics and use utilization fallback 2026-09-17 02:43:32 +05:30
Palash Debnath 1fea27543b Harden release authorization and preserve demo playback across languages 2026-09-17 02:38:09 +05:30
Palash Debnath c0a17f4272 Merge pull request #2157 from debpalash/fix/linux-electron-setup-sidebar
Make Electron primary and integrate desktop and README fixes
2026-09-16 13:44:29 -07:00
Palash Debnath 6cb1200b4e Merge branch 'velixio_dev' into fix/linux-electron-setup-sidebar 2026-09-17 01:55:33 +05:30
Palash Debnath 50d01567d4 Keep read-only telemetry outside worker drain and shutdown gates 2026-09-17 01:55:23 +05:30
Palash Debnath a91193d29d Merge branch 'review/gpu-2149' into fix/linux-electron-setup-sidebar 2026-09-17 01:47:57 +05:30
Palash Debnath 38d588fe93 Verify eager fallback remains usable on all desktop platforms 2026-09-17 01:47:52 +05:30
Palash Debnath 6896a19f91 Apply required frontend formatting to existing Electron transition changes 2026-09-17 01:47:18 +05:30
Palash Debnath 9fdb2b2fc9 Derive branding version checks from the canonical app version 2026-09-17 01:46:10 +05:30
Palash Debnath 835a2a9964 Gate Electron publication on signing checks or an explicit unsigned exception 2026-09-17 01:43:52 +05:30
Palash Debnath 03460c89b7 Merge branch 'velixio_dev' into fix/linux-electron-setup-sidebar 2026-09-17 01:42:08 +05:30
Palash Debnath 2ee455cc1f Prevent telemetry admission while the worker is retiring 2026-09-17 01:41:59 +05:30
Palash Debnath 715bbbde70 Merge branch 'velixio_dev' into fix/linux-electron-setup-sidebar 2026-09-17 01:41:14 +05:30
Palash Debnath 9ca9f74255 Merge branch 'review/gpu-2149' into fix/linux-electron-setup-sidebar 2026-09-17 01:41:08 +05:30
Palash Debnath b67aed2794 Resolve GPU stability review privacy, localization and test isolation findings 2026-09-17 01:40:58 +05:30
Palash Debnath 8804ddac9a Own telemetry probes through shutdown and localize Electron metrics 2026-09-17 01:40:58 +05:30
Palash Debnath 2faaafa9fc Merge remote-tracking branch 'origin/main' into review/gpu-2149 2026-09-17 01:39:08 +05:30
Palash Debnath 002601072e Merge branch 'review/readme-2153' into fix/linux-electron-setup-sidebar
# Conflicts:
#	CHANGELOG.md
2026-09-17 01:36:38 +05:30
Palash Debnath 3709517bf6 Credit the README transition contribution 2026-09-17 01:36:22 +05:30
Palash Debnath 8b5b44d5ea Validate manual Tauri draft assets without premature publication 2026-09-17 01:34:52 +05:30
Palash Debnath 016670c6ad Scope legacy launcher checks to Tauri and protect Electron backend reuse 2026-09-17 01:33:33 +05:30
Palash Debnath 39d66d5299 Prepare v0.5.3 for Electron and the final Tauri update 2026-09-17 01:30:17 +05:30
Palash Debnath 18e54e2352 Sequence the Electron transition after the final Tauri draft 2026-09-17 01:29:01 +05:30
Palash Debnath 1616f9dcab Verify compact README catalogs remain linked and complete 2026-09-17 01:24:47 +05:30
Palash Debnath 623d00eebf Document the integrated Electron transition and CI repairs 2026-09-17 01:24:31 +05:30
Palash Debnath 867cd3c4ac Merge branch 'fix/dubbing-demo-editor' into fix/linux-electron-setup-sidebar 2026-09-17 01:23:57 +05:30
Palash Debnath 0a556e78ba Cover pause and sample-switch playhead continuity 2026-09-17 01:23:53 +05:30
velixio 719ef8a698 fix(workers): harden remote telemetry 2026-09-17 01:23:37 +05:30
Palash Debnath d3f2d06a26 Merge branch 'fix/macos-sidebar-titlebar' into fix/linux-electron-setup-sidebar 2026-09-17 01:22:42 +05:30
Palash Debnath cab5766d28 Merge branch 'review/readme-2153' into fix/linux-electron-setup-sidebar
# Conflicts:
#	README.md
2026-09-17 01:22:42 +05:30
Palash Debnath 7098c3f809 Test offline notification state without hiding desktop updates 2026-09-17 01:22:23 +05:30
Palash Debnath ba954645b3 Clarify Electron transition without blocking desktop contributions 2026-09-17 01:22:06 +05:30
Palash Debnath 9af0fcbd89 Merge branch 'fix/macos-sidebar-titlebar' into fix/linux-electron-setup-sidebar
# Conflicts:
#	CHANGELOG.md
2026-09-17 01:21:38 +05:30
Palash Debnath fc2d94d266 Merge main and resolve sidebar and disabled notification review findings 2026-09-17 01:21:27 +05:30
Palash Debnath a06ef166aa Merge branch 'fix/dubbing-demo-editor' into fix/linux-electron-setup-sidebar
# Conflicts:
#	CHANGELOG.md
2026-09-17 01:21:27 +05:30
Palash Debnath 1f978ef3ee Merge main and preserve A/B playback position when switching samples 2026-09-17 01:20:22 +05:30
Palash Debnath 853331f3ec Merge branch 'feat/tauri-native-parity' into fix/linux-electron-setup-sidebar
# Conflicts:
#	frontend/src-tauri/src/watch_folder_core.rs
#	frontend/src-tauri/src/wayland_shortcut_core.rs
2026-09-17 01:19:59 +05:30
Palash Debnath fe35bf6737 Merge main and resolve native bridge review findings 2026-09-17 01:19:40 +05:30
Palash Debnath 22637a7954 Align CI fixtures with safe dubbing and compact Electron documentation 2026-09-17 01:18:48 +05:30
Palash Debnath 2b0e43da05 Refresh and brand VoiceStudio installable agent skills 2026-09-17 01:10:37 +05:30
velixio 10d086359c feat(workers): show remote telemetry 2026-09-17 01:08:38 +05:30
Palash Debnath fd9a9aa3df Improve VoiceStudio Original dark theme accent contrast 2026-09-16 22:17:00 +05:30
Palash Debnath 5af8fe6967 Make Electron the default desktop and refresh setup documentation 2026-09-16 22:11:59 +05:30
Palash Debnath 559f3fc0c1 Respect Electron matrix architecture for macOS packaging 2026-09-16 21:47:27 +05:30
Palash Debnath 4ce42decdd Keep Linux Electron installer names consistent with updater feeds 2026-09-16 21:45:15 +05:30
Palash Debnath 7af1181fc2 fix(desktop): complete shared native bridge contracts 2026-09-16 21:36:14 +05:30
cyberspace-cs 30cda178de docs: add Electron rewrite status note to README
Clarify that the desktop shell is being rebuilt and desktop-app-related
issues/PRs should be avoided until the rewrite ships.
2026-09-17 00:04:58 +08:00
Palash Debnath 3470f5b2d8 Make runtime tests honor macOS paths and CUDA availability 2026-09-16 21:32:11 +05:30
Palash Debnath 28d29f7c71 Handle detached navigation failures in integration and support screens 2026-09-16 21:28:15 +05:30
Palash Debnath 0b8de98783 Prepare Electron desktop releases and final Tauri sunset workflow 2026-09-16 21:24:27 +05:30
Palash Debnath 8a1e6eb058 Merge remote-tracking branch 'origin/main' into fix/linux-electron-setup-sidebar
# Conflicts:
#	README.md
2026-09-16 17:27:32 +05:30
Palash Debnath f21fab4e03 Polish Electron navigation and theme, refresh README and agent skills 2026-09-16 17:25:00 +05:30
Palash Debnath ae82468ec3 Update README.md 2026-09-16 17:23:24 +05:30
Shivendra-CoherentandClaude Opus 5 64f74e0dd2 fix: keep the backend alive on pre-Ampere NVIDIA GPUs (#2135)
On a Tesla T4 the backend exited during the first /generate with no
traceback and no HTTP response, leaving the client with
RemoteDisconnected and every later call with ConnectionRefused. Three
separate defects combined, which is why none of the reporter's
workarounds helped.

1. torch.compile(mode="reduce-overhead") captures CUDA graphs. T4
   (sm_75) passed the existing arch gate, so capture was attempted and
   aborted the process from inside the native CUDA library — below the
   interpreter, where neither the #278 eager-fallback wrapper nor any
   except clause can see it. The compile mode is now resolved per GPU:
   Ampere (sm_80) and newer keep the cudagraph mode, older cards drop to
   the non-cudagraph "default" mode and keep their compiled Inductor
   kernels. Fails open on any probe error, so no GPU that works today
   loses the optimization. OMNIVOICE_FORCE_CUDAGRAPH=1 restores it.

2. should_torch_compile() never read TORCH_COMPILE_DISABLE. main.py sets
   it on win32, build_engine_env injected it into subprocesses, and
   docs/install/windows.md tells users to export it — but the in-process
   gate ignored it, so the reporter exported the documented variable and
   still got "torch.compile applied". The gate now honours
   TORCH_COMPILE_DISABLE / TORCHDYNAMO_DISABLE / TORCHINDUCTOR_DISABLE on
   every platform, and an env opt-out on the parent propagates to engine
   subprocesses. The settings DB path is logged alongside the toggle:
   the reporter had three omnivoice.db files and edited one the backend
   never opened.

3. Settings -> Performance -> "Disable torch.compile" was rendered
   disabled outside Windows in both the Tauri and Electron UIs, so the
   one control that would have stopped this was unreachable for the
   affected Linux user. The toggle is now live on every platform, and
   build_engine_env honours it everywhere rather than only on win32.

Also arms faulthandler before torch is imported, so a fatal native
signal writes the faulting thread's Python stack to backend_err.log
instead of the process vanishing silently. This does not prevent a
crash; it makes one diagnosable. OMNIVOICE_DISABLE_FAULTHANDLER=1 skips
it.

Tests fail before / pass after, verified by stashing the source and
running the new tests against unfixed code. The crash test kills a real
child interpreter with a real SIGSEGV and requires a named Python frame
in the output. test_torch_compile_path_gate's fixture now clears the
compile-disable env vars: main.py setdefaults them on win32, so on a
Windows runner they leaked into os.environ and decided those tests.

Not verified on real hardware — no Turing GPU available. The sm_80 floor
is inferred from the crash report and from docs/hardware-notes-tesla-t4.md,
which already flagged cudagraphs on T4 as attempted by default and never
evaluated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 17:20:09 +05:30
Palash Debnath a966ae20b6 docs: restore trending badge 2026-09-16 13:15:18 +05:30
Palash Debnath 3f1f9bf9d0 docs: refresh Electron showcase and packaging prep 2026-09-16 13:09:32 +05:30
Palash Debnath 1d5d3d873e feat: refresh support and integrations experience 2026-09-16 12:38:16 +05:30
Palash Debnath 61d52944f2 fix(dub): correct long timeline rendering and repeated ASR context 2026-09-15 20:02:18 +05:30
Palash Debnath 36ccec5e61 feat(dub): show translation output and live logs in shared agent footer 2026-09-15 19:38:20 +05:30
Palash Debnath 730c7f377c feat(dub): save custom agent translation style instructions 2026-09-15 19:10:24 +05:30
Palash Debnath 7758ccc500 fix(dub): preserve original sound outside dialogue intervals 2026-09-15 19:02:25 +05:30
debpalash 49618ea893 chore: record React DOM type dependency in Bun lockfile 2026-09-15 17:55:00 +05:30
Palash Debnath 0a6ea976cb fix(dub): preserve complete speech and reject silent partial output 2026-09-15 15:44:57 +05:30
debpalash 2e7edb091d docs: note dubbing demo improvements 2026-09-15 15:36:28 +05:30
debpalash d47884ed14 fix(electron): improve dubbing demo playback and editing 2026-09-15 15:36:12 +05:30
debpalash 332d09e1a0 fix(electron): register notification hit region after titlebar drag regions 2026-09-15 15:25:03 +05:30
Palash Debnath c524ea3235 fix(electron): widen persistent secondary sidebars and fit video controls 2026-09-15 15:14:22 +05:30
Palash Debnath f137333abd fix(electron): queue early video playback and pass compositing launch flag 2026-09-15 15:06:26 +05:30
Palash Debnath 3569ac2d24 fix(electron): render video preview posters and document Linux compositing workaround 2026-09-15 14:59:58 +05:30
Palash Debnath d6ca24a6bb docs: link Linux onboarding and sidebar notes to PR 2129 2026-09-15 14:48:54 +05:30
Palash Debnath a4cfc49af2 fix(electron): polish Linux onboarding and workspace sidebar controls 2026-09-15 14:46:33 +05:30
debpalash 4c3766c0c9 docs: note macOS shell improvements in changelog 2026-09-15 12:13:39 +05:30
debpalash a0e3347639 fix(electron): refine macOS sidebar and titlebar controls 2026-09-15 12:13:12 +05:30
Palash Debnath 975799d4dd docs: note shared native desktop contracts 2026-09-14 21:03:24 -07:00
Palash Debnath bc8f563987 fix(desktop): complete shared native bridge contracts 2026-09-14 21:01:55 -07:00
Palash Debnath 4e55180f70 fix(electron): isolate packaged runtime setup 2026-09-14 11:25:07 -07:00
Palash Debnath dad59c1318 style: format shared frontend contracts 2026-09-14 11:05:42 -07:00
Palash Debnath 141b42f7a6 fix: restore desktop integration contracts 2026-09-14 10:53:55 -07:00
Palash Debnath f832d616f7 feat(electron): add full VoiceStudio desktop app 2026-09-14 10:22:23 -07:00
Palash Debnath eaf8bb9538 Update README.md 2026-09-10 22:50:20 -07:00
Palash Debnath 8932b09339 Merge pull request #2043 from debpalash/fix/mcp-timeout-follows-backend
fix(mcp): tools wait for the backend's own budget, not a fixed 120 s
2026-09-10 16:37:46 -07:00
Palash Debnath a35c7babdf docs(mcp): the first generation also downloads the model inside the budget 2026-09-10 16:16:01 -07:00
Palash Debnath bd9be744eb Merge remote-tracking branch 'origin/main' into fix/mcp-timeout-follows-backend
# Conflicts:
#	CHANGELOG.md
2026-09-10 16:15:59 -07:00
Palash Debnath 14778aaba4 Merge pull request #2044 from debpalash/fix/pytorch-whisper-vram-budget
fix(asr): PyTorch Whisper budgets the VRAM its model needs
2026-09-10 16:15:17 -07:00
Palash Debnath 5ca94b7d72 Merge remote-tracking branch 'origin/main' into fix/pytorch-whisper-vram-budget
# Conflicts:
#	CHANGELOG.md
2026-09-10 15:52:34 -07:00
Palash Debnath 7247cef1b6 Merge remote-tracking branch 'origin/main' into fix/mcp-timeout-follows-backend
# Conflicts:
#	CHANGELOG.md
2026-09-10 15:52:21 -07:00
Palash Debnath ff5bad5a92 Merge pull request #2042 from debpalash/fix/transcribe-m4a
fix(asr): PyTorch Whisper transcribes M4A through the ffmpeg fallback
2026-09-10 15:51:53 -07:00
Palash Debnath e7911e3e97 fix(asr): reduced VRAM budgets only for exact OpenAI Whisper ids
Review follow-up. The budget matched model names by substring, so a custom
repo whose name contains "turbo", "small" or "base" (or a word such as
"database") got a reduced budget and could be admitted to CUDA without
enough memory. Reduced budgets now apply only to the exact OpenAI
checkpoint ids, .en variants included. Any other repository, fine-tunes
included, keeps the conservative 5.0 GB, as the engine doc says.
2026-09-10 15:33:11 -07:00
Palash Debnath d37da79b34 fix(mcp): generation waits also cover the backend's queue clock
Review follow-up. A generation first waits in the GPU pool's queue, on its
own clock (GPU_QUEUE_TIMEOUT_S, 1800 s by default), before its execution
budget starts, so a 630 s wait could still give up before the backend
returned its queue error. The generate wait now adds the queue budget.
Transcription is unchanged: run_transcribe_guarded starts its 300 s clock
at submission, so queue time already counts against it.

Parity tests pin each tool's wait above the backend's own worst case,
from ASR_TRANSCRIBE_TIMEOUT_S, GPU_QUEUE_TIMEOUT_S and generate_timeout_s
on cpu, cuda and mps. A later backend change that outgrows the MCP wait
now fails a test. A new case covers a CPU budget larger than the GPU one.
2026-09-10 15:31:14 -07:00
Palash Debnath b4cfd7aac0 fix(asr): keep readable audio at its native rate; cover the M4A route
Review follow-up. The shared loader resamples soundfile-readable audio to
16 kHz with a linear interpolation, which can alias 44.1/48 kHz input into
Whisper's band. PyTorch Whisper now reads such audio at its native rate,
as before, so the pipeline's band-limited resampler converts it. Only
files soundfile can't open (MP4/M4A) take the ffmpeg path, which resamples
properly. New tests: a 48 kHz stereo WAV reaches the pipeline at 48 kHz
without ffmpeg, a POST /transcribe with an .m4a returns 200, and the MCP
tool names M4A bytes .m4a.
2026-09-10 15:30:23 -07:00
Palash Debnath e501a596d3 docs(changelog): note the PyTorch Whisper VRAM budget fix (#2044) 2026-09-10 15:06:52 -07:00
Palash Debnath 033f8f2f64 fix(asr): PyTorch Whisper budgets the VRAM its model needs
The CUDA preflight demanded 5.0 GB of free VRAM for every model, sized for
full large-v3. The default model is large-v3-turbo, which the pipeline
loads in fp16: about 1.6 GB of weights, not the 3.2 GiB fp32 figure in the
old comment. So a 6 GB card with nothing else resident reported 5.0 GB
free and was sent to CPU every time, although CUDA ran the same audio in
37 s against minutes on CPU.

The budget is now fp16 weights + 1.5 GB workspace (batch 16) + 0.5 GB
headroom, per model and capped at the old 5.0 GB. That is 3.6 GB for turbo
and 5.0 GB for full large-v3 and any unrecognised model. Both CPU-fallback
warnings name the model and the OMNIVOICE_ASR_VRAM_PREFLIGHT=0 opt-out.
The engine doc lists the budgets.

Fixes #2041
2026-09-10 15:06:45 -07:00
Palash Debnath 39c348af26 docs(changelog): note the MCP timeout fix (#2043) 2026-09-10 15:05:25 -07:00
Palash Debnath 58d590a686 fix(mcp): tools wait for the backend's own budget, not a fixed 120 s
The MCP tools gave up on a backend POST after OMNIVOICE_MCP_TIMEOUT_S,
default 120 s, while the backend's own budgets run longer: 300 s for ASR,
and 300 s or more for generation. A transcription the backend would have
finished came back as an empty client-side timeout, and the abandoned
job kept holding the worker. The docstring said the timeout followed
OMNIVOICE_GENERATE_TIMEOUT_S, but it never read it.

Unset, each tool now waits for the backend's budget plus 30 s, never less
than 120 s. transcribe follows OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S, and
generate_speech follows the larger of the GPU and CPU generation budgets,
scaled by text length like the backend. An explicit
OMNIVOICE_MCP_TIMEOUT_S still wins. docs/mcp.md says so.

Fixes #2040
2026-09-10 15:05:18 -07:00
Palash Debnath 6ff9ab80f8 docs(changelog): note the M4A transcribe fix (#2042) 2026-09-10 15:05:01 -07:00
Palash Debnath aef275a041 fix(asr): PyTorch Whisper transcribes M4A through the ffmpeg fallback
PyTorchWhisperBackend.transcribe read uploads with soundfile, and
libsndfile cannot open MP4/M4A (AAC). /transcribe and the MCP tool both
accept .m4a, so every such upload failed with "Format not recognised"
before ASR ran. It now uses _load_audio_16k_mono_f32, the loader the
sherpa backend already uses: soundfile first, then the validated ffmpeg
path, returning the 16 kHz mono float32 Whisper expects.

Fixes #2039
2026-09-10 15:04:53 -07:00
Palash Debnath 0525a078e0 Merge pull request #2035 from debpalash/fix/cosyvoice-advisory-floors
test(cosyvoice): raise the advisory floors to GitHub's fixes
2026-09-10 14:57:12 -07:00
Palash Debnath 083e7965f8 Merge pull request #2037 from debpalash/fix/sidecar-ready-diagnostics
fix(sidecar): a failed ready handshake names its cause
2026-09-10 14:37:51 -07:00
Palash Debnath df81bcd1ac fix(deps): raise protobuf and transformers bounds in pyproject too
The floors test checked only uv.lock, and pyproject still allowed
transformers>=5.5.0 and protobuf>=4.25. A fresh resolution, or an install
without the lock, could still pick a release below the advisory fixes.
The manifest now requires transformers>=5.10.0 and protobuf>=5.29.6. The
lock already resolved 5.15.1 and 7.36.0, so no package version changes;
only the recorded requirements do.
2026-09-10 14:35:31 -07:00
Palash Debnath aa98bb6974 Merge pull request #2038 from debpalash/fix/worker-control-teardown-flake
test(workers): give Control teardown waits a generous bound
2026-09-10 14:25:11 -07:00
Palash Debnath 1f386817cd Merge remote-tracking branch 'origin/main' into fix/sidecar-ready-diagnostics
# Conflicts:
#	CHANGELOG.md
2026-09-10 14:16:28 -07:00
Palash Debnath edc7199ab8 Merge pull request #2036 from debpalash/fix/youtube-bot-check-guidance
fix(dub): YouTube's bot check points at the app's cookie import
2026-09-10 14:15:26 -07:00
Palash Debnath 942561bbb3 test(sidecar): repair the literals in the new drain and error-frame tests 2026-09-10 14:12:04 -07:00
Palash Debnath 2e2b9fedd9 fix(sidecar): bind each stderr drain to its own process
The drain thread took its buffer from spawn but read self._proc when it
started, so a drain that started late could read a replacement process's
stderr into the old buffer. Spawn now passes the process and its buffer
together. Tests cover a late drain and the error-frame branch of the ready
diagnostics.
2026-09-10 14:09:20 -07:00
Palash Debnath 837baa922a Merge remote-tracking branch 'origin/main' into fix/cosyvoice-advisory-floors
# Conflicts:
#	CHANGELOG.md
2026-09-10 14:07:43 -07:00
Palash Debnath 0688c37f04 Merge pull request #2020 from debpalash/feat/engine-list-detail
feat(catalogue): engine list + detail panel, weights under their engine
2026-09-10 14:06:26 -07:00
Palash Debnath 973f44c0a5 docs(changelog): note the worker teardown flake fix (#2038) 2026-09-10 14:01:58 -07:00
Palash Debnath acae6bea4e test(workers): give Control teardown waits a generous bound
test_blocked_reconnect_persistence_does_not_stall_another_worker failed on
a Windows runner (#2020's Smoke job) when its final wait for the Control
stream to end expired at 2s. That wait only proves the stream ends: after
both releases, the real disconnect path persists all 64 queued tasks, and
a slow runner took longer. The assertions that matter, another worker's
heartbeat and prewarm staying under 0.2s while reconciliation is blocked,
are untouched.

The three waits that only bound a Control stream's exit now share
_CONTROL_EXIT_TIMEOUT_S = 10s. A stream that never ends still fails.
2026-09-10 14:01:50 -07:00
Palash Debnath fc9bd8c8d8 test(dub): import app modules at test time in the bot-check tests 2026-09-10 13:52:50 -07:00
Palash Debnath 9e4d67325f fix(sidecar): each spawn quotes only its own stderr
A shared deque cleared on spawn could still receive late lines from the
previous process's drain thread, and a new start-up failure would quote
them. Each spawn now gets a fresh buffer passed to its own drain thread.
2026-09-10 13:52:41 -07:00
Palash Debnath bb7eb06697 test(deps): raise the app's protobuf and transformers floors too
The app's own lock already resolves transformers 5.15.1 and protobuf
7.36.0, but PYTHON_FLOORS still allowed transformers 5.5.0 and did not
pin protobuf. A later lock change could have slipped below the fixes for
the advisories Dependabot raised on CosyVoice's manifest (#2030, #2031).
2026-09-10 13:49:32 -07:00
Palash Debnath 3a90be329f docs(changelog): note the sidecar ready-failure diagnostics (#2037) 2026-09-10 13:46:38 -07:00
Palash Debnath 6a923f5a10 fix(sidecar): a failed ready handshake names its cause
A sidecar that failed its ready handshake reported "did not signal ready:
None" for the two most likely causes. _recv returns None on EOF, and a
deadline kill closes stdout just like a crash, so the report could not
tell a slow start from a start-up crash. The exit code was never read,
and stderr went only to the log.

The error now names which it was:

- no ready frame within the deadline, so it was stopped;
- it exited with code N before signalling ready;
- it sent the wrong op, or reported an error frame (with its message).

Each ends with the sidecar's last stderr lines, scrubbed of home paths
and secrets and capped at 800 characters. The prefix is unchanged, so
existing matching still works. The echo sidecar gets test-mode hooks to
exit, stall or send the wrong op before ready.

Fixes #2026
2026-09-10 13:46:31 -07:00
Palash Debnath 8bad7a316b docs(changelog): note the YouTube bot-check guidance (#2036) 2026-09-10 13:44:29 -07:00
Palash Debnath 437593bfbb fix(dub): YouTube's bot check points at the app's cookie import
A YouTube link that hits "Sign in to confirm you're not a bot" showed
yt-dlp's raw advice to pass --cookies-from-browser or --cookies, which are
CLI flags nobody using VoiceStudio can pass. It now has its own failure
class, VIDEO_DOWNLOAD_BOT_CHECK. Its hint says to export signed-in cookies
as a cookies.txt file and attach it in Dub, or to upload the video
instead. The class is terminal, so it is neither retried as a network drop
nor escalated through player clients like a 403.

Fixes #2034
2026-09-10 13:44:21 -07:00
Palash Debnath 686d9c5fcf fix(catalogue): repair the guidance the Weights-list rename broke
The rename to "the engine's Weights list in Model Catalogue" left several
messages without a verb, and pointed others at the wrong place:

- The offline and create-voice messages say what to do again.
- pyannote has no owning engine, so diarization points at Other weights.
- The Hugging Face mirror moved to Settings → Network, and voice previews
  moved to Settings → Storage.
- Unloading and switching engines happen in the engine list, not a
  Weights list.
- A bad saved path points at Settings → Storage or the env file.
- Docstrings that read "the the" are fixed.

The dub stream-drop fallback goes through i18n in all 21 locales. A
Dictation pick on a row that is already downloading no longer starts a
second install: the row's radio is disabled while it works, and
useModelDownloads refuses a second mutation for a repo already in flight.
The Supertonic-3 license test checks for the Accept wording.
2026-09-10 13:34:54 -07:00
Palash Debnath 01d503f64f Merge remote-tracking branch 'origin/main' into fix/cosyvoice-advisory-floors 2026-09-10 13:31:49 -07:00
Palash Debnath c94585a339 Merge pull request #2032 from debpalash/fix/darwin-eperm-exit-settle
fix(lifecycle): wait for a Darwin exit to register after EPERM
2026-09-10 13:28:34 -07:00
Palash Debnath 40139bb2a2 Merge remote-tracking branch 'origin/main' into fix/cosyvoice-advisory-floors 2026-09-10 13:20:21 -07:00
Palash Debnath 045f75c30e Merge pull request #2031 from debpalash/dependabot/pip/backend/engines/cosyvoice_subprocess/transformers-5.10.1
chore(deps): bump transformers from 4.57.6 to 5.10.1 in /backend/engines/cosyvoice_subprocess
2026-09-10 13:20:16 -07:00
Palash Debnath 4a4fa6a9e0 test(cosyvoice): raise the advisory floors to GitHub's fixes
The floors came from OSV alone, which did not list the protobuf and
transformers advisories GitHub flagged once the manifest reached main
(Dependabot #2030, #2031). Raise them to the first fixed releases so a
later downgrade fails the test. Note that transformers 5 was checked
against CosyVoice's tokenizer and cached decoding.
2026-09-10 13:04:49 -07:00
Palash Debnath 5f3f35f134 Merge remote-tracking branch 'origin/main' into fix/darwin-eperm-exit-settle
# Conflicts:
#	CHANGELOG.md
2026-09-10 13:04:31 -07:00
Palash Debnath 32393613be Merge branch 'main' into feat/engine-list-detail 2026-09-10 13:03:00 -07:00
Palash Debnath 2c83f5c16c Merge pull request #2013 from debpalash/feat/catalogue-one-page
feat(catalogue): one page — setup summary over per-family engines and weights
2026-09-10 13:02:39 -07:00
Palash Debnath a76e743415 docs(changelog): note the Darwin exit-settle fix (#2032) 2026-09-10 13:02:08 -07:00
Palash Debnath 8cf7bebb61 fix(lifecycle): wait for a Darwin exit to register after EPERM
XNU stops signalling a process as soon as it starts exiting but posts
NOTE_EXIT later in the same exit. A KILL that lands in that window gets
EPERM while the exit probe still reads "alive", so stopping a child that
TERM had just ended could fail with "Operation not permitted". This
flaked the macOS run of contained_exit_probe_preserves_a_live_child.

The EPERM branch now re-probes for up to 250 ms before treating the error
as a live, unsignalable root. Live roots, reaped roots and probe failures
are still errors.
2026-09-10 13:02:00 -07:00
Palash Debnath 05df8159eb Merge branch 'main' into dependabot/pip/backend/engines/cosyvoice_subprocess/transformers-5.10.1 2026-09-10 12:58:26 -07:00
Palash Debnath 5980fce852 Merge pull request #2030 from debpalash/dependabot/pip/backend/engines/cosyvoice_subprocess/protobuf-5.29.6
chore(deps): bump protobuf from 4.25.9 to 5.29.6 in /backend/engines/cosyvoice_subprocess
2026-09-10 12:58:02 -07:00
dependabot[bot] 6b98b1bded chore(deps): bump transformers in /backend/engines/cosyvoice_subprocess
Bumps [transformers](https://github.com/huggingface/transformers) from 4.57.6 to 5.10.1.
- [Release notes](https://github.com/huggingface/transformers/releases)
- [Commits](https://github.com/huggingface/transformers/compare/v4.57.6...v5.10.1)

---
updated-dependencies:
- dependency-name: transformers
  dependency-version: 5.10.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-10 19:35:13 +00:00
dependabot[bot] b597007152 chore(deps): bump protobuf in /backend/engines/cosyvoice_subprocess
Bumps [protobuf](https://github.com/protocolbuffers/protobuf) from 4.25.9 to 5.29.6.
- [Release notes](https://github.com/protocolbuffers/protobuf/releases)
- [Commits](https://github.com/protocolbuffers/protobuf/commits)

---
updated-dependencies:
- dependency-name: protobuf
  dependency-version: 5.29.6
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-10 19:35:06 +00:00
Palash Debnath 0eaa1bdc3e Merge branch 'main' into feat/catalogue-one-page 2026-09-10 12:34:20 -07:00
Palash Debnath 44d3639c02 Merge pull request #2025 from debpalash/feat/cosyvoice3-isolated
feat(engines): CosyVoice 3 installs in one click into its own venv
2026-09-10 12:33:01 -07:00
Palash Debnath ba8f6cfd15 Merge feat/catalogue-one-page (with main) into feat/engine-list-detail
Conflicts: RecoBanner (deleted here: the recommendation card is gone),
ModelStoreTab + EngineCompatibilityMatrix (this branch's rewrite kept),
supertonic3 backend (main's own-venv check kept, message without the
retired "→ Engines" step), CHANGELOG (base layout + the #2020 line).

Carried over from main and the review:
- The list row hides Install while only the license review is left
  (main's #2017 rule, now in the row; Accept lives in the panel).
- Engine action aria labels go through i18n (engines.aria*, all 21
  locales) instead of hardcoded English.
- Every "Model Catalogue → Engines/Models" path main added, plus the
  frontend strings that still named the retired panes, now point at the
  one-page catalogue.
- test_engine_unavailable_reason_1866 reads the license matcher from its
  new home, engines/engineDisplay.js.
2026-09-10 12:17:02 -07:00
Palash Debnath 80ec0d06c9 Merge origin/main into feat/catalogue-one-page; harvest remaining review threads
Conflicts: CHANGELOG (main cut 0.5.2 and opened a new Unreleased; the
#2013 lines move there), pockettts/supertonic3 docs (main's new one-click
install text kept, with the retired "Model Catalogue → Engines →" step
dropped, as in the rest of docs).

Review fixes on top:
- Bulk installs name the right repository on failure: every "install
  several" button now pairs allSettled results with their request before
  filtering (shared failedInstalls/installFailureMessage + unit test).
- confucius4-tts docs: "The first synthesis triggers".
2026-09-10 12:09:39 -07:00
Palash Debnath 85ce7206b5 fix(engines): keep the literal data filter so tar extraction is visibly validated
Bandit B202 flagged the custom filter callable. Links are now left out by
choosing the members, and extractall keeps filter="data", which validates
everything extracted. Same behaviour: the Matcha-TTS tarball's absolute
link is skipped instead of aborting the fetch.
2026-09-10 12:02:33 -07:00
Palash Debnath 180371910e fix(cosyvoice): patched dependencies, no wetext download, Windows long-path hint
Built the real environment on Windows to check the one-click install.
- The pins upstream carries with published advisories (diffusers,
  hydra-core, lightning, modelscope, onnx, protobuf, transformers) are
  raised to fixed releases; that exact set installs and passes the
  installer's import probe. pyarrow, not needed, is dropped. A test keeps
  every pin at or above its advisory fix.
- wetext is dropped with its install-time fetch. Its data exists only on
  ModelScope, which rate-limits downloads; a throttled fetch left files
  missing while reporting success, so the normaliser failed silently.
  CosyVoice reads text as written without it, and nothing reaches
  ModelScope at install or synthesis time. The unused post-install hook
  goes with it.
- openai-whisper builds from source, and under a long cache path the build
  hits Windows' 260-character limit; the install failure now says to turn
  on long-path support.
2026-09-10 11:39:50 -07:00
Palash Debnath efdec50081 fix(engines): a source tarball with a link in it still extracts
The stdlib data filter raises on a link to an absolute path, and the
pinned Matcha-TTS tarball ships one (data -> its author's training
folder), which aborted the fetch and with it the CosyVoice install. Found
by building the real environment. Links are now skipped on every
interpreter, as the pre-3.11.4 path already did.
2026-09-10 11:34:03 -07:00
Palash Debnath 655db866d1 Merge pull request #2029 from debpalash/fix/release-checksum-race
fix(release): one job writes the notes and publishes, after every platform
2026-09-10 11:31:19 -07:00
Palash Debnath 7c1599b73a test(release): pin the missing-checksum gate; docs drop the manual publish step
The guard now proves the job exits on a missing platform before it
rewrites the notes or publishes, so removing that exit fails CI.
RELEASING.md's end-to-end updater test still said to publish the draft by
hand; the workflow publishes it.
2026-09-10 11:08:10 -07:00
Palash Debnath 399b9ec3f6 Merge remote-tracking branch 'origin/main' into feat/cosyvoice3-isolated 2026-09-10 11:06:04 -07:00
Palash Debnath b804a6f23f Merge pull request #2022 from debpalash/feat/moss-tts-nano-isolated
feat(engines): MOSS-TTS-Nano installs into its own venv, pinned to a working upstream
2026-09-10 11:05:21 -07:00
Palash Debnath 9561d41fb6 fix(cosyvoice): a missing model-folder override is an error, not a silent swap
When OMNIVOICE_COSYVOICE_MODEL pointed at a folder that no longer existed,
the sidecar quietly loaded the installed CosyVoice 3 model instead, so it
spoke with a model and voice the user had not chosen while
model_identity() still named theirs. It now fails with a message naming
the path and how to fix or clear it.
2026-09-10 10:51:13 -07:00
Palash Debnath 6305f577e1 docs(changelog): link the release-publish entry to #2029 2026-09-10 10:49:16 -07:00
Palash Debnath dd1dce75cb fix(release): one job writes the notes and publishes, after every platform
Each build leg appended its checksums to the shared release notes with
softprops/action-gh-release. Two consequences, both seen on v0.5.2:

- The appends were concurrent read-modify-writes, so a leg that read the
  notes before another wrote them lost its section. v0.5.1 and v0.5.2
  both shipped without the macOS Apple Silicon checksums in the notes,
  though the SHA256SUMS file was attached.
- softprops defaults to draft: false, so the first leg to finish
  published tauri-action's draft while the others were still building.
  v0.5.2 went public at 17:27; its complete latest.json landed at 17:38.

Legs now only attach their SHA256SUMS file with gh release upload. A new
release-notes-checksums job runs once after the matrix, the updater-manifest
repair and the uninstall scripts, writes all four platforms' checksums in
matrix order, fails if one is missing (leaving the release a draft), and
then publishes it. contributors-strip edits the notes after it, so the
notes have one writer at a time. A test pins all of that.

docs/RELEASING.md described a manual publish and .exe / .deb / .nsis.zip
artifacts the workflow no longer builds; it now matches what ships.
2026-09-10 10:49:08 -07:00
Palash Debnath 9c925b17a0 Merge remote-tracking branch 'origin/main' into feat/moss-tts-nano-isolated 2026-09-10 10:38:43 -07:00
Palash Debnath 30c64dc44b Merge pull request #2021 from debpalash/feat/voxcpm2-isolated
feat(engines): VoxCPM2 installs into its own venv and runs in a sidecar
2026-09-10 10:37:43 -07:00
Palash Debnath e8be51f751 Merge remote-tracking branch 'origin/feat/moss-tts-nano-isolated' into feat/cosyvoice3-isolated
# Conflicts:
#	CHANGELOG.md
2026-09-10 10:13:53 -07:00
Palash Debnath 919116dfa5 Merge remote-tracking branch 'origin/feat/voxcpm2-isolated' into feat/moss-tts-nano-isolated
# Conflicts:
#	CHANGELOG.md
2026-09-10 10:13:21 -07:00
Palash Debnath 9dd005eab5 docs(changelog): list VoxCPM2 under Unreleased, not the released 0.5.2
Merging main placed the entry beside its old neighbour, which now sits in
the released [0.5.2] section.
2026-09-10 10:11:47 -07:00
Palash Debnath 8a497aa4db Merge remote-tracking branch 'origin/main' into feat/voxcpm2-isolated 2026-09-10 10:10:37 -07:00
Palash Debnath 38c2405fa1 Merge pull request #2007 from debpalash/release/v0.5.2
docs(changelog): cut the 0.5.2 release section
2026-09-10 10:09:20 -07:00
Palash Debnath 97b153df3a test(changelog): count versions the way the app's changelog parser reads them
The one-section-per-version check matched raw heading text, so [v0.5.2]
and [0.5.2] would pass as two versions while core.changelog, which serves
the in-app notes, strips the v and whitespace and reads them as one. The
check now uses parse_changelog itself.
2026-09-10 09:47:54 -07:00
Palash Debnath 6b62866fb5 test(cosyvoice): its fake venv carries the completion marker 2026-09-10 09:46:21 -07:00
Palash Debnath ab565e9b5c Merge remote-tracking branch 'origin/feat/moss-tts-nano-isolated' into feat/cosyvoice3-isolated 2026-09-10 09:46:02 -07:00
Palash Debnath 594daf9251 test(moss-tts-nano): its fake venv carries the completion marker 2026-09-10 09:45:53 -07:00
Palash Debnath 8114132a72 Merge remote-tracking branch 'origin/feat/voxcpm2-isolated' into feat/moss-tts-nano-isolated 2026-09-10 09:45:33 -07:00
Palash Debnath 63a1c1791a fix(engines): an own-venv engine needs a finished install, and VoxCPM2 PCM is 48 kHz
engine_venv_python() accepted any venv interpreter, so a reinstall that
failed partway made the resolver pick the sidecar and report it ready,
hiding the working in-process engine. It now also requires the completion
marker, which the import probe writes only after the engine imported in
that venv and which a new dependency step removes. That covers VoxCPM2
here and PocketTTS / Supertonic-3 from #2016.

The parent reads sidecar PCM at the engine's fixed 48 kHz, so the VoxCPM2
sidecar resamples a model that reports another rate instead of passing its
samples through mislabelled. The retry test now proves a permanent error
runs once.
2026-09-10 09:45:20 -07:00
Palash Debnath 4e37e21745 chore(release): v0.5.2
Cut the Unreleased notes as 0.5.2 (2026-09-10) and fold in the untagged
0.5.2 section prepared on 2026-09-02, so release.yml publishes one
complete section. A test now requires one section per version.
2026-09-10 09:42:08 -07:00
Palash Debnath 3bfdeeddba Merge pull request #2024 from debpalash/fix/uv-no-app-config
fix(engines): engine installs never inherit VoiceStudio's own uv config
2026-09-10 09:42:02 -07:00
Palash Debnath be722ff608 test: register the CosyVoice sidecar with the stdout guard and STRUCTURE.md 2026-09-10 09:22:24 -07:00
Palash Debnath 0b7117b833 Merge remote-tracking branch 'origin/feat/moss-tts-nano-isolated' into feat/cosyvoice3-isolated 2026-09-10 09:22:06 -07:00
Palash Debnath a88c39eac1 test: register the MOSS-TTS-Nano sidecar with the stdout guard and STRUCTURE.md 2026-09-10 09:21:58 -07:00
Palash Debnath 3f19b2e9a7 Merge remote-tracking branch 'origin/feat/voxcpm2-isolated' into feat/moss-tts-nano-isolated 2026-09-10 09:21:40 -07:00
Palash Debnath 172a142825 test: register the VoxCPM2 sidecar with the stdout guard and STRUCTURE.md
The #1428 guard pins the exact set of sidecars it scans, and
docs/STRUCTURE.md must list every engine adapter; both failed CI on the
new engines/voxcpm2_subprocess.
2026-09-10 09:21:32 -07:00
Palash Debnath ac35351000 Merge remote-tracking branch 'origin/main' into fix/uv-no-app-config
# Conflicts:
#	CHANGELOG.md
2026-09-10 09:11:59 -07:00
Palash Debnath f9da18842a docs(changelog): link the CosyVoice entry to #2025 2026-09-10 09:10:09 -07:00
Palash Debnath c4206ee1aa Merge remote-tracking branch 'origin/feat/moss-tts-nano-isolated' into feat/cosyvoice3-isolated 2026-09-10 09:09:41 -07:00
Palash Debnath ec1529479d feat(engines): CosyVoice 3 installs in one click into its own venv
CosyVoice ran only in-process, which needed it importable from the app's
own interpreter; upstream's setup (its own Python 3.10 environment, pins
that clash with the app's) never gives that, so the engine was unusable
from an installer build. The one-click installer now clones a reviewed
commit (074ca6dc) and its Matcha-TTS submodule (dd9105b3) into
DATA_DIR/engines/cosyvoice/, builds a Python 3.10 venv, downloads the
CosyVoice 3 weights, and a self-contained sidecar runs the model there.

Upstream's requirements cannot be used as they are: they add a
third-party Azure DevOps index, pin torch 2.3.1 (CUDA 12.1 only, no
RTX 50-series), pull TensorRT and DeepSpeed on Linux, and do not even
resolve together (fastapi). A trimmed list ships with the app
(engines/cosyvoice_subprocess/requirements.txt, which says what was
dropped and why); torch comes from the per-host 2.7.0 pins. Every package
installs from a wheel except two pure-Python ones, so nothing needs a
compiler, and SoX is not used.

The installer gains what CosyVoice needs, generally:
- ExtraSource fetches a pinned submodule tree, which neither a depth-1
  clone nor GitHub's tarball includes;
- weights_allow_patterns downloads only the files the model loads
  (5.4 of 9.8 GB);
- post_install_code runs an optional fetch in the engine's venv: here,
  wetext's normalisation data, which it otherwise downloads from
  ModelScope on every model load. A failure only leaves that to first use.
- The completion marker now applies to engines with weights too, since
  weights left by an earlier run do not prove the dependency step
  finished. IndexTTS keeps its weights check, so no existing install is
  asked to reinstall.

The sidecar formats prompts as upstream's v3 examples do, speaks in
upstream's sample voice when there is no reference clip (v3 has no
built-in speakers), and never hands AutoModel a missing folder, which
would start a ModelScope download. The gateway keeps offering the
CosyVoice weights as a plain download to remote workers that run the
in-process engine.
2026-09-10 09:09:37 -07:00
Palash Debnath 61a5418c6a Merge remote-tracking branch 'origin/feat/voxcpm2-isolated' into feat/moss-tts-nano-isolated 2026-09-10 09:08:38 -07:00
Palash Debnath b28390af30 test(voxcpm2): patch the base class the engine really inherits
Other suites purge services from sys.modules, so importing
subprocess_backend inside the test can yield a different module than the
one VoxCPM2SubprocessBackend subclassed, and the patch then misses: the
real generate() ran and failed. Patch VoxCPM2SubprocessBackend's own base.
2026-09-10 09:07:59 -07:00
Palash Debnath 764fed09bc Merge pull request #2018 from debpalash/fix/license-reason-passthrough
fix(engines): say when a license or the platform is what blocks an engine
2026-09-10 09:04:07 -07:00
Palash Debnath 4fdde15079 Merge pull request #2023 from debpalash/fix/translation-uninstall-guard
fix(translation): uninstalling an engine never removes what others need
2026-09-10 09:00:07 -07:00
Palash Debnath 14a8a9cd46 docs(changelog): link the uv config entry to #2024 2026-09-10 08:46:57 -07:00
Palash Debnath 1af1ecc5bd fix(engines): engine installs never inherit VoiceStudio's own uv config
The backend runs inside the app's tree, and the uv processes it starts for
engine installs had no working directory of their own. uv therefore
discovered VoiceStudio's pyproject.toml and applied its [tool.uv]
constraint-dependencies, torch==2.8.0 among them, to each engine's venv.

Resolved that way, MOSS-TTS-v1.5 (torch==2.9.1+cu128) and Confucius4
(torch==2.7.0) are unsatisfiable, so their one-click installs and
bootstraps could never succeed. An engine that pins no torch got the app's
instead of its own, so its venv was not really independent.

uv_subprocess_env, which every one-click install step and every engine
bootstrap already uses, now always sets UV_NO_CONFIG=1. It used to return
None in two cases and let the call inherit the environment; it now always
returns a copy. Region and custom mirrors still apply, because they reach
uv as UV_INDEX_URL. The app's own `uv sync` and the translation installer,
which install into the app's environment on purpose, are unchanged.
2026-09-10 08:46:49 -07:00
Palash Debnath e0879027d5 Merge remote-tracking branch 'origin/main' into fix/license-reason-passthrough
# Conflicts:
#	CHANGELOG.md
#	tests/test_engine_unavailable_reason_1866.py
2026-09-10 08:41:45 -07:00
Palash Debnath 5298c8e5e1 Merge pull request #2016 from debpalash/feat/isolated-engine-installs
feat(engines): isolated one-click installs for five engines
2026-09-10 08:38:39 -07:00
Palash Debnath 1de4fedad7 docs(changelog): file the translation guard under Fixed, not Highlights 2026-09-10 08:33:55 -07:00
Palash Debnath c5f32d020a fix(translation): uninstalling an engine never removes what others need
The translation uninstall route runs `pip uninstall -y <package>` on the
app's own environment. Two entries made that break something else:

- The LLM engine's package is openai, a core dependency that Settings →
  LLM Providers also uses. Unlike Argos, it was not marked builtin, so
  uninstalling that engine removed it from the app.
- Google, DeepL, Microsoft and MyMemory share deep_translator. Uninstalling
  any one removed it for all four.

The route now asks uninstall_blocker() first. It refuses (400) a package
VoiceStudio itself depends on, read from the installed package metadata so
there is no second list to keep in step, and refuses (409) a package
another engine shares, naming the engines that would stop working. The
openai entry is marked builtin as well, and a test requires every entry
backed by an app dependency to be.
2026-09-10 08:26:53 -07:00
Palash Debnath f5ea363724 fix(engines): name the missing MPS on Apple Silicon, and hold the MLX test on every host
mlx_supported() fails three ways. Only the non-Apple branch is a platform
gap; Apple Silicon whose PyTorch cannot use MPS was left on the generic
line, and the live-host test asserted a platform reason even there, so it
would fail on such a Mac. That case now has its own owned sentence, and
the live test asserts only the branch its host can produce, with literal
cases covering the rest on every machine.
2026-09-10 08:22:20 -07:00
Palash Debnath 4280157407 docs(changelog): link the MOSS-TTS-Nano entry to #2022 2026-09-10 08:21:06 -07:00
Palash Debnath 240a64f086 feat(engines): MOSS-TTS-Nano installs into its own venv, pinned upstream
The in-process MOSS-TTS-Nano engine needs moss_tts_nano installed into the
app's own environment, with upstream's exact pins (torch 2.7.0,
transformers 4.57.1) landing there too. It also cannot work with today's
upstream: moss_tts_nano now exports only __version__, and the model class
the engine looks for is gone. The entry point is the top-level
moss_tts_nano_runtime.NanoTTSService.

The one-click installer now clones a reviewed commit (8b7bcc93,
2026-09-06) into DATA_DIR/engines/moss-tts-nano/ with its own venv. A
self-contained sidecar drives NanoTTSService there: it preloads the model,
heartbeats through the downloads (the model on load, the audio tokenizer on
the first synthesis) and never after, downmixes to mono as the in-process
engine did, resamples to the reported 48 kHz if needed, and reuses one
output file instead of leaving one per call.

The resolver becomes a table (_OWN_VENV_SIDECARS) now that two engines use
it, with a test that each entry's env var is the one its installer sets.
The three identical Intel-Mac gates become one factory.
2026-09-10 08:20:58 -07:00
Palash Debnath 388d554ba0 docs(changelog): link the VoxCPM2 entry to #2021 2026-09-10 08:13:57 -07:00
Palash Debnath 21105b24a5 feat(engines): VoxCPM2 installs into its own venv and runs in a sidecar
VoxCPM2 ran only in-process, so using it meant installing voxcpm, and a
torch of its choosing, into the app's own environment. It now has a
one-click install into DATA_DIR/engines/voxcpm2/.venv and a self-contained
sidecar (engines/voxcpm2_subprocess) that imports nothing from the app.
tts_backend resolves the voxcpm2 id to the sidecar once that venv exists,
and to the in-process class otherwise, so an existing pip install keeps
working.

voxcpm leaves torch unpinned. Resolved with the CUDA index, that paired
PyPI's newest torch (CPU-only on Windows) with a +cu128 torchaudio, and
uv's --torch-backend fell back to voxcpm 1.5.0 on Windows. A new
torch_pins spec field pins the pair, and the host picks the build: +cu128
on CUDA hosts, +cpu on other Windows and Linux hosts, plain on macOS. Each
was resolved with voxcpm==2.0.3. Not offered on Intel Macs, where torch
2.11 has no build.

The parent keeps the reference-clip preparation and the trailing-silence
trim, so output matches the in-process engine. The sidecar retries a
transient weight download like the app's loader does.
2026-09-10 08:13:57 -07:00
Palash Debnath ae68dd5bfd fix(engines): a half-finished install is repaired, not reported installed
An engine with no weights download counted as installed once its venv
interpreter existed, so a dependency install that died halfway made the
next attempt answer already_installed and the engine failed at its first
import. The import probe now writes a completion marker, and a fresh
dependency step removes the old one. IndexTTS keeps its weights check, so
no existing install is asked to reinstall.

The MOSS bootstrap no longer blames a non-CUDA host for an install
failure; the index is always supplied, and uv's error says what failed.
The PocketTTS and Supertonic guides now say where the sidecar runs, and
the three repository-engine guides say what to do if the first weight
download outlasts the compute-time budget.
2026-09-10 08:12:17 -07:00
Palash Debnath 0aafbf4421 fix(catalogue): sherpa dictation weights keep cancel/remove/reinstall — the dictation pick is a radio on the row 2026-09-10 08:09:37 -07:00
Palash Debnath 341b475226 feat(catalogue): icons on the setup summary rows 2026-09-10 08:01:42 -07:00
Palash Debnath 09b7e77629 docs(changelog): engine list + detail panel (#2020); drop the unused weights heading key 2026-09-10 08:00:40 -07:00
Palash Debnath 3008f89919 feat(catalogue): engine list + detail, weights under their engine
The engine matrix (five columns, three-line rows, every chip on every row)
becomes a shadcn table with three columns — Engine · Runs on · Status — and
one primary action per row (Use / Install). Everything else lives in a
detail panel for the selected row: GPU compatibility chips, isolation,
hints and reasons, health and self-test probes, one-click install
progress, setup snippet, disk usage, docs, license, the curated-model
picker, and now the engine's downloadable WEIGHTS.

Weights belong to their engine: every models.yaml entry names the backend
ids that load it (`engines:`), the detail panel lists and installs them
(EngineWeights, on the model store's install/cancel/remove flow via the
extracted useModelDownloads hook), and the sherpa-onnx engine shows its
dictation-model picker there. The page's "Downloaded weights" list and
recommendation card are gone; only weights no engine owns (speaker
diarisation) remain in a small "Other weights" list. A backend test pins
the mapping: every entry has an `engines` list and every id is a real
backend.

- useEngineInventory: the matrix's state machines extracted verbatim
  (shared/local fetch, residency, health/self-test cooldowns, install
  poller with overlap guard + epoch, disk-usage generations, license).
- Row status phrases: GPU active / CPU fallback / CPU / Available /
  Needs setup / Installing… / failed; routing "unavailable" never reads
  Ready. Group captions keep "Ready to use" / "Add more engines".
- Engine titles read "Engines" (each locale's own word); backend
  "Model Catalogue → Engines/Models" messages and docs updated to the
  new structure.
- Dead matrix CSS (phone-tier grid) removed; scopeReco and RecoBanner gone.
2026-09-10 07:59:47 -07:00
Palash Debnath 2fd4cb3caf test(engines): cover dots.tts's own Windows reason in the platform category 2026-09-10 07:53:10 -07:00
Palash Debnath e532c9d8dd fix(engines): say when an engine can't run on this platform
The public reason sanitizer had no category for a host the engine cannot
run on at all, the same gap that hid the license button. "MLX requires
Apple Silicon" and PocketTTS's Intel-Mac reason became the generic
"check installation" line, and "not supported on this platform" became
"isn't installed yet", which an existing test asserted. Each sent people
after an install that could never work.

A platform category, matched after the license and before the install
and file checks, now says the engine doesn't run on this platform and
points at its guide. mlx-audio's "Apple Silicon only" wording is left
out of the markers: it also appears on an M-series Mac when the package
is simply missing, where installing does help.
2026-09-10 07:52:46 -07:00
Palash Debnath 5516750a4f fix(engines): accept Confucius4's requirements.txt-only source layout
Source validation demanded a pyproject.toml in every checkout, and
Confucius4 ships none, so its install could never get past fetching the
source. The manifest file is now per spec. The regression test fabricates
each pinned upstream's real root files.
2026-09-10 07:48:52 -07:00
Palash Debnath c079b721ed feat(engines): Supertonic-3 and PocketTTS install into their own venvs
Both engines ran with the app's interpreter, installed as optional extras
into the app's own environment (`uv sync --extra`). They now get one-click
installs like the sidecar engines: a PyPI-only spec (no source to fetch)
creates DATA_DIR/engines/<id>/.venv and installs the app's own pinned wheel
there, so nothing they install can touch the app or another engine.

Each engine prefers its own venv and falls back to the app's interpreter,
so an existing `uv sync --extra` install keeps working and is never
provisioned over: the spec counts a package found in the app environment
as installed.

PocketTTS installs from PyTorch's CPU index: it never uses a GPU, and
PyPI's Linux torch pulls ~15 NVIDIA packages. It stays unoffered on Intel
Macs, where no usable torch exists.

Supertonic's sidecar loads its constants by path when the revision env var
is absent, instead of importing the engines package, whose __init__
imports the app backend that its own venv does not have.

The Install button is hidden once only the license review stands between
the user and the engine. The installer tests' autouse fixture now removes
every spec's env var on teardown: a bare delenv of an unset var restored
nothing, and a persisted path leaked into later suites.
2026-09-10 07:44:42 -07:00
Palash Debnath 664c9ea4b5 fix(engines): keep the license reason so Supertonic-3 and PocketTTS can be enabled
The Model Catalogue shows an engine's license Accept button only when its
reason matches /license not accepted/i. public_backends() replaces probe
text with owned sentences, and no category covered a license gate, so the
reason arrived as the generic "Engine unavailable" line and the only way
to enable Supertonic-3 or PocketTTS never rendered (#2017).

A license category, matched first, keeps those words. The test reads the
regex out of EngineCompatibilityMatrix.jsx, so a wording change on either
side fails CI instead of silently hiding the button.
2026-09-10 07:34:06 -07:00
Palash Debnath faa3d39836 feat(engines): isolated one-click installs for MOSS-TTS-v1.5, Confucius4 and dots.tts
Three engines that shipped as terminal-only setups now install from Model
Catalogue → Engines with the existing sidecar installer, which is
generalised to take a per-engine venv interpreter, install target, import
probe and host gate.

Each engine gets DATA_DIR/engines/<id>/ with its own checkout and .venv;
every uv pip install passes --python for that venv, never the app's
interpreter. Switching the active engine only changes a pref, so moving
between engines and back cannot corrupt a working one, and uninstalling one
removes only its own folder. Tests pin both invariants for every spec.

MOSS-TTS-v1.5's [torch-runtime] extra pins torch==2.9.1+cu128, which exists
only on PyTorch's index, so its manual install and its bootstrap could never
resolve (#2015). core.torch_indexes defines the index once for the
installer and the bootstrap, and a test ties it to the app's own
pytorch-cuda index.

Install buttons appear only where the install can work: MOSS on CUDA hosts,
dots.tts off Windows (upstream publishes no Windows install). A direct POST
on an unsupported host gets a 409 with the reason. An engine with no
one-click install now points at its guide, not at a page with no Install
button.
2026-09-10 07:28:42 -07:00
Palash Debnath b0cae840c4 Merge pull request #2010 from debpalash/fix/pill-visibility-desync
fix(dictation): pill window, card and stale model on Windows (#2009, #2012)
2026-09-10 07:18:22 -07:00
Palash Debnath a7cfe288cb fix(catalogue): harvest review findings on #2013
- SetupSummary: an installed engine whose routing is "unavailable" reads
  Needs setup, not Ready (select is refused for it too); a failed /engines
  or /dictation/models fetch renders as an error with Retry instead of
  posing as "Off" / "Needs setup".
- Bulk installs (summary, model store, recommendation card) wait for every
  request to settle before re-enabling, and report which repos failed —
  one early rejection can no longer re-arm the button mid-flight.
- The weights list stays mounted across family switches (hidden under LLM)
  so download progress and Retry/Dismiss state survive navigation.
- Settings search: Hugging Face mirror terms route to Network; the legacy
  "models" tab id resolves to Storage. "Manage models" opens the TTS tab.
- Locales: uk "Рушії", zh-TW "引擎", vi "Engine" for the Engines heading.
- Docs name the family tab wherever the instruction depends on it.
2026-09-10 07:08:43 -07:00
Palash Debnath 566aad7bba fix(dictation): mirror the widget's shadow setting in the macOS overlay
Tauri applies platform config as a JSON Merge Patch, so the windows array in
tauri.macos.conf.json REPLACES the base one rather than merging window objects.
desktopWindowConfig.test.js pins that every base window property survives on
macOS, and it caught the base gaining "shadow": false without the overlay.

The programmatic builder already sets shadow(false) on every platform; this
keeps the macOS declaration in step with it. The Linux and Windows overlays
declare no widget window, so the base entry applies there unchanged.
2026-09-10 06:53:49 -07:00
Palash Debnath 204a7294ca docs(changelog): one-page Model Catalogue (#2013) 2026-09-10 06:46:49 -07:00
Palash Debnath 3cae853440 feat(catalogue): one page, one axis — setup summary over per-family engines and weights
The Model Catalogue put the same decision on two axes: an Engines pane with
TTS/ASR/LLM tabs and a Models pane with TTS/ASR/Dictation/Diarisation
sections, dictation shown in both, plus storage stats, the HF token and the
voice-preview toggle parked on the model list. Settings → Voice still carried
Engines and Models entries that only pointed back here.

Now the page reads top-down: a SetupSummary (speech, transcription,
dictation, language model — engine, device, one status word, Change), the
engine list for one family, and that family's downloadable weights under it
(TTS under TTS; offline ASR, streaming dictation and diarisation under ASR;
nothing for LLM, whose engines bring their own). One storage line points at
Settings → Storage.

- ModelStoreTab takes a `family` and scopes sections and the recommendation
  preset to it (scopeReco); stats strip, HF-token toolbar and previews
  panel removed from it.
- Settings: Engines/Models categories and CataloguePointer removed; models
  directory → Storage, HF mirror → Network (both restart-flagged), voice
  previews → Storage. "Manage models" in disk usage opens the catalogue.
- Store: openCatalogue takes a family (pane key tolerated, ignored);
  pendingCatalogueTab gone.
- Engine matrix title is now the locale's plain "Engines".
- i18n: catalogue.* summary keys in all 21 locales; pane/pointer keys dropped.
- Docs: "Model Catalogue → Engines" is "Model Catalogue"; "→ Models" is
  "→ Downloaded weights".
2026-09-10 06:46:09 -07:00
Palash Debnath e9965b2cb4 fix(dictation): never show the pill natively when Tauri's show fails
Greptile on #2010: if win.show() fails while hwnd() succeeds, the native
SW_SHOWNOACTIVATE show still ran, putting an always-on-top window on screen
that Tauri believes is hidden. dismiss() and the idle reconcile then cannot
remove it — the stranded-window bug this PR fixes, reached by a different
door.

The previous commit degraded to the native show on purpose, reasoning that a
visible pill beats none. That was the wrong trade: the tray's red dot already
tells the user they are being recorded, while an unhidable always-on-top window
is left behind for the rest of the session. The native show now runs only
after Tauri's show has succeeded, and the test that pinned the fallback is
flipped to pin its absence.
2026-09-10 06:36:06 -07:00
Palash Debnath 52887ae44d fix(dictation): no card around the pill, and the pill uses your model (#2009, #2012)
Two more defects from the same live report.

The card. Tauri's default window shadow on Windows gives an undecorated
window a 1px white border and, on Windows 11, rounded corners — drawn around
the whole 460x164 pill window, which is far wider than the pill (at most
284px). That is the bordered card framing empty space, visible whether or not
the pill is showing. The widget window now sets shadow(false), in the builder
and in its tauri.conf.json declaration; the capsule draws its own edge in CSS.

The stale model (#2012). The pill runs in its own window with its own store,
created at app start, usually before the backend listens. CaptureWidget
hydrated the dictation prefs once and memoized the promise whether or not the
load worked, and loadDictationPrefs swallowed the failure and marked itself
loaded. So the widget kept the store's seed, sherpa-whisper-tiny, for the
session. The main window checked the model actually picked (Parakeet,
installed) and said ready; the widget asked the server for the seed, and the
server correctly answered that it was not installed.

loadDictationPrefs now reports whether the backend answered. Only a successful
load is kept; a failed one is retried, and a capture start re-reads so a model
chosen in the main window reaches the widget. An in-flight load is shared.

On the same path, the missing-model install toast was called from the widget,
whose window has no <Toaster> in the desktop app, so it rendered nowhere. The
install recommendation now rides the dictation notice to the main window,
which shows the one-click download. The browser build, where the widget lives
inside the main window, keeps its local toast. The pill labels a missing model
as that rather than "Transcription failed: ...", and clamps error text to two
lines instead of spilling a paragraph past the capsule.

Tests: a failed first load is retried instead of pinning the seed, and a model
changed elsewhere is picked up at the next capture — both fail against main's
widget. The notice routes a missing model to the install toast. The
setup-race test now asserts no local toast in Tauri and the notice payload
instead. 91 frontend tests and 256 Rust lib tests pass; typecheck:ci passes.
2026-09-10 06:34:26 -07:00
Palash Debnath b0f1558db0 fix(dictation): tell Tauri the pill is visible, not just Windows (#2009)
Closing the dictation pill on Windows left an empty dark rectangle on screen,
always on top, removable only by quitting the app.

show_pill_noactivate called the raw Win32 ShowWindow(hwnd, SW_SHOWNOACTIVATE)
and never Tauri's own win.show(). The flag is there for a good reason (#982: a
pill that takes foreground makes the dictated text paste into the pill instead
of the user's document), but going straight to Win32 puts the window on screen
behind Tauri's back. Tauri went on believing it was hidden, and every mechanism
that could have removed it was disabled by that one desync:

  - isVisible() answered false while the user was looking at the window;
  - hide() was a no-op on a window Tauri thought was already hidden, so
    dismiss() in CaptureWidget could not remove it;
  - the idle reconcile — the backstop that exists precisely to clean up a
    stranded pill — asks isVisible() first, and concluded there was nothing
    to clean up.

win.show() now runs first, then the native flag. The no-activation behaviour is
carried by the WS_EX_NOACTIVATE style bit that mark_pill_noactivate applies at
creation, which is what makes the Tauri show safe: the style bit, not the show
flag, is what refuses activation. The flag stays as a second line of defence,
since hwnd() can fail and the bit might not have been applied.

Windows only. macOS and Linux already took the win.show() branch.

The ordering is now a function with both shows as parameters, so two tests can
pin it: Tauri's show runs and runs first, and a failing Tauri show still puts
the pill on screen — degrading to the old behaviour beats not showing the user
that they are being recorded.

256 Rust lib tests green; the 10 stranded-pill frontend tests unchanged and
still passing.
2026-09-10 05:57:22 -07:00
Palash Debnath dea884a22d Merge pull request #2006 from debpalash/fix/worker-artifact-id-posix
fix(worker): identify a staged input the same way on every OS (#2005)
2026-09-10 05:31:39 -07:00
Palash Debnath bee9fa25ff Merge remote-tracking branch 'origin/main' into fix/worker-artifact-id-posix
# Conflicts:
#	CHANGELOG.md
2026-09-10 05:02:02 -07:00
Palash Debnath 06c15ce37f fix(worker): identify a staged input the same way on every OS (#2005)
A staged task input's artifact id was built with os.path.join, so a Windows
control plane produced `inputs\<sha256>.wav`. That id is not a local path. It
is persisted into remote_tasks.params_json, shipped to remote workers over
gRPC as the identifier for the input they must fetch, and compared against a
later disk sweep to decide whether a staged file is still referenced.

So a Windows host hands a Linux worker `inputs\abc.wav`, where the backslash is
an ordinary filename character and no such file exists. Remote GPU workers are
a shipped feature; this broke them for every Windows control plane. The same
ids also stop matching when an omnivoice_data/ directory moves between
operating systems.

artifact_id_for() makes it canonical POSIX — resolve_within already treats both
separators as structural, so resolution is unchanged. normalize_artifact_id()
covers the upgrade: rows written by the old code carry a backslash, and the
sweeper decides "unreferenced" by comparing ids, so without it an upgraded
install reads every legacy row as garbage and deletes inputs that surviving
tasks still point at.

Two other tests in this run asserted POSIX-only behaviour rather than product
behaviour, and are corrected here too:

  - the durability-barrier test required a directory fsync, which
    _fsync_parent_directory deliberately skips without os.O_DIRECTORY. It now
    gates on that same attribute rather than on the OS name, so the test and
    the code it checks cannot drift apart.
  - the read-only-cache test built its scenario with chmod(0o500), which on
    Windows only toggles a read-only FILE attribute and does not stop a file
    being created inside the directory. It verifies its premise by probing and
    skips when the host writes anyway — which also covers root and anything
    holding CAP_DAC_OVERRIDE, replacing a geteuid check that named only one of
    them.

Then the reason none of this was visible: CI runs tests/ on Linux only. The two
worker suites join the existing Windows step in the smoke matrix. They need no
ffmpeg, so they cost seconds. Verified green on Windows first — 244 tests
across the four suites in that step.

Fails before, passes after, both directions: a staged id containing a
backslash, and a legacy-id input deleted by the sweeper.
2026-09-10 04:51:36 -07:00
Palash Debnath bb9149ce1f Merge pull request #2004 from debpalash/fix/grpc-loop-budget
test(worker): size the loop-responsiveness budget against what it measures
2026-09-10 04:40:54 -07:00
Palash Debnath 4c10e802ab test(worker): size the loop-responsiveness budget against what it measures
main is red. Smoke (Windows) failed on
test_upload_durability_barrier_does_not_block_the_grpc_loop with

    assert 0.20299999999997453 < 0.2

Three milliseconds of scheduling noise on a shared runner, and a red build
that says nothing about the product.

The three tests here prove a blocking filesystem call does NOT stall the gRPC
event loop: they park the call on a barrier and check the loop still ran their
own coroutine promptly. That is a wall clock on shared hardware, so the two
numbers have to be chosen against each other. The discriminator was a 0.5 s
watchdog — a stalled loop could not proceed until it fired — while the budget
sat at 0.2 s. Responsive measured ~0.2. The line was drawn exactly where the
noise lives.

Both numbers are named constants now, with the reasoning next to them:
a 1.5 s hold, a 0.75 s budget. Responsive lands near 0.2, stalled lands at 1.5,
and the line sits between them with room on both sides. The waiter's own cap
moved above the hold too, so a genuinely stalled loop is reported by the budget
assertion that names the problem rather than by a bare TimeoutError.

Verified the assertion is still worth having: with the blocking call made to
stall the loop for real, the test fails. It is a wider net, not a hole.

55 tests in the file, three runs in a row.
2026-09-10 04:15:06 -07:00
Palash Debnath c021fac1d8 Merge pull request #2003 from debpalash/land/2002-inert-badge
feat(pronunciation): badge a stored-but-inert entry in the list (#1949)
2026-09-10 04:05:28 -07:00
Palash Debnath 24e0737e0d fix(pronunciation): badge only the entries the backend calls inert
Both review bots found the same thing independently: `e.type !== 'respelling'`
badged rows the user had switched OFF. A disabled entry is indeed not applied,
but for a reason the toggle already shows — labelling it "not applied yet"
reads as a defect rather than their own choice. `inert_entries_for_language`
excludes disabled rows for exactly that reason, so the badge now matches it.

Not taken: the suggestion to name `ipa` and `cmu` explicitly instead of testing
against respelling. The backend's rule is that everything which is not
respelling is inert today, and mirroring it keeps the two in step. An explicit
list would silently stop badging a notation added later — an entry that saves,
validates, toggles on and quietly does nothing, which is the invisibility #1949
exists to remove. A test pins that direction with an unfamiliar type.

Two tests, one per direction. The disabled case fails without the enabled gate.
2860 vitest tests green.
2026-09-10 03:38:44 -07:00
Palash Debnath 06f7c4af7b feat(pronunciation): badge a stored-but-inert entry in the list (#1949)
Takes the part of #2002 by @utkarsha741 that #1984 did not already cover.

An IPA or CMU row saves, validates and toggles on, and is then dropped before
term matching — Phase 1 only substitutes respelling. #1984 made that visible in
the "Test a sentence" preview, which the user sees only if they run a test. The
entry LIST is where they look at what they have saved, and there it still
looked like every other working row.

So the row carries the same fact: a warning badge next to the type and scope
badges, on IPA and CMU only. Badging a respelling row would be the opposite
lie — those do take effect.

The rest of #2002 is already on main under a different name: it re-added the
backend skip detection and the test-preview line as `skipped_terms`, where
`inert_entries_for_language` and `inert_entries` have shipped since #1984.
Landing that half would have been a second implementation of one behaviour with
two response fields for it.

String added to all 21 locales, not just en, so the badge is not an English
island in a translated panel. Fails before, passes after: with the badge
condition disabled the new test cannot find it. 2858 vitest tests, 529 locale
and CJK guard tests green.
2026-09-10 03:26:44 -07:00
Palash Debnath c9a587ea75 Merge pull request #2000 from debpalash/land/1998-powershell-key
docs(docker): make the PowerShell key URL-safe too (#1998)
2026-09-10 03:21:39 -07:00
Palash Debnath cef2b73e02 Merge remote-tracking branch 'origin/main' into land/1998-powershell-key 2026-09-10 02:52:35 -07:00
Palash Debnath 50b1184df3 Merge pull request #1996 from debpalash/fix/1933-port-holder
fix(backend): name who actually holds port 3900 (#1933)
2026-09-10 02:52:16 -07:00
Palash Debnath c48e8c4ff5 Merge remote-tracking branch 'origin/main' into land/1998-powershell-key 2026-09-10 02:22:57 -07:00
Palash Debnath 0c9f1bd7c4 Merge remote-tracking branch 'origin/main' into fix/1933-port-holder
# Conflicts:
#	CHANGELOG.md
2026-09-10 02:22:49 -07:00
Palash Debnath b6a0f96e90 Merge pull request #1992 from debpalash/fix/1900-attempt-id
fix(bootstrap): emit an attempt id so the splash stops guessing (#1900)
2026-09-10 02:22:28 -07:00
Palash Debnath 93771a0201 Merge pull request #2001 from debpalash/fix/worker-barrier-flake
test(worker): stop spinning the loop the awaited work needs
2026-09-10 02:22:22 -07:00
Palash Debnath a4f87bad6e Merge remote-tracking branch 'origin/main' into fix/1933-port-holder
# Conflicts:
#	CHANGELOG.md
#	frontend/src-tauri/src/backend.rs
2026-09-10 01:53:46 -07:00
Palash Debnath 703c665c6c Merge remote-tracking branch 'origin/main' into fix/worker-barrier-flake 2026-09-10 01:52:50 -07:00
Palash Debnath f301db1b02 Merge remote-tracking branch 'origin/main' into land/1998-powershell-key 2026-09-10 01:52:41 -07:00
Palash Debnath c90d32845a Merge remote-tracking branch 'origin/main' into fix/1900-attempt-id
# Conflicts:
#	CHANGELOG.md
2026-09-10 01:52:33 -07:00
Palash Debnath 9a7f5fd7ed Merge pull request #1994 from debpalash/fix/1850-crash-tail
fix(crash): capture the dying backend's last words, not the log so far (#1850)
2026-09-10 01:52:06 -07:00
Palash DebnathandClaude Opus 5 25393b4654 test(worker): stop spinning the loop the awaited work needs
Smoke (Windows) fails intermittently in
test_revocation_during_result_barrier_cannot_ack_published_bytes with a bare
TimeoutError. It hit two PRs in a row today, one of them documentation-only,
which rules out any change under review.

The wait is a busy loop:

    while not barrier_finished.is_set():
        await asyncio.sleep(0)

asyncio.sleep(0) yields to the loop but never sleeps, so this runs the loop
flat out on the one thread the upload task also needs to reach
_durable_replace and set the event. On a loaded Windows runner the waiter
starves the worker it is waiting for, and the 1 s cap fires with nothing
actually wrong — a failure with no signal in it, which is worse than no test.

_await_event parks the wait on a worker thread with asyncio.to_thread, leaving
the loop free. Deterministic, and faster: the test drops from a full second of
spinning to the time the work actually takes.

Deliberately narrow. The other spin-waits in this file sit inside
"assert elapsed < 0.2" blocks that exist to prove the gRPC loop stayed
RESPONSIVE during a blocking call — spinning is the measurement there, and
converting them would delete the assertion's meaning.

123 tests across both worker files pass; the target test passes three runs in
a row.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 01:42:15 -07:00
Palash DebnathandClaude Opus 5 e11270990b fix(backend): identify the port holder by the marker header, not a body sniff
CodeRabbit: running_backend_version accepts a /system/info body that merely
CONTAINS "model_checkpoint" or "data_dir". That is a substring sniff, and it is
fine for the decision it was written for — whether to attach to a healthy
same-version backend. It is not fine for this one, which ends in a message
naming a process for the user to kill. Any service can serve that body.

port_holder now requires the x-omnivoice-backend header that backend/main.py
stamps on every response, the same gate startup_progress already applies for
the same reason: a foreign process on our port must not narrate our UI, and it
certainly must not be the thing we point a user's kill command at. Unmarked
means Foreign, which keeps the conservative wording and offers no command.

Two tests against a real one-shot loopback responder: a spoofed body with no
marker is Foreign and gets no terminal command, and the marker is what makes a
responder ours. The first fails with the header check disabled.

244 lib tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 01:13:28 -07:00
Palash DebnathandClaude Opus 5 f21ad9a27d fix(crash): close the far end of a dead run's slice, and settle one at a time
Two more review findings, both real.

CodeRabbit: pinning the START of the dying run's slice was only half the fix. A
start with an unbounded end still does not identify one run — the replacement
writes BELOW those lines, and a tail reads the last N of the file, so the newer
run's healthy startup is exactly what the dead run's crash marker would get.
read_dead_run_tail closes the range: it is called after the settle, and takes
the end from wherever the current run now begins, which is either still the
pinned start (nothing replaced it) or the replacement's own offset — precisely
where this run's slice ends. An end that is not a usable boundary degrades to
the rest of the file, matching how an unusable start already degrades.

CodeRabbit: settle_err_log moved every handle out of the list and then waited
without that lock, so two callers could interleave — the second found an empty
list, concluded there was nothing to wait for, and read the log while the first
was still waiting for exactly the drainer it needed. A settlement lock makes
each caller's return mean the waiting is genuinely done.

Two tests: a dead run's slice stops where the replacement begins (and the
unbounded read really does return the newer run, so the assertion is not
vacuous), and an unusable end degrades rather than capturing nothing.

241 lib tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 01:11:46 -07:00
Palash Debnath fc45ec3b5b Merge remote-tracking branch 'origin/main' into land/1998-powershell-key 2026-09-10 01:08:11 -07:00
Palash Debnath 1eda03827d Merge remote-tracking branch 'origin/main' into fix/1900-attempt-id 2026-09-10 01:08:03 -07:00
Palash Debnath 555210bbdb Merge remote-tracking branch 'origin/main' into fix/1933-port-holder 2026-09-10 01:07:55 -07:00
Palash Debnath 78d6dbe802 Merge remote-tracking branch 'origin/main' into fix/1850-crash-tail 2026-09-10 01:07:46 -07:00
Palash Debnath e3efe4a39a Merge pull request #1999 from debpalash/fix/windows-backend-tests
test: make the isolated backend session pass on Windows, and gate it there
2026-09-10 01:07:27 -07:00
Palash DebnathandClaude Opus 5 36a9515320 docs(docker): make the PowerShell key URL-safe too
Lands #1998 by @yangfan-yf-yf. Correct finding: the PowerShell example I added
in #1993 generated the administrator key with `python -c`, and the whole point
of the Docker path is that the host does not need Python. On a Windows host
without it, the very first line of the setup fails.

One thing on top. The key is also accepted as an `?api_key=` query parameter
(core/auth.py), and raw Base64 carries `+`, `/` and `=`. A `+` in a query
string decodes to a space, so a user who pasted such a key into a URL would get
a silent mismatch with nothing to explain it. The Bash line next to it uses
`secrets.token_urlsafe` and never had this shape, so the two now agree:
trim the padding, map `+` to `-` and `/` to `_`.

Verified in Windows PowerShell 5.1 (5.1.26100): the block parses and runs, and
the key is 43 URL-safe characters — the same shape `secrets.token_urlsafe(32)`
produces. validate-install-docs.py and both docker/changelog test files pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 01:00:14 -07:00
Palash Debnath 9618f26f6e Merge remote-tracking branch 'origin/main' into fix/1900-attempt-id
# Conflicts:
#	CHANGELOG.md
2026-09-10 00:58:27 -07:00
Palash Debnath 9e12c7b181 Merge remote-tracking branch 'origin/main' into fix/1933-port-holder
# Conflicts:
#	CHANGELOG.md
2026-09-10 00:58:18 -07:00
Palash Debnath baaef34a4a Merge remote-tracking branch 'origin/main' into fix/1850-crash-tail
# Conflicts:
#	CHANGELOG.md
2026-09-10 00:58:09 -07:00
Palash Debnath 5b842280c8 Merge pull request #1997 from debpalash/fix/1931-blackwell-docs
docs+test: finish the RTX 50-series story (#1931)
2026-09-10 00:56:32 -07:00
Palash DebnathandClaude Opus 5 9d71c7cedd test: pin the port diagnosis to the matcher, not to one sentence
The backend-lifecycle harness asserted a literal — "is already in use, so the
backend could not" — while its own comment said the point was "the exact
phrasing BootstrapSplash.detectHints localizes". Those are not the same thing,
and the gap showed: rewording the message by who actually holds the port kept
the matcher firing and still failed the test.

It now asserts the real contract, the same one the Rust unit tests pin: the
message mentions a port, says it is in use after that, and names the port
number. Any wording that satisfies detectHints satisfies this; any that does
not, fails — which is the failure worth catching, because it silently costs
the user the localised hint.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 00:45:34 -07:00
Palash Debnath 770fe35da9 Merge pull request #1995 from debpalash/fix/1927-crash-hint
fix(crash): say what the exit code means in the crash details (#1927)
2026-09-10 00:36:16 -07:00
Palash DebnathandClaude Opus 5 db48c7851c test: make the isolated backend session pass on Windows, and gate it there
Four tests in backend/tests/ cannot pass on a stock Windows checkout. The
`test` job runs that session on Linux only, so all four were invisible to CI
and hit every Windows contributor on their first `pytest` run — with failures
that have nothing to do with whatever they changed. Same class as #1990.

  - test_contained_subprocess_waitid_fallback.py simulates macOS by deleting
    os.waitid, then drives the fallback with os.waitpid/os.WNOHANG and
    start_new_session. Windows has none of those; os.WNOHANG is an
    AttributeError before the first assertion. The module is POSIX-only by
    premise, so it says so.
  - test_invalid_or_missing_desktop_drain_fd_fails_safe asserts a RuntimeError
    that cannot be raised off POSIX: backend_drain_fd returns None there before
    it reads the environment. The file already had this skipif on its sibling.
  - test_mps_proxy_survives_fatal_child_exit_and_recovers raced the OS. The
    child calls os._exit and the parent raises the moment its pipe hits EOF —
    before the process is reaped. Asserting poll() on the next line is a race
    Linux won and Windows lost every time. It waits for the death now, which is
    what the test actually claims.

Then the reason all four survived: nothing runs this session on Windows. The
smoke matrix already does a full `uv sync` there, so the session costs forty
seconds and now runs as a step in it. Verified green on Windows before adding
the gate — 355 passed, 8 skipped — so this cannot break main.

Kept to Windows deliberately: that is the platform I can verify here, and a
gate added blind on macOS would be a guess about a host I cannot run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 00:34:58 -07:00
Palash DebnathandClaude Opus 5 512c2fa7d6 fix(bootstrap): bind the attempt id to the work it labels
Four review findings on the attempt id, all real.

Greptile, P1 — the status snapshot could mismatch. The bump and the stage write
were separate, so a restart landing between them returned the PREVIOUS attempt's
stage stamped with the new attempt's id: the splash then recorded
`installing_deps` as work this attempt did, which is precisely the #1894
fabrication the id exists to remove. `begin_attempt_with` now takes the stage
lock across both writes, and `bootstrap_status` reads stage and attempt under
that same lock. The pair a reader observes is always self-consistent.

CodeRabbit, major — an output pump is a thread reading a pipe, and it outlives
the run it drains. It stamped each line with the counter's value at read time,
so a restart relabelled the dying run's trailing output as the new attempt's
evidence. `emit_log_for_attempt` takes the attempt explicitly, and all four
pumps (the backend's stdout and stderr, and both sides of `run_streaming`)
capture theirs when they start. Every other call site runs inside the attempt
it describes and keeps reading the counter.

CodeRabbit, minor — the tests that advance the process-global counter raced
each other under cargo's threaded runner, so one could read a value another had
just moved. They serialize on a lock now, like the env-var tests above them.

CodeRabbit, minor — the backfill-to-live seam deduplicated on stage plus text,
and installer output repeats itself constantly. Across a restart that is not a
replayed line, it is the new attempt's own evidence, and dropping it can remove
the only proof for a stage the poll never samples. The attempt is part of the
identity now.

Two new tests: a new attempt never carries the previous stage, and a repeated
line belonging to a different attempt is kept. The second fails against the
previous dedup key. 241 Rust lib tests and 2853 vitest tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 00:15:33 -07:00
Yang Fan 55e6d9a07b docs(docker): generate PowerShell keys without Python 2026-09-10 15:14:00 +08:00
Palash DebnathandClaude Opus 5 192c1f5127 fix(crash): pin the run slice before settling, and keep every drainer
Two review findings on the settle, both real.

Greptile, P1: settling can take up to two seconds, and a Retry arriving in that
window installs a new run and moves ERR_LOG_RUN_START past the dying run's
output. Reading "the current run" after the wait would then hand the dead
process's crash marker the REPLACEMENT's healthy startup — the cross-run
attribution #1510 exists to prevent, reintroduced through the wait added to fix
the tail. Every death path now pins the offset BEFORE settling and reads from
it, via read_error_log_tail_from.

CodeRabbit: a single drainer slot loses a timed-out handle the moment a new run
installs its own. Dropping a JoinHandle detaches the thread, so nothing can
ever wait for that run's output again and both guarantees quietly stop holding.
The slot becomes a list: a settle drains it, joins what finished, and puts back
what is still running, ahead of anything a concurrent spawn pushed.

Two tests: a pinned offset still names the dying run's slice after a respawn
moved the current one, and an unfinished drainer survives another run
installing its own. 239 lib tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 00:10:53 -07:00
Palash DebnathandClaude Opus 5 7610a0d86e fix(backend): show the listener before ending it, never a blind kill
Greptile, security, on the reclaim guidance: the convenient one-liner does not
preserve the identity port_holder established.

  - `lsof -ti tcp:3900` matches CONNECTED CLIENTS as well as the listener, so
    piping it into kill can end a process that merely talks to VoiceStudio.
  - Windows `findstr :3900` matches `:39001` and established connections too.

And the identity itself is a fact about the moment the message was written. By
the time a user runs a command it has to be re-established, and only they can
do that.

So both platforms now get two steps: a lookup restricted to the LISTENING
socket that prints the pid and process name, and a kill of that pid once the
user has confirmed what it is. A test pins that the guidance never pipes a
lookup into kill and always shows something to confirm.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 00:08:22 -07:00
Palash Debnath 2afea02c94 Merge remote-tracking branch 'origin/main' into fix/1933-port-holder
# Conflicts:
#	CHANGELOG.md
2026-09-10 00:05:45 -07:00
Palash Debnath 46b1fb2f57 Merge remote-tracking branch 'origin/main' into fix/1927-crash-hint
# Conflicts:
#	CHANGELOG.md
2026-09-10 00:05:39 -07:00
Palash Debnath ac427fb821 Merge remote-tracking branch 'origin/main' into fix/1850-crash-tail
# Conflicts:
#	CHANGELOG.md
2026-09-10 00:05:35 -07:00
Palash Debnath 924172e012 Merge pull request #1986 from debpalash/fix/1960-name-the-language
fix(dub): name the source language code that was rejected
2026-09-10 00:04:53 -07:00
Palash DebnathandClaude Opus 5 14d6b90836 docs+test: finish the RTX 50-series story (#1931)
The code half of #1931 landed already: `torchaudio.set_audio_backend()` is
guarded, so the torch 2.9.x upgrade a Blackwell card needs no longer trades one
`ml_imports` crash for another. Two things were still missing.

The changelog said the upgrade was documented. It was not — nothing in docs/
mentions sm_120, Blackwell, or the 50-series at all, so a user hitting a native
access violation inside `import torch` had the issue thread and nothing else.
troubleshooting.md now carries it: why the pinned torch 2.8.0 cannot work
(no sm_120 kernels in the wheel — not a setting, not a workaround), the trio
that has to move together, the verification command that proves the kernels
arrived, and the fact that the change is to the repo's own pins so a later pull
will undo it.

The part most likely to be missed is that there are TWO pin lists.
`constraint-dependencies` governs `uv sync`/`uv lock`/`uv run`;
deploy/torch-constraints.txt governs the `uv pip install` paths, which ignore
project-level uv settings. Editing one leaves the other behind, which is what
`RuntimeError: operator torchvision::nms does not exist` looks like from the
outside. Both are named.

The second gap: nothing protected the guard. CI runs the pinned torch 2.8.0,
where `set_audio_backend` still exists, so deleting the `hasattr` as a
"simplify this no-op" cleanup would pass every test in the suite and restore a
hard startup crash for every RTX 50-series user. tests/ now walks the backend
AST and fails on any reach for a torchaudio API that 2.9 removed unless
something proves it is there — a `hasattr`/`getattr` check or a `try`. Fails
with the guard removed, passes with it.

Not addressed here, because it is already fixed: the reporter's third
observation, that launching through the desktop shell hung inside `import
torch`'s native init, is the OpenBLAS/stdin-pipe deadlock closed under #1952.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-10 00:03:56 -07:00
Palash Debnath 2cde7663a1 Merge remote-tracking branch 'origin/main' into fix/1900-attempt-id
# Conflicts:
#	CHANGELOG.md
2026-09-10 00:00:42 -07:00
Palash Debnath 2bf922d9e5 Merge pull request #1993 from debpalash/land/1987-followup
docs(docker): finish the ARM64 Compose guidance (#1987)
2026-09-10 00:00:11 -07:00
Palash Debnath d52a4c89a7 Merge pull request #1991 from debpalash/fix/windows-symlink-tests
test: let a stock Windows checkout run the symlink tests (#1990)
2026-09-09 23:59:58 -07:00
Palash DebnathandClaude Opus 5 6be80a4b3d fix(backend): name who actually holds port 3900 (#1933)
The port-conflict failure asserted "already in use by another application"
without ever asking who held the port. In the reports behind #1933 (and its
duplicates #1935, #1936, #1937 — same machine, same 45-minute window) the
holder was the user's OWN orphaned backend from an earlier run. So the app
told them to quit a copy of VoiceStudio that has no window, and there was no
action in the message that could have worked.

@Chang-Jin-Lee diagnosed this precisely on #1936, including the observation
that the identity check already exists: `running_backend_version(port)` asks
`/system/info` who is there, and is already trusted for the more consequential
decision of whether to attach to a healthy same-version backend. It simply was
not consulted on this path.

So it is now. `port_holder()` returns one of three answers, and
`port_conflict_message()` words the failure from it:

  - our own backend at this version (or one too old to report one) — say so,
    say it has no window to quit, and give the terminal command that ends it;
  - our own backend at a different version — name the version, which is what
    identifies it, and give the same command;
  - anything else — the existing wording, now actually justified.

The terminal command is only ever offered for a listener that identified
itself as ours. An unidentified one keeps the conservative wording: a user must
never be told to go kill a process that may not be theirs. A listener that
accepts a connection but does not answer `/system/info` counts as
unidentified, which is the reading that cannot do harm.

All three sites that reported this — take-ownership, respawn, and the
early-exit path on EXIT_PORT_IN_USE — go through the one builder now.

What is deliberately unchanged: `kill_orphan_on_port` still refuses to signal
a PID discovered through lsof/netstat. That refusal is correct — the reuse race
is real, and a matching foreign service must never be terminated. This changes
what the user is told, not what the app is willing to kill.

Five Rust tests, one per branch plus the suffix behaviour, and one that pins
every wording against the `detectHints` matcher — that regex is what turns
these English strings into the localised `bootstrap.hint_port`, and an earlier
draft of one message silently lost the translation by saying "is held by". The
frontend test that pinned the old literals now pins the new ones.

241 Rust lib tests and 2851 vitest tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:59:09 -07:00
Palash DebnathandClaude Opus 5 f67e5da289 fix(crash): say what the exit code means in the crash details (#1927)
The report in #1927 is a Windows access violation 13 seconds after "Loading
VoiceStudio model on device: cuda" — a native fault inside the compute stack,
which produces no Python traceback because the process is executing bad machine
code. What the user was shown was "Backend died (exit code -1073741819)", a
timestamp, an uptime, and a log ending mid-startup. The issue they filed has an
empty description, which is the honest response to being handed a number and no
next step.

The classification already existed and is good: `crashCauseHint` distinguishes
a native fault (a GPU driver disagreeing with the bundled CUDA runtime, or a
partially downloaded weight file), an exit 78 port conflict, an OOM kill and a
half-built Python environment, and names concrete actions including the
crash-isolated engines. It just never reached this surface — the only place it
rendered was the message on a stream dropped by a crash, and a crash with no
request in flight has no stream to drop.

So the details dialog renders it. A sentinel marker is deliberately excluded:
it cannot know a crash happened at all (sleep, force-quit and a stopped VM
leave the same trace), so it has no cause to explain, and asserting one would
be the #1375 fabrication in a new place.

Three tests: the access violation gets the compute-stack guidance, a port
conflict gets its own rather than the GPU one, and a sentinel gets none. The
first two fail against the previous component.

2853 vitest tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:51:50 -07:00
Palash Debnath b945ddd163 Merge remote-tracking branch 'origin/main' into tmp/1986
# Conflicts:
#	CHANGELOG.md
2026-09-09 23:48:36 -07:00
Palash Debnath 525a2809eb Merge pull request #1989 from debpalash/land/1981-structure
docs: refresh STRUCTURE.md and pin its counts to the tree (#1981)
2026-09-09 23:48:11 -07:00
Palash DebnathandClaude Opus 5 1b7fcf92c3 fix(crash): capture the dying backend's last words, not the log so far (#1850)
The crash report in #1850 carries a stderr tail that stops 58 seconds before
the death it is meant to explain. That is not a quiet backend — it is a race.

`wait()` returns the moment the child exits, but stderr is drained by a
separate Rust thread reading a pipe and appending to backend_err.log. The two
crash-marker sites read that file immediately on detecting death, so the
drainer's in-flight lines — the traceback that names the cause — land after the
tail is taken. The report then shows a log that simply stops, and the crash is
undiagnosable no matter how good the rest of the capture is. Every silent
"exit code 1" report is a candidate for this.

The machinery to wait already existed for a different reason: #1510 joins the
drainer before a respawn records its start offset, so a dying run's buffered
tail cannot be attributed to the new run. It was just never applied to the
death paths. `join_previous_err_drainer` becomes `settle_err_log`, called from
the crash-marker sites in both the startup and supervisor paths as well as
before a respawn.

One behaviour change while it moves: when the bound expires the handle is now
handed back rather than dropped. Dropping detaches the thread, and every later
caller — including the respawn that #1510 protects — silently loses the ability
to wait for that run's output at all. Bound stays at 2 s, so a wedged pipe still
cannot stall crash recording.

Regression tests: a drainer that writes a traceback 120 ms after death (the
tail contains it now, contains only "steady state" before), and a wedged
drainer that outlives the bound (the slot still holds it). The three tests that
install into the process-global drainer slot now serialize on a lock — they
were racing each other under cargo's threaded runner.

238 lib tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:47:18 -07:00
Palash DebnathandClaude Opus 5 78a6c489c9 docs(docker): finish the ARM64 Compose guidance (#1987)
@yangfan-yf-yf pushed two more commits to #1987 after the first pass landed.
Two things in them were worth taking:

  - an explicit `compose pull` step, so the platform override is proven before
    `up -d` rather than discovered when the pull inside it fails; and
  - a PowerShell form. An ARM64 Windows host cannot use `export`, and the
    surrounding page only ever shows Bash — so the guidance did not actually
    reach the users most likely to need it.

Not taken: the same commits also moved `--platform linux/amd64` into the
default `docker pull` / `docker run` quick start. That is a no-op for the
amd64 majority and contradicts the Architecture section directly above, which
introduces the flag as the conditional ARM64 step. The canonical command stays
the one almost everyone should run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:41:44 -07:00
Palash Debnath 5e4f863c4d Merge pull request #1988 from debpalash/land/1987-docker-arch
docs(docker): document the amd64-only images and the ARM64 workaround (#1987)
2026-09-09 23:39:41 -07:00
Palash Debnath 12794f3dae Merge pull request #1985 from debpalash/fix/1773-classic-error-class
fix(errors): carry the backend error class onto a classic 500
2026-09-09 23:39:30 -07:00
Palash DebnathandClaude Opus 5 0ce5cca1b5 fix(bootstrap): emit an attempt id so the splash stops guessing (#1900)
The first-run splash decided which bootstrap attempt a piece of evidence
belonged to by inferring attempt boundaries from the `bootstrap_status` stage,
which is sampled about once a second. Inference from a sampled signal cannot be
airtight, and two routes slipped through it:

  - a retry that goes failed -> checking -> starting_backend inside one sample
    window, where the poll sees no restart stage at all; and
  - the supervisor's own venv rebuild, which re-enters `checking` with no
    `failed` stage and no click behind it. If the poll samples the same stage
    name on either side of it, the sequence is `installing_deps` ->
    `installing_deps` — literally no signal that anything restarted, and the
    previous attempt's completed steps stayed on screen as this attempt's work.
    That is the #1894 fabrication arriving by a route stage inference cannot
    close.

The producer knows the answer exactly, so it now says so. `ATTEMPT` is a
monotonic counter bumped wherever the bootstrap really restarts —
`respawn_backend`, which both retry commands and the scoped reset funnel
through, and the automatic venv rebuild. `bootstrap_status` returns it beside
the stage (a flattened `BootstrapStatus`, so the wire shape the frontend
already reads is unchanged), and every `bootstrap-log` line carries it too.

The splash scopes stage evidence by equality on that id and the boundary
heuristics are gone: `RESTART_STAGES`, the leaving-`failed` rule, the
wall-clock `attemptStart`, and the `selfInitiatedRef` guard that existed only
to stop the poll re-stamping a boundary the UI had already opened. `beginAttempt`
is now presentation only — it clears the visible log for a retry the user asked
for.

Two new tests cover what only an id can carry: a Rust-side restart the poll
cannot see at all, and a log line from the previous attempt that must not count
toward this one. Both fail against the previous component and pass now. Rust
side: 240 lib tests green, including four on the counter and the status shape.
Frontend: 2847 vitest tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:38:39 -07:00
Palash Debnath 744a0550d3 Merge remote-tracking branch 'origin/main' into tmp/1986
# Conflicts:
#	CHANGELOG.md
2026-09-09 23:25:50 -07:00
Palash DebnathandClaude Opus 5 56e9e42ac2 test: let a stock Windows checkout run the symlink tests (#1990)
Creating a symlink on Windows needs SeCreateSymbolicLinkPrivilege, which a
normal account does not hold unless Developer Mode is on. GitHub's hosted
Windows runners hold it, so seven unguarded call sites passed in CI and failed
only on a contributor's own machine, with WinError 1314 and no connection to
whatever they were working on:

  tests/backend/services/test_audiocpp_backend.py  (5)
  tests/test_exports_api.py                        (1)
  tests/test_storage_report.py                     (1)

The repo already knew about this — tests/test_hf_cache_repair.py carries a
private _symlink_or_skip helper whose docstring describes exactly this failure.
The pattern simply never reached the other files, which is the whole class of
the bug: a convention that lives in one module's private helper gets rewritten
from scratch, or forgotten, at every new call site.

So the helper is now a `symlink_or_skip` fixture in tests/conftest.py, and
tests/test_symlink_guards.py walks the AST of every test module and fails on a
raw symlink_to / os.symlink that has no way to skip. Guarded means the fixture,
a try, a skipif marker (module-level pytestmark included), or a test that has
already run a skipping helper — the three legitimate existing patterns, which
it recognises rather than forcing a rewrite.

Coverage is unchanged: the full pytest job runs on Linux, where nothing skips.
Fails before (7 errors, then the guard reports the offending files), passes
after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:25:16 -07:00
Palash Debnath c2f27c39d2 Merge remote-tracking branch 'origin/main' into tmp/1985
# Conflicts:
#	CHANGELOG.md
2026-09-09 23:24:18 -07:00
Palash Debnath 60b696f37f Merge pull request #1984 from debpalash/fix/1949-phoneme-visible
fix(pronunciation): say when an IPA/CMU entry is stored but not applied
2026-09-09 23:23:18 -07:00
Palash DebnathandClaude Opus 5 5237f7849a docs: keep STRUCTURE.md's README annotation in English
tests/test_no_hardcoded_cjk.py rejects CJK outside frontend/src/i18n/, and the
refreshed tree annotated README_CN.md with the characters themselves. The file
name already says which language it is; the annotation does not need to be in
it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:22:58 -07:00
Palash DebnathandClaude Opus 5 a5f4d61eda docs: keep STRUCTURE.md's counts honest with a test
Lands #1981 by @Dawcraft, which refreshes docs/STRUCTURE.md to match the tree
as it actually is — the old file still described a root-level layout that the
2026-07-12 cleanup removed, and pointed at a tests/services/ mirror that has
not existed since the tests/backend/ reorganisation.

Verified every path, directory and CI claim in the refreshed file against the
repo: the router auto-include list, the isolated backend/tests/ pytest step,
the smoke-matrix job and its HF_HUB_OFFLINE guard, and every file the tree
names. One number was off — backend/services/ holds 78 modules, not 79.

Off-by-one in a doc is the symptom; the class is a count nothing checks, which
is wrong the week after it is written. tests/test_structure_doc.py now pins
the router count, the service count and the engine-adapter list to the tree,
so the next module to land fails the suite with the line to update instead of
quietly aging the doc. Fails before the fix (79 != 78), passes after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:06:59 -07:00
Palash Debnath c7d0e314d9 Merge branch 'pr/1981' into land/1981-structure 2026-09-09 23:03:14 -07:00
Palash DebnathandClaude Opus 5 c4f214858b docs(docker): cover Compose in the ARM64 guidance
Lands #1987 by @yangfan-yf-yf, which closes #1921.

The published images are linux/amd64 only, and the quick start reached image
resolution before saying so — an ARM64 user met "no matching manifest for
linux/arm64/v8" with no explanation. Verified against docker.yml, which says
so in its own comment: "only building linux/amd64".

One gap in the original: the platform override was documented for docker pull
and docker run, but Compose has no per-command --platform flag, so the
recommended Compose command still resolved the missing ARM64 manifest and
failed exactly as before. DOCKER_DEFAULT_PLATFORM covers it, with the same
caveat the rest of the section makes — emulation, not native support, and only
the CPU profile makes sense under it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 23:00:44 -07:00
Palash Debnath 7aa2333823 Merge remote-tracking branch 'origin/pr/1987' into land/1987-docker-arch 2026-09-09 23:00:14 -07:00
Palash Debnath fe012ad3a5 Merge remote-tracking branch 'origin/main' into fix/1960-name-the-language
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:58:31 -07:00
Palash Debnath 1b7e858886 Merge remote-tracking branch 'origin/main' into fix/1773-classic-error-class
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:58:26 -07:00
Palash Debnath fd74865f7c Merge remote-tracking branch 'origin/main' into fix/1949-phoneme-visible
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:58:22 -07:00
Palash Debnath 0f237a9b25 Merge pull request #1983 from debpalash/fix/1847-bootstrap-log-file
fix(bootstrap): keep the first-run install log after setup finishes
2026-09-09 22:57:53 -07:00
Yang Fan ae25a6a594 docs: clarify Docker image architecture requirements 2026-09-10 13:51:58 +08:00
Palash DebnathandClaude Opus 5 aa984075a7 test(dub): pair each rejected code with its own response
The two requests in this case send different bad codes; asserting one string
against both bodies passed on whichever happened to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 22:46:12 -07:00
Palash DebnathandClaude Opus 5 80e0b91ce8 test(dub): assert the substance of the rejection, not its exact wording
The existing case pinned the literal string "Invalid source language code",
which the #1960 fix replaces with a message that names the offending code. It
now asserts what the test is actually about — a 400 that identifies the code —
so improving the guidance again does not fail it for the wrong reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 22:44:48 -07:00
Palash DebnathandClaude Opus 5 d1e4c847cb fix(dub): name the source language code that was rejected
Closes #1960.

The report was "400 Bad Request: Invalid source language code" and nothing
else. That cannot be acted on or triaged: it does not say which of the ninety
or so codes was wrong, so neither the user nor a maintainer reading the
auto-filed issue can tell whether the picker offered something the backend does
not accept, or a stale preference from an older build is still being sent.

I could not determine the cause from the report, which is exactly the problem.
Naming the code makes the next one answerable instead of guessing at this one.

The value is a language code chosen from a menu, not private data, and the
engine validator a few lines away already echoes its input the same way.

Also adds the check I actually wanted while investigating: a test that reads
the picker's own LANG_CODES and asserts the backend accepts every one of them,
so a code added to the menu cannot silently become a 400. It passes today —
the menu and the allowlist do agree — which is how I ruled that out as the
cause rather than assuming it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 22:42:27 -07:00
Palash Debnath 91515310f0 Merge remote-tracking branch 'origin/main' into fix/1773-classic-error-class
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:40:35 -07:00
Palash Debnath 6087085d43 Merge remote-tracking branch 'origin/main' into fix/1949-phoneme-visible
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:40:30 -07:00
Palash Debnath 2f66e95958 Merge remote-tracking branch 'origin/main' into fix/1847-bootstrap-log-file
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:40:26 -07:00
Palash Debnath f905230fd8 Merge pull request #1982 from debpalash/fix/1974-dev-port-ownership
fix(dev): reclaim the port from a backend the app itself left running
2026-09-09 22:39:58 -07:00
Palash DebnathandClaude Opus 5 b6ed598770 fix(errors): carry the backend error class onto a classic 500
Closes #1773.

The 500 handler has always put error_class in the response body, but nothing
lifted it onto the Error object — and the auto bug reporter reads the Error. So
every unclassified 500 filed "VoiceStudio hit an internal error; check the
backend log for details." and nothing else: identical reports, none of them
triageable, with the distinguishing datum sitting unused in the payload that
produced them.

#1956 did exactly this for the streaming path. The classic path had been
carrying the field on the wire the whole time; it just never survived the hop
onto the exception.

Only a string is kept. A 404 or a validation error has no class, and an empty
one would put a blank line in every report; a non-string is ignored rather than
stringified. Both pinned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 22:25:56 -07:00
Palash Debnath ac09da3c24 Merge remote-tracking branch 'origin/main' into fix/1949-phoneme-visible
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:22:37 -07:00
Palash Debnath ef8636ee78 Merge remote-tracking branch 'origin/main' into fix/1847-bootstrap-log-file
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:22:31 -07:00
Palash Debnath 72249a8d8e Merge remote-tracking branch 'origin/main' into fix/1974-dev-port-ownership
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:22:26 -07:00
Palash Debnath 8e2928ca56 Merge pull request #1980 from debpalash/fix/1858-shortcut-registration
fix(dictation): say when another app already owns the shortcut
2026-09-09 22:22:01 -07:00
Palash DebnathandClaude Opus 5 de962d164a fix(pronunciation): say when an IPA/CMU entry is stored but not applied
Closes #1949.

Settings offers three notations. Only Respelling substitutes text today; IPA
and CMU rows save cleanly, are validated, get a badge and can be toggled on,
then get dropped before term matching and are never read again.

That much is Phase 1 behaving as designed. The defect is that it was INVISIBLE:
"Test a sentence" answered "No entries match — spoken as written" for a term
that does match. Not a degraded answer, a wrong one — and it sent the user off
to re-type an entry that was already correct, or to convert it to Respelling,
where a phoneme string is then read as graphemes.

docs/specs/01-expressive-tts.md asked for exactly the opposite: such entries
"passed through and flagged 'phoneme not honored on this engine' (parity-rule:
visible degradation)". That flag was never implemented. This is it.

The dry run reports inert entries separately, and the panel names them. The
substitution path is deliberately untouched — this does NOT start feeding raw
phoneme strings into the grapheme stream, which is the thing Phase 1 refuses on
purpose, and a test pins that it still refuses.

Not Phase 2. Lowering IPA/CMU to engine markup is a real feature per engine and
stays open; what changes here is that the gap is now honest rather than silent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 22:11:05 -07:00
Palash Debnath 3d76f1ca26 Merge remote-tracking branch 'origin/main' into fix/1847-bootstrap-log-file
# Conflicts:
#	CHANGELOG.md
#	frontend/src-tauri/src/bootstrap.rs
2026-09-09 22:06:20 -07:00
Palash Debnath 6203fe41e9 Merge remote-tracking branch 'origin/main' into fix/1974-dev-port-ownership
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:04:53 -07:00
Palash Debnath 8e87755845 Merge remote-tracking branch 'origin/main' into fix/1858-shortcut-registration
# Conflicts:
#	CHANGELOG.md
2026-09-09 22:04:48 -07:00
Palash Debnath 3013cdc241 Merge pull request #1979 from debpalash/fix/1898-windows-quit
fix(windows): stop reporting every deliberate quit as a crash
2026-09-09 22:04:19 -07:00
Palash DebnathandClaude Opus 5 939ed4b12d fix(bootstrap): keep the first-run install log after setup finishes
Closes #1847.

The splash is the only surface with a Show/Copy affordance for these lines, and
it unmounts the moment the stage flips to ready — so on a successful first run
the whole install log was gone for good, with no completion pause and nowhere
to retrieve it. A user who wanted to check what had just been installed, or
attach it to a bug report, had nothing.

The lines are written to bootstrap.log beside backend.log now, so everything
about a run is in one directory and a bug report does not have to hunt in two.

Truncated once per process rather than appended forever: a bootstrap is a
single episode and the useful question is always "what happened this time".
That also bounds the file across repeated retries without needing a hook on
every restart path. The docs say so, and say to copy it first if you need a
superseded attempt.

Best effort throughout — a log that cannot be written must never take the
bootstrap down with it, and a test pins that it does not.

The counter half of this issue (Activity frozen at 200) was already fixed on
main by #1918; I verified that before starting rather than assuming the whole
issue was open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 21:48:53 -07:00
Palash Debnath 9e768c9552 Merge remote-tracking branch 'origin/main' into fix/1974-dev-port-ownership
# Conflicts:
#	CHANGELOG.md
2026-09-09 21:45:27 -07:00
Palash Debnath 18d2d1e076 Merge remote-tracking branch 'origin/main' into fix/1858-shortcut-registration
# Conflicts:
#	CHANGELOG.md
2026-09-09 21:45:22 -07:00
Palash Debnath 8d0b826c92 Merge remote-tracking branch 'origin/main' into fix/1898-windows-quit
# Conflicts:
#	CHANGELOG.md
2026-09-09 21:45:17 -07:00
Palash Debnath 939bdc7248 Merge pull request #1978 from debpalash/fix/1857-reduced-motion
feat(a11y): add an in-app Reduce motion switch
2026-09-09 21:44:50 -07:00
DawcraftandClaude Opus 5 71d4114354 docs: correct the CI and mirror-path claims in STRUCTURE.md
Both points from the review are right:

- The three test homes do not each get their own CI job. `ci.yml` runs all
  three as steps of the single `test` job (`Run pytest`, `Run pytest
  (backend/tests, isolated)`, `Run Vitest`); what makes `backend/tests/`
  separate is the pytest session, not the job.
- `tests/backend/services/test_dub_pipeline*.py` does not exist — that
  regression test is flat, at `tests/backend/test_dub_pipeline_wav.py`.
  The mirroring example now uses a path that exists
  (`backend/services/ffmpeg_utils.py` ->
  `tests/backend/services/test_ffmpeg_utils.py`) and says that
  backend-wide and cross-cutting suites stay flat.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LDyC6prbjFydox9XQGhyny
2026-09-10 13:32:07 +09:00
Palash DebnathandClaude Opus 5 5f7df6ac52 fix(dev): reclaim the port from a backend the app itself left running
Closes #1974.

The dev launcher only treated a port holder as ours when it ran out of the git
checkout. A backend the Tauri shell spawned lives under a per-app directory
named after the bundle id instead, so the launcher saw its OWN orphaned backend
as a stranger, refused to free port 3900, and aborted the run with "Refusing to
stop unrelated process" and no way forward but Task Manager.

Ownership now also accepts the app's reverse-DNS identifier in the executable
path or the command line. A bundle id is specific enough to be safe: nothing
else on the machine carries it, which is the point of the namespace.

The guard itself is unchanged in spirit — a foreign listener on the port is
still refused, and a test pins that widening ownership did not widen it to
everything, including a process from some other vendor's bundle.

Known limit, since I hit it in this repo: on Windows the check is given the
command line and executable path but not the working directory, so a backend
started by hand from an arbitrary interpreter — a bare `uvicorn` whose only
link to the checkout is a relative --app-dir — is still not recognised. That is
a different shape from the reported one and needs the cwd, which this code path
does not currently have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 21:31:19 -07:00
Palash Debnath 12216ecc9a Merge remote-tracking branch 'origin/main' into fix/1858-shortcut-registration
# Conflicts:
#	CHANGELOG.md
2026-09-09 21:27:39 -07:00
Palash Debnath 1339fa6814 Merge remote-tracking branch 'origin/main' into fix/1898-windows-quit
# Conflicts:
#	CHANGELOG.md
2026-09-09 21:27:38 -07:00
Palash Debnath 4b5b19ba47 Merge remote-tracking branch 'origin/main' into fix/1857-reduced-motion
# Conflicts:
#	CHANGELOG.md
#	frontend/src/i18n/locales/ar.json
#	frontend/src/i18n/locales/de.json
#	frontend/src/i18n/locales/es.json
#	frontend/src/i18n/locales/fr.json
#	frontend/src/i18n/locales/hi.json
#	frontend/src/i18n/locales/id.json
#	frontend/src/i18n/locales/it.json
#	frontend/src/i18n/locales/ja.json
#	frontend/src/i18n/locales/ko.json
#	frontend/src/i18n/locales/nl.json
#	frontend/src/i18n/locales/pl.json
#	frontend/src/i18n/locales/pt.json
#	frontend/src/i18n/locales/ru.json
#	frontend/src/i18n/locales/sv.json
#	frontend/src/i18n/locales/th.json
#	frontend/src/i18n/locales/tr.json
#	frontend/src/i18n/locales/uk.json
#	frontend/src/i18n/locales/vi.json
#	frontend/src/i18n/locales/zh-CN.json
#	frontend/src/i18n/locales/zh-TW.json
2026-09-09 21:27:37 -07:00
Palash Debnath 7d1f44f8bd Merge pull request #1977 from debpalash/land/1975-light-theme
fix(a11y): land the light theme, with one token raised to clear WCAG AA
2026-09-09 21:26:39 -07:00
Palash DebnathandClaude Opus 5 647ddc842b fix(dictation): say when another app already owns the shortcut
Closes #1858.

Whichever app registers a global shortcut first wins, and the default collides
with 1Password Quick Access on macOS — so for a large share of installs the
hotkey the onboarding screen advertises silently does nothing.

Registration failure was a Rust-side log line and nothing else. There was no
publish on the error path, so the frontend kept reporting whatever accelerator
had been REQUESTED, with no way for any screen to know the OS had refused it.
The failure is published now, carrying the outcome in `backend` and still
naming the accelerator so the UI can say WHICH combination is taken.

Surfaced as its own state rather than folding into the existing "no hotkey
registered" badge. That one means "not checked yet"; this means "this exact
combination belongs to another app, pick a different one" — different
situations needing different actions.

Detection rather than a new default, deliberately. Any default can collide with
something, so changing the value would move the problem rather than remove it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 21:20:32 -07:00
DawcraftandClaude Opus 5 638392994a docs: refresh STRUCTURE.md to match the tree as it is today
STRUCTURE.md still described the April layout: it was missing
backend/engines, worker, mcp_shim, speech_client, migrations, plugins,
hooks and config; the frontend e2e suites, i18n and src-tauri packaging
inputs; and the bin, skills, .agents/skills, notebooks, omnivoice-gallery
and .github/workflows top-level entries. Stale docs are bugs.

Three corrections beyond the missing entries:

- "all tests live here, no exceptions" was wrong. There are three homes
  (tests/, backend/tests/, co-located vitest) and the split is deliberate:
  pyproject testpaths, a separate ci.yml job, and the sys.modules-stub
  hazard documented in backend/tests/conftest.py. Replaced the claim with
  a table that records why each home exists.
- .env.example does not exist and the app never reads a repo-local .env;
  the durable user env file is ~/.config/omnivoice/env
  (backend/core/user_env.py), written by the Settings panel.
- .agents/ was listed as deleted, but it is back with a different job:
  the canonical skill copies pinned by skills-lock.json.

Also fixes the dead blob/main/STRUCTURE.md URL in the backlink script --
the file has lived in docs/ since the cleanup pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LDyC6prbjFydox9XQGhyny
2026-09-10 13:19:44 +09:00
Palash DebnathandClaude Opus 5 dba3074376 fix(windows): stop reporting every deliberate quit as a crash
Closes #1898.

On Windows the shell terminates the backend's job object with no graceful
phase — a console-less GUI child has no reliable control event — so the
backend never runs its lifespan shutdown and never clears its own run
sentinel. Every deliberate quit came back on the next launch as "The backend
did not shut down cleanly last run — it likely crashed or was killed". The
backend-side fix in #1895 only helps platforms where teardown actually begins.

A process about to be killed cannot record its own intent, so the shell records
it: the sentinel is retired immediately before the tree is terminated, on the
one path that knows the stop is deliberate. Anything that dies WITHOUT passing
through that path still leaves its sentinel behind and is still reported as a
crash, which is the property worth keeping.

Best effort by design. The data directory comes from the running backend, so if
it cannot be reached the file stays and the next launch reports a crash — the
same behaviour as before, never worse. A test pins that specifically: silently
erasing evidence of a real crash would be worse than a false positive.

Split into a pure file-level half and the port lookup so the behaviour is
testable without standing up a stub server. Verified on Windows with a real
toolchain: 233 Rust tests pass, including the three new ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 21:16:13 -07:00
Palash DebnathandClaude Opus 5 d0422b8913 feat(a11y): add an in-app Reduce motion switch
Closes #1857.

The CSS honours prefers-reduced-motion in about a dozen separate blocks, but
that is the OS switch and nothing else. Someone who wants a calm app without
turning motion off system-wide had no way to ask for it, and someone whose OS
setting is not respected by their environment had no recourse at all.

Settings → Appearance → Reduce motion sets data-motion="reduce" on the root,
and one blanket rule covers the whole tree including pseudo-elements. That
shape is deliberate: a per-component list is what let the header status dot
keep pulsing under Reduce Motion (Part B of the same issue, fixed separately),
and a single rule cannot have that gap.

Additive by design. The media query is left untouched and keeps working on its
own, so turning this off never re-enables motion for someone whose system asked
for less. A test pins that the two stay independent.

Durations go to 0.01ms rather than none: a zero duration skips animationend /
transitionend, which strands anything waiting on them. Imperceptible, still
fires. Also pinned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 21:11:13 -07:00
Palash DebnathandClaude Opus 5 a31b50c5cb fix(a11y): land the light theme, with one token raised to clear WCAG AA
Closes #1973. Lands #1975 by @CoDe-ReDz.

The app shipped six themes, all dark, and "auto" stayed dark on a light-mode
OS — so a user looking for a light mode found nothing. Light text on a dark
background causes halation for people with astigmatism, which makes this an
accessibility gap rather than a preference.

One correction to the contributor's palette: --chrome-fg-muted at #586e75 gives
4.39:1 against --chrome-bg #eee8d5, just under the 4.5:1 AA threshold for
normal text. Raised to #4d5f66 (5.45:1) in both the explicit light block and
the prefers-color-scheme mirror, keeping it in the Solarized family.

For the record, the two contrast failures the review bot flagged as P1 are not
real: --color-fg-subtle measures 6.66:1 and --chrome-fg-dim 5.86:1, both
comfortably AA. The token that actually failed was one it did not mention.

Everything else the bots raised was already handled on the branch: both theme
labels go through t() with real translations in all 21 locales, and the
header's white-to-grey gradient is overridden for the explicit light theme and
the auto mirror alike.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 21:06:19 -07:00
Palash Debnath cb4aaa1580 Merge remote-tracking branch 'origin/pr/1975' into land/1975-light-theme 2026-09-09 21:04:07 -07:00
Palash Debnath d0ab15c376 Merge pull request #1972 from debpalash/fix/1826-input-too-short
fix(errors): explain a too-short input instead of quoting torch's conv error
2026-09-09 20:59:29 -07:00
CoDe-ReDz 212a1afb69 fix(i18n): replace english placeholder labels with localized translations 2026-09-09 19:55:46 -07:00
CoDe-ReDz 58a4674316 fix(a11y): address bot review feedback (contrast, header, i18n, tests) 2026-09-09 19:43:30 -07:00
CoDe-ReDz bcc47882e9 a11y: Add Solarized Light theme and wire OS auto-sync
Resolves #1973. Addresses eye strain halation by adding a WCAG AA compliant light theme and exposing the Auto preference.
2026-09-09 19:08:39 -07:00
Palash Debnath 355f7465c8 Merge remote-tracking branch 'origin/main' into fix/1826-input-too-short
# Conflicts:
#	CHANGELOG.md
2026-09-09 16:29:07 -07:00
Palash Debnath 9426006d3c Merge pull request #1971 from debpalash/fix/1849-uiscale-order
fix(setup): offer the text-size control before first run, not after it
2026-09-09 16:28:44 -07:00
Palash Debnath f82ec3f579 Merge remote-tracking branch 'origin/main' into fix/1826-input-too-short
# Conflicts:
#	CHANGELOG.md
#	backend/core/failure.py
2026-09-09 16:12:00 -07:00
Palash Debnath b4fc4cb53f Merge remote-tracking branch 'origin/main' into fix/1849-uiscale-order
# Conflicts:
#	CHANGELOG.md
2026-09-09 16:11:23 -07:00
Palash Debnath c6d7a19f42 Merge pull request #1970 from debpalash/fix/1879-clone-reference
fix(errors): say 'no reference clip' instead of naming library parameters
2026-09-09 16:10:58 -07:00
Palash DebnathandClaude Opus 5 c902a42f6b fix(errors): explain a too-short input instead of quoting torch's conv error
Closes #1826.

A degenerately short generation reaches a convolution whose kernel is wider
than the tensor it was handed, and torch reports that in its own terms —
"Calculated padded input size per channel: (1). Kernel size: (2). Kernel size
can't be greater than actual input size". It arrived doubly wrapped in
"Underlying error:" and named nothing the user could change, when the fix on
their side is simply to type more than one character.

It is worth classifying for a second reason: this is not transient. The generic
wrapper told the user to "retry once", and this class fails identically on
every retry, so the advice actively wasted their time. The new remedy says so.

Matched on torch's own wording, which nothing else produces, so it is safe on
the context-free surfaces — and it needs to be, because that is exactly how it
reaches the user, through the generic 500 and the streaming error frame.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 15:55:38 -07:00
Palash Debnath 7a24b459bc Merge remote-tracking branch 'origin/main' into fix/1849-uiscale-order
# Conflicts:
#	CHANGELOG.md
2026-09-09 15:52:55 -07:00
Palash Debnath 0862b3efed Merge remote-tracking branch 'origin/main' into fix/1879-clone-reference
# Conflicts:
#	CHANGELOG.md
2026-09-09 15:52:51 -07:00
Palash Debnath ac63eff33d Merge pull request #1969 from debpalash/fix/1931-torchaudio-backend
fix(startup): guard the torchaudio API removed in 2.9, and document the Blackwell path
2026-09-09 15:52:29 -07:00
Palash DebnathandClaude Opus 5 f92e33f8e3 fix(setup): offer the text-size control before first run, not after it
Closes #1849.

UiScaleSetup is a client-side zoom — it makes no backend calls at all — but it
was gated on backendReady. So on a clean install the user watched the entire
bootstrap, and answered the macOS Accessibility prompt, at whatever size the
app had guessed, and was offered the size control only once all of that had
finished. Someone who cannot comfortably read the UI had to get through the
least readable part of the product first.

The gate now runs as soon as the store has hydrated, which is its only real
prerequisite: uiScaleConfigured lives in the store, and reading it earlier
would flash the screen at someone who had already chosen a scale.

Pinned at the source level. Rendering App in jsdom to observe the ordering
would need the whole backend, store and Tauri surface mocked — a far more
fragile test than the two facts it pins.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 15:41:03 -07:00
Palash Debnath 70aecf7ad0 Merge remote-tracking branch 'origin/main' into fix/1879-clone-reference
# Conflicts:
#	CHANGELOG.md
2026-09-09 15:36:52 -07:00
Palash Debnath 164534016a Merge remote-tracking branch 'origin/main' into fix/1931-torchaudio-backend
# Conflicts:
#	CHANGELOG.md
2026-09-09 15:36:47 -07:00
Palash Debnath 7b7b491ba0 Merge pull request #1968 from debpalash/fix/small-batch-1
fix(errors): point a generation timeout at Settings, not an environment variable
2026-09-09 15:36:25 -07:00
Palash DebnathandClaude Opus 5 0a019b5d7c fix(errors): say 'no reference clip' instead of naming library parameters
Closes #1879.

mlx-audio raises a bare ValueError in its own vocabulary — "No conditionals
available. Either provide audio_prompt/audio_prompt_sr for voice cloning, or
ensure conds.safetensors is in the model directory." — and the generate route
passed it straight through as the 400 detail. The user was told to supply an
argument they have no way to name and to check for a file they have never heard
of, when what happened is simply that they asked to clone with nothing to clone
from.

Classified now, with a remedy in the user's terms: pick a profile that has a
saved reference clip, or record one. It also notes that a designed voice with
no saved reference cannot be cloned from, which is the case that produces this.

The route still passes through every ValueError it cannot classify. Most are
VoiceStudio's own validation messages and are exactly what the user should
read, so replacing them wholesale would have been a regression — tests pin four
of them as untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 15:24:12 -07:00
Palash Debnath 4da62e664d Merge remote-tracking branch 'origin/main' into fix/1931-torchaudio-backend
# Conflicts:
#	CHANGELOG.md
2026-09-09 15:20:05 -07:00
Palash Debnath 508943ae71 Merge remote-tracking branch 'origin/main' into fix/small-batch-1
# Conflicts:
#	CHANGELOG.md
2026-09-09 15:20:00 -07:00
Palash Debnath 42b4bf856f Merge pull request #1967 from debpalash/fix/1866-engine-reason
fix(engines): say why an engine is unavailable instead of reporting a failed check
2026-09-09 15:19:40 -07:00
Palash DebnathandClaude Opus 5 c119afb023 fix(startup): guard the torchaudio API removed in 2.9, and document the Blackwell path
Refs #1931.

torchaudio 2.9 removed set_audio_backend(). soundfile has been the only backend
since 2.0, so the call was already a no-op — but unguarded it raises
AttributeError inside the ml_imports startup phase, and a failure there takes
the whole backend down: the desktop app sits on "starting backend" forever and
/health stays 503.

The group hitting it is not hypothetical. RTX 50-series (Blackwell, sm_120)
cards have no kernels in the pinned torch 2.8.0, so those users MUST move to
torch 2.9.x, which brings torchaudio 2.9 with it. Being forced to upgrade and
then crashing on a line that does nothing is the whole defect.

This does NOT raise the torch pin. Doing that changes the CUDA build on every
platform, in Docker and in CI, so it is the owner's call rather than something
to slip into a bug fix — the issue stays open for it. What lands here is the
half that is safe: the guard, plus a troubleshooting section with the exact
upgrade recipe and the command to confirm the card is visible, so an affected
user has a supported path today.

The guard is tested at the source level: reproducing it needs a real torchaudio
2.9 in the environment, which the pinned test env does not have. One test also
pins that the guard actually WRAPS the call, since a hasattr elsewhere in the
file would satisfy a naive substring check while the real call stayed bare.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 15:06:18 -07:00
Palash Debnath 0ba91e6d82 Merge remote-tracking branch 'origin/main' into fix/small-batch-1
# Conflicts:
#	CHANGELOG.md
2026-09-09 15:03:10 -07:00
Palash Debnath 28d9d2fbe8 Merge remote-tracking branch 'origin/main' into fix/1866-engine-reason
# Conflicts:
#	CHANGELOG.md
2026-09-09 15:03:05 -07:00
Palash Debnath 5069872752 Merge pull request #1966 from debpalash/fix/1845-setup-pill
fix(dictation): stop the Accessibility prompt owning the screen indefinitely
2026-09-09 15:02:45 -07:00
Palash DebnathandClaude Opus 5 030436a549 fix(errors): point a generation timeout at Settings, not an environment variable
Closes #1808.

#1797 moved the compute-time budget into Settings → Performance & Device, but
three branches of _timeout_guidance still told the user to raise
OMNIVOICE_GENERATE_TIMEOUT_S. That sends someone to set an environment variable
for a value the app now exposes as a control — and on Windows, setting one
durably is the trap this project's own docs warn against.

Nothing about the mechanism changed: the variable still works and still takes
precedence over the setting. Only which of the two the message names.

Two existing tests asserted the env var appears in that text. They predate
#1797 and were pinning the behaviour this issue reports as wrong, so they now
assert the control instead. A third test guards the whole class rather than the
three instances, so a branch added later cannot quietly reintroduce it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 14:50:58 -07:00
Palash Debnath b9324b84b3 Merge remote-tracking branch 'origin/main' into fix/1866-engine-reason
# Conflicts:
#	CHANGELOG.md
2026-09-09 14:45:56 -07:00
Palash Debnath 353cc40fdc Merge remote-tracking branch 'origin/main' into fix/1845-setup-pill
# Conflicts:
#	CHANGELOG.md
2026-09-09 14:45:52 -07:00
Palash Debnath a4cb4afe37 Merge pull request #1965 from debpalash/fix/1856-dictation-step
fix(setup): stop the last onboarding step failing three times with no model
2026-09-09 14:45:32 -07:00
Palash DebnathandClaude Opus 5 075ab7b68b fix(engines): say why an engine is unavailable instead of reporting a failed check
Closes #1866.

Model Catalogue → Engines showed "Engine unavailable. Check installation and
configuration." and "Last error: A previous engine check failed." for engines
the user had simply never installed. Neither names a missing package, a missing
step, or a next action, and the second reads like a crash or a poisoned cache
rather than "you have not installed this yet" — so a normal, expected state
looked like a fault.

The probe's own sentence still cannot cross the boundary: it carries exception
text, local paths and sometimes credentials, which is why it was replaced in
the first place. What changed is that the private diagnostic is now CLASSIFIED
into a VoiceStudio-owned category — package not installed, needs configuring,
file missing or unreadable — exactly the shape _public_routing_reason already
uses for routing. Anything unrecognised keeps the old generic sentence rather
than asserting a cause the probe never gave.

test_docs_url_survives_the_public_metadata_scrub pinned the generic wording
while testing something else; it now asserts what it is actually about, that no
private text survives.

Also skips the exec-bit placeholder test on Windows, where os.access(X_OK) is
true for any existing file so the assertion cannot fail — it errored the whole
module on a Windows checkout. Pre-existing, unrelated to this change, and in
the way of running these tests at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 14:37:27 -07:00
Palash Debnath e5a0a553e9 Merge remote-tracking branch 'origin/main' into fix/1845-setup-pill
# Conflicts:
#	CHANGELOG.md
2026-09-09 14:27:16 -07:00
Palash Debnath b7eb5428c7 Merge remote-tracking branch 'origin/main' into fix/1856-dictation-step
# Conflicts:
#	CHANGELOG.md
2026-09-09 14:27:12 -07:00
Palash Debnath bb4e99a336 Merge pull request #1964 from debpalash/fix/1957-untrusted-mount
fix(errors): explain a Windows untrusted-mount failure instead of echoing WinError 448
2026-09-09 14:26:49 -07:00
Palash DebnathandClaude Opus 5 accef57865 fix(dictation): stop the Accessibility prompt owning the screen indefinitely
Closes #1845. Closes #1886.

The widget window is created always-on-top, and the setup state had no time
limit at all. On a clean macOS install the pill sat over the first-run setup
window — covering the disk-space line and the Start installation button — and
over every other application, until Accessibility was granted or the user
dismissed it by hand. There was no cap and no safety net: the stranded-pill
reconcile only runs while idle, and this state is not idle.

A permission the user has not granted yet does not outrank what they are
actually doing, and mid-setup they usually cannot grant it yet anyway. The
prompt now gets a bounded claim on the screen and then steps aside.

Polling deliberately continues after the window hides, so granting
Accessibility later still returns the widget to idle on its own — what expires
is the pill's claim on the screen, not the reconciliation. The hide is latched
so it fires once rather than fighting anything that legitimately shows the
window again; both properties have a test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 14:07:53 -07:00
Palash Debnath d07fa1ffff Merge remote-tracking branch 'origin/main' into fix/1957-untrusted-mount
# Conflicts:
#	CHANGELOG.md
2026-09-09 14:04:34 -07:00
Palash DebnathandClaude Opus 5 d313045b13 fix(setup): stop the last onboarding step failing three times with no model
Closes #1856.

The mandatory-only install path ships no speech-to-text model, and the
dictation step rendered its three script cards regardless. Every card came up
red with "No speech-to-text model is installed", and the step's own copy
invited the user to press the hotkey or hit Replay, neither of which can
transcribe anything. That is the final screen of first-run setup, so the last
thing a new user saw was three failures they were told to cause.

The step now checks readiness the same way the component already checks for
its bundled sample WAVs, and when no model is installed it offers the model
chooser in place of the cards — the same picker the Transcriptions page uses,
so the user installs one and continues rather than reading an error three
times. A model already on disk can be selected without a download.

`checking` deliberately keeps the cards: the probe resolves in well under a
second, and flashing the install panel first would be worse than the wait.
A test pins that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 14:03:03 -07:00
Palash Debnath 6a44733297 Merge pull request #1962 from debpalash/land/queue2
Land the reviewed PR queue, part two: log panel, bootstrap mirror, setup diagnostics
2026-09-09 13:58:18 -07:00
Palash DebnathandClaude Opus 5 fd7d2d94dd fix(errors): explain a Windows untrusted-mount failure instead of echoing WinError 448
Closes #1957.

A download failed with nothing but the OS sentence: "[WinError 448] The path
cannot be traversed because it contains an untrusted mount point". That is a
Windows rule about the VOLUME — Dev Drives, mounted VHD/ReFS volumes and
junctions into another user profile all trigger it — so retrying the same link
can never work, and the message names nothing the user can change.

Classified now, with a remedy that points at Settings → Storage and gives the
fsutil escape hatch for a folder that has to stay put. Matched on the numeric
code first, since Windows translates the sentence, with the English phrase as a
fallback. Allowlisted for context-free surfaces because it arrives through the
global 500 handler, which otherwise attaches no hint at all — and its trigger
is unmistakable, so it cannot land on an unrelated failure.

A test pins that the offending path never comes back in the payload.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 13:47:51 -07:00
Palash Debnath f82f35ee97 Merge pull request #1963 from debpalash/land/queue3
Land the reviewed PR queue, part three: zh-CN locale and null capture timings
2026-09-09 13:43:45 -07:00
Palash Debnath 1c37c479be Merge remote-tracking branch 'origin/main' into land/queue2
# Conflicts:
#	CHANGELOG.md
2026-09-09 13:43:21 -07:00
Palash DebnathandClaude Opus 5 b55098f044 fix: repair the Colab cell edit and point the clear test at the split resolver
Two CI failures, both mine to fix.

The warning I added to the Colab ASR cell used \n escapes inside the notebook
JSON, and they landed as real newlines, so the cell's Python had an
unterminated string and tests/test_colab_asr_setup.py could not exec it. The
block prints line by line now, with no escapes to get wrong.

test_tauri_log_clear_reports_truncate_failure patched _tauri_log_candidates,
but #1925 moved Clear onto _tauri_plugin_log_candidates, so the patch no longer
reached the code under test and the real resolver was consulted instead. It
passed on a machine with a shell log on disk and failed on a clean runner.
Patches both halves, matching the fixture in test_tauri_log_clear.py.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 13:43:17 -07:00
Palash Debnath 854609e459 Merge pull request #1952 from debpalash/feat/local-workspace-dictation-polish
Improve dictation controls and creative workspace layouts
2026-09-09 13:41:19 -07:00
Palash DebnathandClaude Opus 5 7d46563ea8 fix(i18n): finish the zh-CN locale and make its own parity suite pass
#1877 completes the zh-CN translation and drops its ratchet to zero, but the
PR never ran today's gates — it has been conflicting, so CI reported nothing —
and the file it lands does not pass tests/test_locale_parity.py.

Three things fixed here:

- Twelve keys were declared twice inside the same object (timing_concise,
  autofit_quality, the plan_* set, the role_* set). Python's parser rejects a
  duplicate key outright, so the whole suite errored rather than failing one
  assertion. Deduped keeping the first occurrence, which is the block #1877
  actually translated.
- The `player` section appeared twice: the complete new one and an older
  two-key stub. JSON keeps the LAST, so the stub silently won and six keys
  vanished at runtime. The stub is gone.
- `settings.hf_source_*_label` appeared twice with slightly different wording.

The file is rewritten as canonical JSON (indent 2, non-ASCII preserved), which
is byte-identical to how en.json already serialises, so the format matches the
other locales exactly. zh-CN now has zero keys missing and zero beyond en.

Also fixes the review finding on #1959: the capture route picks its engine from
a `mode` form field, not an `accurate` flag, so parametrising on `accurate`
sent a field the route ignores and ran the default fast path twice. Both
engines are exercised now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 13:24:06 -07:00
Palash Debnath b6143cce08 Merge remote-tracking branch 'origin/pr/1959' into land/queue3 2026-09-09 13:18:49 -07:00
Palash Debnath 512aad5c48 Merge remote-tracking branch 'origin/main' into land/1952
# Conflicts:
#	frontend/src/pages/AudiobookTab.jsx
2026-09-09 13:17:13 -07:00
Palash Debnath 8526ff56f7 Merge remote-tracking branch 'origin/main' into land/queue2
# Conflicts:
#	CHANGELOG.md
2026-09-09 13:15:23 -07:00
Palash Debnath a68f85c51b Merge pull request #1958 from debpalash/land/queue1
Land the reviewed PR queue: audiobook, engines, bootstrap, setup and download fixes
2026-09-09 13:13:25 -07:00
Palash DebnathandClaude Opus 5 df7dab43b6 fix(i18n): keep the China-mirror comments out of the CJK guard
tests/test_no_hardcoded_cjk.py fails on any non-English text outside the
translation layer, allowlist aside, and #1892 put the mirror region's Chinese
label into two Rust doc comments. The comments only quote what the UI shows, so
naming the region in English says the same thing and keeps the guard green
without widening the allowlist for a comment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 13:12:22 -07:00
Palash DebnathandClaude Opus 5 f48a33edcf fix: address the review findings on the PRs landed here
Bot review raised one blocking and several real findings across these PRs.
Each is fixed here rather than merged and followed up.

#1892 — apply_pypi_index_env() now runs for EVERY uv invocation, and with the
default region "auto" it called the UNCACHED auto_detect_region(), racing two
live network probes with a 4s timeout per uv call. On a blocked or offline
network that is a repeated multi-second stall, and it multiplies the outbound
calls a local-first app makes unasked. The probe is memoised for the life of
the process. It also now clears UV_INDEX_URL before setting it, so a stale
ambient value cannot outrank the region the user picked.

#1925 — backend.rs trimmed OMNIVOICE_LOG_DIR for its emptiness guard but built
the path from the RAW value, while the Python reader strips it. A padded value
therefore had the writer and the reader looking at different directories, which
is the divergence the PR exists to close.

#1920 — the rotation walk caught bare OSError, so a PermissionError or a real
I/O failure was swallowed and the panel silently rendered less. Only the race
the guard exists for (a file that rolled away, and on Windows the handler's own
sharing violation) is skipped now; anything else surfaces.

#1951 — the fix was right but shipped no tests and no changelog entry. Both
added, including a case pinning that the wizard preflight and the diagnostic
route the same host the same way, since they carry separate copies of the
branch.

#1923 — the cell hardcodes the CTranslate2 model, but if the cuDNN 8 step
failed the backend falls back to PyTorch Whisper and downloads a second
multi-gigabyte model. The cell now says so while the download it just spent is
still on screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 13:09:12 -07:00
Palash DebnathandClaude Opus 5 791a46d9dd fix(errors): match WebKit's space-labelled frames so the filter works on Safari
The frame matcher required non-whitespace from line start up to the `@`,
which is right for rejecting a V8 header posing as a frame but wrong for
JSC: it labels top-level frames `global code@url`, `eval code@url` and
`module code@url`. Those are exactly the frames an injected extension
script throws from, so on WKWebView — the macOS desktop shell — and Safari
no frame matched, the origin came back unknown, and the extension's error
still offered "Report this bug". #1901 was fixed on Chromium only.

The three labels are enumerated rather than allowing spaces generally, so
the header false positive the anchoring exists for stays closed; a test
pins that. Each WebKit case uses a distinct message because shouldShow()
throttles by message text and a shared one would pass on the throttle
instead of the frame match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 13:02:24 -07:00
Palash Debnath 61eb27dbba Merge remote-tracking branch 'origin/pr/1892' into land/queue2 2026-09-09 13:02:15 -07:00
Palash Debnath 19d371f9d6 Merge remote-tracking branch 'origin/pr/1951' into land/queue2 2026-09-09 13:02:15 -07:00
Palash Debnath c760347339 Merge remote-tracking branch 'origin/pr/1925' into land/queue2
# Conflicts:
#	CHANGELOG.md
#	backend/api/routers/system.py
2026-09-09 13:02:15 -07:00
Palash Debnath ee8a8fb928 Merge remote-tracking branch 'origin/pr/1924' into land/queue2
# Conflicts:
#	CHANGELOG.md
2026-09-09 13:01:46 -07:00
Palash Debnath d7177be66c Merge remote-tracking branch 'origin/pr/1923' into land/queue2
# Conflicts:
#	CHANGELOG.md
2026-09-09 13:01:46 -07:00
Palash Debnath 88e59326fd Merge remote-tracking branch 'origin/pr/1920' into land/queue2
# Conflicts:
#	CHANGELOG.md
2026-09-09 13:01:45 -07:00
Palash Debnath f02880cc5c Merge remote-tracking branch 'origin/pr/1897' into land/queue2
# Conflicts:
#	CHANGELOG.md
2026-09-09 13:01:45 -07:00
Palash Debnath f9ff185a22 Merge remote-tracking branch 'origin/main' into land/1952
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:58:49 -07:00
Palash Debnath cc96de27b8 Merge remote-tracking branch 'origin/main' into land/queue1 2026-09-09 12:58:32 -07:00
Palash DebnathandClaude Opus 5 e8af4fae12 fix(engines): give audiocpp its docs link and pin docs_url in the registry shape
#1917 adds a Learn more link to the unavailable-engine row, driven by
_ENGINE_DOCS, and a guard that every registered engine has a doc page. It was
written before audiocpp landed on main, so against today's main the guard
failed on audiocpp and the registry shape test failed on the new docs_url key
— PR-green under an older base, main-red on merge.

docs/engines/audio-cpp.md already existed; only the id-to-path mapping was
missing. The shape test now expects docs_url, with a note pointing at the
guard so the next engine added without a doc fails loudly rather than
quietly dropping its link.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 12:58:29 -07:00
Palash Debnath e0309be21f Merge pull request #1955 from debpalash/fix/windows-backend-watchdog-hang
fix(backend): stop Windows desktop launches freezing at "Loading ML runtime"
2026-09-09 12:56:20 -07:00
Chang-Jin-LeeandClaude Opus 5 2e4465a3f4 fix(capture): pass a null segment end through instead of rounding it
`round(s.get("end", 0), 2)` does not defend against a stored None: the key is
present, so `.get` returns the None rather than the default, and `round` raises

    TypeError: type NoneType doesn't define __round__ method

`max(s.get("end", 0) for s in segments)` on the line above raises first when any
other segment is timed:

    TypeError: '>' not supported between instances of 'NoneType' and 'float'

Two engines reach these builders with end=None. `_sherpa_result` sets
duration=None when it cannot derive one from the sample rate, and sherpa is the
first capture engine. `OpenAICompatASRBackend._adapt_response` emits end=None for
every plain-text response, which is what a server that rejects verbose_json
returns, and that backend is selectable as the active one used by accurate mode.

Measure the duration from the segments that carry a number, and pass the nulls
through. That is the shape the segment list already renders since #1904 — it
shows whichever half of the range is known — and it keeps the honest null the
producers deliberately write instead of inventing a zero.

capture_ws.py has the same two lines and gets the same treatment; it also emits
end=None itself in five of its own streaming payloads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 04:33:18 +09:00
Palash Debnath 9ffc9d52d7 Merge remote-tracking branch 'origin/main' into land/1952
# Conflicts:
#	CHANGELOG.md
#	bun.lock
#	frontend/package.json
#	frontend/src/components/CaptureWidget.jsx
#	frontend/src/pages/Transcriptions.jsx
2026-09-09 12:31:06 -07:00
Palash Debnath cea6678ea7 Merge pull request #1956 from debpalash/fix/1800-stream-error-class
fix(errors): make an unclassified failure report say something true and specific
2026-09-09 12:30:09 -07:00
Palash Debnath 40199e39cf Merge remote-tracking branch 'origin/main' into work/1955
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:29:36 -07:00
Palash DebnathandClaude Opus 5 c21696b9ae style: format the two files #1930 left unformatted
`bun run format:check` is a CI gate and SetupWizard.jsx plus its test came in
unformatted. No behaviour change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 12:28:05 -07:00
Palash Debnath 67817b89bf Merge remote-tracking branch 'origin/pr/1942' into land/queue1
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:25:35 -07:00
Palash Debnath 20bc24c312 Merge remote-tracking branch 'origin/pr/1930' into land/queue1 2026-09-09 12:25:35 -07:00
Palash Debnath 810abf9187 Merge remote-tracking branch 'origin/pr/1919' into land/queue1
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:25:35 -07:00
Palash Debnath cfb318c50b Merge remote-tracking branch 'origin/pr/1918' into land/queue1
# Conflicts:
#	CHANGELOG.md
#	frontend/src/components/BootstrapSplash.jsx
2026-09-09 12:25:34 -07:00
Palash Debnath 2e1e52f26e Merge remote-tracking branch 'origin/pr/1917' into land/queue1
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:24:12 -07:00
Palash Debnath 2f663011e5 Merge remote-tracking branch 'origin/pr/1916' into land/queue1
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:24:11 -07:00
Palash Debnath 773c4728c8 Merge remote-tracking branch 'origin/pr/1915' into land/queue1
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:24:11 -07:00
Palash Debnath 6b68c4d9fe Merge remote-tracking branch 'origin/pr/1914' into land/queue1
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:24:11 -07:00
Palash Debnath 5f5bb04f41 Merge remote-tracking branch 'origin/pr/1912' into land/queue1
# Conflicts:
#	CHANGELOG.md
2026-09-09 12:24:10 -07:00
Palash Debnath cea00a4889 Merge pull request #1908 from psiberfunk/codex/fix-synthesis-progress-indeterminate
fix(studio): show honest synthesis progress
2026-09-09 12:22:07 -07:00
Palash Debnath 6268eacf35 Merge pull request #1903 from psiberfunk/codex/fix-transcriptions-dictation
fix(dictation): make Transcriptions capture actionable
2026-09-09 12:22:01 -07:00
Palash Debnath a2653824ce Merge pull request #1896 from psiberfunk/fix/bootstrap-splash-observed-stages-1894
fix(bootstrap): stop showing first-run install steps on warm starts
2026-09-09 12:21:55 -07:00
Palash Debnath 71056bcdb4 Merge pull request #1890 from psiberfunk/fix/widget-pill-shadow-clip
fix(dictation): stop clipping the widget pill's drop shadow into a rectangle
2026-09-09 12:21:49 -07:00
Palash Debnath b5441bea36 Merge pull request #1888 from psiberfunk/fix/macos-dock-reopen
fix(desktop): handle Dock-icon reopen on macOS
2026-09-09 12:21:44 -07:00
Palash Debnath d7751ad7ce Merge pull request #1885 from psiberfunk/fix/capture-pill-setup-tooltip
fix(dictation): give the setup-state pill label a hover tooltip
2026-09-09 12:21:38 -07:00
Palash Debnath 2b027cb54a Merge branch 'main' into fix/1859-launchpad-scroll-shrink 2026-09-09 12:20:27 -07:00
Palash Debnath d2fb292217 Merge pull request #1938 from debpalash/dependabot/npm_and_yarn/frontend/vitest-4.1.11
chore(deps-dev): bump vitest from 4.1.9 to 4.1.11 in /frontend
2026-09-09 12:18:24 -07:00
Palash DebnathandClaude Opus 5 7afb82bb43 style: format the four files this branch left unformatted
`bun run format:check` is a CI gate and these four came in unformatted. No
behaviour change — running the formatter in write mode touches nothing else
in the tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 12:14:47 -07:00
Palash DebnathandClaude Opus 5 d630db29f5 fix(errors): stop a context-free failure borrowing another stage's remediation
Closes #1943.

A macOS mlx-audio text-to-speech failure returned a 500 advising the user
that "the connection to the video server dropped mid-download". No video was
involved. VIDEO_DOWNLOAD_NETWORK triggers on bare phrases — "timed out",
"connection reset", "broken pipe" — so any unrelated failure carrying one
is handed a confidently wrong next step, which is worse than no hint at all.

failure._CONTEXT_FREE_HINT_CLASSES already existed for exactly this, and its
own comment names VIDEO_DOWNLOAD_NETWORK as the class that must never appear
on a stageless surface. Only append_hint honoured it; public_exception_response
took over the 500 path without carrying the rule across, and the streaming
error frame then inherited the same gap through it.

The filter now lives in public_exception_response, so every context-free
caller gets it. MODEL_CACHE_CORRUPT joins the allowlist — its trigger is a
VoiceStudio-authored sentence, no library can produce it, and the 500 handler
is the surface a corrupt cache actually reaches.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 12:13:25 -07:00
Palash DebnathandClaude Fable 5.1 f49d31d5b3 test(backend): only accept positive startup evidence in the Windows watchdog test
The poll loop treated any status other than "starting" as success, so a
backend that stayed alive but reported a failed startup would pass the
very test meant to catch a broken start (CodeRabbit + Greptile on #1955).
Succeed only when the ML import step is done or status is ready; fail
loudly on any other terminal status or error. Also give the child an
empty HF_HUB_CACHE so it never reads the developer's populated cache.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCDZcBpP6QQa4dzUa8z6rh
2026-09-09 12:06:46 -07:00
Palash DebnathandClaude Opus 5 4f4d21ff5b fix(errors): name the backend error class on an unclassified streaming failure
Closes #1800.

Every engine failure the taxonomy cannot classify renders one floor message,
"Generation failed. Check the selected engine and try again." The auto bug
reporter puts that message and a stack of minified bundle frames into the
issue, so unrelated faults arrive as byte-identical reports — roughly a dozen
of the open issues are that same report filed again, and none of them can be
told apart, let alone triaged.

The streaming error frame now carries the exception's TYPE NAME, the frontend
keeps it on StreamingPreviewError, and the report prints it as "Backend error
class: …". A MemoryError and a FileNotFoundError stop being the same issue.

Only the class name — no substring of the exception message is copied, so the
response-safety contract still holds and a test pins that a path in the
exception never reaches the payload. This is the same datum the dub routes
already put on the wire as error_class and the analytics allowlist already
treats as content-free.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 12:06:40 -07:00
Palash DebnathandClaude Opus 5 b82205fbab fix(deps): regenerate the workspace lockfile for the vitest bump
frontend/ is a bun workspace, so its package.json is locked by the
repo-root bun.lock. Dependabot bumped only the manifest, so
`bun install --frozen-lockfile` — which CI and deploy/Dockerfile both
run — rejected the tree and the Tests job never got past install.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 12:01:45 -07:00
Palash DebnathandClaude Opus 5 7601f0f1a8 fix(dictation): drop the token border the picker's nesting rail reintroduced
tests/test_no_literal_borders.py guards the app-wide border removal: a
`border-[var(--chrome-border…)]` renders a stray hairline the moment that
token stops resolving transparent. The picker's indent rail used one, which
failed the guard. The indent and padding already carry the nesting, so the
rail keeps its width as border-transparent and shows nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 12:00:42 -07:00
Palash DebnathandClaude Fable 5.1 5c7f8a8e7a test(backend): Windows integration regression for the desktop stdin watchdog hang
Spawns the real backend the way the desktop shell does (containment marker
plus a piped stdin) and asserts startup gets past the ML import. On the
pre-fix watchdog it times out after 180 s; on the fix it passes in ~4 s.
Windows-only, since the deadlock is a Windows loader-lock interaction and
CI's backend job runs on Linux.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCDZcBpP6QQa4dzUa8z6rh
2026-09-09 11:51:17 -07:00
Palash DebnathandClaude Opus 5 7dcb88452c fix(deps): regenerate the root lockfile for this branch's refreshed pins
frontend/package.json moved but the workspace-root bun.lock did not, so
`bun install --frozen-lockfile` — what CI and deploy/Dockerfile both run —
rejected the tree. Plain `bun install` tolerates the drift, so a green local
run said nothing about it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 11:49:29 -07:00
Palash Debnath db50111c54 fix(sidebar): restore project rename, lost when the Dub landing became history-only
The Dub landing used to carry a WorkspaceProjects panel, and that panel was
the only caller of the project rename endpoint. Making the landing
history-only removed the panel and left rename unreachable from the entire
UI, while the rename route, the project list and the inline-rename CSS all
stayed. App.jsx's renameProject became an unused variable, which is what
failed CI lint — the lint error was the symptom, the lost capability was the
bug.

Projects now live only in the sidebar rail, so the affordance moves there:
inline rename on each project row, commit on Enter or Save, abandon on
Escape, empty and unchanged names ignored, and the button hidden when no
handler is wired. The orphaned WorkspaceProjects component is deleted.

Also fixes a Windows-only failure in initialLoadRetry.test.js: it took
.pathname off a file:// URL, which on Windows yields "/C:/..." and made
readFileSync resolve "C:\C:\...", so the file ENOENT'd on every Windows
checkout. Uses fileURLToPath instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
(cherry picked from commit 0a164add119335c3a150725cfe4c4ce8959c50a5)
2026-09-09 11:48:22 -07:00
Palash Debnath 6b7c2bac8f feat(transcriptions): let the user pick which dictation model to install
The missing-model empty state offered exactly one action: download the
recommended Whisper Tiny. The six other catalogue models — the more accurate
English Parakeet, the 25–44 MB streaming models that show text while you
speak, the bilingual zh/en ones — were only reachable through Settings, and
a user who already had one on disk was still told to download Whisper Tiny.

The page now lists the whole sherpa-onnx catalogue grouped by the trade-off
the user is actually choosing between (best accuracy vs lowest latency),
with languages and download size on every row. Any model can be installed
in one click, an installed one can be switched to without a download, and
the progress bar names the model that was picked. If the catalogue cannot
be read the single recommended-download button remains as the fallback.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
(cherry picked from commit 9c4ba2bab2a2253fbf4f82bf082db530267ddadc)
2026-09-09 11:47:38 -07:00
Palash DebnathandClaude Opus 5 48d1b22fd9 feat(dictation): model picker in the engine quick-switch, and a recoverable Windows dev stack
Adds the sherpa-onnx dictation model picker under the Transcription engine
row so the model the hotkey loads is switchable without opening Settings,
routes the Sherpa transcription path through that same preference, and makes
the Windows desktop dev stack recover instead of demanding Task Manager.
Refreshes the Tauri and npm dependency pins that went with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ypcgSsh5j2PEonSJiAU1S
2026-09-09 11:47:31 -07:00
Palash DebnathandClaude Fable 5.1 6a6dd83efa docs(changelog): note the Windows backend startup hang fix (#1955)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCDZcBpP6QQa4dzUa8z6rh
2026-09-09 11:40:13 -07:00
Palash DebnathandClaude Fable 5.1 09a260bb26 fix(backend): stop Windows desktop launches freezing at "Loading ML runtime"
Every backend spawned by the Windows desktop shell hung forever in the
startup worker's `import torch`, inside the loader for numpy's OpenBLAS
DLL. The desktop parent-liveness watchdog (0a20aeb0) parks a synchronous
read on the stdin pipe the shell hands the backend, and that pending read
deadlocks the DLL initializer. The identical command from a terminal, with
no stdin pipe and no watchdog, starts in seconds — which is why it only
reproduced under the app.

Bisected outside the app by spawning the backend with the shell's exact
env, pipes, creation flags and job object: a watchdog thread that merely
sleeps is harmless; a pending ReadFile, via the C runtime or straight to
the kernel, hangs it every time. Native stacks (py-spy --native) show the
watchdog in NtReadFile and the importer waiting on a critical section from
inside the OpenBLAS initializer.

Fix: on Windows the watchdog polls PeekNamedPipe and reads only bytes that
are already buffered, so no I/O is ever outstanding on the pipe. It still
exits the instant the desktop closes its end (ERROR_BROKEN_PIPE), and a
non-pipe stdin keeps the shared blocking reader. Verified: the app-style
spawn goes from an indefinite hang to ready in ~3 s, and the desktop-prod
build boots and loads the model.

Not in v0.5.1; the watchdog landed 2026-08-30 on main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCDZcBpP6QQa4dzUa8z6rh
2026-09-09 11:39:50 -07:00
Palash Debnath a754faec04 docs: link dictation changes to PR 1952 2026-09-09 18:58:27 +05:30
Palash Debnath 8fc003ad7c Merge remote-tracking branch 'origin/main' into feat/local-workspace-dictation-polish 2026-09-09 18:55:28 +05:30
Palash Debnath 99a534882c feat: improve dictation controls and creative workspaces 2026-09-09 18:55:19 +05:30
Chang-Jin-LeeandClaude Opus 5 81f889f3ff fix(errors): anchor the JSC frame pattern so a header cannot pose as one
Greptile on #1924: the second alternative in FRAME_LINE was unanchored, so a
V8 HEADER whose message happens to read `... user@chrome-extension://...`
matched as a JSC frame. FRAME_URL then took the message's URL as the throw
site and suppressed the report -- the same false positive the previous commit
fixed, one layer down.

The round-1 test missed it because its message carried a chrome-extension://
URL with no `@` before it, so the header never matched either alternative.

A JSC frame is `fn@url` and a function name has no spaces, so the `@` must be
reachable from the line start through non-whitespace only: `^\s*\S*@`. A V8
header is `Name: message`, so the space after the colon stops the match. Both
stack dialects still work, including an anonymous Firefox frame that begins
with the `@`.

Two tests: our error whose MESSAGE contains `user@chrome-extension://` with
our frames below is still reported (red before), and a bare
`Y@chrome-extension://...` with no header line at all is still recognised as a
frame, so the anchor does not cost the Safari/Firefox shape.

Full suite 328 files / 2740 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-09 22:13:00 +09:00
michaelhuamanfloresandClaude Sonnet 5 693a43b26c fix(setup): stop misreporting low-VRAM caveat as kernel-launch risk
The GPU routing "accelerated" caveat branch in preflight and diagnose
always showed the driver/arch "may fail at kernel launch" fix hint,
even when the actual reason was a low-VRAM advisory unrelated to
drivers or torch. Gate that message on KERNEL_RISK_MARKER and show an
accurate VRAM-appropriate hint otherwise.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 07:47:18 -05:00
Chang-Jin-Lee d848c2eb3b Merge remote-tracking branch 'upstream/main' into fix/1847-bootstrap-log-retention 2026-09-09 12:55:59 +09:00
Chang-Jin-Lee d5d28ddabf Merge remote-tracking branch 'upstream/main' into fix/1866-engine-unavailability-reason 2026-09-09 12:55:59 +09:00
Chang-Jin-Lee 3622102ace Merge remote-tracking branch 'upstream/main' into fix/1848-bidi-start-truncation 2026-09-09 12:55:59 +09:00
Chang-Jin-Lee 31d7b4d251 Merge remote-tracking branch 'upstream/main' into fix/1859-launchpad-scroll-shrink 2026-09-09 12:55:59 +09:00
Chang-Jin-Lee 5e3af17d34 Merge remote-tracking branch 'upstream/main' into fix/1901-extension-error-reports 2026-09-09 12:55:59 +09:00
Chang-Jin-Lee 67ef5e22eb Merge remote-tracking branch 'upstream/main' into fix/tauri-clear-preserves-backend-stderr 2026-09-09 12:55:06 +09:00
Chang-Jin-Lee a84dc9e8ef Merge remote-tracking branch 'upstream/main' into fix/system-logs-rotation 2026-09-09 12:54:56 +09:00
marreiradigital b15fa1209a fix(download): log nao promete um fallback que ainda nao aconteceu
Achado do CodeRabbit (fora do diff) no PR #1942. Depois de desacoplar os dois
sinais, a mensagem passou a ser escolhida so por `_segmented_off` — mas na
penultima tentativa o plano devolve (True, True): o acelerador esta esgotado E
a tentativa re-lanca, entao o `snapshot_download` so entra na PROXIMA. O log
dizia "falling back to snapshot_download" enquanto na verdade ia retentar.

Sao tres estados distintos e ler so um sinal funde dois deles. A frase virou o
helper puro `_segmented_retry_note(disable, reraise)`, testado nos tres ramos —
inclusive o (True, True), que e o que a revisao pediu para cobrir.
2026-09-08 23:47:42 -04:00
marreiradigital 8a2c9a4828 Merge branch 'main' into fix/windows-dev-stack-and-download-resume
Conflito unico em `install_model`: a `main` (#1926) acrescentou
`and not allow_patterns` a condicao do acelerador, e este branch trocou
`_attempt == 1` por `not _segmented_off`. As duas guardas valem e foram
mantidas juntas.

O comentario acima da condicao ainda dizia que qualquer falha cai no
snapshot_download; atualizado para a regra atual (falha nao-transitoria, ou a
ultima tentativa), com ponteiro para `_segmented_retry_plan`.
2026-09-08 23:28:28 -04:00
marreiradigital adebe98e0f test(download): resolve tambem o import pre-existente em runtime
O `_segmented_enabled` no topo do arquivo viola a mesma instrucao de caminho
que o CodeRabbit apontou nos testes novos. Deixar metade do arquivo fora da
regra so garante que a proxima revisao aponte de novo.
2026-09-08 23:25:47 -04:00
marreiradigital 97e0e70bf2 fix(download): entrega o caminho simples so na ultima tentativa
Achado P1 do Greptile no PR #1942: com `disable` e `reraise` amarrados um ao
outro, a tentativa 4 de 5 desligava o acelerador E ja caia no
`snapshot_download` na mesma iteracao. Resultado: o acelerador ficava com 3
tentativas em vez de 4, e uma nova queda abandonava o manifesto reaproveitavel
uma tentativa antes do necessario, recomecando por um arquivo separado — que e
exatamente o que este helper existe para evitar.

Os dois sinais agora sao independentes: a tentativa que esgota o acelerador
ainda re-lanca, entao o caminho simples comeca na ULTIMA tentativa. Acelerador
fica com 1-4, `snapshot_download` com a 5.

Tambem blindei o caso de o acelerador falhar ja na ultima tentativa: ali nao ha
para onde re-lancar, entao a decisao vira "caminho simples agora" em vez de
estourar o laco sem nunca ter tentado o fallback.

Os testes do helper passaram a resolver o modulo da app em tempo de execucao,
como pede a instrucao de caminho para tests/**/*.py (import de modulo da app no
topo fica velho se um teste anterior sujar o sys.modules) — apontado pelo
CodeRabbit no mesmo round.
2026-09-08 23:25:30 -04:00
marreiradigital 66f7ef8cfe fix(download): reentra no acelerador na proxima tentativa apos queda
Achados do CodeRabbit no PR #1942.

O mais grave: com o erro classificado como transitorio, o codigo mantinha o
acelerador ligado mas caia direto no `snapshot_download` na MESMA tentativa. Se
esse download desse certo, o laco terminava e o manifesto do `.part` nunca era
reusado — exatamente o recomeco-do-zero que a correcao existe para impedir.

Agora o erro transitorio e propagado para o retry externo, cuja proxima
tentativa reentra no `_segmented_snapshot` e retoma do manifesto. A decisao
virou o helper puro `_segmented_retry_plan`, testavel direto (o laco mora dentro
de `install_model`, uma rota de ~200 linhas). A ultima tentativa fica reservada
para o caminho simples, entao o acelerador continua sem poder ser o motivo de um
install falhar de vez.

Tambem deste round de revisao:

- `Invoke-CimMethod ... Terminate` tinha o retorno descartado com `$null =`. O
  Win32_Process.Terminate reporta falha pelo ReturnValue, nao lancando: um kill
  negado por permissao era reportado como sucesso e a porta seguia presa. Agora
  o ReturnValue e validado, com exit 4 proprio e a mensagem carregando o codigo.
- O teste de concorrencia era vazio: o handler sincrono do MockTransport retorna
  antes de qualquer outra task rodar, entao `peak` nunca passava de 1 e a
  asserção `peak <= 4` passava sem exercitar o semaforo. Passou a segurar as
  requisicoes abertas com um asyncio.Event e a exigir `peak == 4` (verificado:
  com o semaforo afrouxado para 1000, o teste acusa 31).
- A doc dizia que OMNIVOICE_DOWNLOAD_MAX_WORKERS limita as faixas e que origem
  sem Range cai no snapshot_download. Nenhum dos dois: `_segmented_snapshot` nao
  passa `num_connections` (usa as 8 padrao) e origem sem Range vira stream unico
  dentro do proprio acelerador.
- Entradas de Highlights do CHANGELOG sem o `(#NNNN)` exigido.
2026-09-08 23:12:14 -04:00
marreiradigital d01fb5e7cc docs(changelog): aponta as entradas para as issues corretas
As issues #1940 (downloader segmentado sem progresso em conexao instavel) e
#1941 (stack de dev irrecuperavel no Windows) foram abertas para estas
correcoes; substitui os refs emprestados de #1224 e #1690.
2026-09-08 22:49:58 -04:00
marreiradigital bc22276fec docs(changelog): registra as correcoes de download e de dev no Windows
Entradas referenciadas a #1224 (truncamento de corpo no download) e #1690
(supervisor do backend de dev), que sao as issues que estas correcoes
estendem. Nao ha issue propria aberta para elas ainda.
2026-09-08 22:46:56 -04:00
marreiradigital db9e9d7fa1 fix(dev): permite destravar porta de dev presa no Windows
`canStop: !windows` fazia o script recusar qualquer parada no Windows com
"stop it in Task Manager and retry". O motivo original é legítimo: `taskkill
/pid` mira um PID reutilizável, e um PID reciclado entre o inspect e o kill
derrubaria um processo alheio.

Só que isso deixava o `bun run dev` permanentemente travado sempre que um
backend ficasse órfão — exatamente o cenário do commit anterior sobre a árvore
de processos. O predev falhava e não havia caminho de recuperação automático.

A parada agora é presa à INSTÂNCIA do processo: um único PowerShell busca a
instância CIM, confere o CreationDate contra a identidade já inspecionada e só
então chama Terminate NAQUELA instância. O terminate age sobre o objeto que a
checagem validou, não sobre um PID buscado de novo depois — a corrida some.
PID reciclado devolve exit 3 e é deixado em paz, em vez de falhar a execução.
2026-09-08 22:46:48 -04:00
marreiradigital e098d1280c fix(dev): normaliza caminho POSIX sem vazar a semantica do host
`belongsToCheckout(..., windows = false)` respeitava a flag na hora de montar
a string, mas normalizava o caminho com `resolve()` do host. Rodando no
Windows, "/work/VoiceStudio" virava "C:\work\VoiceStudio" e não casava com
nada numa linha de comando POSIX — o mesmo valia para o separador `sep`.

Efeito prático: o teste "command ownership requires a checkout path boundary"
já falhava na `main` limpa em qualquer máquina Windows, passando só no CI
Linux. Passa a usar `path.posix` quando a flag diz POSIX.
2026-09-08 22:46:29 -04:00
marreiradigital 2127aa7716 fix(dev): mata a arvore de processos do backend no Windows
O supervisor faz `spawn("uv", ...)` e o uv sobe o uvicorn como filho dele.
Windows não tem sinais: `child.kill()` vira TerminateProcess só no filho
DIRETO, então matar o `uv` deixava o uvicorn neto vivo segurando a porta 3900.
O spawn seguinte falhava com `[Errno 10048]`, o supervisor contava como crash,
e três desses derrubavam a stack inteira de dev — inclusive o Vite, via
`--kill-others-on-fail`.

`killProcessTree` usa `taskkill /T` no win32 e mantém o envio de sinal no
POSIX. Como o kill forçado devolve exit não-zero e sinal nulo, o reload que nós
mesmos pedimos passaria por crash; isso é tratado olhando se o tree-kill de
fato aconteceu, e não a plataforma — um crash de verdade durante um reload
continua indo para a recuperação de crash (coberto por teste que já existia).
2026-09-08 22:45:44 -04:00
marreiradigital 0d3fb07c1f fix(download): segmenta em blocos limitados e retoma o acelerador
O downloader segmentado gravava progresso no manifesto apenas quando um
segmento INTEIRO terminava, e dimensionava os segmentos como
tamanho/num_connections. Num blob de 806 MB isso dava 8 segmentos de ~100 MB:
numa conexão que cai a cada ~50 MB nenhum segmento jamais completava, o
manifesto nunca era escrito e cada tentativa recomeçava do zero.

Pior, o acelerador só rodava na PRIMEIRA tentativa (`_attempt == 1`), então
depois da primeira queda todas as retentativas iam para o `snapshot_download`
e o `.part` acumulado ficava órfão para sempre.

Agora os segmentos são limitados a 16 MB e a concorrência passa a ser
controlada por semáforo (antes vinha da própria contagem de segmentos), e o
acelerador é preservado entre tentativas quando o erro é de rede — reusando
`_is_retryable_download_error`, que já é a fonte única dessa classificação.
Ele só é desligado de vez quando a falha NÃO é transitória, ou seja, quando o
acelerador de fato não serve naquele host.

Reproduzido em rede real: `peer closed connection without sending complete
message body (received 54260979, expected 100708200)`.
2026-09-08 22:45:31 -04:00
dependabot[bot] 3ad31feccf chore(deps-dev): bump vitest from 4.1.9 to 4.1.11 in /frontend
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.9 to 4.1.11.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md)
- [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.11/packages/vitest)

---
updated-dependencies:
- dependency-name: vitest
  dependency-version: 4.1.11
  dependency-type: direct:development
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-09 00:18:48 +00:00
Palash Debnath 18f56a7940 Merge pull request #1926 from debpalash/feat/audiocpp-gpu-routing
feat(audiocpp): route native GPU backends
2026-09-08 23:53:34 +05:30
Palash Debnath 870b6f93ba fix(audiocpp): harden GPU runtime routing 2026-09-08 23:33:40 +05:30
Som 4c0bf2e010 fix(i18n): correct grammatical agreement in de/es/fr/pt/ru/sv consent.desc
CodeRabbit review on this PR caught real agreement errors in the
translations added for consent.desc (and de's try_dictation_desc
register):

- de: try_dictation_desc used informal du/deine while every other
  setup.* string in this locale uses formal Sie/Ihre — switched to match.
- es/fr/pt: adjectives after the anonymous-stats noun (feminine plural
  in each language) didn't agree in gender/number.
- ru: predicate adjectives didn't agree with the feminine noun
  "статистика".
- sv: adjectives didn't agree with "statistik" (en-gender noun).
2026-09-08 15:38:20 +00:00
Som 7d6f5c5b0c fix(setup): give the dictation step its own distinct rail label
Removing the redundant SectionHead (previous commit) left one more
duplicate on the dictation onboarding step: the step rail and
DictationDemo's own heading both said "Try dictation"/"Try Dictation" —
flagged by Greptile review on this PR.

Give the rail a genuinely distinct short label
(setup.dictation_step_label, "Dictation"), mirroring the consent step's
already-correct pattern (rail label vs. card title are different
strings). Translated into all 21 locales to hold the locale-parity
ratchet. Added a new test that renders the real DictationDemo (not
mocked to null, unlike the existing consent test) to prove the rail
label and the card's own title are distinct and each render exactly
once — the gap that let this ship undetected.
2026-09-08 15:36:54 +00:00
Palash Debnath 4f4dc7e068 fix(audiocpp): budget unknown GPU memory safely 2026-09-08 21:02:44 +05:30
Som 56a420a898 fix(setup): stop repeating the consent and dictation step titles
STEP_SUBTITLES.consent and STEP_SUBTITLES.dictation reused the exact
same i18n key as the step's title/label, and each step's body then
rendered a SectionHead with that same key again, on top of the step's
own content component rendering its own heading — showing the same
phrase 2-3 times on screen.

Give both steps a genuine short description (consent.desc,
setup.try_dictation_desc) for the header subtitle, matching the
system/models steps' existing pattern, and drop the now-redundant
SectionHead in each step's body since the content component
(AnalyticsConsentCard, DictationDemo) already renders its own title.

The two new keys are translated into all 21 locales to hold the
locale-parity ratchet (tests/test_locale_parity.py) at its current
baseline.

Closes #1855
2026-09-08 15:26:11 +00:00
Palash Debnath d5c0fa6f75 fix(worker): derive every engine capacity 2026-09-08 20:52:24 +05:30
Palash Debnath fe53e02693 fix(audiocpp): harden async and worker routing 2026-09-08 19:45:48 +05:30
Palash Debnath 56f8b616cc test(audiocpp): make CPU thread assertion portable 2026-09-08 19:10:44 +05:30
Palash Debnath f1016cdedd fix(audiocpp): budget low-memory Vulkan GPUs 2026-09-08 18:42:32 +05:30
Palash Debnath 68194c004d docs(changelog): note audio.cpp GPU routing 2026-09-08 18:32:46 +05:30
Palash Debnath 23dd8728d1 fix(audiocpp): preserve GPU routing across workers 2026-09-08 18:30:12 +05:30
Chang-Jin-LeeandClaude Opus 5 22ab31e0c7 fix(system): correct the Linux log path in the docstring, and follow the writer's override
CodeRabbit: the docstring said $XDG_STATE_HOME/VoiceStudio where the code says
OmniVoice. Checked against the writer rather than guessing which side was
wrong -- backend.rs::backend_log_path() joins "OmniVoice" on Linux, so the
code was right and the docstring was a pre-existing error. It matters because
that docstring is what gets read when telling a Linux user where the file is.

Reading backend_log_path() to settle it turned up something worth fixing in
the resolver this PR introduced: Rust checks OMNIVOICE_LOG_DIR before any
per-OS default, and nothing on the Python side knew about it. The backend is a
child of the shell, so an ambient override reaches both processes -- a
resolver that ignored it would look in the per-OS default while the writer
wrote somewhere else. That is the same divergence class as the desktop Logs
panel in #1782, and leaving a newly added resolver knowingly wrong was not an
option.

Two tests: the override moves both candidates, and a whitespace-only value
falls back to the default, matching the writer's !dir.trim().is_empty() guard.
The platform parity test now also clears OMNIVOICE_LOG_DIR so it stays a
statement about the defaults.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 21:46:35 +09:00
Palash Debnath fc237d3929 fix(audiocpp): harden native GPU routing 2026-09-08 18:10:51 +05:30
Palash Debnath 260519d4ee Merge branch 'main' into feat/audiocpp-gpu-routing
# Conflicts:
#	CHANGELOG.md
#	backend/engines/audiocpp/__init__.py
#	backend/engines/audiocpp/bootstrap.py
2026-09-08 17:29:25 +05:30
psiberfunk 9b1f2149ec Merge upstream main into fix/bootstrap-splash-observed-stages-1894 2026-09-08 07:23:38 -04:00
psiberfunk 311ff43de0 Merge upstream main into fix/run-sentinel-clear-before-slow-shutdown-1895 2026-09-08 07:22:34 -04:00
Palash Debnath 8f140e550d Merge pull request #1891 from debpalash/work/local-audiocpp-catalogue
feat(audiocpp): add CPU Breeze-TTS-2 backend and responsive catalogue rows
2026-09-08 16:43:12 +05:30
Palash Debnath 25c41e630c fix(audiocpp): gate readiness on model 2026-09-08 16:29:28 +05:30
Palash Debnath ffc5c07949 fix(audiocpp): require explicit model install 2026-09-08 16:21:08 +05:30
Palash Debnath 503407272e feat(audiocpp): route native GPU backends 2026-09-08 15:05:23 +05:30
Palash Debnath 6b2683eb32 fix(audiocpp): replace stale model aliases 2026-09-08 14:51:42 +05:30
Chang-Jin-LeeandClaude Opus 5 4999bac7a8 fix(system): stop the Tauri tab's Clear from wiping the backend's stderr
/system/logs/tauri/clear iterated every entry in _tauri_log_candidates() and
truncated each one, including backend_err.log -- the spawned backend's stderr.
Three things make that data loss rather than a tidy-up.

The tab that owns the button does not show it. On desktop the Frontend/Tauri
panel goes through the Rust read_log_tail command, whose tauri_log_path()
resolves tauri.log and nothing else, so the user truncates a file they were
never shown.

backend.rs::open_err_log_for_run() opens it APPEND-ONLY so "a respawn must not
destroy the previous run's evidence" (#1510) and rotates it to .1 rather than
truncating. It manages its own size; clearing it from here only undoes that
design. The same file's spawn diagnostics are described there as "retained in
backend_err.log across runs and lands verbatim in bug reports".

A native death -- a Windows access violation, a SIGSEGV -- writes nothing to
the Python log by construction, so this file is the only record it happened.
#1777 and #1782 are both threads where the maintainer had to ask a reporter
for it by hand.

Clear is narrowed to the shell's own log. The READ path is unchanged: the
candidate list was split into two halves and recomposed, and a parametrized
test pins that /system/logs/tauri still reaches all four files in the same
order on darwin, linux and win32 -- the recompose is where a slip would
silently hide a log.

Refs #1510

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 18:04:05 +09:00
Palash Debnath 132fdf313e fix(audiocpp): support cross-filesystem model aliases 2026-09-08 14:34:00 +05:30
Palash Debnath 9bfa3ec751 fix(audiocpp): ship verified CPU runtime 2026-09-08 14:27:19 +05:30
Chang-Jin-LeeandClaude Opus 5 c637166f02 fix(errors): read the origin off the first stack FRAME, not the header
Both bots caught the same defect from opposite sides, and both were right.
`err.stack` starts with a header line carrying the message, and the regex
scanned the whole string for the first URL -- so a URL in the MESSAGE was
mistaken for the throw site.

Greptile's half: an extension error reading "Failed to fetch
https://example.com" reported the message's URL as its origin and escaped the
filter. The bug this PR exists to fix, surviving inside the fix.

CodeRabbit's half, and the worse one: one of OUR failures that happens to
quote a chrome-extension:// URL in its text was suppressed as if an extension
had thrown it. A false positive here silences a real bug, which is strictly
worse than the noise it saves.

Frame detection now covers both stack dialects -- V8's "    at fn (url:1:2)"
after a header, and JSC/SpiderMonkey's "fn@url:1:2" with no header at all --
and reads the URL off the FIRST frame only. When that frame names no URL (a
native or anonymous throw site) the origin is unknown and the report IS
offered: walking deeper would attribute the error to a frame that did not
throw it.

Three tests added, all three red against the previous commit; against
upstream/main the two original extension cases and Greptile's are red, while
the two "still reports" ones pass there because upstream offers a report for
everything.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 16:58:27 +09:00
Chang-Jin-LeeandClaude Opus 5 658269d42f fix(errors): stop offering to report a browser extension's exception
A browser extension injected into the page throws into the page's own error
channel, so window.onerror surfaced it with a "Report this bug" action and a
user filed it. #1901 is one such report: the stack is entirely
chrome-extension://eppiocemhmnlbhjplcgkofciiegomcon/executors/200.js with no
VoiceStudio frame in it, and the message ("Cannot read properties of
undefined") gives a maintainer nothing to tell it apart from a real bug.

IGNORE_PATTERNS matches on the message and cannot help here -- an extension's
TypeError reads exactly like one of ours. The existing `Script error.` entry
covers only the opaque cross-origin case; an extension's script is not opaque,
so it arrives with a full stack and goes straight through.

Filter on the THROW SITE: `e.filename` when the event carries one, else the
first stack frame naming a URL. Deliberately not "any frame mentions an
extension" -- an extension that patches a built-in leaves its frame in the
middle of a stack whose fault is genuinely ours, and dropping those would
silence real bugs, which is worse than the noise it saves. A test pins that
case.

consoleBuffer still records these into Settings -> Logs -> Frontend. What is
suppressed is only the offer to file them against this project.

Closes #1901

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 16:50:18 +09:00
nidhi-singh02 ae85dd36e4 fix(colab): install ASR model before transcription and dubbing 2026-09-08 13:17:49 +05:30
Chang-Jin-LeeandClaude Opus 5 17fda0d79f fix(system): survive a rollover racing the read, and stop trusting the scan
Three review findings, all applied.

Greptile P1, read race: a rollover can rename a candidate between the
existence check and the open, and the handler exposes no lock a route can
take. Per-file OSError now skips that file instead of 500ing the whole panel
-- which is what the single-file version did in the same situation, so this is
strictly better than before rather than a new guarantee. A roll landing
mid-walk can still shift which chunk a file holds, so a tail taken at that
instant may repeat or miss a block; the panel re-polls every 5s and the next
read is clean. Buying strict consistency would mean reaching into logging's
internals from a route.

Greptile P1, clear race: enumerating first left a window where a rollover
created a backup after the scan and its history survived a Clear that
reported success. Clear now works off the fixed name set -- every name the
handler can write is known up front, so there is nothing to enumerate and no
snapshot to go stale.

CodeRabbit: the CHANGELOG lines ended in (#1782), which reads as "this fixes
#1782" when the desktop path defect that thread is about is untouched. Now
(#1920).

Two tests added, both red before: a candidate vanishing mid-walk still fills
the request from the next file, and a Clear whose scan reported nothing still
empties the backups.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 16:46:18 +09:00
Palash Debnath c6fbedbecb fix(audiocpp): resolve security review findings 2026-09-08 13:09:56 +05:30
Palash Debnath e2b5b3d6ed fix(audiocpp): validate the real pinned runtime 2026-09-08 13:02:35 +05:30
Chang-Jin-LeeandClaude Opus 5 15c6cc698d fix(launchpad): constrain the readiness checklist too
CodeRabbit was right: the invariant had a hole. `<ReadinessChecklist compact />`
is a fifth direct child of `.launchpad`, rendered when profiles or studio
projects exist -- which is the state the #1859 reporter was in. The first pass
gave shrink-0 to the three unconditional blocks and missed it, and the test
could not have caught it: the fixture rendered the EMPTY page, where that
branch does not mount.

Wrapped rather than passing the class down. ReadinessChecklist takes no
className, is mounted twice (nested inside the empty state as well as here),
and shrink-0 is a fact about this parent's flex column, not about the
component.

The test now runs the invariant over BOTH page states, because they render
different direct children -- empty gives the flex-1 empty state, populated
gives the checklist -- so checking one leaves the other unconstrained. The
ReadinessChecklist stub renders a marker node instead of null so the wrapper
the page owns is still findable.

Against upstream/main 3 of the 4 in this file are red; against the previous
commit, the two populated-page ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 14:42:19 +09:00
Chang-Jin-LeeandClaude Opus 5 3beebc6d57 fix(system): tail the backend log across a rollover, and clear the backups
main.py rolls omnivoice.log at 2 MB into .1/.2/.3, and /system/logs read only
the current file. For the minutes after a rollover the Backend tab showed a
handful of lines while up to 6 MB of history sat in omnivoice.log.1. Measured
with 3 lines in the current file and 500 in each of two backups: tail=200
returned 3 lines and reported total_lines: 3.

That is the panel CONTRIBUTING and the engine guides tell a reporter to paste
from, so the gap costs a round trip on every bug report that lands near a
roll.

The tail now reaches into the rotated siblings, but only when the current file
cannot satisfy the request -- the panel polls every 5s and opening 6 MB of
backups on each call would be a bad trade for a case that only matters right
after a roll. The response gains a `paths` list so a report can say whether
its tail crossed a boundary.

Clear is in the same commit because the two are coupled: it truncated only
omnivoice.log, so it freed almost nothing, and once the tail can see the
backups a Clear that leaves them looks like it did nothing at all.

Found while reading #1782, and it does NOT close it. That thread's blank panel
is the desktop path, which never reaches this route -- details in a comment
there.

Refs #1782

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 14:39:49 +09:00
Chang-Jin-LeeandClaude Opus 5 a9d90936ce fix(launchpad): let the page scroll instead of compressing its own hero
`.launchpad` is grid row 2 of `.app-container` and is itself a flex column
with `overflow-y: auto`. Opening the log panel grows grid row 4 and shrinks
row 2, which the shell is built for -- row 2 is supposed to hand the shortfall
to its own scrollbar.

It did not, because a flex item shrinks before its container scrolls. The
three content blocks under `.launchpad` carried the default flex-shrink: 1, so
they absorbed the shortfall: the hero is overflow-hidden, so it clipped its
own heading mid-line, and the deck moved up into the hero's artwork. That is
exactly what #1859 describes.

Measured in headless Chromium 153, 720px window, same nesting as the shell,
log panel at 560px:

  variant                       hero box/natural  clipped  deck top vs h1 bottom
  footer collapsed (28px)            145/145        no            +59
  current                             64/145       YES            -22
  + min-height:0 on .launchpad        64/145       YES            -22
  + shrink-0 on the children         145/145        no            +59

The third row answers the lead #1859 flagged as unconfirmed. The missing
min-height: 0 on `.launchpad` -- its two siblings, .app-container >
.main-content and .studio-panel, both set it -- is inert: overflow-y: auto
already zeroes a grid item's automatic minimum size, so the track was
shrinking correctly all along. Not adding it, because a rule that looks like
the fix and is not would mislead the next reader.

The test states the invariant rather than the three class names: a
`.launchpad` child either protects a height (shrink-0) or is deliberately
elastic (flex-1, which is what the empty state is). A fourth block cannot be
added unlabelled.

Not touched: the bell-icon-opens-a-log-console labelling mismatch the report
mentions as a separate observation.

Closes #1859

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 14:27:54 +09:00
Chang-Jin-LeeandClaude Opus 5 4b601ee6c1 fix(bootstrap): hold the run in a ref, and dedup only the handover seam
Two CodeRabbit Major findings, both applied.

O(n^2) retention was mine, introduced in the previous commit. Keeping the full
run in `logs` state meant `prev.concat([entry])` copied the whole array on
every event, so a 5000-line install did ~12.5M element copies and could
stutter the splash on exactly the verbose runs that need it. The run now lives
in `allLogsRef` and is pushed to (O(1)); `logs` state is back to holding only
the rendered tail, so the one array copied per event is bounded by
VISIBLE_LOG_LINES again. `totalLines` is what re-renders on a new line. Copy
reads the ref, and the two full-run scans (detectHints, isUnrecoverableFailure)
read it behind `isFailed`, when nothing is arriving any more.

Permanent dedup was pre-existing and does lose real output: installer text
repeats constantly, and any line matching one of the last five was dropped, so
the counter undercounted and Copy lost lines. The overlap it guards can only
happen on the first live event after backfill, so it now runs until the first
accepted line and never again. Telling a true repeat from a replayed one for
the whole run would need a sequence number from the Rust side; narrowing the
window to where the ambiguity actually is does not.

Three tests added: a repeated line survives, the backfill seam is still
deduped (regression guard for the narrowing), and the rendered state stays at
200 across a 5000-line stream, which is the invariant that keeps the append
cheap. Against upstream main 6 of the 7 in this file are red; the seam test is
a guard, not a fail-before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 14:15:58 +09:00
Chang-Jin-LeeandClaude Opus 5 03c1396b14 fix(bootstrap): keep the whole first-run log, render only the tail
MAX_LOG_LINES = 200 capped the `logs` state itself, so on a cold install every
line past the 200th destroyed an earlier one. Four consumers read that array
and all four degraded once a real install ran past the cap:

- the Activity heading renders logs.length, so the counter sat pinned at 200
  for the rest of a multi-minute bootstrap while lines kept streaming. Live
  progress read as stalled when it was not. This is what #1847 reported.
- handleCopyLogs serializes the same array, so Copy could only ever return the
  newest 200 lines -- and this splash is the only place in the app with a
  copy-log affordance at all.
- detectHints(message, logs) scans for actionable failure markers. A failure
  early in a long install lost its marker, so the card fell back to
  hint_default and told the user nothing specific. This is the severe one: the
  screen still looks helpful while saying nothing.
- isUnrecoverableFailure(message, logs) runs off the same scan.

Only the <pre> needed the cap -- it is a DOM budget, not a retention policy.
Keep the run in state, slice at the one place that writes DOM, and rename the
constant to VISIBLE_LOG_LINES so the next reader cannot make the same mistake.

The array is bounded by one bootstrap: both retry paths clear it and App.jsx
unmounts the splash when the stage flips to 'ready'. The two full-array scans
are behind `isFailed`, so they never run while lines are streaming.

Not fixed here: the other half of #1847, that the splash vanishes on success
with no completion state and the log is then unrecoverable. That needs either
a lifecycle change in App.jsx or a Rust-side persisted stream, and the report
frames the two as separable.

Refs #1847

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 14:03:04 +09:00
Chang-Jin-LeeandClaude Opus 5 276bf626da docs(changelog): match the narrowed health-log wording
The entry said the log records "what kind of failure it was", which is the
overclaim Greptile flagged on the field itself. It records whether the probe
raised, and nothing about the cause.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 13:44:36 +09:00
Chang-Jin-LeeandClaude Opus 5 09a582c65a fix(engines): narrow the health log field to what the probe did
Greptile P1: `failure=unavailable` reads as a classification of the cause, but
SubprocessBackend.health_check() swallows its own exceptions by contract, so a
dead sidecar and a package that was never installed both return (False, msg)
and land in the same bucket. Rename to `probe=`, with `raised:<Class>` and
`returned-unavailable` as the two values, so the field states what the probe
did and claims nothing about why. The limitation and what it would take to fix
it properly (structured failure metadata from the probes) are named in the
comment and the test docstring.

Greptile P2: drop the trailing arrow glyph from the Learn more button. It sat
outside t(), and a bare "→" points the wrong way once the app switches to an
RTL locale. InfoHint hardcodes the same glyph and would want the same
treatment, but that is not this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 13:43:36 +09:00
Chang-Jin-LeeandClaude Opus 5 adfe7b8588 fix(engines): give an unavailable engine row somewhere to send the user
Model Catalogue -> Engines can only ever say "Engine unavailable. Check
installation and configuration." and "Last error: A previous engine check
failed." for an engine whose package is not importable. That is by design:
public_backends() replaces reason and last_error because an availability probe
can carry exception text, a local path or a credential, and two of the shipping
is_available() implementations do interpolate an exception into their message.

So the row cannot explain itself. docs/engines/<engine>.md can -- accurate,
and for cosyvoice CI-guarded against the installer registry -- but nothing in
frontend/src referenced docs/engines at all, so the point of failure was a dead
end. #1746 is that dead end reaching the tracker.

Add a registry-authored docs_url next to install_hint and setup_snippet. It is
a VoiceStudio-owned constant keyed on the engine id, not probe output, so the
public scrub leaves it intact by construction rather than by classification --
which is what keeps the security boundary where the maintainer put it. The row
renders it with the same "Learn more" affordance MCPBindingsPanel and
RemoteBackendPanel already use, so no new i18n key is needed.

The health log line said "Engine health check failed; details withheld" and
named neither the engine nor the kind of failure, while the response tells the
user to check the backend log and docs/engines asks them to copy that engine's
lines. Log the registry id and a stable exception class -- the same class=
shape core.public_errors.public_failure() already logs. The diagnostic text
stays out and the id is flattened to one token, so
tests/test_response_safety.py's existing log-injection test passes unchanged.

Not touched: publishing the computed reason itself. That is the product
decision the reporter flagged, and the log line is deliberate, not an
oversight -- test_engine_health_route_logs_but_does_not_return_private_
diagnostic pins it.

Refs #1866

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 13:29:57 +09:00
Chang-Jin-LeeandClaude Opus 5 d3eb502395 fix(setup): isolate start-truncated paths from the bidi algorithm
The preflight detail line and the storage-path row set dir="rtl" purely to
move the ellipsis to the start of the line, so a long path keeps its tail
visible. That also makes the line an RTL paragraph, and the Unicode bidi
algorithm places a leading run of European numbers or neutrals at the visual
right edge of one. Every detail starting with a digit or a "/" therefore
rendered with its head at the end:

  48.0 GB total                 -> GB total 48.0
  /home/user/.cache/huggingface -> home/user/.cache/huggingface/

Letter-first details were unaffected, which is the signature of bidi
reordering rather than bad data; the backend emits these strings correctly
ordered.

Wrap the text in <bdi> at both call sites. The isolate is dir="auto", so the
run is ordered by its own content while the box keeps direction: rtl for the
ellipsis side. Measured in headless Chromium 153: "48.0" moves from x=49 to
x=0 and "total" from x=20 to x=46, and an overflowing path still has its head
clipped off the start (x=-142), so the start-side ellipsis survives.

jsdom resolves no bidi, so the regression test pins the DOM shape instead, and
a source scan fails any future dir="rtl" element that owns its text directly.

Closes #1848

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chang-Jin-Lee <ckdwls525@gmail.com>
2026-09-08 12:28:28 +09:00
psiberfunk a88ad00bdf fix(audiobook): show the recovery manifest path 2026-09-07 23:02:36 -04:00
psiberfunk ceef6d66bd docs: place OmniVoice fix in highlights 2026-09-07 22:37:19 -04:00
psiberfunk 035add4c5c fix(audiobook): keep recovery actions available 2026-09-07 22:36:25 -04:00
psiberfunk 31abce74b6 test(engines): cover hidden MPS compatibility 2026-09-07 22:33:41 -04:00
psiberfunk 7668d045fa fix(audiobook): refresh resumable jobs after resume 2026-09-07 22:03:02 -04:00
psiberfunk 7157fd0fbe fix(engines): preserve hidden MPS compatibility state 2026-09-07 21:56:46 -04:00
psiberfunk fb98bbfaf6 fix(audiobook): surface interrupted renders 2026-09-07 21:55:40 -04:00
psiberfunk a935c08d6a fix(audiobook): scale chapter timeouts 2026-09-07 21:45:55 -04:00
psiberfunk 8d60929236 fix(engines): hide redundant OmniVoice MPS sidecar 2026-09-07 21:43:58 -04:00
psiberfunk 6e651ecbbd test(studio): cover progress prop forwarding 2026-09-07 20:16:38 -04:00
psiberfunk 4a53fd2668 fix(studio): show honest synthesis progress 2026-09-07 20:08:50 -04:00
psiberfunk 955c62e671 test(dictation): pin timeout completion race 2026-09-07 18:26:10 -04:00
psiberfunk 89c5d46660 test(dictation): cover acknowledgement failure 2026-09-07 18:15:40 -04:00
psiberfunk 70091bf97c fix(dictation): close delivery races 2026-09-07 18:11:12 -04:00
psiberfunk 0189b2a84f fix(dictation): report capture acceptance 2026-09-07 18:06:15 -04:00
psiberfunk 0d7625ee1a fix(dictation): make transcription capture actionable 2026-09-07 17:41:04 -04:00
psiberfunkandClaude Opus 5 5d99271865 style(test): satisfy oxfmt on the widget pill shadow test
"Tests (backend + frontend)" was failing on this branch: the backend suite
passed (7059 tests) and the frontend suite passed, but `bun run format:check`
flagged src/test/widgetPillShadowClip.test.jsx. Formatting only — no
behavioural change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 15:20:16 -04:00
psiberfunkandClaude Opus 5 b5879343a0 fix(bootstrap): treat leaving failed as a new attempt
Review finding on 5e9538a0. The reset keyed only off arriving at a restart
stage, so a retry whose restart stage the poll never sampled did not reset at
all: failed -> checking -> starting_backend inside one ~1s window surfaces as
failed -> starting_backend, leaving the FAILED attempt's stages in
polledStages and rendering its install chrome as this attempt's completed
work. Not merely a late boundary — no boundary.

retry_bootstrap / clean_and_retry_bootstrap are the only exits from `failed`,
so leaving that stage is itself proof a new attempt began. Keying off it as
well as off restart stages closes the case without new producer state.

Regression test fails against 5e9538a0, passes here. Suite green (25 tests).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 14:05:19 -04:00
psiberfunkandClaude Opus 5 5e9538a049 fix(bootstrap): open the attempt boundary when a retry starts, not when polled
Review finding on ece08bd7. The attempt boundary was stamped when the ~1s
status poll first reported `checking` — which lands after the Rust side has
already emitted the new attempt's first log lines. Those lines were then
filtered out as belonging to the previous attempt, so a fast stage that the
poll also missed stayed pending: the exact evidence loss the log union was
added to prevent.

Retries we initiate now open the attempt in beginAttempt(), at initiation, so
their boundary is exact. A guard stops the stage-transition effect from
re-stamping a boundary we already set a poll interval earlier — without it
the fix would have been undone one tick later.

The effect remains the fallback for restarts begun on the Rust side, where
the poll is the only signal available. That window is documented rather than
hidden: it fails toward showing a step pending (conservative and honest)
rather than done (the fabrication this PR removes). Closing it entirely needs
a Rust-provided attempt id — a new IPC surface, deliberately out of scope.

Regression test fails against ece08bd7, passes here. Suite green (24 tests).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 14:02:15 -04:00
psiberfunkandClaude Opus 5 ece08bd790 fix(bootstrap): derive observed stages from logs, and reset them on retry
Two review findings on #1896, both real:

1. Polling misses completed stages. `bootstrap_status` is sampled ~1/s, so a
   stage that starts and finishes between samples was never recorded — on a
   fast disk `creating_venv` routinely does — and would render pending
   forever even though it ran. The stage-tagged `bootstrap-log` stream is
   emitted as the work happens, so a line tagged with a stage is independent
   proof that stage ran. The observed set is now the union of the two; the
   comment claiming the poll "guarantees" a stage lands at least once was
   wrong and is gone.

2. Retry kept stale stages. The observed set was add-only and the splash
   stays mounted across a Retry, so a stage the failed attempt reached would
   still render done in the new attempt even when that attempt skipped it —
   the exact fabrication this change exists to remove. Arriving back at a
   restart stage from anywhere else now starts a fresh attempt. Keyed off the
   stage transition rather than our own Retry buttons, so a Rust-side restart
   resets it too. Log evidence is filtered to the current attempt; the
   visible log is deliberately left alone (clearing it would destroy the
   user's context, cf. #1847).

Regression tests added for both; each fails against 1c344bf9 and passes here.
Full BootstrapSplash suite green (6 files, 23 tests).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 13:57:02 -04:00
psiberfunkandClaude Opus 5 e3e8e80888 fix(lifecycle): scope the sentinel-clear comment honestly for Windows
Greptile P1 on #1897: the comment claimed the moved clear_sentinel() sits
"comfortably inside any shutdown deadline, including Windows' effectively-zero
one". That is false. On Windows tools.rs terminates the job object with no
graceful phase, so lifespan teardown never begins and this line is never
reached — a deliberate quit is still misreported as a crash there.

The code change is unaffected and still correct for the platforms where
teardown does begin. Only the claim was wrong, so only the claim changes.
Windows needs the shell to signal deliberate intent before the hard kill,
which is a Rust-side change tracked separately.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 13:52:42 -04:00
psiberfunkandClaude Opus 5 4ad5aa2831 fix(lifecycle): clear the run sentinel before the slow shutdown work
Measured on macOS build 0.5.2-153: a backend given SIGTERM directly
completes graceful shutdown in 5.25s and clear_sentinel() runs
correctly. But the desktop shell's quit path
(frontend/src-tauri/src/bootstrap.rs:388) grants only a 2s grace
before force_terminate()+kill(), and Windows
(frontend/src-tauri/src/tools.rs:522-528) grants no graceful phase at
all. clear_sentinel() used to be the LAST statement of lifespan
shutdown (backend/main.py:1213), behind ~50s of bounded waits and
model unload/free_vram()/gc.collect()/httpx close — so a deliberate,
clean quit routinely got SIGKILLed before reaching it, leaving
run_sentinel.json behind for the next launch to misreport as a crash.

Move the sentinel clear to the TOP of the shutdown block, immediately
after `yield`: once uvicorn has begun graceful shutdown the exit is
deliberate by definition, so the sentinel has already done its job.
One os.remove is comfortably inside any shutdown deadline, including
Windows' effectively-zero one. The later clear_sentinel() call is
removed (not duplicated) so a later failure in this function can't
mask the early result; the truthful "Shutdown: done."/degraded log at
the end now reads that earlier return value instead of re-clearing.

Adds a regression test that forces a later shutdown step
(model_loads_begin_shutdown) to raise, simulating the kill hitting
mid-teardown, and asserts the sentinel is already gone and
detect_unclean_shutdown() reports no crash. Confirmed fail-before /
pass-after against this change.

Fixes #1895

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 13:48:37 -04:00
psiberfunkandClaude Opus 5 1c344bf9ec fix(bootstrap): stop showing first-run install steps on warm starts
BootstrapSplash derived step "done" state purely from STEPS.indexOf(stage),
so a warm start (bootstrap.rs finds the venv healthy and jumps straight from
Checking to StartingBackend) rendered downloading_uv/creating_venv/
installing_deps as fabricated green DONE ticks, including the "first run,
5-10 min." label. A repair sync (venv exists, only InstallingDeps runs) hit
the same fabrication, and JourneyRail hardcoded Setup=done/Installing=active
regardless of stage.

Track which stages are actually observed (sticky, via the ~1s
bootstrap_status poll) and derive doneness and journey-chrome visibility from
that instead of list position. Journey rail, the "Installing" heading, the
step list, and the resume note stay hidden until a genuine install stage
(downloading_uv/creating_venv/installing_deps/awaiting_setup) is observed;
the masthead, live stage label, progress meter, and activity log are
unaffected. No new user-facing strings.

Fixes #1894

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 13:48:18 -04:00
杨月政andCursor 37aea507a4 fix(setup): lengthen HF connectivity probe timeouts to 8s
China / high-latency paths often need 3–5s just for TCP to
huggingface.co or hf-mirror.com; the previous 2s/3s probes
falsely reported Unreachable on port 443 and could hard-fail
older preflight builds.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-07 23:18:33 +08:00
杨月政andCursor 9979f1a84e fix(bootstrap): apply China PyPI mirror on repair uv sync
Repair previously called apply_uv_env without UV_INDEX_URL, so China
region installs still hit pypi.org for hatchling and failed with
tls handshake eof while the UI showed 中国 (镜像). Fold index selection
into apply_uv_env so first-run, drift, repair, and pip repairs share it.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-07 23:04:24 +08:00
Palash Debnath af216d8411 docs: reference local work draft PR in changelog 2026-09-07 20:21:47 +05:30
Palash Debnath 65ea0dc376 feat: preserve audio.cpp backend and responsive catalogue work 2026-09-07 20:20:54 +05:30
psiberfunkandClaude Opus 5 3105647afb test(capture): assert the error detail, not just a differing title
CodeRabbit: the previous assertions passed for any non-empty title that
differed from the visible text, so a fallback to the label would have
satisfied them — they did not prove `errorInfo.message` takes precedence.

Match on "audio group", which lives only in capture.mic_hint_linux and never
in the label, so a fallback or an unrelated tooltip fails the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 10:07:14 -04:00
psiberfunkandClaude Opus 5 50be1944b4 test(capture): cover the error branch of the pill title
CodeRabbit: only the setup title was covered. Add the error branch that
carries a message, asserting the title is strictly more than the visible
clipped text. The no-message branch falls through to `label` and is covered
by the setup-state test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 10:02:23 -04:00
psiberfunkandClaude Opus 5 e014d63253 test(widget): pin the pill's content width instead of bounding it
CodeRabbit: the max-width assertion used `<=`, so a future cap of 200px
would pass while silently narrowing the pill and clipping more of the label
— the truncation #1884 is about. Assert equality with the real content box
(WINDOW_WIDTH - GUTTER * 2) so the required width is pinned, not just
bounded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 10:00:58 -04:00
psiberfunkandClaude Opus 5 55bcfbdb6a test(macos): make the pill-visible reopen case actually discriminating
CodeRabbit correctly flagged the regression test as tautological: with a
single `main_window_visible` argument, `should_restore_on_reopen(false)`
passes under the aggregate `has_visible_windows` rule too, so the test gave
no protection against the bug it was written for.

Take Cocoa's aggregate flag as a second, deliberately-ignored parameter and
pass it from the event arm. That lets the test state the contract that
matters — main window hidden, pill visible, still restore — and it now fails
if the body ever reverts to deciding on the aggregate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 10:00:26 -04:00
psiberfunkandClaude Opus 5 12b398b662 fix(dictation): stop clipping the widget pill's drop shadow into a rectangle
The `widget` Tauri window is exactly 300x64 with an 8px body padding on
every edge around the 48px pill, leaving only an 8px gutter before
`overflow: hidden` clips anything painted outside it. `.capture-pill`'s
shared shadow (`0 8px 32px` + `0 2px 8px`, ~40px of needed clearance) had
nowhere to go there, so it hard-clipped into a straight edge at the
window boundary — a rounded capsule sitting inside a hard-edged dark
rectangle instead of floating free over the desktop. The recording/
transcribing state shadows had the same problem and fully override the
base shadow, so the clip would reappear the instant dictation started.

Scope a tighter shadow to `html[data-window='widget']` for the base pill
and both state variants, each layer verified to keep
`|y-offset| + blur + spread <= 8` (the gutter), and cap `max-width` to
the window's 284px content box (the shared 340px value is wider than the
window itself). The main-window `.capture-pill-host` path is untouched —
it has real clearance and its shadow already renders correctly there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 09:43:32 -04:00
Palash Debnath 9790d28922 Merge pull request #1883 from debpalash/codex/fix-windows-nonadmin-smoke
fix(ci): prepare hosted Windows policy for non-admin MSI smoke
2026-09-07 18:43:14 +05:30
psiberfunkandClaude Opus 5 7736e9cc60 fix(macos): key Reopen on the main window, not any visible window
The first revision of this handler gated the Dock-icon restore on Cocoa's
`has_visible_windows`. That is too coarse for this app: the dictation pill
is a second, always-on-top window (`widget`) shown and hidden independently
of the main window, and its Accessibility-setup state stays on screen until
the permission is granted.

So with the main window closed and the pill up, Cocoa reports
`has_visible_windows: true`, the guard returns false, and clicking the Dock
icon still does nothing — the exact failure this handler was added to fix.

Ask the main window directly instead. `is_visible()` fails only if the
window is gone, and a redundant show is harmless next to a Dock icon that
stays dead, so an error is treated as "not visible".

Adds a regression test naming the pill case so the coarse predicate cannot
come back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 09:09:34 -04:00
psiberfunkandClaude Opus 5 fdeb8a40fc fix(desktop): handle Dock-icon reopen on macOS (#1887)
Closing the main window with the traffic light hides it (CloseRequested
never destroys it, by design) but the app.run() loop only handled
ExitRequested — there was no RunEvent::Reopen arm anywhere in the
crate, so clicking the Dock icon with no visible windows did nothing.
The window was only recoverable via the tray's "Show VoiceStudio" item
or a full quit/relaunch.

Add a macOS-gated RunEvent::Reopen arm that runs the same
show/unminimize/focus sequence as the tray's "show" handler, now
factored into a shared show_and_focus_main_window() so both recovery
paths can't drift apart. The restore decision (skip when a window is
already visible) is split into a pure should_restore_on_reopen() helper
so it's unit-testable outside the real Cocoa event loop, following the
file's existing with_noactivate_style/is_app_origin pattern.

Scope: macOS-only implementation of a macOS-only platform convention.
Does not touch CloseRequested's hide-instead-of-close behavior, and
does not add multi-window support (out of scope per #1887).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 09:06:09 -04:00
psiberfunkandClaude Opus 5 ebbe74b14b fix(dictation): give the setup-state pill label a hover tooltip
The pill's label <span> only ever set `title` for state === 'error';
every other state, including 'setup', rendered `title={undefined}`.
The 300x64 widget window leaves ~284px for the label after padding and
the status dot/button, so `capture.a11y_setup` ("Allow Accessibility
so dictation can type for you") clips to "Allow Acc…" with nothing to
reveal the rest on hover.

Reuse the already-localized `label` as the tooltip for every state,
keeping the richer error message where one exists. No new strings, no
locale changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 09:00:12 -04:00
Palash Debnath 6a733a4c52 docs: reference installer smoke fix in changelog 2026-09-07 18:21:21 +05:30
Palash Debnath 5cde34c113 fix(ci): prepare hosted Windows policy for non-admin MSI smoke 2026-09-07 18:20:43 +05:30
Palash Debnath bd84169ff2 Merge pull request #1881 from debpalash/codex/fix-per-user-resource-snapshot
fix(windows): preserve frontend resources during per-user MSI build
2026-09-07 17:13:27 +05:30
Palash Debnath 4486b479a1 test: run generated asset hook inside MSI fixture 2026-09-07 16:44:14 +05:30
Palash Debnath 74aa444485 docs: reference installer resource fix PR 2026-09-07 16:43:01 +05:30
Palash Debnath e1e96ae5a2 fix: preserve frontend resources during per-user MSI build 2026-09-07 16:41:20 +05:30
Palash Debnath 0d39a5b283 test: exercise MSI build hooks with changing resource filenames 2026-09-07 16:40:12 +05:30
Palash Debnath 33ce88e465 Merge pull request #1874 from debpalash/codex/pr-queue-integration
chore(integration): land reviewed app, lifecycle and installer fixes
2026-09-07 15:51:52 +05:30
Palash Debnath 10a0fc27f0 Merge commit '63ade08b' into codex/pr-queue-integration 2026-09-07 15:30:30 +05:30
Palash Debnath 63ade08b6f fix(i18n): reuse translated token cleanup error 2026-09-07 15:30:09 +05:30
Palash Debnath fc7021e2b4 Merge commit '1038a487' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 15:26:02 +05:30
Palash Debnath 1038a487a8 fix(auth): gate onboarding replacement on known token state 2026-09-07 15:25:01 +05:30
Palash Debnath 19b9b4a136 Merge commit 'd9bf2516' into codex/pr-queue-integration 2026-09-07 15:09:06 +05:30
Palash Debnath d9bf251665 test(auth): exercise token paths on native backend hosts 2026-09-07 15:08:24 +05:30
Palash Debnath 5bc9f7856a Merge commit '8428517d' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 15:04:50 +05:30
Palash Debnath 8428517d95 fix(auth): preserve local Hugging Face token paths and cleanup 2026-09-07 15:03:12 +05:30
Palash Debnath 6caabf3b0b Merge commit '6b802bea' into codex/pr-queue-integration
# Conflicts:
#	tests/test_locale_parity.py
2026-09-07 14:29:30 +05:30
Palash Debnath 6b802beaf9 fix: complete compact engine family locale labels 2026-09-07 14:28:45 +05:30
yearthmain e059c4f330 docs(changelog): note the zh-CN locale completion (#1877)
Signed-off-by: yearthmain <yearthmain@gmail.com>
2026-09-07 16:56:13 +08:00
yearthmain 8998faa98b fix(i18n): translate all 486 missing zh-CN keys, tighten parity ratchet to 0
zh-CN.json was missing 486 of en.json's 3,057 leaf keys (baseline 486 in
_MISSING_BASELINE), so Settings / Models / Engines / Dictation surfaces
rendered English fallback. Translate every missing key following the ko
overhaul in #1776: brand and technical terms verbatim (Tauri, Discord,
Hugging Face, LLM, FFmpeg, torch.compile, DELETE), {{placeholders}}
preserved on all 70 keys that carry them, i18next tags (<1>, <code>,
<issueLink>) intact, existing translations and key order untouched.

Tighten _MISSING_BASELINE['zh-CN'] from 486 to 0 — zh-CN now matches
ko at full parity, verified with tests/test_locale_parity.py (248
passed).

Signed-off-by: yearthmain <yearthmain@gmail.com>
2026-09-07 16:47:30 +08:00
Palash Debnath 2ed38c476e Merge commit '99f93164' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
#	tests/test_locale_parity.py
2026-09-07 14:16:01 +05:30
Palash Debnath 078bc06428 Merge commit 'af245a45' into codex/pr-queue-integration 2026-09-07 14:15:11 +05:30
Palash Debnath 99f93164be fix: translate Hugging Face token source labels 2026-09-07 14:14:29 +05:30
Palash Debnath af245a45ee fix: complete workspace playback and engine translations 2026-09-07 14:14:28 +05:30
Palash Debnath dbbad3d58e Merge commit 'bcd2e531' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
#	docs/install/macos.md
2026-09-07 13:48:19 +05:30
Palash Debnath 8667b34ca5 Merge commit 'f18a7add' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
#	docs/install/macos.md
2026-09-07 13:47:50 +05:30
Palash Debnath bcd2e53175 fix(lifecycle): verify root exit after Darwin signal permission errors 2026-09-07 13:46:59 +05:30
Palash Debnath f18a7adddf fix(lifecycle): verify root exit after Darwin signal permission errors 2026-09-07 13:46:06 +05:30
Palash Debnath a968bc57b3 Merge commit 'f2cfdaa4' into codex/pr-queue-integration 2026-09-07 13:31:45 +05:30
Palash Debnath 7ac5e5ab3d Merge commit 'a768e84f' into codex/pr-queue-integration 2026-09-07 13:31:35 +05:30
Palash Debnath f2cfdaa430 test: assert first-sound Chinese taxonomy normalization 2026-09-07 13:31:12 +05:30
Palash Debnath a768e84ffe test: enforce device-neutral Confucius catalog label 2026-09-07 13:31:11 +05:30
Palash Debnath 4648c23256 Merge commit '24252712' into codex/pr-queue-integration 2026-09-07 13:19:19 +05:30
Palash Debnath 1e96748402 Merge commit '30307ae8' into codex/pr-queue-integration
# Conflicts:
#	.gitignore
2026-09-07 13:19:19 +05:30
Palash Debnath 169e36d5b6 Merge commit '0e7ee3d9' into codex/pr-queue-integration 2026-09-07 13:18:56 +05:30
Palash Debnath 135dd0a3a7 Merge commit '9dbf45ca' into codex/pr-queue-integration 2026-09-07 13:18:49 +05:30
Palash Debnath ef51c6b75d Merge commit 'e32bc942' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 13:18:49 +05:30
Palash Debnath 24252712d2 fix(moss): use declared device identifiers in install hint 2026-09-07 13:18:27 +05:30
Palash Debnath b49d2954d5 Merge commit '4008499e' into codex/pr-queue-integration 2026-09-07 13:18:20 +05:30
Palash Debnath bf8da941d1 Merge commit '3b64692e' into codex/pr-queue-integration 2026-09-07 13:18:11 +05:30
Palash Debnath ca3bf8367f Merge commit 'fc8db259' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 13:18:11 +05:30
Palash Debnath 0c255dfabd Merge commit '9afc9db2' into codex/pr-queue-integration 2026-09-07 13:17:45 +05:30
Palash Debnath 4008499ea5 fix(confucius): list supported device codes in the install hint 2026-09-07 13:17:37 +05:30
Palash Debnath e32bc94248 fix: preserve conversion state and finish workspace labels 2026-09-07 13:17:00 +05:30
Palash Debnath fc8db259a9 fix(bootstrap): prevent stale timeout after retry invalidation 2026-09-07 13:16:43 +05:30
Palash Debnath 0921292319 fix(confucius): keep catalogue hardware metadata accurate 2026-09-07 13:15:07 +05:30
Palash Debnath 9afc9db287 test(budget): resolve application modules at test runtime 2026-09-07 13:14:14 +05:30
Palash Debnath 3b64692eed fix(moss): align catalog install hint with accelerator routing 2026-09-07 13:13:49 +05:30
Palash Debnath 30307ae883 chore: ignore generated Windows MSI diagnostic artifacts 2026-09-07 13:13:36 +05:30
Palash Debnath 9dbf45ca2c test(onboarding): resolve the current runtime validator 2026-09-07 13:12:52 +05:30
Palash Debnath 0e7ee3d912 test(release): isolate cleanup module instances per test 2026-09-07 13:12:52 +05:30
Palash Debnath 140e9262b9 Merge commit '18a0c11c' into codex/pr-queue-integration 2026-09-07 12:47:40 +05:30
Palash Debnath 18a0c11c23 test(devices): clear mocked live probe cache 2026-09-07 12:44:32 +05:30
Palash Debnath a0b84a6b61 Merge commit '8c67b109' into codex/pr-queue-integration 2026-09-07 12:42:34 +05:30
Palash Debnath 8c67b10942 test(devices): patch the active capability module after reloads 2026-09-07 12:42:13 +05:30
Palash Debnath 760e9e5ac7 Merge commit '25498dd1' into codex/pr-queue-integration 2026-09-07 12:31:58 +05:30
Palash Debnath 25498dd17b docs(devices): explain optional DirectML fallback 2026-09-07 12:31:33 +05:30
Palash Debnath 16020c2178 Merge commit 'a41e66e0' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 12:28:46 +05:30
Palash Debnath e4471008b4 Merge commit 'df46b71b' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 12:28:27 +05:30
Palash Debnath 2823d6a8ea Merge commit '57150ede' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 12:28:06 +05:30
Palash Debnath 57150ede64 fix(moss): recover from accelerator probe failures 2026-09-07 12:26:04 +05:30
Palash Debnath df46b71be2 fix: tolerate Confucius and DOTS accelerator probe failures 2026-09-07 12:24:42 +05:30
Palash Debnath a41e66e0f0 docs: note capture widget hide permission repair 2026-09-07 12:23:48 +05:30
Palash Debnath 2b2bfe1a7f fix(capture): allow the widget window to hide after recording 2026-09-07 12:23:32 +05:30
Palash Debnath f728071819 Merge commit '59a487e7' into codex/pr-queue-integration 2026-09-07 12:13:44 +05:30
Palash Debnath 59a487e742 test: start upload watchdog at the blocked write 2026-09-07 12:13:29 +05:30
Palash Debnath ef9bda2e17 Merge commit 'e8a69508' into codex/pr-queue-integration 2026-09-07 12:09:10 +05:30
Palash Debnath e8a6950898 fix: allow exact nonsecret dubbing pane storage key 2026-09-07 12:08:47 +05:30
Palash Debnath 43d47f4d36 Merge commit '93bbfd64' into codex/pr-queue-integration 2026-09-07 12:04:25 +05:30
Palash Debnath 93bbfd64ea docs(device): explain optional accelerator probe fallbacks 2026-09-07 12:03:38 +05:30
Palash Debnath c973b1c983 Merge commit 'b4b10c14' into codex/pr-queue-integration 2026-09-07 12:03:21 +05:30
Palash Debnath c4530b4919 Merge commit '99218e2c' into codex/pr-queue-integration
# Conflicts:
#	frontend/src/components/Header.jsx
2026-09-07 12:03:21 +05:30
Palash Debnath b4b10c1431 fix(header): satisfy native Mac detection lint gate 2026-09-07 12:02:35 +05:30
Palash Debnath 99218e2c34 fix(header): satisfy native Mac detection lint gate 2026-09-07 12:02:15 +05:30
Palash Debnath b759ae7749 Merge commit 'ee3e1474' into codex/pr-queue-integration 2026-09-07 11:58:24 +05:30
Palash Debnath b2f5de328e Merge commit 'af7f1ceb' into codex/pr-queue-integration 2026-09-07 11:58:16 +05:30
Palash Debnath f2dd6c95db Merge branch 'codex/fix-per-user-msi' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:58:16 +05:30
Palash Debnath ee3e147401 test: prove upload revocation without scheduling threshold 2026-09-07 11:57:04 +05:30
Palash Debnath af7f1ceb3d test: construct heartbeat backend with lifecycle state 2026-09-07 11:56:16 +05:30
Palash Debnath 0cc22ad803 test: reject unresolved MSI authoring and invalid component identity 2026-09-07 11:56:13 +05:30
Palash Debnath 62b0e05c1d fix(windows): preserve registry separators during template expansion 2026-09-07 11:55:42 +05:30
Palash Debnath aed50121b4 docs: reference per-user MSI repair PR (#1873) 2026-09-07 11:54:25 +05:30
Palash Debnath 079be2fbe5 docs: describe automatic Windows MSI authoring validation 2026-09-07 11:52:41 +05:30
Palash Debnath f463276ca5 ci: validate both Windows MSI scopes with a tiny payload 2026-09-07 11:52:22 +05:30
Palash Debnath a438f5db0e Merge branch 'codex/fix-sidecar-timeout-reap' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:45:27 +05:30
Palash Debnath f72ab3a3b2 fix(sidecar): quarantine timeout owners until bounded cleanup succeeds 2026-09-07 11:41:18 +05:30
Palash Debnath 1d30a50c1c docs: reference sidecar recovery PR (#1872) 2026-09-07 11:30:21 +05:30
Palash Debnath c36df0ddf0 Merge branch 'codex/fix-sidecar-timeout-reap' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:29:26 +05:30
Palash Debnath 85cb7242ea fix(sidecars): finish timeout cleanup before allowing recovery 2026-09-07 11:28:13 +05:30
Palash Debnath c547da351c Merge branch 'codex/consolidate-1809' into codex/pr-queue-integration 2026-09-07 11:22:48 +05:30
Palash Debnath 602be02a17 Merge branch 'codex/review-1862' into codex/pr-queue-integration 2026-09-07 11:22:48 +05:30
Palash Debnath f395c90aa1 Merge branch 'codex/review-dialog-resolution' into codex/pr-queue-integration 2026-09-07 11:22:48 +05:30
Palash Debnath 01cb810c65 test(bootstrap): release launch gate before timeout assertion 2026-09-07 11:22:18 +05:30
Palash Debnath f4e5d7b9ab Merge branch 'codex/review-1806' into codex/pr-queue-integration 2026-09-07 11:22:11 +05:30
Palash Debnath caedcde0c6 Merge branch 'codex/review-logs-state' into codex/pr-queue-integration 2026-09-07 11:22:11 +05:30
Palash Debnath 2a36a24612 Merge branch 'codex/consolidate-1831' into codex/pr-queue-integration 2026-09-07 11:22:11 +05:30
Palash Debnath 610303ec11 Merge branch 'codex/consolidate-1830' into codex/pr-queue-integration 2026-09-07 11:22:11 +05:30
Palash Debnath 3a114d62dd test(header): isolate visual fixture system polling 2026-09-07 11:21:55 +05:30
Palash Debnath 65e6c275c6 fix(workers): retain conservative budgets for legacy missing workers 2026-09-07 11:21:39 +05:30
Palash Debnath 9a90fe41be test(vite): exercise conditional dialog alias configuration 2026-09-07 11:21:32 +05:30
Palash Debnath d55838542e fix(confucius): preserve legacy torch device selection 2026-09-07 11:21:13 +05:30
Palash Debnath 9c0a5469cc fix(moss): retain older manually provisioned torch environments 2026-09-07 11:21:12 +05:30
Palash Debnath c7a8d24af6 fix(logs): report failed refresh after clearing logs 2026-09-07 11:19:51 +05:30
Palash Debnath 7d963cf297 Merge branch 'fix/release-rerun-asset-collisions' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:19:07 +05:30
Palash Debnath b9cb6c204f Merge branch 'codex/review-1865' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
#	frontend/src/components/Header.jsx
2026-09-07 11:18:50 +05:30
Palash Debnath 6157c1717b Merge branch 'codex/review-1863' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:18:28 +05:30
Palash Debnath c3b54393c4 docs: record release retry recovery fix (#1871) 2026-09-07 11:17:30 +05:30
Palash Debnath f0bb3fb8b2 Merge branch 'codex/consolidate-1810' into codex/pr-queue-integration 2026-09-07 11:17:24 +05:30
Palash Debnath 157f2987dc fix(api): retain Node ESM compatibility for abortable delay imports 2026-09-07 11:17:24 +05:30
Palash Debnath 426bd1ecea fix(desktop): preserve macOS window behavior with native controls 2026-09-07 11:17:17 +05:30
Palash Debnath d34490fd15 docs: explain retrying partially published releases 2026-09-07 11:16:29 +05:30
Palash Debnath 9b9d4d6579 fix(header): reserve traffic-light space only in native Mac windows 2026-09-07 11:16:24 +05:30
Palash Debnath eee35ab771 Merge branch 'codex/consolidate-1831' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:16:00 +05:30
Palash Debnath 0095c4a089 Merge branch 'codex/consolidate-1830' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:15:51 +05:30
Palash Debnath ef30dbece7 fix(release): clear target asset collisions before retry uploads 2026-09-07 11:15:40 +05:30
Palash Debnath a9062c2fa1 Merge branch 'codex/review-logs-state' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
#	frontend/src/test/LogsFooterNotifications.test.jsx
2026-09-07 11:15:38 +05:30
Palash Debnath b7ebd6e54d Merge branch 'codex/review-hf-onboarding' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:15:16 +05:30
Palash Debnath cd58bbded0 fix(moss): align accelerator detection and routing status 2026-09-07 11:13:33 +05:30
Palash Debnath b9183bb2a6 fix(engines): align accelerator metadata and DOTS runtime precision 2026-09-07 11:12:38 +05:30
Palash Debnath e9495beb8b fix(auth): inspect token presence locally until explicit validation 2026-09-07 11:12:26 +05:30
Palash Debnath 409dd016ce fix(logs): require successful current snapshots before all-clear 2026-09-07 11:11:46 +05:30
Palash Debnath db10e4b70f Merge branch 'codex/review-1862' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
#	frontend/src/test/visual/specs.jsx
2026-09-07 11:10:42 +05:30
Palash Debnath 3613b3d39b Merge branch 'codex/review-1861' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath 4277859a44 Merge branch 'codex/review-ui-1841' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath eefc45ec49 Merge branch 'codex/review-1821' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath 2f5a52f9eb Merge branch 'codex/review-1819' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath 729cad8912 Merge branch 'codex/review-1806' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath 86f9eda65f Merge branch 'codex/consolidate-1810' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath 32c96eeee4 Merge branch 'codex/consolidate-1809' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath f5d52e479e Merge branch 'codex/review-dialog-resolution' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath eafd94d0a6 Merge branch 'codex/review-oom-order' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath 2b5e3cdbc6 Merge branch 'codex/consolidate-1799' into codex/pr-queue-integration
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:10:11 +05:30
Palash Debnath 2229a68cfc fix(gallery): calibrate preview guard against shipped speech fixtures 2026-09-07 11:07:16 +05:30
Palash Debnath 0581fb69fd Merge current main into PR #1819 2026-09-07 11:07:16 +05:30
Palash Debnath 85a368a08b test(header): verify reduced motion and responsive status visibility 2026-09-07 11:06:37 +05:30
Palash Debnath a8c57c816b test(onboarding): validate first-sound request against engine taxonomy 2026-09-07 11:05:30 +05:30
Palash Debnath 360d1d29c8 docs: record connection diagnostics fix under Unreleased 2026-09-07 11:04:51 +05:30
Palash Debnath 7366555c95 docs: record startup fix under Unreleased 2026-09-07 11:04:50 +05:30
Palash Debnath b8ff2721a5 fix(tts): normalize complete signed ranges without rewriting chains 2026-09-07 11:04:44 +05:30
Palash Debnath da26a240e2 Merge current main into PR #1821 2026-09-07 11:04:44 +05:30
Palash Debnath fb3d9b7139 fix(bootstrap): preserve retry cancellation across lifecycle acquisition 2026-09-07 11:03:42 +05:30
Palash Debnath 6c47ae7b21 fix(api): honor cancellation throughout transport diagnostic waits 2026-09-07 11:03:05 +05:30
Palash Debnath fda73384f0 fix(workers): retain granted deadlines across disconnects and restart 2026-09-07 11:02:45 +05:30
Palash Debnath 84add8b279 Merge current main into PR #1806 2026-09-07 11:02:27 +05:30
Palash Debnath 8b510d73db fix(ui): address workspace review findings and restore CI 2026-09-07 11:02:18 +05:30
Palash Debnath b1194c6223 test: cover nested and hoisted dialog dependency resolution 2026-09-07 11:02:06 +05:30
Palash Debnath 75f8924222 Merge remote-tracking branch 'origin/main' into codex/review-dialog-resolution
# Conflicts:
#	CHANGELOG.md
2026-09-07 11:00:31 +05:30
Palash Debnath 2a6d089f4a Merge remote-tracking branch 'origin/main' into codex/consolidate-1810 2026-09-07 11:00:17 +05:30
Palash Debnath 5d134f22ec test: enforce VRAM reclaim before clone prompt retry 2026-09-07 10:59:09 +05:30
Palash Debnath 08af0a971e Merge remote-tracking branch 'origin/main' into codex/review-oom-order 2026-09-07 10:58:32 +05:30
Palash Debnath ff0ce4d37d docs: keep pending transcription fix under Unreleased 2026-09-07 10:58:08 +05:30
Palash Debnath 1ab03067cf Merge remote-tracking branch 'origin/main' into codex/consolidate-1809 2026-09-07 10:58:01 +05:30
Palash Debnath 7a2f86066b fix(transcriptions): consolidate safe clipboard handling and regression tests 2026-09-07 10:57:08 +05:30
Palash Debnath d91beef0fd docs: credit Windows console help fix (#1815) 2026-09-07 10:56:19 +05:30
Palash Debnath 574b634688 Merge remote-tracking branch 'origin/main' into codex/review-cp1252 2026-09-07 10:55:45 +05:30
Palash Debnath 62ab62fc57 Merge remote-tracking branch 'origin/main' into codex/consolidate-1799 2026-09-07 10:55:05 +05:30
电车司机小李 ecbd152c43 fix(logs): avoid false all-clear state
Signed-off-by: 电车司机小李 <39351936+motodriver@users.noreply.github.com>
2026-09-07 11:51:02 +08:00
psiberfunkandClaude Sonnet 5 fec2e7b5f3 docs(changelog): credit the contributor for #1864
Greptile flagged the Unreleased entry as missing the contributor
credit the changelog convention requires for community PRs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 21:02:33 -04:00
psiberfunkandClaude Sonnet 5 04bc9cedac fix(header): hide the custom window controls on macOS
macOS draws its own native traffic-light cluster even with
decorations:false (tauri.conf.json's titleBarStyle:"Overlay" still
overlays it), but Header.jsx's showWindowControls only checked
whether the app was running under Tauri, not which OS — so the
custom Windows-style minimize/maximize/close row rendered on macOS
too, duplicating the native controls.

Gate it on platform using the same navigator.platform check already
used in HotkeyTab.jsx / SettingsSearch.jsx. Windows/Linux keep the
custom row since decorations:false gives them no chrome otherwise.

Fixes #1864.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 20:51:38 -04:00
psiberfunkandClaude Opus 5 e397e10d64 test(header): restore navigator.platform after mac-inset tests
CodeRabbit flagged that setPlatform() redefined navigator.platform as
an own property but nothing ever restored it, so after this file's
tests run the global stays pinned to whichever platform ran last
('Linux x86_64') — order-dependent and able to leak into any later
test in the same environment that reads navigator.platform. Capture
the original descriptor (undefined, since it's an inherited jsdom
getter) and restore it — or delete the own-property override — in
afterEach.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 20:33:58 -04:00
psiberfunkandClaude Opus 5 8e1da3c22f fix(firstrun): send a voice-design-safe instruct instead of omitting it
Greptile flagged that omitting `instruct` from the first-sound request
fixes OmniVoice (whose `_resolve_instruct` rejected the old free-text
prose) but breaks a different engine: mlx-audio's Qwen3 VoiceDesign
backend requires a truthy `instruct` and raises ValueError without one
(`_is_voice_design()` in backend/services/tts_backend.py), so a user who
picked that engine during onboarding would still get silent first-sound
failure — same bug class, different engine.

Send 'middle-aged, low pitch' instead of omitting the field: it's the
exact taxonomy string the backend's own "Narrator" personality preset
uses (backend/core/personalities.py), so it's valid vocabulary for
OmniVoice's `_resolve_instruct` and a non-empty description for any
voice-design engine. Rewrote firstSoundInstruct.test.js, which
previously asserted instruct was absent entirely (passing for the wrong
reason); it now asserts a non-empty, taxonomy-only instruct is sent and
cross-checks its value against personalities.py's narrator preset so the
two can't silently drift apart.

Also credited the community contributor in CHANGELOG.md per the repo's
own convention (Greptile P2) and promoted the entry to Highlights,
matching every other credited entry in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 20:32:25 -04:00
psiberfunkandClaude Opus 5 b92cc3bb79 fix(setup): keep the overwrite warning visible on narrow screens
CodeRabbit flagged that the hf_token_replace_warning text shared the
same `max-[560px]:hidden` class as the dismissable "add a token" pitch,
so a user replacing an already-active token on a narrow viewport (mobile
width, or a small first-run window) never saw the warning that doing so
clobbers the working token. Only the pitch should hide at that width —
the overwrite warning is safety copy and must always render. Added a
regression test asserting the warning's className never carries the
responsive-hide class.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 20:26:33 -04:00
psiberfunkandClaude Opus 5 d757f160fa fix(header): inset the breadcrumb clear of macOS traffic lights
tauri.conf.json sets decorations:false + titleBarStyle:"Overlay" on
every platform, so on macOS the native traffic-light cluster is drawn
on top of the web content instead of getting its own row; Windows and
Linux draw nothing there. The header's left block (status dot +
kicker) had no inset at all for that zone, so on macOS the traffic
lights sat on top of it.

Fix, macOS-only: detect macOS the same way HotkeyTab.jsx /
SettingsSearch.jsx already do (navigator.platform), and apply a new
.header-area__left--mac-inset class that completes header-area's own
16px left padding to the same flat 64px-from-window-edge total that
.header-area--tabs already reserves for the identical cluster. Windows
and Linux get no inset, so no space is wasted where nothing is
overlaid.

Fixes #1860.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 19:53:18 -04:00
psiberfunkandClaude Opus 5 a9ac8b1b46 fix(header): stop the status dot pulsing under Reduce Motion
Header.jsx's status dot animates via the hqPulse keyframes, applied as
a Tailwind arbitrary [animation:...] utility. None of index.css's
twelve @media (prefers-reduced-motion: reduce) blocks named hqPulse
(it isn't a stable CSS class, so those selector-based blocks can't
reach it), so the purely decorative pulse kept running with OS Reduce
Motion on.

Fix: append motion-reduce:[animation:none] to the dot's className -
the same mechanism LogsFooter.jsx already uses for its own
arbitrary-utility pulses (heart-glow, donate-pop-in).

Part B only, per the issue split - Part A, an in-app motion toggle, is
a product decision and stays open.

Fixes #1857 (Part B).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 19:49:32 -04:00
psiberfunkandClaude Opus 5 f88bfecabc fix(firstrun): stop sending free-text instruct prose on first sound
App.jsx's post-onboarding "first sound" request appended a hardcoded
narrator prose string as `instruct`. Every engine's instruct is a
controlled vocabulary (OmniVoice's `_resolve_instruct` rejects
anything outside a fixed token list), so this 400ed on every first
run — silently, since the surrounding catch is deliberately silent
by design (a first impression must never surface an error).

Omit `instruct` entirely instead of swapping in valid vocabulary:
it matches every other call site in the app (`if (instruct)
fd.append('instruct', ...)`), matches the seeded demo profile's
empty stored instruct, and every engine backend already treats a
missing/empty instruct as "no styling" rather than a required field
— so this can't regress no matter which TTS engine is active.

Fixes #1853.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 19:46:15 -04:00
psiberfunkandClaude Opus 5 753891adc9 fix(setup): stop pitching an HF token when one is already active
HfTokenCard.jsx unconditionally rendered the "add a free Hugging Face
token" pitch in first-run's Models & engines step, even when the
backend had already resolved and validated one (app/env/hf-cli). Since
Save persists via huggingface_hub.login(), which overwrites the
canonical $HF_HOME/token file outright, complying with the unnecessary
prompt could silently clobber an already-working token.

The card now checks GET /system/hf-token/state (the same resolver the
Settings -> API Keys panel already consumes) before rendering:
- an active, validated token shows the source + masked value instead
  of the pitch
- replacing it requires an explicit "Replace..." click plus an inline
  overwrite warning, rather than one blind paste-and-Save
- a still-loading check shows a neutral placeholder
- a failed check falls back to the pre-fix pitch rather than hiding
  the card

Fixes #1851.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 16:54:26 -04:00
Palash Debnath 4efde9ce4e feat(ui): polish dubbing workspace layout and controls 2026-09-06 20:26:43 +05:30
li-lizhe ae2fa83b02 fix: handle None from current_accelerator(); exclude MPS in dots_tts
current_accelerator() returns None on CPU-only builds (no accelerator
compiled in), so .type would crash. Use check_available=True and fall back
to 'cpu' when None. For dots_tts, also select fp32 on MPS since
DotsTtsRuntime is untested on MPS.

Addresses greptile P1 + coderabbit Functional Correctness review comments.
2026-09-06 21:06:33 +08:00
li-lizhe a671004cb7 fix: handle None from current_accelerator() on CPU-only builds
current_accelerator() returns None on CPU-only PyTorch builds (no
accelerator compiled in), so accel.type would crash. Use check_available=True
and fall back to 'cpu' when None.

Addresses greptile P1 + coderabbit Stability review comments.
2026-09-06 21:06:19 +08:00
li-lizhe ed6fca46a9 fix(tts-engines): select device via torch.accelerator in confucius4 and dots_tts
Both Confucius4 and DOTS-TTS engine sidecars hardcode device selection to
`torch.cuda.is_available()`, which returns False on Ascend NPU, Intel XPU,
and other non-CUDA accelerators — causing the models to silently run on CPU
(in fp32) instead of the available accelerator.

Replace with `torch.accelerator.current_accelerator().type`, the
device-agnostic API that auto-detects CUDA, NPU, XPU, MPS, and CPU. MPS is
excluded for Confucius4 (upstream untested on Apple Silicon). dtype stays
bf16 for any GPU-class accelerator and fp32 on CPU.

Also verified on Ascend 910B (torch 2.14, torch_npu, 4 NPU):
  before: cuda_available=False → device "cpu", precision "float32"
  after:  accelerator → confucius4 device="npu", dots_tts precision="bfloat16"
2026-09-06 09:24:39 +08:00
li-lizhe 65d37288ce fix(moss_tts_v15): select device via torch.accelerator instead of CUDA hardcode
MOSS-TTS-v1.5 engine hardcoded device selection to
`device = "cuda" if torch.cuda.is_available() else "cpu"`. On an Ascend
NPU (torch_npu) host, `torch.cuda.is_available()` is False, so the whole
model silently runs on CPU in fp32 — never using the accelerator — even
though torch.accelerator reports `npu` and bf16 is supported.

Replace with the device-agnostic `torch.accelerator.current_accelerator()`
so any backend (CUDA / NPU / XPU / MPS) is picked up automatically. MPS is
still excluded (MOSS's upstream trust_remote_code modelling code is untested
on Apple Silicon); dtype is bf16 for any GPU-class accelerator and fp32 on
CPU.

Verified on Ascend 910B (torch 2.14, torch_npu, 4 NPU):
  before: torch.cuda.is_available()==False -> device "cpu", dtype float32
  after:  accelerator -> device "npu", dtype bfloat16
2026-09-06 09:23:22 +08:00
Palash Debnath 293b0812f4 feat(ui): organize dubbing controls and export drawer 2026-09-05 23:17:50 +05:30
Palash Debnath b441a486cc feat(ui): refine voice controls and expandable navigation 2026-09-05 22:37:02 +05:30
Palash Debnath d4bf1fe9a9 feat(ui): polish voice interactions and fix notification badge 2026-09-05 20:43:42 +05:30
Palash Debnath 5574cbbc16 fix(ui): name OmniVoice correctly and cycle active engine labels 2026-09-05 20:06:24 +05:30
Palash Debnath 13eae6ff02 fix(ui): show selected engine on title bar 2026-09-05 19:44:55 +05:30
Palash Debnath fd5e78cdab refactor(ui): simplify engine quick access with family tabs 2026-09-05 19:38:19 +05:30
Palash Debnath 67c316e81b feat(studio): consolidate engine controls and pin mode actions 2026-09-05 18:52:43 +05:30
Palash Debnath 81fb585557 fix(studio): anchor engine menu to header trigger 2026-09-05 17:51:57 +05:30
Palash Debnath 3675d750d6 fix(studio): reuse rich language picker for cloning 2026-09-05 17:46:08 +05:30
Palash Debnath bbbf185cf3 fix(studio): keep sticky language picker inside viewport 2026-09-05 17:11:24 +05:30
Palash Debnath 0c4ed0546f fix(studio): constrain sticky controls on short screens 2026-09-05 16:38:43 +05:30
Palash Debnath 17df9210cd fix(studio): pin synthesis controls below scrolling form 2026-09-05 16:26:13 +05:30
Palash Debnath aedca15f7e docs: record voice workspace tabs 2026-09-05 15:23:59 +05:30
Palash Debnath 7a14c31a8e feat(studio): promote voice modes to workspace tabs 2026-09-05 15:22:44 +05:30
flutterkage2kandClaude Opus 5 ed746bab57 fix(tts): speak the tilde in digit ranges instead of mashing the numbers
"20~30초" is read aloud as a single number — OmniVoice says "이십삼" (23).
The separator never reaches the listener, so any written range is heard as
the wrong figure.

`normalize_text` only ran its number pass behind `_num2words_lang`, which
returns None for ko/ja/zh/th/vi (those scripts read digits natively and are
deliberately outside num2words). Nothing else looked at the range mark, so
the tilde went to the engine untouched and the two numbers ran together.

Rewrite `N~M` into the spoken form before the engine sees it, outside the
num2words gate so the CJK languages are covered too. Verified by rendering
each candidate and transcribing it back (ko, OmniVoice, cloned voice):

    "대략 20~30초짜리"      heard "23초"           WRONG
    "대략 20-30초짜리"      heard "23초"           WRONG (reproduces it)
    "대략 20에서 30초짜리"   heard "20에서 30초짜리"  correct
    "20〜30分ぐらい" → "20から30分" heard "20〜30分くらい"  correct

Deliberately narrow:

* Only the tilde family (U+007E, U+301C, U+FF5E). Japanese and Korean IMEs
  emit the latter two. An ASCII hyphen is left alone — between digits it
  also spells dates, phone numbers and product codes, where "to" is wrong
  (`tests` already pin "pages 3-5" as unchanged).
* Only languages with a verified spoken form (ko/ja/zh/en). Anything else
  keeps its tilde, matching how `_PERCENT_WORD` is scoped.
* Spacing belongs to the form, not the caller: a Korean postposition binds
  to its numeral ("20에서 30"), Japanese and Chinese set no spaces, English
  needs them on both sides.
* Neighbour guards block digits and ASCII letters but allow CJK, because
  CJK writes the unit hard against the digits ("20~30초"); a `\w` guard
  rejects exactly the cases the rule exists for.

`ko`/`ja`/`zh` join `_FULL_NAME_TO_CODE` so the new resolver can see them.
They stay out of `_NUM2WORDS_LANGS`, so this does not open a num2words path
for them — the same inert-entry pattern the file already documents for
"vietnamese".

`backend/services/text_normalization.py` joins the functional-CJK allowlist
in tests/test_no_hardcoded_cjk.py, under the text-processing group and by
the procedure that file documents: the range words are engine input, not
user-facing UI strings.

Tests: 7 new change-cases and 8 new leave-unchanged cases (hyphen, date,
phone number, product code, decimals, a non-numeric tilde, an unverified
language, and no language at all). All 7 change-cases fail against the
previous implementation.

Full suites before and after: the same 17 failures, none of them touched by
this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 17:06:15 +09:00
flutterkage2kandClaude Opus 5 847ca6d44c fix(desktop): resolve plugin-dialog wherever the package manager put it
`bun run desktop` dies before the window opens on a fresh clone:

    Error: ENOENT: no such file or directory, open
    '.../frontend/node_modules/@tauri-apps/plugin-dialog/dist-js/index.js'

The alias hardcoded `frontend/node_modules/...`, but this is a bun
workspace: bun hoists the package to the workspace root and leaves
`frontend/node_modules` empty, so the path the alias names does not
exist. Vite's dep optimizer reads it directly and throws, taking
`beforeDevCommand` — and the whole desktop shell — down with it.

Probe both layouts and fall through to Vite's own resolution when
neither is present, so a missing package degrades to normal resolution
instead of crashing the dev server.

Verified on macOS 26.6 (Apple Silicon), bun 1.2.22, fresh clone: the
window now opens and the backend serves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 17:06:12 +09:00
flutterkage2kandClaude Opus 5 784494dc3c fix(gallery): measure spectral flatness per frame so real speech passes
Most gallery previews fail with "the voice engine returned no audible
audio for this archetype". The renders are fine — the guard is not.

Two problems, both in the degenerate-buzz check:

1. `_spectral_flatness` took ONE FFT of the whole clip. Spectral
   flatness is defined over short frames; a full-length transform gets
   finer frequency resolution the longer the clip is, so voiced
   harmonics carve deeper and deeper nulls and the geometric mean
   collapses. The number tracked clip length, not timbre.

2. `_DEGENERATE_FLATNESS = 0.015` was calibrated against
   `_speech_like()` in the unit test — a synthetic harmonics+noise
   stand-in that is far flatter than real speech. Real renders measure
   well below it, so the threshold sat inside the speech range.

Measured on this engine's own output (framed, per this patch):

    pure tone 80 Hz        2.6e-10    two-tone buzz    3.3e-09
    quietest real speech   2.0e-04    (VoxCPM2 ko)

Frame the measurement (1024/512, skipping inter-word frames at the
noise floor) and move the threshold to 1e-5 — ~3000x above the tonal
cases, ~20x below the quietest real render.

Before: 6 of 8 renders rejected; ml_japanese_explainer,
ml_japanese_companion and feat_23_the_explainer all 503 through
GET /archetypes/{id}/preview.
After: 0 false positives across 27 real clips (Japanese, Korean and
English archetypes, cloned voices, human reference recordings), and
those three previews return 200. Every accepted clip was confirmed as
real speech by transcribing it with the app's own ASR.

Not addressed: a render that collapses toward NOISE rather than a tone
still passes (one observed at flatness 0.073, ASR returns a
hallucination). The old threshold missed it too, so this is not a
regression — calibrating an upper bound needs more than one sample.

Tests: frame-based measurement must be clip-length invariant, and the
threshold must sit between the measured tonal ceiling and the measured
real-speech floor. Both fail against the previous implementation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 17:05:59 +09:00
dajiaohuang 6c42b50fe3 test(ci): document cp1252 help regression 2026-09-05 05:21:45 +08:00
dajiaohuang 2594abebf8 fix(ci): keep install-docs help cp1252-safe 2026-09-05 05:17:34 +08:00
Palash DebnathandClaude Opus 5 e4f00bc564 fix(desktop): let Retry preempt the readiness wait
Greptile's P1 on #1809, and it is right. `launch_backend_and_wait` holds
`BackendState::lifecycle` around the entire launch, including the readiness
wait — which this branch just made unbounded for as long as the backend
answers `/startup/progress`. Retry, Clean & Retry, reset and uninstall all
need that same lock, so on a slow start the user's own escape hatch would
block behind the wait instead of interrupting it: an app with no way out,
which is worse than the early kill the branch set out to remove.

Every flow that is about to take lifecycle ownership now bumps a generation
counter first, before reaching for the lock. The waiting loop snapshots that
counter once its caller holds ownership — so a bump that predates it is not
mistaken for a preemption — and stands down within one 500 ms poll when it
changes, releasing the lock for whoever asked.

That also settles what happens at the splash's six-minute stall budget: it
flips to failed and offers Retry and the logs, and Retry now actually works,
while its /health recovery poll still walks straight into the app if the slow
start finishes first. Either way the user gets out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HR6J9zKQop9TGGVUwjypnF
2026-09-04 18:15:30 +05:30
Palash DebnathandClaude Opus 5 6a2ada2f11 fix(tts): don't answer a GPU OOM by repeating the same allocation
`_get_clone_prompt` catches everything and returns None so synthesis falls
back to `generate()`'s inline reference path. For a device OOM that is not a
fallback at all: the inline path runs the SAME encode on the SAME device —
producing identical output is the entire point of the precompute — so it is
guaranteed to hit the same wall moments later, on a GPU with even less
headroom than the first attempt found. Two reporters' backends died with a
Windows access violation (exit code -1073741819) seconds after this fallback
logged, mid-generation, on a card that had just refused an 86 MiB
allocation.

An OOM here is also the most recoverable kind. The allocator is typically
sitting on reserved-but-unallocated blocks — #1790's own log reports 90 MiB
reserved against that 86 MiB request — so drop them and try once more. If it
still will not fit, raise: the failure layer turns a device OOM into "close
other GPU-heavy apps or unload models, then retry", which is a far better
answer than walking into a native fault.

Every other failure still falls back silently, since for a non-memory fault
the inline path may genuinely succeed.

Fixes #1790.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HR6J9zKQop9TGGVUwjypnF
2026-09-04 18:11:20 +05:30
Palash DebnathandClaude Opus 5 9ac1045917 test: reject a backslash in a tracked path too
CodeRabbit's catch on #1799: git stores paths with `/` separators, so a `\`
that survives into a path component is part of a NAME. It is legal to commit
one from Linux or macOS and impossible to check out on Windows, where git
refuses it under `core.protectNTFS` — the same checkout-time failure, before
any test runs, that the stray `:memory:.ses` caused.

The rule moves into a pure `windows_hostile_reason` so it can be exercised
directly: the repo cannot carry a fixture for each hostile shape without
becoming the very thing the test rejects. Both directions are pinned — every
shape Windows refuses, and ordinary paths that merely resemble one (a file
called `console.md`, `com10.py`, a component containing but not ending in a
dot).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HR6J9zKQop9TGGVUwjypnF
2026-09-04 18:04:38 +05:30
Palash DebnathandClaude Opus 5 31d90def83 fix(errors): don't claim the backend crashed with no evidence that it did
Two Apple Silicon reporters were told "it most likely crashed or was killed
mid-request" while generating. Neither bug report carried a crash marker,
because none had been recorded — the app had no evidence for the one thing
it asserted, and the advice that follows that sentence is Retry and Clean &
Retry, which rebuilds the whole Python environment to fix a backend that had
not died.

Two causes, both fixed here.

The desktop shell learns the backend died from a ~2 s poll: it has to notice
the child exit before it can write the marker. `apiFetch` asked for that
marker exactly once, at the instant the transport gave up, so it raced the
poll and lost either way round — a backend that really died was reported
with the vague sentence instead of its exit code and crash notice, and one
that never died was reported as dead anyway. `streamDropError` already waits
that poll out (#1119); the request path never did. The loop is now a shared
`awaitBackendCrashMarker`, used by both, with a shorter budget here because
the transport cascade has already cost the user a few seconds.

And the copy itself overshot what it could know. By construction it is
reached only once a crash has been looked for and not found, so it no longer
names one: it says the backend stopped answering with no crash recorded, and
that a heavy job holding the engine is the likelier story — which on a
memory-pressured Mac mid-generation it is. Updated in all 21 locales, since
a translation still asserting a crash would be the same bug in another
language.

The #1337 test that required the crash wording is updated with it: #1337
established that the backend had answered seconds earlier, not what silenced
it, and requiring the stronger claim is what pinned this in place.

Fixes #1802.
Fixes #1805.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HR6J9zKQop9TGGVUwjypnF
2026-09-04 18:03:23 +05:30
Palash DebnathandClaude Opus 5 d5c33eda27 fix(desktop): don't kill a backend that is still starting
The launcher waited a flat five minutes from spawn for the backend to
report ready, then killed it and tried again. On a host where the cold
start genuinely takes longer — the reporter's project lived on a mapped
network drive, and `import torch` off one is slow the first time, as is a
first CUDA load or a cold spinning disk — that deadline expired *while the
backend was still importing*. The respawn threw away the warm page cache
and raced the same clock, so the app could never start, and it blamed the
backend: "the backend never reported ready". Launching that same backend by
hand reached ready in well under a minute once the cache was warm.

A backend answering `/startup/progress` with `status: "starting"` is not
one we have to guess about: it bound its socket, it is serving HTTP, and it
is naming the step it is on. Killing it cannot make the retry faster, and
the launcher knows nothing the user doesn't. So keep waiting while it
answers, and keep narrating each step. The budget still governs silence —
nothing answering, or a self-reported `failed` — where a slow backend and a
wedged one really are indistinguishable and the existing stderr-tail
failure is the right answer.

The splash needed the same correction. Its stall watchdog keys on
`bootstrap_status`, which sits on `starting_backend` for the whole of a slow
start, so it would have called the launch stuck at six minutes anyway; the
proof of life arrives on the separate `bootstrap-log` stream. Output now
counts as activity, and a genuinely silent backend still trips the watchdog
so the info-less spinner of #879 stays fixed.

Fixes #1791.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HR6J9zKQop9TGGVUwjypnF
2026-09-04 17:50:23 +05:30
Palash DebnathandClaude Opus 5 36a38bdbb9 fix(ci): don't ship a tracked path Windows cannot check out
A stray sqlite session artifact named `:memory:.ses` was committed by
accident on this branch. Git on Windows rejects a path containing `:` with
`error: invalid path` and exits 128 during **checkout** — so both Windows
jobs went red before a single build or test step ran, pointing at a file
nobody had edited, while Linux and macOS stayed green.

Drop the file, ignore the `*.ses` artifact class, and add a guard that scans
the index on every platform for paths Windows cannot represent: illegal
characters, components ending in a space or dot, and reserved DOS device
names. The failure now surfaces as a named test on every runner instead of
as a checkout crash on one leg of the matrix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HR6J9zKQop9TGGVUwjypnF
2026-09-04 17:41:02 +05:30
VishvakR 02c3a651ac test(generate): drive the scheduler itself, and restore env before reloading
Third review round on the PR, both findings in the new test file.

CodeRabbit: the disconnect regression called deadlines.for_task directly, so it
would have passed even if Scheduler._budget_for stopped coercing a missing
worker to the CPU budget -- the very thing it exists to pin. It now builds a
real WorkerPool and Scheduler, assigns the task to the 4 GB worker, asserts the
bound budget is the CPU one, disconnects the worker and asserts the
recomputation is not shorter. Forcing under_provisioned=False in _budget_for
fails it with `assert 300 == 600`.

CodeRabbit: the two env-var tests deleted the variables and reloaded
model_manager inside a finally, which runs BEFORE pytest restores them -- so on
a machine that already exports either var, the module constants would describe
an environment pytest was about to put back, and every later test would read
the mismatch. Both use monkeypatch.context() now, so the environment is restored
before the reload.
2026-09-04 15:36:18 +05:30
VishvakR fa148ecb55 fix(generate): tighten the VRAM-floor tests and document budget precedence
Second review round on the PR.

CodeRabbit: the repo-wide dispatch assertion accepted any nested min_vram_gb
keyword, so a budget computed with 0 or another engine's floor would pass while
the guard used the right one. It now compares the two expressions.

CodeRabbit: the awaiting-side deadline test restated gpu_gateway's formula
instead of calling it, so it would not have noticed that function starting to
select a shorter ceiling. It calls _default_deadline now, on cuda and rocm.

CodeRabbit: the docs said an explicit OMNIVOICE_GENERATE_TIMEOUT_S is honoured
"everywhere" while also saying the CPU var governs under-provisioned cards --
the two cannot both be true. Verified against the code (both vars set, 4 GB
cuda, engine floor 6 GB -> 200s, the accelerated value) and documented as a
precedence table rather than prose. The accelerated var deliberately wins on an
under-provisioned host: that is what keeps "lower it to fail fast everywhere"
working. Pinned by a test so the table cannot drift from the behaviour.

CodeRabbit also flagged that Scheduler._budget_for recomputes with no worker
after a disconnect, dropping under_provisioned to False. That cannot shorten
anything: no worker means no execution_device, which _base_execution_seconds
already coerces to "cpu" -- the same budget the floor raises an
under-provisioned card to. Added a test pinning that rather than persisting a
dispatch-time budget on the attempt. The residual case it describes -- an
operator who raised the accelerated budget ABOVE the CPU one sees a shorter
recomputation once the worker is gone -- predates this change and applies to
every GPU worker, not just under-provisioned ones, so it belongs in its own fix.
2026-09-04 13:52:23 +05:30
VishvakR f172d0c3be fix(generate): apply the VRAM-floor budget to remote workers and /convert
Review findings on the PR, fixed here rather than left for a fourth report.

Greptile (P1): the control plane sets a remote attempt's deadline, so the same
inversion reached remote workers. Its suggested fix -- thread the engine floor
into generate_timeout_s() -- would read the wrong machine: that function probes
THIS host, so a Mac control plane dispatching to a 4 GB Windows worker learns
nothing (MPS is excluded by design), and a 4 GB box dispatching to a 24 GB
worker would wrongly get the longer budget. The worker already advertises both
figures it takes -- free_memory_bytes and min_memory_bytes, both set in
worker/capabilities.py -- so ConnectedWorker.under_provisioned() decides from
those, and deadlines.for_task() floors the execution budget at what the same job
would get on a CPU. The task-level ceiling in gpu_gateway._default_deadline is
computed before a worker is bound and already asks for the CPU budget, so it
still covers the raised lease; a test pins that.

CodeRabbit (major): /convert had the identical split -- min_vram_gb to the
guard so a timeout could name the card, and a budget computed without it.

CodeRabbit (minor): the docs promised the CPU-class floor for any GPU, while
the code scopes it to dedicated-VRAM families. Reworded to say CUDA/ROCm and to
say why MPS is excluded.

CodeRabbit (minor): the call-site assertion compared global occurrence counts,
so one dispatch could drop both arguments while another gained an extra and the
total still matched. It now walks the AST and checks each dispatch on its own,
and the pairing is additionally enforced repo-wide across backend/api/routers:
a dispatch that knows the engine's floor well enough to explain a timeout must
know it well enough to set the budget.

Three inline capability-selection loops in ConnectedWorker collapse into one
_capability_for(), so the new predicate cannot select a different capability
than execution_device() does.
2026-09-04 13:33:00 +05:30
VishvakR fcac8e1bae fix(generate): budget an under-provisioned GPU like the CPU it performs like
A GPU with less VRAM than the engine declares it needs pages to system RAM
over PCIe, so it renders slower than the same machine's CPU. The compute-time
budget picked its value from the device family alone, so that card was treated
as fast hardware and given 300s -- half the 600s a plain CPU host gets. It is
the slowest configuration the app supports and it had the shortest watchdog.

Everything else already acted on the verdict. resolve_routing() raises the
caveat, the synth preflight warns before the user waits, and _timeout_guidance()
names the card in the failure. Each TTS generate dispatch even hands the guard
the engine's floor on the line above the timeout that ignored it. #1226 and
#1222 were the same 4 GB cards on the same engine; both were closed by making
the app explain the timeout better, never by correcting the budget behind it.

generate_timeout_s() now floors an under-provisioned accelerator at the CPU
budget. The length scaling is unchanged, and an explicitly configured
OMNIVOICE_GENERATE_TIMEOUT_S is still honoured verbatim, so an operator who
lowered the watchdog to fail fast keeps that. The floor is a max(), never an
assignment, so a raised accelerated budget is never cut down. Engines that
declare no floor, a failed VRAM probe, and MPS (whose vram_gb is a unified-
memory heuristic, not a dedicated pool) are all untouched.

The three-clause "is this host under-provisioned" test was written out inline
in the caveat and in the timeout message, which is how the budget came to
disagree with the warning printed beside it; it is now one predicate,
under_provisioned_vram(), that all three read.

Reported on a GTX 1650 (4 GB) running the omnivoice engine, whose breadcrumbs
show the budget ending the job on the dot: 372s and 301s are exactly
300 + max(0, len - 1200) / 40 for the two takes.

Fixes #1804.
2026-09-04 13:02:14 +05:30
Palash DebnathandClaude Opus 5 4e6c36848c fix(transcriptions): render segments that have no timings
An OpenAI-compatible ASR answering in json/text format returns no
timestamps, and services/asr_backend.py records that honestly as
`end: None` rather than inventing a number. The segment list called
`.toFixed()` on it unconditionally, so the render threw and the whole
Transcriptions view went blank — a transcript that merely lacked timings
became one the user could not read at all.

Show whichever bound is known and nothing when neither is, so the text
stays readable either way. Non-finite values are treated as unknown too,
so a bad timing prints nothing rather than NaN.

Fixes #1798.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HR6J9zKQop9TGGVUwjypnF
2026-09-04 03:26:51 +05:30
1436 changed files with 204930 additions and 14952 deletions
+29
View File
@@ -0,0 +1,29 @@
# VoiceStudio compatibility entry
The current cross-agent package is [voicestudio](../../../../skills/voicestudio/SKILL.md).
For new installations use `npx skills add debpalash/VoiceStudio --skill voicestudio`.
Use the running backend at the user's configured address (default
`http://localhost:3900`). Check `/health`, discover `/openapi.json` and
`/v1/audio/voices`, then use the installed schema for speech, transcription,
profiles, and jobs. The HTTP MCP endpoint is `/mcp`; discover tools from the
connected server instead of assuming this older package's tool inventory.
Launch the installed Electron app if the backend is unavailable. For source
development follow the checkout's Electron README. Existing helpers in
`scripts/` support legacy source installations; inspect their environment
and dependency assumptions before running them.
Model downloads and remote services require the user's choice. Never silently
install models, promise fixed latency, or treat compatibility voice names as
real provider voices. Validate saved audio and asynchronous job completion
before reporting success. Protected backends require configured credentials;
never disable authentication to make an example work.
Source and current setup documentation:
https://github.com/debpalash/VoiceStudio
This archived entry is not an installable skill. Existing installations should
remove the old `omnivoice` / `oss-maintainer` entries and install `voicestudio` /
`voicestudio-maintainer` from the canonical repository. Legacy helpers remain
for existing users; the Electron supervisor is the preferred launcher.
-172
View File
@@ -1,172 +0,0 @@
---
name: omnivoice
description: "Local TTS, voice cloning, voice design, and video dubbing via the VoiceStudio MCP server (open-source ElevenLabs alternative; nothing leaves the machine, runs on MPS/CUDA/CPU). Use when: (1) generating speech from text in any of 646 languages, (2) cloning a voice from a 3-second reference clip, (3) designing a voice by gender/age/accent/pitch/style, (4) dubbing a video into another language, (5) listing voice profiles or personality presets, (6) producing narration where privacy, cost, or absent API keys matter, (7) non-English narration where Edge TTS/kokoro fall short, (8) batch audio for blog posts or content pipelines. Triggers: 'omnivoice', 'voice clone', 'clone this voice', 'tts', 'narrate', 'generate speech', 'voice synthesis', 'dub video', 'voice design', 'local tts', 'multilingual voice', 'narrate this post', 'elevenlabs alternative'."
---
# VoiceStudio
The canonical cross-agent package lives at `skills/omnivoice/SKILL.md`. This
Claude-specific package retains the MCP lifecycle helpers and references.
## Overview
Generate audio locally via the VoiceStudio MCP server. Tools: `generate_speech`, `list_voices`, `list_personalities`, `list_languages`, `check_health`. Resources: `voice://{id}`, `history://recent`.
## Prerequisites — Backend Must Be Running
The MCP tools all hit `$OMNIVOICE_API_URL` (default `http://localhost:3900`). If the backend is down, every tool returns a connection error. Install + boot:
```bash
git clone https://github.com/debpalash/VoiceStudio.git "$OMNIVOICE_HOME"
cd "$OMNIVOICE_HOME"
uv sync
VIRTUAL_ENV="$(pwd)/.venv" uv pip install 'mcp[cli]'
```
Then:
```bash
scripts/check-health.sh # exit 0 if up
scripts/start-backend.sh # boot in background (MPS/CUDA auto-detected)
```
First synthesis call lazy-downloads the `k2-fsa/OmniVoice` model (~2.4 GB) from HuggingFace — cached on subsequent boots.
## Task Index — Pick the Right Tool
| Task | Tool | Notes |
|---|---|---|
| Verify backend is up | `check_health` | Returns `{"status":"ok","device":"mps|cuda|cpu"}` |
| Text → audio with a saved voice | `generate_speech(text, profile_id)` | Returns base64 WAV. `profile_id="demo0001"` is the bundled demo voice |
| Text → audio without a clone (voice design) | `generate_speech(text, instruct="…")` | Omit `profile_id`; pass an `instruct` like `"warm middle-aged female narrator, calm pace"` |
| Multilingual narration | `generate_speech(text, language="es")` | Any ISO 639 code or `"Auto"` |
| List existing voices | `list_voices` | Returns id, name, type, personality |
| List personality presets | `list_personalities` | Returns narrator / casual / news-anchor / etc. with their `instruct` strings |
| List supported languages | `list_languages` | 646 total; returns 20 popular + the full count |
For non-trivial decisions (which engine to use, when to pick VoiceStudio over kokoro / Edge TTS / ElevenLabs), see [references/engines-comparison.md](references/engines-comparison.md).
For MCP wiring details, backend lifecycle, troubleshooting, and a clean teardown, see [references/mcp-setup.md](references/mcp-setup.md).
## Common Workflows
### 1. One-shot narration with the demo voice
```python
# As called through the MCP client (your agent will do this for you):
result = generate_speech(
text="Hello — this is VoiceStudio generating speech locally.",
profile_id="demo0001",
language="English",
steps=16, # 8 = fast/draft · 16 = balanced · 32 = quality
)
# result is JSON with audio_id, generation_time_s, audio_duration_s, format, wav_base64
```
Benchmark: 4.2 s of audio in ~24 s server-side on Apple Silicon MPS at 16 diffusion steps.
### 2. Save the WAV to disk and play
Tool returns base64 PCM WAV (16-bit, mono, 24 kHz). Decode + write:
```python
import base64, json
payload = json.loads(result_text) # parse JSON the tool returns
open("out.wav","wb").write(base64.b64decode(payload["wav_base64"]))
```
On macOS: `afplay out.wav`. Convert to MP3 with `ffmpeg -i out.wav -codec:a libmp3lame -b:a 128k out.mp3`.
### 3. Voice clone — end-to-end recipe
Cloning needs a 3-10 second reference clip the model will use as a speaker embedding. The MCP server does NOT expose profile creation — it only reads existing profiles. Two paths to create one:
**Path A — bundled helper (macOS, recommended for fresh clones):**
```bash
scripts/record-reference.sh ~/Downloads/my-ref.wav 12 1
# args: output_path raw_duration_sec mic_index
# Default mic_index=1 (MacBook built-in); list devices via:
# ffmpeg -f avfoundation -list_devices true -i ""
```
The script gives **audible** countdown + start/stop cues via macOS `say` + `/System/Library/Sounds/Ping.aiff` so the user knows when to speak (terminal stdout is buffered — text "speak now" prompts arrive too late). It records a longer raw window, then trims to ~10 seconds of speech via `silenceremove + atrim`, plays back for verification, and prints the next-step `curl` command.
**Path B — manual:**
```bash
# 1. Record (mono, 24 kHz native — matches model's internal rate)
ffmpeg -f avfoundation -i ":1" -t 12 -ac 1 -ar 24000 raw.wav
# 2. Trim leading silence + take first 10 sec of speech
ffmpeg -i raw.wav \
-af "silenceremove=start_periods=1:start_silence=0.05:start_threshold=-40dB,atrim=end=10" \
-ac 1 -ar 24000 ref.wav
# 3. Verify
ffmpeg -i ref.wav -af volumedetect -f null - 2>&1 | grep volume # max should be > -20 dB
afplay ref.wav
```
**POST to /profiles** (multipart/form-data — required fields: `name`, `ref_audio`):
```bash
curl -X POST http://127.0.0.1:3900/profiles \
-F "name=carlos-clone" \
-F "ref_audio=@ref.wav" \
-F "ref_text=The exact text spoken in the clip" \
-F "language=English" \
| python3 -m json.tool
# returns { "id": "abc12345", "name": "carlos-clone" }
```
Once created, pass `profile_id` to `generate_speech` (via MCP) or directly via `POST /generate`. Profiles persist in SQLite + reference-audio files at `~/Library/Application Support/OmniVoice/voices/<id>.<ext>` (the backend preserves the uploaded extension — `.wav` if you uploaded a WAV, `.mp3` if MP3, etc.). State persists across backend restarts.
**Reference clip tips that materially affect quality:**
| Factor | Why it matters |
|---|---|
| Single speaker | Mixed speakers blur the embedding |
| Clean speech, no music/noise | Model embeds the noise too |
| Natural prosody (avoid pangrams) | Diffusion samples replicate prosody, not just timbre |
| 3-10 sec is the sweet spot | < 3 s lacks information; > 10 s adds compute without quality gain |
| Match `ref_text` to what's spoken | Improves alignment, especially on noisy refs |
| `language` correct | Wrong language → cross-lingual transfer artifacts |
| Loudness peak ≥ -15 dB | Quiet refs work but normalize poorly |
### 4. Voice design (no reference clip)
Skip `profile_id`; provide an `instruct` string describing the desired voice:
```python
generate_speech(
text="Welcome to the future of agentic systems.",
instruct="warm middle-aged female narrator, calm authoritative pace, documentary style",
)
```
Get pre-made instructs via `list_personalities` and copy the one matching the brief (narrator, casual, news-anchor, etc.).
### 5. Video dubbing (web UI only)
The MCP server does not expose the dubbing endpoint. The full transcribe → translate → re-voice → mux pipeline lives behind the desktop UI (`bun run desktop` in `$OMNIVOICE_HOME`) and the `/dub/*` REST routes. When the user asks to dub a video, point them to the UI; surface this skill only for the synthesis primitives above.
## When NOT to use VoiceStudio
- **Fast English-only narration on weak hardware** → `kokoro-tts` is ~10× smaller and 2× realtime on CPU (see [references/engines-comparison.md](references/engines-comparison.md))
- **Lowest-friction one-off TTS** → Edge TTS needs no install or backend
- **Highest possible quality regardless of cost** → ElevenLabs still wins on English narration polish; VoiceStudio ties or wins on multilingual + cloning
- **Real-time streaming dictation** → use the VoiceStudio desktop widget (`⌘+⇧+Space`), not the MCP server
## Resources
- [references/engines-comparison.md](references/engines-comparison.md) — Decision tree across VoiceStudio / kokoro / Voicebox / Edge TTS / ElevenLabs / cloud APIs
- [references/mcp-setup.md](references/mcp-setup.md) — MCP wiring, backend lifecycle, env vars, troubleshooting
- [scripts/check-health.sh](scripts/check-health.sh) — `curl /health`, exit 0/1
- [scripts/start-backend.sh](scripts/start-backend.sh) — Start uvicorn on 127.0.0.1:3900 with health probe
- [scripts/stop-backend.sh](scripts/stop-backend.sh) — Clean shutdown via `kill -TERM` on the bound PID
- [scripts/record-reference.sh](scripts/record-reference.sh) — macOS-only: record + trim + verify a reference clip for cloning, with audible cues (`say` + system beeps) that bypass terminal output buffering
Backend Swagger / OpenAPI: `http://127.0.0.1:3900/docs` (when backend is up).
Upstream: github.com/debpalash/VoiceStudio. The app uses AGPL-3.0-only; optional engines and downloaded models retain their own licenses. See `LICENSE-NOTICE.md` in the repository.
+26 -10
View File
@@ -56,7 +56,17 @@ bun install
bun run dev
```
This starts both services:
This launches Electron with hot reload. Its runtime supervisor manages backend setup
and startup; do not launch a second backend. See [Electron setup](../electron/README.md).
```bash
bun run build # build Electron
bun run start # launch the built Electron app
bun run dist # package locally without publishing
bun run dev:web # legacy browser UI + backend
```
The legacy browser command starts both services:
| Service | URL | What it does |
|---------|-----|---|
@@ -71,28 +81,34 @@ cause doesn't scroll away with the terminal. The same death is also reported
as a crash notice in the UI the next time the backend starts (see
[docs/install/troubleshooting.md §14c](docs/install/troubleshooting.md)).
### Desktop App (Tauri)
### Legacy Desktop App (Tauri)
```bash
bun run desktop # dev: hot-reload Tauri shell + backend
bun run desktop-prod # production: builds, bundles the backend, then launches
bun run tauri # legacy dev: hot-reload Tauri shell + backend
bun run tauri:desktop-prod # legacy production: builds, bundles the backend, then launches
```
Both run `uv sync` first (so the Python backend env is set up) and start the
backend automatically — you do **not** start it separately. Use the exact script
names: there is no `desktop=prod` (note the **hyphen** in `desktop-prod`).
`desktop-prod` is Windows-aware (auto-detects bash/git; see `scripts/desktop-prod.mjs`).
names: there is no `desktop=prod` (note the **hyphen** in `tauri:desktop-prod`).
`tauri:desktop-prod` is Windows-aware (auto-detects bash/git; see `scripts/desktop-prod.mjs`).
Requires [Rust](https://rustup.rs/) and platform-specific Tauri dependencies — see the [Tauri prerequisites](https://v2.tauri.app/start/prerequisites/).
After installing Rust with rustup on macOS/Linux, either open a new terminal or
load Cargo into the current one before starting the desktop app:
After installing Rust with rustup (or `uv` with its installer), a terminal that
was already open still has the old `PATH`. The desktop launchers (`bun tauri`,
`bun tauri:desktop-prod`, `bun tauri:desktop-fresh`) detect this and add `~/.cargo/bin` /
`~/.local/bin` for that run, printing a one-line note; to make it permanent,
open a new terminal, or on macOS/Linux load Cargo into the current one:
```bash
source "$HOME/.cargo/env"
bun desktop
bun run tauri
```
If Rust is genuinely not installed, the launchers stop up front with the
install command instead of failing later inside `cargo metadata`.
On Linux, errors such as `Package gdk-3.0 was not found`, `pango.pc` missing,
or `javascriptcoregtk-4.1` missing mean the native packages above were not
installed; changing `PKG_CONFIG_PATH` does not fix libraries that are absent.
@@ -304,7 +320,7 @@ that — the agent recalls the architecture, conventions, and your past findings
instead of re-reading the tree each time. [**memxt**](https://github.com/debpalash/memxt)
(100% local, MCP-based, built by this project's maintainer) exists for exactly
this; any MCP memory server works. Pair it with the repo's agent skill —
`npx skills add debpalash/omnivoice-studio` — so your agent knows the project's
`npx skills add debpalash/VoiceStudio` — so your agent knows the project's
hard rules from the first prompt.
## Quality gates your PR must pass
+89 -2
View File
@@ -11,6 +11,11 @@ on:
push:
branches: [main]
workflow_dispatch:
inputs:
windows_wix_diagnostic:
description: Run only the tiny nonpublishing Windows MSI authoring diagnostic
type: boolean
default: false
permissions:
contents: read
@@ -21,6 +26,7 @@ env:
jobs:
test:
if: ${{ !inputs.windows_wix_diagnostic }}
name: Tests (backend + frontend)
runs-on: ubuntu-22.04
env:
@@ -166,6 +172,12 @@ jobs:
working-directory: frontend
run: node --experimental-strip-types --no-warnings --test ../tests/frontend/*.test.mjs
# Electron used to be built only after a release started, so renderer,
# preload and packaging regressions could pass the required PR gate.
# Keep this command shared with release.yml through the root script.
- name: Electron typecheck, tests and production contract
run: bun run check:electron
# Production-bundle blank-screen gate. Everything above runs UN-minified
# (dev server + Vitest/jsdom), so a crash that exists ONLY in the minified
# release bundle — a TDZ reorder that throws before React mounts — passes
@@ -180,6 +192,28 @@ jobs:
- name: Production-bundle smoke — no blank screen
working-directory: frontend
run: bun run test:prod-bundle
- name: Electron renderer workflow smokes
shell: bash
run: |
set -euo pipefail
export OMNIVOICE_PORT=3999
export VOICESTUDIO_UI_URL=http://localhost:3912
export PLAYWRIGHT_CHANNEL=chromium
bun run --cwd electron smoke:server > /tmp/voicestudio-electron-smoke.log 2>&1 &
server_pid=$!
trap 'kill "$server_pid" 2>/dev/null || true' EXIT
for _ in {1..60}; do
if curl --fail --silent --show-error "$VOICESTUDIO_UI_URL" >/dev/null; then
break
fi
sleep 0.25
done
curl --fail --silent --show-error "$VOICESTUDIO_UI_URL" >/dev/null || {
cat /tmp/voicestudio-electron-smoke.log
exit 1
}
node electron/tests/playback-smoke.mjs
node electron/tests/dub-smoke.mjs
# ── Cross-platform Tauri shell check ────────────────────────────────────
# Catches platform-specific Rust regressions on PR (cfg(target_os=...)
@@ -189,6 +223,7 @@ jobs:
# and `cargo test --lib` runs the shell's unit tests natively on each OS.
# Full bundling stays in release.yml on tag push.
tauri-cross-platform:
if: ${{ !inputs.windows_wix_diagnostic }}
name: Tauri shell check (${{ matrix.label }})
needs: test
strategy:
@@ -289,6 +324,7 @@ jobs:
# job above misses. Narrow scope (tests/smoke/ only) — full pytest stays
# on Linux until Phase 1's INST-01 lands setuptools for WhisperX.
smoke-matrix:
if: ${{ !inputs.windows_wix_diagnostic }}
name: Smoke (${{ matrix.label }})
needs: test
strategy:
@@ -428,17 +464,68 @@ jobs:
PY
- name: Run smoke tests
# Exercise credential paths on native Windows as well as POSIX hosts.
if: matrix.backend_supported
run: uv run --no-sync pytest tests/smoke/ -q --tb=short
run: uv run --no-sync pytest tests/smoke/ tests/test_hf_token_cache_paths.py -q --tb=short
env:
HF_HUB_OFFLINE: "1" # same no-silent-downloads guard as the main pytest job
HF_HUB_CACHE: ${{ runner.temp }}/pockettts-empty-hf-cache
# The isolated backend session, on Windows. The `test` job runs it on
# Linux only, which is how four tests that CANNOT pass on Windows shipped
# unnoticed: two reach for os.WNOHANG and os.waitid (POSIX-only, an
# AttributeError before the first assertion), one asserts a RuntimeError
# that `backend_drain_fd` returns None instead of raising off POSIX, and
# one raced the OS reaping a crashed child — a race Linux won and Windows
# lost every time. All four were invisible to CI and hit every Windows
# contributor on their first `pytest` run. Forty seconds closes the class.
- name: Isolated backend session (Windows)
if: runner.os == 'Windows' && matrix.backend_supported
run: uv run --no-sync pytest backend/tests/ -q --tb=short
env:
HF_HUB_OFFLINE: "1"
# Artifact commits depend on native Windows rename/replace semantics;
# Linux emulation cannot exercise sharing rules or path parsing.
# test_worker_task_store and test_worker_inbound_transport joined this
# step after a Windows run found a real portability bug the Linux-only
# `test` job could not see: a staged input's artifact id was built with
# os.path.join, so a Windows control plane persisted and shipped
# `inputs\<sha>.wav` — which a Linux worker cannot resolve. These suites
# need no ffmpeg, so they cost seconds here.
- name: Remote-worker artifact paths (Windows)
if: runner.os == 'Windows' && matrix.backend_supported
run: uv run --no-sync pytest tests/test_worker_upload_server.py tests/test_worker_server_integrity.py -q --tb=short
run: >-
uv run --no-sync pytest
tests/test_worker_upload_server.py
tests/test_worker_server_integrity.py
tests/test_worker_task_store.py
tests/test_worker_inbound_transport.py
-q --tb=short
env:
HF_HUB_OFFLINE: "1"
HF_HUB_CACHE: ${{ runner.temp }}/worker-artifact-empty-hf-cache
windows-wix-diagnostic:
name: Windows MSI authoring (no publishing)
needs: test
if: ${{ !cancelled() && (inputs.windows_wix_diagnostic || needs.test.result == 'success') }}
runs-on: windows-2022
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
- uses: oven-sh/setup-bun@v1
- name: Bundle canonical system and per-user templates with a tiny payload
shell: pwsh
run: ./scripts/diagnose-windows-wix.ps1
- name: Restore hosted Installer policy after failed standard-user installation
shell: powershell
run: ./scripts/test-msi-policy-cleanup.ps1
- name: Preserve verbose linker output and rendered authoring
if: always()
uses: actions/upload-artifact@v4
with:
name: windows-wix-diagnostic
path: wix-diagnostic-artifacts/
if-no-files-found: warn
retention-days: 3
+115
View File
@@ -0,0 +1,115 @@
name: Electron packaging rehearsal
# Explicitly artifact-only: no tag, schedule, release, or publishing permission.
on:
workflow_dispatch:
permissions:
contents: read
concurrency:
group: electron-rehearsal-${{ github.ref }}
cancel-in-progress: true
jobs:
package:
runs-on: ${{ matrix.runner }}
timeout-minutes: 60
strategy:
fail-fast: false
matrix:
include:
- runner: ubuntu-24.04
platform: linux
arch: x64
target: x86_64-unknown-linux-gnu
flags: --linux --x64
- runner: windows-2022
platform: win32
arch: x64
target: x86_64-pc-windows-msvc
flags: --win --x64
- runner: macos-15
platform: darwin
arch: arm64
target: aarch64-apple-darwin
flags: --mac --arm64
- runner: macos-15-intel
platform: darwin
arch: x64
target: x86_64-apple-darwin
flags: --mac --x64
defaults:
run:
shell: bash
env:
VOICESTUDIO_RUST_TARGET: ${{ matrix.target }}
VOICESTUDIO_UPDATE_CHANNEL: electron-preview-${{ matrix.platform }}-${{ matrix.arch }}
CSC_IDENTITY_AUTO_DISCOVERY: 'false'
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22'
- uses: oven-sh/setup-bun@v2
with:
bun-version: '1.4.2'
- uses: dtolnay/rust-toolchain@stable
with:
targets: ${{ matrix.target }}
- uses: Swatinem/rust-cache@v2
with:
workspaces: native/desktop-bridge -> target
key: electron-${{ matrix.target }}
- uses: astral-sh/setup-uv@v6
with:
version: '0.12.13'
enable-cache: false
- name: Linux native dependencies
if: runner.os == 'Linux'
run: |
sudo apt-get update
sudo apt-get install -y libasound2-dev libxdo-dev libxtst-dev libx11-dev libxkbcommon-dev libwayland-dev libssl-dev pkg-config xvfb
- name: Bundle pinned uv for the host architecture
run: |
node --input-type=module <<'NODE'
import { execFileSync } from 'node:child_process';
import { mkdirSync, copyFileSync, chmodSync } from 'node:fs';
import { join } from 'node:path';
const expected = process.env.VOICESTUDIO_RUST_TARGET;
const targets = { 'linux-x64': 'x86_64-unknown-linux-gnu', 'win32-x64': 'x86_64-pc-windows-msvc', 'darwin-arm64': 'aarch64-apple-darwin', 'darwin-x64': 'x86_64-apple-darwin' };
if (targets[`${process.platform}-${process.arch}`] !== expected) throw new Error('Runner architecture does not match package target');
const source = execFileSync(process.platform === 'win32' ? 'where.exe' : 'which', ['uv'], { encoding: 'utf8' }).trim().split(/\r?\n/)[0];
const dir = 'frontend/src-tauri/binaries';
mkdirSync(dir, { recursive: true });
const destination = join(dir, `uv-${expected}${process.platform === 'win32' ? '.exe' : ''}`);
copyFileSync(source, destination);
if (process.platform !== 'win32') chmodSync(destination, 0o755);
NODE
- name: Install locked dependencies
run: bun install --frozen-lockfile
- name: Validate and build Electron
run: bun run check:electron
- name: Package without publishing
working-directory: electron
run: |
bun x electron-builder --config electron-builder.config.mjs ${{ matrix.flags }} --publish never
node tests/packaging-contract.mjs --artifact
node tests/update-package-contract.mjs --platform ${{ matrix.platform }} --arch ${{ matrix.arch }}
- name: Packaged startup smoke test
working-directory: electron
run: |
if [ "$RUNNER_OS" = Linux ]; then
xvfb-run -a node tests/packaged-smoke.mjs --setup
else
node tests/packaged-smoke.mjs --setup
fi
- name: Save installers and updater metadata for review
uses: actions/upload-artifact@v4
with:
name: electron-rehearsal-${{ matrix.platform }}-${{ matrix.arch }}
retention-days: 14
if-no-files-found: error
path: |
electron/release/VoiceStudio-Electron-*
electron/release/electron-*.yml
+218
View File
@@ -0,0 +1,218 @@
name: Electron desktop release
# Builds are safe by default. Only an explicit publish dispatch exposes a release.
on:
push:
tags: ['v*']
workflow_dispatch:
inputs:
publish:
description: "Publish the tagged Electron release after all platforms pass"
type: boolean
default: false
allow_unsigned:
description: "Explicitly accept unsigned/unnotarized Electron installers and documented updater limitations"
type: boolean
default: false
permissions:
contents: read
concurrency:
group: electron-release-${{ github.ref }}
cancel-in-progress: false
jobs:
validate:
# The transition tag is assembled after the manual Tauri draft succeeds.
if: github.event_name == 'workflow_dispatch' || github.ref_name != vars.TAURI_SUNSET_TAG
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Require an exact version tag
env:
REF: ${{ github.ref }}
ALLOW_UNSIGNED: ${{ inputs.allow_unsigned }}
DISPATCH_ACTOR: ${{ github.actor }}
RERUN_ACTOR: ${{ github.triggering_actor }}
OWNER: ${{ github.repository_owner }}
run: |
if [ "$ALLOW_UNSIGNED" = true ]; then
test "$DISPATCH_ACTOR" = "$OWNER" && test "$RERUN_ACTOR" = "$OWNER" || {
echo "Only the repository owner may accept unsigned installers"; exit 1;
}
fi
VERSION=$(node -p "require('./frontend/package.json').version")
test "$REF" = "refs/tags/v$VERSION" || { echo "Dispatch on the exact version tag"; exit 1; }
package:
needs: validate
runs-on: ${{ matrix.runner }}
timeout-minutes: 60
strategy:
fail-fast: false
matrix:
include:
- runner: ubuntu-24.04
platform: linux
arch: x64
target: x86_64-unknown-linux-gnu
flags: --linux --x64
- runner: windows-2022
platform: win32
arch: x64
target: x86_64-pc-windows-msvc
flags: --win --x64
- runner: macos-15
platform: darwin
arch: arm64
target: aarch64-apple-darwin
flags: --mac --arm64
- runner: macos-15-intel
platform: darwin
arch: x64
target: x86_64-apple-darwin
flags: --mac --x64
defaults:
run:
shell: bash
env:
VOICESTUDIO_RUST_TARGET: ${{ matrix.target }}
VOICESTUDIO_UPDATE_CHANNEL: electron-stable-${{ matrix.platform }}-${{ matrix.arch }}
CSC_IDENTITY_AUTO_DISCOVERY: 'false'
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22'
- uses: oven-sh/setup-bun@v2
with:
bun-version: '1.4.2'
- uses: dtolnay/rust-toolchain@6bed0761d98439e5a578e2877258200ad565ba87 # stable
with:
targets: ${{ matrix.target }}
- uses: Swatinem/rust-cache@v2
with:
workspaces: native/desktop-bridge -> target
key: electron-${{ matrix.target }}
- uses: astral-sh/setup-uv@v6
with:
version: '0.12.13'
enable-cache: false
- name: Linux native dependencies
if: runner.os == 'Linux'
run: |
sudo apt-get update
sudo apt-get install -y libasound2-dev libxdo-dev libxtst-dev libx11-dev libxkbcommon-dev libwayland-dev libssl-dev pkg-config xvfb
- name: Bundle pinned uv for the host architecture
run: |
node --input-type=module <<'NODE'
import { execFileSync } from 'node:child_process';
import { mkdirSync, copyFileSync, chmodSync } from 'node:fs';
import { join } from 'node:path';
const expected = process.env.VOICESTUDIO_RUST_TARGET;
const targets = { 'linux-x64': 'x86_64-unknown-linux-gnu', 'win32-x64': 'x86_64-pc-windows-msvc', 'darwin-arm64': 'aarch64-apple-darwin', 'darwin-x64': 'x86_64-apple-darwin' };
if (targets[`${process.platform}-${process.arch}`] !== expected) throw new Error('Runner architecture does not match package target');
const source = execFileSync(process.platform === 'win32' ? 'where.exe' : 'which', ['uv'], { encoding: 'utf8' }).trim().split(/\r?\n/)[0];
const dir = 'frontend/src-tauri/binaries';
mkdirSync(dir, { recursive: true });
const destination = join(dir, `uv-${expected}${process.platform === 'win32' ? '.exe' : ''}`);
copyFileSync(source, destination);
if (process.platform !== 'win32') chmodSync(destination, 0o755);
NODE
- name: Install locked dependencies
run: bun install --frozen-lockfile
- name: Validate and build Electron
run: bun run check:electron
- name: Package without publishing
env:
CSC_LINK: ${{ secrets.ELECTRON_CSC_LINK }}
CSC_KEY_PASSWORD: ${{ secrets.ELECTRON_CSC_KEY_PASSWORD }}
working-directory: electron
run: |
bun x electron-builder --config electron-builder.config.mjs ${{ matrix.flags }} --publish never
node tests/packaging-contract.mjs --artifact
node tests/update-package-contract.mjs --platform ${{ matrix.platform }} --arch ${{ matrix.arch }}
- name: Verify macOS signing and notarization before publication
if: inputs.publish == true && inputs.allow_unsigned != true && matrix.platform == 'darwin'
run: |
APP=$(find electron/release -maxdepth 2 -name VoiceStudio.app -type d -print -quit)
test -n "$APP"
codesign --verify --deep --strict "$APP"
spctl --assess --type execute --verbose=2 "$APP"
- name: Verify Windows installer signature before publication
if: inputs.publish == true && inputs.allow_unsigned != true && matrix.platform == 'win32'
shell: pwsh
run: |
$installers = @(Get-ChildItem electron/release/VoiceStudio-Electron-*.exe)
if ($installers.Count -eq 0) { throw "No installer to verify" }
foreach ($installer in $installers) {
$signature = Get-AuthenticodeSignature $installer.FullName
if ($signature.Status -ne 'Valid') { throw "Installer signature is not trusted: $($installer.Name)" }
}
- name: Packaged startup smoke test
working-directory: electron
run: |
if [ "$RUNNER_OS" = Linux ]; then
xvfb-run -a node tests/packaged-smoke.mjs --setup
else
node tests/packaged-smoke.mjs --setup
fi
- name: Save installers and updater metadata for review
uses: actions/upload-artifact@v4
with:
name: electron-release-${{ matrix.platform }}-${{ matrix.arch }}
retention-days: 14
if-no-files-found: error
path: |
electron/release/VoiceStudio-Electron-*
electron/release/electron-*.yml
release:
needs: package
runs-on: ubuntu-latest
permissions:
contents: write
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
TAG: ${{ github.ref_name }}
SUNSET_TAG: ${{ vars.TAURI_SUNSET_TAG }}
PUBLISH: ${{ inputs.publish }}
steps:
- uses: actions/checkout@v4
- uses: actions/download-artifact@v4
with:
pattern: electron-release-*
merge-multiple: true
path: release-assets
- name: Validate all platforms before creating a release
run: |
python3 scripts/prepare_electron_release.py --assets release-assets --tag "$TAG"
- name: Preserve the final Tauri updater feeds
run: |
test -n "$SUNSET_TAG" || { echo "Set TAURI_SUNSET_TAG before releasing"; exit 1; }
# The transition tag already holds its own final Tauri feeds.
# Later releases carry copies pointing to the immutable sunset payloads.
gh release download "$SUNSET_TAG" --pattern latest.json --dir release-assets
gh release download "$SUNSET_TAG" --pattern latest-user.json --dir release-assets
python3 scripts/prepare_electron_release.py --assets release-assets --tag "$TAG" --sunset-tag "$SUNSET_TAG"
- name: Disclose explicitly accepted unsigned artifacts
if: inputs.allow_unsigned == true
run: |
cat >> release-assets/RELEASE_NOTES.md <<'EOF'
### Electron installer trust
These Electron installers are unsigned or ad-hoc signed and are not Apple-notarized.
Windows/macOS may show trust warnings. macOS automatic updates are unverified;
use manual installer updates. Tauri updater signatures remain independently verified.
EOF
- name: Create or update draft
run: |
if ! gh release view "$TAG" >/dev/null 2>&1; then
gh release create "$TAG" --verify-tag --draft --title "$TAG — VoiceStudio" --notes-file release-assets/RELEASE_NOTES.md
fi
test "$(gh release view "$TAG" --json isDraft --jq .isDraft)" = true || { echo "Refusing to replace a published release"; exit 1; }
gh release edit "$TAG" --notes-file release-assets/RELEASE_NOTES.md
find release-assets -maxdepth 1 -type f ! -name RELEASE_NOTES.md -print0 | xargs -0 gh release upload "$TAG" --clobber
- name: Publish only when explicitly requested
if: github.event_name == 'workflow_dispatch' && inputs.publish == true
run: gh release edit "$TAG" --draft=false --latest
+243 -29
View File
@@ -27,27 +27,22 @@
# each to surface PyInstaller/Tauri issues that never showed up locally on
# macOS — iterate on CI.
name: Desktop Release
name: Tauri sunset (manual only)
# Legacy workflow: run once on the final Tauri version tag.
# Electron releases are owned by electron-release.yml.
on:
push:
tags: ['v*']
schedule:
# 07:00 UTC daily — rolling `preview` prerelease from `main`. The
# preview-gate job no-ops the matrix when main hasn't moved in a day.
- cron: '0 7 * * *'
workflow_dispatch:
inputs:
draft:
description: "Create as draft release (tag push only)"
required: false
description: "Keep the final Tauri release draft until Electron artifacts are ready"
default: "true"
publish_preview:
description: "Publish a rolling 'preview' prerelease (updater Preview channel). Previews ALWAYS build from main — dispatching from any other branch fails the preview-gate."
required: false
description: "Legacy compatibility input; previews are retired"
type: boolean
default: false
permissions:
contents: write # needed to attach artifacts + updater manifest to GH Release
@@ -76,6 +71,16 @@ jobs:
name: Tests (backend + frontend)
runs-on: ubuntu-22.04
steps:
- name: Require the designated final Tauri tag
env:
SUNSET_TAG: ${{ vars.TAURI_SUNSET_TAG }}
REF: ${{ github.ref }}
PREVIEW: ${{ inputs.publish_preview }}
run: |
test -n "$SUNSET_TAG" || { echo "Set TAURI_SUNSET_TAG to the final v* tag first"; exit 1; }
test "$REF" = "refs/tags/$SUNSET_TAG"
test "$PREVIEW" != "true"
- uses: actions/checkout@v4
- name: Setup Python 3.11
@@ -142,6 +147,9 @@ jobs:
working-directory: frontend
run: node --experimental-strip-types --no-warnings --test ../tests/frontend/*.test.mjs
- name: Electron typecheck, tests and production contract
run: bun run check:electron
# Decide preview-vs-stable, and for nightly runs whether `main` actually
# moved in the last day. Outputs gate the expensive matrix (`build`) and the
# `preview-notes` job, so a no-commit night costs only this ~30s job.
@@ -311,7 +319,7 @@ jobs:
libwebkit2gtk-4.1-dev \
build-essential curl wget file libxdo-dev libssl-dev \
libayatana-appindicator3-dev librsvg2-dev \
libasound2-dev ffmpeg
libasound2-dev ffmpeg xvfb
# ── Frontend build ─────────────────────────────────────────────────
- name: Cache bun deps
@@ -344,7 +352,7 @@ jobs:
- name: Bundle uv (${{ matrix.rust_target }})
shell: bash
env:
UV_VERSION: "0.11.7"
UV_VERSION: "0.12.13"
TRIPLE: ${{ matrix.rust_target }}
run: |
set -euo pipefail
@@ -473,6 +481,9 @@ jobs:
fi
{
echo 'body<<RELEASE_BODY_EOF'
echo '## Final Tauri update'
echo 'VoiceStudio desktop is moving to Electron. This is the last Tauri release. Back up your data and install Electron separately: https://github.com/debpalash/VoiceStudio/blob/main/docs/electron-migration.md'
echo
echo "$BODY"
echo 'RELEASE_BODY_EOF'
} >> "$GITHUB_OUTPUT"
@@ -604,6 +615,21 @@ jobs:
fi
done < /tmp/stale.txt
# A retried job reuses its version and can collide with installers it
# uploaded before a later step failed. Keep other versions/arches intact;
# macOS versionless updater archives are scoped by release tag and arch.
- name: Clear this target's installer assets on retry
if: github.run_attempt > 1
shell: bash
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
RELEASE_TAG: ${{ (needs.preview-gate.outputs.is_preview == 'true') && 'preview' || github.ref_name }}
RELEASE_TARGET: ${{ matrix.rust_target }}
run: |
VERSION=$(python -c 'import json; print(json.load(open("frontend/package.json"))["version"])')
python scripts/clear-release-rerun-assets.py \
--tag "$RELEASE_TAG" --version "$VERSION" --target "$RELEASE_TARGET"
- name: Build + release (Tauri)
uses: tauri-apps/tauri-action@v0
env:
@@ -648,6 +674,84 @@ jobs:
updaterJsonPreferNsis: false
includeUpdaterJson: true
- name: Build + publish Electron desktop
if: false # Electron is released independently by electron-release.yml.
shell: bash
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
RELEASE_TAG: ${{ (needs.preview-gate.outputs.is_preview == 'true') && 'preview' || github.ref_name }}
IS_PREVIEW: ${{ needs.preview-gate.outputs.is_preview }}
VOICESTUDIO_RUST_TARGET: ${{ matrix.rust_target }}
run: |
set -euo pipefail
case "${{ matrix.rust_target }}" in
aarch64-apple-darwin) ELECTRON_OS=darwin; ELECTRON_ARCH=arm64; FLAGS="--mac --arm64" ;;
x86_64-apple-darwin) ELECTRON_OS=darwin; ELECTRON_ARCH=x64; FLAGS="--mac --x64" ;;
x86_64-pc-windows-msvc) ELECTRON_OS=win32; ELECTRON_ARCH=x64; FLAGS="--win --x64" ;;
x86_64-unknown-linux-gnu) ELECTRON_OS=linux; ELECTRON_ARCH=x64; FLAGS="--linux --x64" ;;
*) echo "Unsupported Electron target: ${{ matrix.rust_target }}"; exit 1 ;;
esac
if [ "$IS_PREVIEW" = "true" ]; then
export VOICESTUDIO_UPDATE_CHANNEL="electron-preview-${ELECTRON_OS}-${ELECTRON_ARCH}"
else
export VOICESTUDIO_UPDATE_CHANNEL="electron-stable-${ELECTRON_OS}-${ELECTRON_ARCH}"
fi
if [ -n "${APPLE_CERTIFICATE:-}" ]; then
export CSC_LINK="$APPLE_CERTIFICATE"
export CSC_KEY_PASSWORD="${APPLE_CERTIFICATE_PASSWORD:-}"
fi
bun install --frozen-lockfile
(
cd electron
bun run build
node tests/packaging-contract.mjs
bun x electron-builder \
--config electron-builder.config.mjs $FLAGS --publish never
node tests/packaging-contract.mjs --artifact
if [ "$RUNNER_OS" = "Linux" ]; then
xvfb-run -a node tests/packaged-smoke.mjs --setup
else
node tests/packaged-smoke.mjs --setup
fi
node tests/update-package-contract.mjs \
--channel "$VOICESTUDIO_UPDATE_CHANNEL" \
--platform "$ELECTRON_OS" \
--arch "$ELECTRON_ARCH"
)
# A rolling preview reuses one release. Remove only this platform /
# architecture's older Electron artifacts before publishing the new
# version; sibling matrix legs own different names and metadata.
if [ "$IS_PREVIEW" = "true" ]; then
gh release view "$RELEASE_TAG" --repo "$GITHUB_REPOSITORY" --json assets \
--jq '.assets[].name' > electron-assets.txt
case "${{ runner.os }}" in
Windows) OS_TOKEN=win ;;
macOS) OS_TOKEN=mac ;;
Linux) OS_TOKEN=linux ;;
esac
while IFS= read -r asset; do
case "$asset" in
VoiceStudio-Electron-*-${OS_TOKEN}-${ELECTRON_ARCH}.*|${VOICESTUDIO_UPDATE_CHANNEL}*.yml)
gh release delete-asset "$RELEASE_TAG" "$asset" --yes --repo "$GITHUB_REPOSITORY"
;;
esac
done < electron-assets.txt
fi
electron_artifact_count=0
while IFS= read -r artifact; do
gh release upload "$RELEASE_TAG" "$artifact" --clobber --repo "$GITHUB_REPOSITORY"
electron_artifact_count=$((electron_artifact_count + 1))
done < <(find electron/release -maxdepth 1 -type f \
\( -name 'VoiceStudio-Electron-*' -o -name "${VOICESTUDIO_UPDATE_CHANNEL}*.yml" \) | sort)
if [ "$electron_artifact_count" -eq 0 ]; then
echo "FAIL — Electron build produced no publishable artifacts"
find electron/release -maxdepth 1 -type f -print || true
exit 1
fi
- name: Build per-user Windows MSI
if: runner.os == 'Windows'
shell: bash
@@ -780,7 +884,7 @@ jobs:
set -euo pipefail
MSI=$(find frontend/src-tauri/target/${{ matrix.rust_target }}/release/bundle/msi -name '*Current*User*.msi' | head -1)
powershell.exe -NoProfile -ExecutionPolicy Bypass \
-File scripts/smoke-per-user-msi.ps1 -MsiPath "$(cygpath -w "$MSI")"
-File scripts/smoke-per-user-msi.ps1 -MsiPath "$(cygpath -w "$MSI")" -PrepareHostedRunner
# linuxdeploy re-links .DirIcon as an ABSOLUTE symlink into the build
# machine AFTER tauri's files-map has placed the real icon bytes — the
@@ -875,10 +979,10 @@ jobs:
# ── Compute SHA-256 checksums (Phase 0 GATE-05) ───────────────────
# Native OS tools: shasum -a 256 (POSIX) / Get-FileHash (Windows).
# Writes SHA256SUMS-<label>.txt for the user-verifiable path AND
# captures the content into $GITHUB_OUTPUT for body append.
# Writes SHA256SUMS-<label>.txt, attached to the release below. The
# release-notes-checksums job puts every leg's file into the notes.
- name: Compute SHA-256 checksums
if: github.event_name == 'push' && startsWith(github.ref, 'refs/tags/v')
if: startsWith(github.ref, 'refs/tags/v') && (github.event_name == 'push' || github.event_name == 'workflow_dispatch')
id: checksums
shell: bash
run: |
@@ -899,6 +1003,10 @@ jobs:
-o -name "*.msi" -o -name "*.msi.sig" \
-o -name "*.AppImage" -o -name "*.AppImage.sig" \
-o -name "*.deb" \) 2>/dev/null | sort)
while IFS= read -r artifact; do
ARTIFACTS+=("$artifact")
done < <(find electron/release -maxdepth 1 -type f \
\( -name 'VoiceStudio-Electron-*' -o -name 'electron-*.yml' \) 2>/dev/null | sort)
if [ ${#ARTIFACTS[@]} -eq 0 ]; then
echo "FAIL — no artifacts found under $BUNDLE_DIR"
@@ -929,15 +1037,77 @@ jobs:
echo "checksums_file=$OUT" >> "$GITHUB_OUTPUT"
- name: Append checksums to release + attach SHA256SUMS file
if: github.event_name == 'push' && startsWith(github.ref, 'refs/tags/v')
uses: softprops/action-gh-release@v2
with:
tag_name: ${{ github.ref_name }}
append_body: true
body_path: ${{ steps.checksums.outputs.checksums_file }}
files: ${{ steps.checksums.outputs.checksums_file }}
fail_on_unmatched_files: true
# Attach only. The notes are one shared text and the publish is one
# decision, so both belong to the single release-notes-checksums job
# that runs after the whole matrix (see there for why).
- name: Attach SHA256SUMS file
if: startsWith(github.ref, 'refs/tags/v') && (github.event_name == 'push' || github.event_name == 'workflow_dispatch')
shell: bash
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
TAG: ${{ github.ref_name }}
FILE: ${{ steps.checksums.outputs.checksums_file }}
run: |
set -euo pipefail
gh release upload "$TAG" "$FILE" --clobber --repo "$GITHUB_REPOSITORY"
# ── Checksums into the notes, then publish (the single writer) ───────────
# Every build leg used to append its checksums to the shared release notes
# with softprops/action-gh-release. Two things went wrong:
# - The appends were concurrent read-modify-writes, so a leg that read the
# notes before another wrote them lost its section. v0.5.1 and v0.5.2
# both shipped without the macOS Apple Silicon checksums in the notes.
# - softprops defaults to draft: false, so the FIRST leg to finish
# published tauri-action's draft while the other installers and the
# complete latest.json were still being built (v0.5.2 went public at
# 17:27; its latest.json was finished at 17:38).
# This job is the only writer of the notes and the only publisher. It runs
# once every leg, the manifest repair and the uninstall scripts are done,
# writes the four platforms' checksums in a fixed order, and fails if one is
# missing, so a failed platform leaves the release a draft.
release-notes-checksums:
needs: [build, repair-updater-manifest, uninstall-scripts]
if: startsWith(github.ref, 'refs/tags/v') && (github.event_name == 'push' || github.event_name == 'workflow_dispatch')
runs-on: ubuntu-22.04
timeout-minutes: 10
permissions:
contents: write
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
TAG: ${{ github.ref_name }}
KEEP_DRAFT: ${{ inputs.draft }}
steps:
- name: Write every platform's checksums into the notes, then publish
shell: bash
run: |
set -euo pipefail
WORK="$(mktemp -d)"
gh release download "$TAG" --repo "$REPO" --pattern 'SHA256SUMS-*.txt' --dir "$WORK"
# The checksum sections and the Contributors strip are always the
# tail of the notes; drop them so a re-run rebuilds rather than
# stacks (contributors-strip re-appends its strip after this job).
gh release view "$TAG" --repo "$REPO" --json body --jq .body \
| awk '/^### .* artifacts$/ || /^## Contributors$/ {exit} {print}' > "$WORK/notes.md"
missing=0
for label in "macOS Apple Silicon" "macOS Intel" "Windows x64" "Linux x64"; do
file="$WORK/SHA256SUMS-${label// /.}.txt"
if [ -f "$file" ]; then
cat "$file" >> "$WORK/notes.md"
else
echo "::error::The release has no checksums for $label"
missing=1
fi
done
[ "$missing" = 0 ] || exit 1
gh release edit "$TAG" --repo "$REPO" --notes-file "$WORK/notes.md"
if [[ "$KEEP_DRAFT" == "true" ]]; then
echo "Final Tauri draft verified; Electron publication owns the transition."
elif [[ "$TAG" == *-* ]]; then
gh release edit "$TAG" --repo "$REPO" --draft=false --prerelease
else
gh release edit "$TAG" --repo "$REPO" --draft=false --latest
fi
# ── Uninstall scripts as release assets (#1089) ───────────────────────────
# The in-app uninstaller (Settings → Storage → Remove all data) is the primary
@@ -957,6 +1127,49 @@ jobs:
# This job runs once after the whole matrix as the single final writer:
# it makes the manifest's linux signature agree with the .sig asset that
# actually shipped, and refuses to leave a mismatch behind.
# The matrix validates each Electron package before upload. This final read-only
# check validates the other half of the contract: GitHub must actually serve
# all four manifests and every payload they name. Without it a green release
# can leave the in-app updater with four 404 feeds.
electron-publish-contract:
needs: [build, preview-gate]
if: false # Electron release workflow owns this contract.
runs-on: ubuntu-22.04
permissions:
contents: read
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
TAG: ${{ (needs.preview-gate.outputs.is_preview == 'true') && 'preview' || github.ref_name }}
CHANNEL: ${{ (needs.preview-gate.outputs.is_preview == 'true') && 'preview' || 'stable' }}
STABLE_TAG: ${{ needs.preview-gate.outputs.stable_tag }}
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- name: Verify published Electron updater assets
shell: bash
run: |
set -euo pipefail
WORK="$(mktemp -d)"
mkdir -p "$WORK/manifests"
gh release view "$TAG" --repo "$GITHUB_REPOSITORY" \
--json tagName,isPrerelease,assets > "$WORK/release.json"
gh release download "$TAG" --repo "$GITHUB_REPOSITORY" \
--pattern "electron-${CHANNEL}-*.yml" --dir "$WORK/manifests"
if [ "$CHANNEL" = "preview" ]; then
VERSION=$(python3 scripts/stamp-preview-version.py \
--package-json frontend/package.json \
--stable-tag "$STABLE_TAG" \
--run-number "${{ github.run_number }}")
else
VERSION=$(python3 -c 'import json; print(json.load(open("frontend/package.json"))["version"])')
fi
python3 scripts/check_electron_release_assets.py \
--release-json "$WORK/release.json" \
--manifest-dir "$WORK/manifests" \
--channel "$CHANNEL" \
--version "$VERSION"
repair-updater-manifest:
needs: [build, preview-gate]
runs-on: ubuntu-latest
@@ -997,7 +1210,7 @@ jobs:
uninstall-scripts:
needs: [build]
if: github.event_name == 'push' && startsWith(github.ref, 'refs/tags/v')
if: startsWith(github.ref, 'refs/tags/v') && (github.event_name == 'push' || github.event_name == 'workflow_dispatch')
runs-on: ubuntu-22.04
permissions:
contents: write
@@ -1034,11 +1247,12 @@ jobs:
# MUST append via `gh release edit` on the EXISTING release (never a second
# softprops publish — that races tauri-action's per-matrix draft and splits
# installers across two releases; see uninstall-scripts). `needs: [build]`
# guarantees the release + all checksum appends already landed, and this job
# guarantees the release exists and release-notes-checksums has written the
# notes, and this job
# is single (no matrix) so there is no write race. Idempotent: it strips any
# prior "## Contributors" block before re-appending, so re-runs don't stack.
contributors-strip:
needs: [build]
needs: [build, release-notes-checksums]
if: >-
github.event_name == 'push'
&& startsWith(github.ref, 'refs/tags/v')
+13
View File
@@ -22,6 +22,8 @@ node_modules
.turbo/
bun.lockb
frontend/src-tauri/target/
electron/.tmp-native-target/
native/**/target*/
# ─────────────────────────────────────────────────────────────────────────
# Secrets & env
@@ -55,6 +57,7 @@ memxt.db-wal
!.claude/agents/**
/.cache*
/.tmp/
/.tmp-*
# ─────────────────────────────────────────────────────────────────────────
# Research clones — upstream repos used as reference, not shipped
@@ -170,3 +173,13 @@ bin/omnivoice-tts-linux-aarch64
# committed (they ship with the app); the per-language source WAVs are just the
# inputs scripts/render_dub_demo_audio.py hands to scripts/build_dub_demo.sh.
backend/assets/samples/demo/dubbing/*.src.wav
# Stray sqlite session artifacts (`<db-path>.ses`). An in-memory DB yields the
# literal name `:memory:.ses`, and a path containing `:` cannot be checked out
# on Windows at all — committing one fails every Windows CI job at the git
# checkout step, before a single test runs. Guarded by
# tests/test_no_windows_hostile_paths.py.
*.ses
# Generated Windows MSI diagnostic logs and installer payloads
/wix-diagnostic-artifacts/
+2
View File
@@ -25,6 +25,8 @@ regexes = [
'''^hf_QWERTYUIOPasdfghjklZXCVBNM0123456789xyzAB$''',
# NLLB generation length argument, not the value of a credential.
'''^max_length=400$''',
# Dubbing pane split-position localStorage key, not a credential.
'''^omnivoice\.dubSplit\.v1$''',
# cryptography's Ed25519 private-key type name, not key material.
'''^Ed25519PrivateKey$''',
]
+5
View File
@@ -33,6 +33,11 @@ Binding for every AI agent (Claude, Codex, Cursor, review bots, …). CLAUDE.md
- `frontend/package.json` dep changes require regenerating root `bun.lock` (Docker runs `--frozen-lockfile`).
- Issues: absorb or decline — never defer to a future version. Check the open-PR queue before implementing community-reported fixes.
## Shared select controls
- Use `frontend/src/components/SearchableSelect.jsx` for all new or redesigned select boxes. Reuse `VoiceSelector` for voice choices. Do not introduce native `<select>` controls.
- Provide a localized `ariaLabel`; use `menuPortal` inside scrolling or clipping containers. Preserve keyboard selection and disabled states.
## Agent skills
Project development skills are pinned in `skills-lock.json` and installed under
+213 -5
View File
@@ -8,25 +8,129 @@ the frozen-backend fallback mirror it for their toolchains.
## [Unreleased]
## [0.5.3] — 2026-09-17
**Highlights**
- Fix current-user Windows installer validation and nested resource cleanup (#730)
- The README is shorter, with a new Electron UI tour and refreshed screenshots (#2129)
- Voice cloning now starts with a clear upload-or-record choice, reveals recording and reference details only when needed, and keeps sampling controls under Production Overrides (#1817)
- Support pages feature cleaner donation cards, with a workspace support shortcut and sponsor footer with hover cards and email inquiries (#2129)
- Integrations has a dedicated sidebar workspace with featured sponsors, searchable AI providers, and smooth sponsor-strip scrolling (#2129)
- Integrations now covers 100+ automation, communications, MCP, agent, developer, data, and productivity tools with config-driven detail pages (#2129)
- Electron now ships as a complete cross-platform VoiceStudio desktop app with local-first cloning, production workspaces, model packs, repair agents, native integrations, updates, parity checks, and the shared backend contracts required by those workflows (#1823)
- The Model Catalogue is one page: what you use now on top, then each family's engines and weights (#2013)
- VoxCPM2 installs in one click into its own environment, with the CUDA build of PyTorch on NVIDIA GPUs (#2021)
- MOSS-TTS-Nano installs in one click into its own environment, pinned to a reviewed upstream commit it works with (#2022)
- CosyVoice 3 installs in one click into its own environment, with a trimmed dependency set that needs no TensorRT, DeepSpeed or third-party package feed (#2025)
### Changed
### Added
- Electron becomes the default source desktop, with artifact-only packaging rehearsals and a separate final Tauri update path (#2157)
- Installable agent skills use current VoiceStudio names and Electron workflows (#2157)
- README clarifies the Electron transition while keeping desktop contributions welcome (#2153) — thanks @cyberspace-cs!
### Docs
- Electron first run uses four simple steps with model packs, optional advanced controls and skippable dictation setup (#2129)
- Model Catalogue is one page: a setup summary (speech, transcription, dictation, language model) on top, one TTS / ASR / LLM switch, and each family's downloadable weights listed under its engines; the separate Models pane and the Settings → Voice → Engines / Models signposts are gone, the models directory and voice previews moved to Settings → Storage and the HF mirror to Network (#2013)
- The engine list is one line per engine (engine, device it runs on, status, one action) with a detail panel for everything else; each engine's weights install from its panel, so the separate weights list and recommendation card are gone (#2020)
- CosyVoice 3 installs patched protobuf and transformers releases, clearing five security advisories (#2030, #2031)
### Fixed
- Keep demo playback aligned across languages, preserve worker GPU metrics, and restrict unsigned releases to owner dispatches (#2157)
## [0.5.2] — 2026-09-02
- Desktop integration checks cover current dubbing safeguards, navigation, and the linked engine catalog (#2157)
- Tauri and Electron now share native dictation, watch-folder, and Wayland shortcut contracts; focused paste stays ordered and first-run uv stays pinned at 0.12.13 (#2122)
- Dubbing demos synchronize playheads without simultaneous playback and let you open a sample in the editor (#2131)
- macOS desktop sidebar clears the traffic lights, uses a narrower collapsed rail, and places notifications and device controls with more space (#2126)
- Dubbing timelines keep short segments proportional, support zoom, and remove timestamp-confirmed duplicate ASR context (#2129)
- Dubbing translation shares the agent footer with live logs, validated output, cancellation and contextual retries (#2129)
- Agent dubbing translation saves a custom tone and adaptation prompt and preserves it during timing rewrites (#2129)
- Dubbing preserves original sound outside dialogue and mixes separated background only beneath replacement speech (#2129)
- Dubbing repairs missing speech caches, rejects incomplete output, avoids oversized speaker references, and fits full speech without early clipping (#2129)
- Workspace sidebars have a working right-edge resize handle, allow 40% more width, remember their size, and keep video controls inside the preview (#2129)
- Pressing Play while a video is loading starts playback when it is ready instead of reporting playback unavailable (#2129)
- Video previews show their thumbnail before playback, including the source video in Dub (#2129)
- Linux and Windows workspace headers consistently expand and collapse the sidebar, with the app logo at the top of the collapsed rail (#2129)
- Stopping a process on macOS no longer fails with "Operation not permitted" when it was already exiting (#2032)
- A YouTube link blocked by its "not a bot" check now says how to attach signed-in cookies in Dub, instead of quoting yt-dlp's command-line flags (#2036, #2034)
- An engine that fails to start now says whether it timed out, crashed (with its exit code and last output) or answered wrongly, instead of "did not signal ready: None" (#2037, #2026)
- Transcribing an M4A file with PyTorch Whisper works, instead of failing with "Format not recognised" (#2042, #2039)
- PyTorch Whisper runs on 6 GB NVIDIA cards instead of falling back to CPU, because its memory check now fits the model it loads (#2044, #2041)
- MCP tools wait as long as the backend does, so a long transcription no longer fails at 120 s with an empty error (#2043, #2040)
- Generating on an older NVIDIA GPU (Tesla T4, and other pre-Ampere cards) no longer kills the backend on the first request — CUDA graphs are not captured below sm_80 (#2135)
- "Disable torch.compile" in Settings → Performance now works on macOS and Linux, not only Windows; it was greyed out on the platforms that needed it (#2135)
- Setting `TORCH_COMPILE_DISABLE=1` in the environment now actually disables torch.compile, for the in-process engine and engine subprocesses alike (#2135)
- A backend killed by a native crash now leaves the faulting thread's stack in `backend_err.log` instead of exiting silently (#2135)
### CI
- A tagged release is published only after every platform's installers and checksums are attached, and its notes list all four platforms' checksums (#2029)
- A worker-transport test no longer fails when a slow Windows runner takes over 2 seconds to tear down (#2038)
## [0.5.2] — 2026-09-10
**Highlights**
- Supertonic-3 and PocketTTS show their license Accept button again, so they can be enabled (#2017)
- An engine that can't run on your platform says so, instead of telling you to install it (#2018)
- MOSS-TTS-v1.5, Confucius4-TTS, dots.tts, Supertonic-3 and PocketTTS install in one click, each in its own environment, so switching engines and back never breaks a working one (#2015, #2016)
- A pronunciation entry that is stored but not applied yet says so, instead of looking like it did not match (#1949)
- A bare 500 report now names the backend error class, so two unrelated faults stop filing the same issue (#1773)
- A rejected dubbing source language now names the code it rejected (#1960)
- The first-run install log is kept on disk instead of vanishing with the setup screen (#1847)
- `bun run desktop` reclaims port 3900 from a backend the app itself left running, instead of refusing to start (#1974)
- A dictation shortcut another app already owns now says so, instead of silently doing nothing (#1858)
- Quitting on Windows is no longer reported as a crash on the next launch (#1898)
- A Reduce motion switch in Settings, for calm without changing your whole system (#1857)
- A light theme, and System Auto now follows a light-mode OS instead of staying dark (#1973) — thanks @CoDe-ReDz!
- Generating from a one-character input now says the input was too short, instead of quoting a convolution error (#1826)
- First run asks about text size before the install, not after it (#1849)
- Cloning without a reference clip now says so, instead of naming library parameters you cannot set (#1879)
- Upgrading torch for an RTX 50-series card no longer trades one startup crash for another, and the upgrade is documented (#1931)
- A generation timeout now points at the compute-time budget in Settings rather than an environment variable (#1808)
- An engine you have not installed now says so, instead of reporting a failed check (#1866)
- The Accessibility prompt no longer floats over first-run setup and every other app until you grant it (#1845, #1886)
- The last onboarding step offers to install a speech-to-text model instead of failing three times when none is installed (#1856)
- A download that fails because the folder sits behind a mount point Windows will not cross now says so, and where to move it (#1957)
- A GPU that is merely short on free memory is no longer told to reinstall its drivers (#1812) — thanks @michaelhuamanflores!
- An error thrown by a browser extension is filtered on Safari and the macOS app too, not only on Chromium (#1901) — thanks @Chang-Jin-Lee!
- Choosing the China mirror no longer re-races the network on every dependency step, which cost seconds per step on blocked connections (#1892) — thanks @yuezheng2006!
- The backend log panel reports a log it cannot read instead of quietly showing less (#1847) — thanks @Chang-Jin-Lee!
- The floating dictation bubble adds pause, resume, stop, close, and a multiline preview (#1952)
- Transcriptions checks model readiness and offers an inline download and shortcut hints (#1952)
- Transcriptions' missing-model prompt lists every dictation model by accuracy vs latency, languages and size, so you install the one that fits — or switch to one already on disk (#1952)
- The Engines menu's Transcription tab picks the dictation model under Sherpa-ONNX, and that choice now also drives Sherpa transcription (#1952)
- A failure with no stage attached no longer borrows another stage's advice, so a text-to-speech error stops telling you the video server dropped the download (#1943)
- A generation failure that the app cannot classify now names the backend error class, so two unrelated faults stop arriving as the same untriageable report (#1800)
- Transcriptions dictation wakes the desktop recorder, presents one contextual start action, and centers its microphone icon with the label (#1902)
- Colab transcription and dubbing now include an explicit ASR model setup step (#1922) — thanks @nidhi-singh02!
- Apple Silicon now shows one canonical OmniVoice choice in the engine picker while retaining its automatic crash-isolated sidecar runtime (#1913)
- Validate current-user Windows installers under a standard account on hosted runners (#1883)
- Model downloads survive a flaky connection instead of restarting from zero (#1940)
- `bun run dev` recovers on Windows instead of demanding Task Manager (#1941)
- The desktop app builds and opens from a fresh clone again (#1818) — thanks @flutterkage2k!
- GPUs with less VRAM than the engine needs no longer get half the compute-time budget a CPU gets (#1806) — thanks @VishvakR!
- Gallery voice previews play again — the quality guard was rejecting good renders as silent (#1819) — thanks @flutterkage2k!
- Tilde-separated number ranges are spoken clearly without running their endpoints together (#1821) — thanks @flutterkage2k!
- Voice modes use themed tabs, with Synthesize and Convert pinned below their scrolling forms (#1823)
- Fix current-user Windows installer validation and nested resource cleanup (#1873)
- Keep generated frontend assets available while building the current-user Windows installer (#1881)
- Voice cloning now starts with a clear upload-or-record choice, reveals recording and reference details only when needed, and keeps sampling controls under Production Overrides (#1817)
- The first-run welcome line uses an instruction accepted by OmniVoice and VoiceDesign engines (#1861) — thanks @psiberfunk!
- audio.cpp joins the engine lineup as an opt-in CPU backend for Breeze-TTS-2 (English + Chinese, clone + voice design, explicit Model Catalogue install, no Python venv) (#1891)
- audio.cpp uses installed native CUDA, HIP, Metal, and Vulkan providers and preserves device routing across remote workers (#1926)
- Show estimated and measured model, dependency, cache, and temporary disk costs in the engine catalogue (#1718)
- Preview builds now stay newer than Stable even when automatic post-release version bumps are disabled (#1762)
- CosyVoice setup guidance now separates downloaded model files from the runtime that makes the engine available (#1761)
@@ -40,10 +144,33 @@ the frozen-backend fallback mirror it for their toolchains.
### Changed
- Tauri 2.11.5 with refreshed plugins (dialog, updater, log, opener, positioner, single-instance), React 19.3, TanStack Query 5.102, lucide 1.43, posthog-js 1.428, and the rest of the npm workspace on current minors; jsdom 30, jest-dom 7, concurrently 10, taze 21 (#1952)
- eslint ignores `src-tauri/`, so a local Tauri build no longer floods `lint:hooks` with parse errors from generated assets (#1952)
- Casting uses responsive SVG voice cards and searchable speaker menus that stay above surrounding panels (#1823)
- Dubbing aligns output settings, brings review status forward, and simplifies transcript and glossary editing; Launchpad files and voices reflow into responsive grids (#1823)
- Transcript segments use three readable rows for text, timing/status and voice controls, with heights that adapt to wrapping (#1823)
- Dragging the waveform pans horizontally while a click still seeks, keeping the timed transcript aligned (#1823)
- Bulk segment editing uses searchable voice and language menus, readable language names and a responsive selection toolbar (#1823)
- Dubbing overlays playback controls on video, combines waveform and transcript in a compact timeline, and removes header/action background fills (#1823)
- Dubbing uses compact casting, translation and output controls with responsive rows to leave more room for editing (#1823)
- Export uses grouped format settings, themed track menus and switches, with a pinned filename summary and download action (#1823)
- Dubbing output settings use icon-labelled switches, themed track and speaker menus, and clearer timing/transcript controls (#1823)
- Casting voice menus use searchable themed options with SVG preset icons instead of native dropdowns (#1823)
- Dubbing groups casting and translation controls with readable labels, SVG icons, searchable menus, and compact timeline spacing (#1823)
- Production Overrides use readable icon-labelled controls and accessible Denoise/Postprocess switches (#1823)
- Expanded navigation uses a theme-accent tint with subtle static wave gradients (#1823)
- Convert groups source audio, target voice, and timing options into clearer controls; design choices include theme-matched SVG icons (#1823)
- The expandable sidebar reveals workspace labels with restrained active states; language menus adapt to multiple columns on wider screens (#1823)
- Voice design and recording use themed, keyboard-accessible selectors with clearer spacing and labels (#1823)
- Voice tabs and upload/record controls have subtle SVG motion; Text adds clipboard paste and the upload area fills available height (#1823)
- The title-bar label cycles through active speech, transcription, and LLM engines; bundled model labels correctly say OmniVoice (#1823)
- The top-bar Engines panel groups Speech, Transcription, and LLM choices into tabs, with compact memory controls and no duplicate pickers (#1823)
- Voice Design simplified: the 12-row fine-grained block collapses to one summary line with a five-field editor, English accent and Chinese dialect merge into a single field, and the starting-point chips now show 5 with an overflow toggle (#1793)
### Added
- Remote-worker metrics distinguish unavailable readings from zero and keep probes off the control loop (#2155)
- The audiobook result is now a synced-lyrics player: chapter text follows playback with the current word highlighted and click-to-seek, timed from the render's own chapter durations with a karaoke-style even split — no ASR pass, fully local (#1766) — thanks @mvanhorn!
- The dub CAST strip expands into a project-level casting board: drag voice chips (clone profiles, design presets, Default) onto speaker rows — or pick from a keyboard listbox — writing the same per-speaker cast fields as the existing dropdowns (#1767) — thanks @mvanhorn!
- Studio's new Convert method turns a dropped or recorded clip into an existing voice profile's voice, with optional source-duration matching (#1765) — thanks @mvanhorn!
@@ -55,6 +182,11 @@ the frozen-backend fallback mirror it for their toolchains.
### Docs
- PowerShell Docker setup now generates the administrator key without requiring Python on the host (#1993) — thanks @yangfan-yf-yf!
- The torch upgrade an RTX 50-series card needs is written down, with the second pin file the resolver checks and the command that proves the kernels are there (#1931)
- Docker quick starts now explain the AMD64-only images and direct Apple Silicon users to the native macOS app (#1921) — thanks @yangfan-yf-yf!
- audio.cpp (Breeze-TTS-2) is now a documented opt-in engine: prebuilt binary install, explicit GGUF download, voice modes, and the weights' research/non-commercial terms (#1891)
- `docs/STRUCTURE.md` describes the tree as it is today, and a test now keeps its counts honest (#1981) — thanks @Dawcraft!
- Local gigastt is now documented as a supported OpenAI-compatible ASR endpoint, with loopback privacy distinguished from remote servers (#1736) — thanks @ekhodzitsky!
- The CosyVoice guide now states that packaged builds have no one-click runtime installer and records the exact readiness checks exposed by [Discussion 1631](https://github.com/debpalash/VoiceStudio/discussions/1631) (#1761)
- A production private-API guide now covers pinned containers, root credentials, network isolation, streaming proxies, health checks, upgrades, and benchmark evidence (#1720)
@@ -62,6 +194,82 @@ the frozen-backend fallback mirror it for their toolchains.
### Fixed
- One-click engine installs no longer inherit VoiceStudio's own PyTorch pin, which made MOSS-TTS-v1.5 and Confucius4 impossible to install (#2024)
- Uninstalling a translation engine no longer removes a package VoiceStudio or another engine still needs (#2019)
- Closing the dictation pill on Windows removes it from the screen: an empty dark rectangle used to stay there, always on top, until the app was quit (#2009)
- The dictation pill on Windows no longer sits inside a bordered card wider than the pill itself (#2009)
- Dictation uses the model you picked instead of one remembered from before the backend started, so it stops reporting no speech-to-text model while one is installed — and when none is, the main window offers the download (#2012)
- The remote-worker loop-responsiveness tests no longer turn a build red over milliseconds of scheduling noise on shared CI hardware (#1990)
- Remote GPU workers work when the machine running VoiceStudio is on Windows: a staged input is now identified the same way on every operating system, instead of with a path only Windows can read (#2005)
- The pronunciation list badges an IPA or CMU entry as not applied yet, so you can see it without running a test (#1949) — thanks @utkarsha741!
- A remote-worker test no longer fails at random on Windows CI: it waited for a background thread by spinning the event loop that thread's work needed (#1990)
- The isolated backend test session passes on a stock Windows checkout, and CI now runs it there so it stays that way (#1990)
- Windows contributors can run the test suite without Developer Mode: tests that create a symlink now skip instead of failing with `WinError 1314` (#1990)
- The crash details dialog now says what the exit code means and what to try, instead of showing a raw number and a log (#1927)
- A crash report now carries the backend's actual last words: the log tail is captured after the dying process's final output lands, not the instant it exits (#1850)
- The first-run setup screen no longer mislabels a step when the bootstrap restarts itself: Rust now says which attempt each stage and log line belongs to, instead of the screen guessing from a once-a-second poll (#1900)
- A port-3900 conflict now names who is actually holding it, and gives the command that ends an orphaned backend, instead of telling you to quit an app that has no window (#1933) — thanks @Chang-Jin-Lee!
- Windows desktop launches no longer freeze at "Loading ML runtime (PyTorch)": the parent-liveness watchdog polls the stdin pipe instead of leaving a read pending, which deadlocked numpy's OpenBLAS initializer (#1952, #1955)
- `bun desktop-prod` and `bun desktop-fresh` find Rust and uv from a terminal opened before they were installed, as `bun desktop` already did; a missing Rust toolchain fails up front with the install steps (#1952)
- Voice synthesis progress no longer races to a fabricated 95%; it stays indeterminate until the active generation path reports real progress (#1907) — thanks @psiberfunk!
- The Backend log tab keeps showing history across a log rollover, instead of going nearly empty until new lines arrive (#1920)
- Clearing the logs now empties the rotated log files too, so it frees the space it appears to (#1920)
- An error thrown by a browser extension no longer offers to file itself as a VoiceStudio bug (#1901)
- Clearing the desktop logs no longer wipes the backend's stderr, which is the only record a native crash leaves behind and is meant to survive a respawn (#1510)
- Long audiobook chapters now use the same device- and text-length-aware synthesis timeout as other TTS routes (#1910) — thanks @psiberfunk!
- Interrupted audiobook renders can resume cached chapters after tab navigation, and their chapter cache is available from the recovery card (#1911) — thanks @psiberfunk!
- System-check details and storage paths beginning with a number or a slash no longer render with their leading text moved to the end of the line (#1848) — thanks @psiberfunk!
- An unavailable engine's row now links to that engine's guide, so the generic "check installation and configuration" message has somewhere to send you (#1866) — thanks @psiberfunk!
- The backend log now records which engine failed a health check and whether its probe raised, instead of a line that identified neither (#1866) — thanks @psiberfunk!
- The first-run Activity log counts every line instead of freezing at 200 while the install is still running, and Copy now hands back the whole run rather than the last 200 lines (#1847) — thanks @psiberfunk!
- A first-run failure that happened early in a long install keeps its specific advice, instead of falling back to the generic retry hint once the log scrolled past 200 lines (#1847) — thanks @psiberfunk!
- Opening the log panel no longer clips the Launchpad's heading and slides the feature cards up over it — the page scrolls instead of squashing itself (#1859) — thanks @psiberfunk!
- Segmented model downloads split files into 16 MB ranges instead of one range per connection, so a dropped connection refetches one range rather than restarting the file (#1940)
- The download accelerator is kept across retries after a transient network failure and resumes from its manifest, instead of falling back to a from-zero `snapshot_download` (#1940)
- `dev-backend.mjs` stops the backend by process tree on Windows, so an orphaned uvicorn no longer holds port 3900 and turns a source reload into three phantom crashes (#1941)
- `clear-dev-ports.mjs` can free a stuck development port on Windows again, bound to the inspected process instance so a recycled pid is never terminated (#1941)
- Checkout-ownership matching no longer resolves POSIX paths with the host's separator, which made the guard's own test fail on Windows (#1941)
- Install documentation help now prints correctly on Windows consoles using legacy encodings (#1815) — thanks @dajiaohuang!
- Saved transcriptions with missing or invalid timestamps now remain readable (#1799) — thanks @yunaremaia and @tvbht!
- Transcribing with an engine that reports no segment end no longer fails with a server error; the null timing is passed through the way the segment list already expects (#1904) — thanks @aeroglu!
- Copying a saved transcription now uses the shared clipboard helper and reports failed copies accurately (#1803) — thanks @tvbht!
- Voice reference preparation reclaims allocator memory before one bounded retry, then reports persistent GPU out-of-memory failures (#1811)
- `bun run desktop` now opens on a fresh clone: the Vite alias for `@tauri-apps/plugin-dialog` no longer assumes a nested `frontend/node_modules`, which bun's workspace hoisting leaves empty (#1818) — thanks @flutterkage2k!
- Slow backend startups remain running with progress updates, and Retry interrupts startup without stale timeout failures (#1809)
- Backend connection errors report crashes only when recorded evidence exists, and diagnostic waits honor cancellation (#1810)
- A CUDA or ROCm GPU with less VRAM than the engine needs now gets the CPU compute-time budget instead of the shorter accelerated one, since it pages to system RAM and renders slower than the CPU would — applied to local generation, voice conversion, and remote worker deadlines alike (#1806) — thanks @VishvakR!
- Gallery previews no longer fail with "the voice engine returned no audible audio" on perfectly good renders: the degenerate-buzz guard measured spectral flatness over the whole clip (so the value tracked clip length) against a threshold calibrated on a synthetic signal, and rejected real speech in every language tested (#1819) — thanks @flutterkage2k!
- Speak tilde separators in integer, signed, and decimal ranges in English, Korean, Japanese, and Chinese (#1821) — thanks @flutterkage2k!
- Keep recording and conversion work safe while switching methods, synchronize dubbing language controls, and localize timeline controls and timing warnings (#1841)
- Audiobook is now a Write → Cast → Produce tab workspace matching the voice workspace, with the warnings/progress/result rail pinned below (#1841)
- Gallery uses a workspace header with zone tabs, hairline section dividers, theme-token cards, and borderless import rows (#1841)
- Gallery cards reset native button faces, cluster icon actions in the header so Use voice never wraps, and use a roomier grid floor (#1841)
- Gallery filters gain name search, removable iconified pills with clear-all, and dimension icons on every facet (#1841)
- Dubbing playback starts before waveform decoding, automatic cast names are readable, and transcript timestamps have more room (#1823)
- The title-bar engine button stays compact and stable while cycling labels, with engine names aligned right (#1823)
- Long dubbing segment errors wrap in a bounded scrollable notice instead of widening the editor (#1823)
- Voice dropdowns match their field width, use theme accents, and show recent voices only once (#1823)
- Language menus no longer show a pale frame around their search header (#1823)
- The notification count stays inside the title bar instead of clipping above the bell (#1823)
- The workspace engine menu opens beside its button instead of at the opposite edge of the page (#1823)
- Cloning reuses the dubbing language picker with flags, search, and single selection, opening above the pinned synthesis controls (#1823)
- The first-run welcome line uses an instruction accepted by OmniVoice and VoiceDesign engines (#1861) — thanks @psiberfunk!
- The header status dot now honors OS Reduce Motion instead of pulsing regardless (#1862) — thanks @psiberfunk!
- Onboarding reads Hugging Face tokens locally, preserves Windows CLI logins, and requires successful discovery before replacing saved credentials (#1852) — thanks @psiberfunk!
- The logs panel no longer reports “All clear” before log retrieval succeeds or while logs contain warnings or errors (#1870) — thanks @motodriver!
- MOSS accelerator routing and status match runtime selection, with CPU fallback when device probing fails (#1830) — thanks @li-lizhe!
- Confucius accelerator routing tolerates failed device probes, and dots.tts keeps safe default precision on non-CUDA hosts (#1831) — thanks @li-lizhe!
- On macOS, the header status dot and kicker no longer render underneath the overlaid traffic lights (#1863) — thanks @psiberfunk!
- The capture widget can hide after recording and recover from being left visible while idle (#1865) — thanks @psiberfunk!
- macOS retains the shared desktop window sizing, resize limits, and file-drop behavior when native chrome is applied (#1865) — thanks @psiberfunk!
- On macOS, the header no longer shows Windows-style minimize/maximize/close buttons alongside the native traffic lights (#1865) — thanks @psiberfunk!
- Release retries replace their own partially uploaded installers without colliding with existing assets (#1871)
- Timed-out voice engines finish process cleanup before retrying, and old timeout callbacks cannot kill replacement engines (#1872)
- Fast macOS process exits no longer turn a completed shutdown into a permission error (#1809)
- The bootstrap splash no longer shows fabricated first-run install steps on a warm start or repair sync — a step now renders done only once it was actually observed (#1894)
- A deliberate, clean quit killed by the desktop shell's short shutdown grace no longer gets reported as a crash on next launch — the run sentinel now clears before the slower shutdown steps instead of after (#1895)
- Model Catalogue engine rows stack into one column on narrow shells instead of clipping actions off-screen (#1891)
- Simplified Chinese locale completed: all 486 missing keys translated and the parity ratchet tightened to zero (#1877) — thanks @yearth!
- The generation compute-time budget is now a Settings control (Performance & Device) instead of an env-var-only setting the timeout error recommended with no UI path — the error copy points there too, and long CPU/MPS renders get an upfront heads-up before they start (#1787)
- Windows: the backend can now start when the install path contains non-English characters (e.g. a CJK username) on a non-UTF-8 system code page — a new or broken Python environment now builds at an ASCII-safe path automatically (a healthy existing one is never relocated), and a specific error message names the cause and a working fix if the interpreter still crashes in `site` (#1783)
- Exports and other native-picker actions no longer 403 with "Invalid or expired desktop authorization" when the desktop app and backend resolve different data directories, e.g. dev mode or a custom data folder (#1781)
+56 -439
View File
@@ -1,484 +1,101 @@
<div align="center">
<p><img src="docs/logo.png" alt="VoiceStudio logo" width="120" height="120" /></p>
<img src="docs/logo.png" alt="VoiceStudio" width="88" />
<h1>VoiceStudio</h1>
<p>
<a href="https://trendshift.io/repositories/28176?utm_source=repository-badge&amp;utm_medium=badge&amp;utm_campaign=badge-repository-28176" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/repositories/28176" alt="VoiceStudio ranking on Trendshift" width="220" height="48" /></a>
</p>
<p><sub>Previously OmniVoice-Studio</sub></p>
<h3>Clone voices, dub video, dictate, and produce long-form audio on your own hardware.</h3>
<p>16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker</p>
<p>No account, API key, subscription, or usage meter for the local workflow.</p>
<p><strong>Open source voice cloning and workflow engine. Build local.</strong></p>
<p>
<a href="#install">Install</a> ·
<a href="#features">Features</a> ·
<a href="#comparison">Compare</a> ·
<a href="#requirements">Requirements</a> ·
<a href="#hardware-recommendations">Hardware</a> ·
<a href="#engines">Engines</a> ·
<a href="#architecture">Architecture</a> ·
<a href="#api">API</a> ·
<a href="https://voicestudio.sh/?utm_source=github&utm_medium=readme&utm_campaign=project">Website</a> ·
<a href="https://github.com/debpalash/VoiceStudio/releases/latest">Download</a> ·
<a href="#get-started">Get started</a> ·
<a href="#documentation">Docs</a> ·
<a href="#faq">FAQ</a> ·
<a href="README_CN.md"><strong>简体中文</strong></a>
<a href="https://discord.gg/bzQavDfVV9">Discord</a> ·
<a href="README_CN.md">简体中文</a>
</p>
<p>
<a href="https://github.com/debpalash/VoiceStudio/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/debpalash/VoiceStudio/ci.yml?branch=main&style=flat-square&label=CI" alt="CI status" /></a>
<a href="https://github.com/debpalash/VoiceStudio/stargazers"><img src="https://img.shields.io/github/stars/debpalash/VoiceStudio?style=flat-square&color=f59e0b" alt="GitHub stars" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases"><img src="https://img.shields.io/github/downloads/debpalash/VoiceStudio/total?style=flat-square&color=8b5cf6&label=downloads" alt="Total downloads" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/github/v/release/debpalash/VoiceStudio?style=flat-square&color=10b981" alt="Latest release" /></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/license-AGPL--3.0-blue?style=flat-square" alt="AGPL-3.0 license" /></a>
<a href="https://discord.gg/bzQavDfVV9"><img src="https://img.shields.io/badge/Discord-Community-5865F2?style=flat-square&logo=discord&logoColor=white" alt="Discord community" /></a>
</p>
<p>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Download-macOS_·_Windows_·_Linux-10b981?style=for-the-badge" alt="Download VoiceStudio" /></a>
<a href="https://github.com/debpalash/VoiceStudio/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/debpalash/VoiceStudio/ci.yml?branch=main" alt="CI" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/github/v/release/debpalash/VoiceStudio" alt="Latest release" /></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/license-AGPL--3.0-blue" alt="AGPL-3.0" /></a>
</p>
</div>
<div align="center">
<img src="docs/media/0.5.0/quick-switch.gif" alt="Switching TTS engines from the VoiceStudio status bar" width="100%" />
</div>
![A tour of the Electron app: voice cloning, voice design, dubbing, and model management](docs/media/electron/voicestudio.gif)
> [!WARNING]
> **Active beta.** Use the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) for stable work. `main` contains the newest fixes and may change between releases. Report problems through [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues).
## Your voice. Your workflow.
## At a glance
| Create | Produce | Connect |
| :--- | :--- | :--- |
| Clone a voice or design your own | Dub videos with timed speech | Local API & MCP for agents |
| Dictate with a floating widget | Stories, audiobooks & batch jobs | Optional remote workers |
| | VoiceStudio |
|---|---|
| **Workflows** | Voice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation |
| **Language catalogue** | 646 TTS languages; actual coverage and quality depend on the selected engine |
| **Engines** | 16 TTS · 11 ASR · switch in Model Catalogue or with <kbd>Ctrl</kbd>/<kbd>Cmd</kbd>+<kbd>E</kbd> |
| **Platforms** | macOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+ |
| **Compute** | CUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers |
| **Interfaces** | Desktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server |
| **Storage** | Voices, projects, settings, and outputs stay on the machine by default |
| **License** | AGPL-3.0 application; downloaded models keep their upstream terms |
Start with **VoiceStudio** (default, powered by k2-fsa/OmniVoice), or choose another engine. [Features & engine catalog](docs/feature-catalog.md).
<a id="install"></a>
Local workflows run on your hardware. Remote services are optional; usage analytics requires consent.
## Install
<details>
<summary><strong>Explore the workspaces</strong> · Clone, dub, design & models</summary>
Download a package from the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest), then follow the platform guide.
<table>
<tr>
<td><img src="docs/media/electron/voice-cloning.png" alt="Electron voice cloning workspace with the bundled demo voice" width="100%" /></td>
<td><img src="docs/media/electron/dubbing.png" alt="Electron video dubbing workspace" width="100%" /></td>
</tr>
<tr><td align="center">Voice cloning</td><td align="center">Video dubbing</td></tr>
<tr>
<td><img src="docs/media/electron/voice-design.png" alt="Describe a voice in the Electron voice design workspace" width="100%" /></td>
<td><img src="docs/media/electron/models.png" alt="Install and manage local speech models" width="100%" /></td>
</tr>
<tr><td align="center">Voice design</td><td align="center">Local models</td></tr>
</table>
| Platform | Package | Guide |
|---|---|---|
| macOS 13.3+ | Apple Silicon DMG | [Install on macOS](docs/install/macos.md) |
| Windows 10/11 | x64 MSI; choose the current-user build when listed to install without admin access | [Install on Windows](docs/install/windows.md#install-pre-built-msi) |
| Linux | AppImage, x86_64 with glibc 2.39+ | [Install on Linux](docs/install/linux.md) |
| Docker | CUDA, ROCm, CPU, and worker-only GPU profiles | [Run with Docker](docs/install/docker.md) |
<img width="2628" height="1950" alt="VoiceStudio desktop workspace" src="https://github.com/user-attachments/assets/b474497d-a453-49a3-a2dd-f023ec6b7659" />
First launch creates a managed Python environment and downloads the default model. Later launches reuse both.
</details>
> [!NOTE]
> On macOS, first launch needs a one-time right-click, then **Open** approval. Intel Macs cannot run the local Python backend; use a [remote backend](docs/install/macos.md) instead.
## Get started
### Quick Docker run
Download from [Releases](https://github.com/debpalash/VoiceStudio/releases/latest), then follow your platform guide:
```bash
docker run -d -p 127.0.0.1:3900:3900 -v omnivoice-data:/app/omnivoice_data --name voicestudio palashdeb/omnivoice-studio:stable
```
**[macOS](docs/install/macos.md) · [Windows](docs/install/windows.md) · [Linux](docs/install/linux.md) · [Docker](docs/install/docker.md)**
### First voice
Open **Voice cloning**, choose a voice or add a clean reference recording, enter your text, and generate. Install the required model when prompted. Hardware needs vary by engine; see [performance](docs/performance.md).
1. Launch VoiceStudio and open **Voice Cloning**.
2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
3. Enter text, choose a language, then select **Generate**.
> [!TIP]
> **Try without installing:** Run VoiceStudio in the cloud via the [Google Colab notebook](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb). Explore audio quality comparisons in [benchmarks](docs/benchmarks.md) and prompt design tips in [expressive speech](docs/expressive-speech.md).
### Audio samples
Listen to sample outputs produced locally with VoiceStudio:
| Workflow | Prompt / Reference Audio | Generated Audio |
|---|---|---|
| **Voice Cloning** | [demo_voice.wav](backend/assets/samples/demo_voice.wav) | [demo_clone_output.wav](backend/assets/samples/demo_clone_output.wav) |
| **Voice Design** (US News Anchor) | *"Clear, authoritative American broadcast tone"* | [demo_voice_design_us_news_anchor.wav](backend/assets/samples/voice_design/demo_voice_design_us_news_anchor.wav) |
| **Voice Design** (UK Audiobook) | *"Warm, expressive British storytelling voice"* | [demo_voice_design_audiobook_uk_narrator.wav](backend/assets/samples/voice_design/demo_voice_design_audiobook_uk_narrator.wav) |
| **Video Dubbing** (Multilingual) | [source.src.wav](backend/assets/samples/demo/dubbing/source.src.wav) | [Spanish](backend/assets/samples/demo/dubbing/dubbed_es.src.wav) · [French](backend/assets/samples/demo/dubbing/dubbed_fr.src.wav) · [Japanese](backend/assets/samples/demo/dubbing/dubbed_ja.src.wav) · [Chinese](backend/assets/samples/demo/dubbing/dubbed_zh.src.wav) |
### Run from source
Install the [development prerequisites](.github/CONTRIBUTING.md#development-setup) (Node 20+/Bun and Python 3.11+), then:
<details>
<summary><strong>Run the Electron preview from source</strong></summary>
```bash
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop
bun run dev
```
The desktop launcher configures Python dependencies on first run via `uv` automatically. Use `bun run dev` for the browser UI. See [Contributing](.github/CONTRIBUTING.md) for services, tests, and platform packages.
See [Electron setup](electron/README.md) for prerequisites and backend configuration.
### If setup fails
</details>
- Run **Settings → About → Run self-check** or `uv run python backend/main.py --diagnose --deep`.
- Check [install troubleshooting](docs/install/troubleshooting.md).
- Save a scrubbed diagnostic bundle from the app when opening an issue.
- For slow generation, compare [measured benchmarks](docs/benchmarks.md) and [performance settings](docs/performance.md).
<a id="features"></a>
## Features
| Area | Included |
|---|---|
| **Voice Cloning** | Zero-shot synthesis from a short reference clip ([guide](docs/engines/README.md)) |
| **Voice Design** | Create a voice from age, accent, pitch, style, and delivery instructions ([expressive speech](docs/expressive-speech.md)) |
| **Video Dubbing** | Transcribe, translate, preserve speakers, synthesize, and export video ([export guide](docs/dubbing/export.md)) |
| **Stories and audiobooks** | Multi-voice scripts · EPUB/PDF import · chapter rendering · `.m4b` export |
| **[Dictation Widget](docs/features/dictation.md)** | System-wide shortcut, live transcription, optional local-LLM cleanup |
| **Vocal Isolation** | Demucs speech/background separation |
| **Speaker Diarization** | Pyannote and WhisperX speaker assignment ([guide](docs/features/diarization.md)) |
| **Batch Queue** | Queue large sets of audio and video jobs with per-job progress, or watch a local folder for new videos |
| **Model Catalogue** | Install, remove, select, and route TTS, ASR, and LLM models ([catalogue](docs/engines/README.md)) |
| **Remote Model Downloads** | Install models on enrolled remote workers with live progress ([guide](docs/downloading-models.md)) |
| **GPU Auto-Detect** | CUDA, MPS, ROCm, and CPU routing with per-engine checks ([performance](docs/performance.md)) |
| **AI Watermark** | AudioSeal embedding and detection |
| **MCP Server** | Synthesis and transcription tools for MCP clients ([guide](docs/mcp.md)) |
| **Diagnostics** | Self-checks, error journal, logs, and scrubbed support bundles ([troubleshooting](docs/install/troubleshooting.md)) |
| **Local-first** | Core creation stays local; network-backed features are explicit opt-ins |
| **Extensible** | Registry-based TTS, ASR, and plugin interfaces ([acceptance](docs/engine-acceptance.md)) |
<table>
<tr>
<td width="50%"><img src="docs/media/0.5.0/catalogue.png" alt="VoiceStudio Model Catalogue" width="100%" /></td>
<td width="50%"><img src="docs/media/0.5.0/gallery-save.png" alt="Saving a gallery voice as a local profile" width="100%" /></td>
</tr>
<tr>
<td align="center"><sub>Model Catalogue: engine, device, and install state</sub></td>
<td align="center"><sub>Gallery: save a shared voice as a local profile</sub></td>
</tr>
</table>
<a id="comparison"></a>
## Comparison
VoiceStudio trades managed cloud compute for local control. This is the practical difference:
| | **VoiceStudio** | **Typical hosted voice service** |
|---|---|---|
| **Best fit** | Private, offline, self-hosted, or high-volume work | Fast setup without local model management |
| **Data path** | Local by default; remote features are opt-in | Audio and text are processed by the provider |
| **Cost model** | Free software; you supply the hardware | Subscription, credits, or metered API use |
| **Setup** | Install the app and model weights | Create an account and use the web app or API |
| **Performance** | Depends on your engine and hardware | Provider manages compute and scaling |
| **Offline use** | Yes, after required models are installed | Usually requires a network connection |
| **Customization** | Source, engines, models, API, and routing are open | Limited to provider options |
| **Maintenance** | You manage updates, disk, and compute | Provider manages infrastructure |
<a id="requirements"></a>
## Requirements
Requirements vary by engine. These values cover the default local workflow.
| | **Minimum** | **Recommended** |
|---|---|---|
| **OS** | Windows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+ | Current supported OS release |
| **RAM** | 8 GB | 16 GB+ |
| **Disk** | 10 GB free | 20 GB+ SSD |
| **GPU** | Optional; CPU mode is supported | NVIDIA CUDA or Apple Silicon |
| **VRAM** | 4 GB when using a GPU | 8 GB+; large optional engines need more |
| **Python from source** | 3.11+ | 3.11 or 3.12 |
ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See [performance](docs/performance.md), [benchmarks](docs/benchmarks.md), and [engine disk usage](docs/engines/disk-usage.md).
<a id="hardware-recommendations"></a>
### Recommended stack by hardware
| Hardware | Recommended TTS | Recommended ASR | Why |
|---|---|---|---|
| **Apple Silicon (M1M4)** | [MLX-Audio](docs/engines/mlx-audio.md) · [OmniVoice](docs/engines/omnivoice.md) (MPS) | [MLX Whisper](docs/engines/mlx-whisper.md) · [Parakeet MLX](docs/engines/parakeet-mlx.md) | Native unified memory, lowest latency on macOS |
| **NVIDIA GPU (8 GB+ VRAM)** | [OmniVoice](docs/engines/omnivoice.md) · [CosyVoice 3](docs/engines/cosyvoice.md) | [WhisperX](docs/engines/whisperx.md) | High-fidelity zero-shot cloning, word timestamps, diarization |
| **Low VRAM / CPU-only** | [PocketTTS](docs/engines/pockettts.md) · [Sherpa-ONNX](docs/engines/sherpa-onnx.md) · [KittenTTS](docs/engines/kittentts.md) | [Moonshine](docs/engines/moonshine.md) · [Faster-Whisper](docs/engines/faster-whisper.md) (`int8`) | Low memory footprint, optimized CPU inference |
<a id="engines"></a>
## Engines
Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: [docs/engines](docs/engines/README.md).
<a id="tts-engines"></a>
### Text to speech
| Engine | Languages | Clone | Instruct | Linux | macOS ARM | Windows | License |
|---|:---:|:---:|:---:|:---:|:---:|:---:|---|
| [**VoiceStudio** (default, powered by k2-fsa/OmniVoice)](docs/engines/omnivoice.md) | 600+ | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0 code, CC-BY-NC weights](https://huggingface.co/k2-fsa/OmniVoice#license)³ |
| [**CosyVoice 3**](docs/engines/cosyvoice.md) | 9 + 18 dialects | Yes | Yes | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| [**GPT-SoVITS**](docs/engines/gpt-sovits.md) | 5 | Yes | No | CUDA/CPU | No | CUDA/CPU | MIT |
| [**VoxCPM2**](docs/engines/voxcpm2.md) | 30 | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | Apache-2.0 |
| [**MOSS-TTS-Nano**](docs/engines/moss-tts-nano.md) | 20 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| [**KittenTTS**](docs/engines/kittentts.md) | English | No | No | CPU | CPU | CPU | MIT |
| [**MLX-Audio**](docs/engines/mlx-audio.md) | Model-dependent | Varies | Varies | No | MLX | No | Varies |
| [**Sherpa-ONNX**](docs/engines/sherpa-onnx.md) | 20+ | No | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| [**IndexTTS 2.5** ⚡](docs/engines/indextts.md) | ZH · EN · JA · ES · AR | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Bilibili model license¹ |
| [**OmniVoice GGUF** ⚡](docs/engines/omnivoice-gguf.md) | 600+ | Yes | Yes | CUDA/CPU | MPS/CPU | CUDA/CPU | [AGPL-3.0](LICENSE) app · [review the derivative model terms](https://huggingface.co/Serveurperso/OmniVoice-GGUF#license)³ |
| [**OmniVoice (subprocess)** ⚡](docs/engines/omnivoice-subprocess.md) | 600+ | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0 code, CC-BY-NC weights](https://huggingface.co/k2-fsa/OmniVoice#license)³ |
| [**PocketTTS** ⚡](docs/engines/pockettts.md) | EN · FR · DE · PT · IT · ES | Yes | No | CPU | CPU | CPU | CC-BY-4.0, gated² |
| [**Supertonic 3** ⚡](docs/engines/supertonic3.md) | 31 | No | No | CPU | CPU | CPU | OpenRAIL-M |
| [**MOSS-TTS-v1.5** ⚡](docs/engines/moss-tts-v15.md) | 31 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| [**dots.tts** ⚡](docs/engines/dots-tts.md) | 24 | Yes | No | CUDA/CPU | CPU | No | Apache-2.0 |
| [**Confucius4-TTS** ⚡](docs/engines/confucius4-tts.md) | 14 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
⚡ Installed or registered on demand.
¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the [model license](https://huggingface.co/IndexTeam/IndexTTS-2.5/blob/main/LICENSE).
² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.
³ The OmniVoice snapshot also includes an audio tokenizer under separate [Boson Higgs Audio 2 and Meta Llama community terms](https://huggingface.co/k2-fsa/OmniVoice/blob/main/audio_tokenizer/LICENSE). VoiceStudio's application license does not replace model or tokenizer terms.
Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.
<a id="asr-engines"></a>
### Speech to text
| Engine | ID | Languages | Best fit |
|---|---|:---:|---|
| [**WhisperX** (default)](docs/engines/whisperx.md) | `whisperx` | ~100 | Dubbing, subtitles, word-level timing |
| [**Faster-Whisper**](docs/engines/faster-whisper.md) | `faster-whisper` | ~100 | General cross-platform transcription |
| [**Faster-Whisper (isolated)**](docs/engines/faster-whisper-isolated.md) | `faster-whisper-isolated` | ~100 | Crash-isolated batch transcription |
| [**MLX Whisper**](docs/engines/mlx-whisper.md) | `mlx-whisper` | ~100 | Apple Silicon |
| [**PyTorch Whisper**](docs/engines/pytorch-whisper.md) | `pytorch-whisper` | ~100 | CUDA, MPS, and CPU fallback |
| [**Parakeet TDT**](docs/engines/nemo-parakeet.md) | `nemo-parakeet` | English + 25 EU | Fast CPU/CUDA transcription |
| [**Parakeet TDT v3 (MLX)**](docs/engines/parakeet-mlx.md) | `parakeet-mlx` | 25 EU | Apple Silicon dictation and word timestamps |
| [**Moonshine**](docs/engines/moonshine.md) | `moonshine` | English | Low-power, low-latency ONNX |
| [**FunASR**](docs/engines/funasr.md) | `funasr` | 50+ | VAD and inline diarization |
| [**sherpa-onnx** (live dictation)](docs/engines/sherpa-onnx-asr.md) | `sherpa-onnx-asr` | Model-dependent | Streaming CPU dictation |
| [**OpenAI-compatible** ⚠️ configured server](docs/engines/openai-compatible-asr.md) | `openai-compat-asr` | Server-dependent | Local gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server |
WhisperX and Faster-Whisper retry with `int8` when efficient `float16` is unavailable. Pin `ASR_COMPUTE_TYPE=int8` or `float32` only if automatic selection still fails.
<a id="architecture"></a>
## Architecture
```text
Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
```
| Layer | Path | Responsibility |
|---|---|---|
| Desktop shell | `frontend/src-tauri/` | Window lifecycle, tray, shortcuts, updater, sidecar bootstrap |
| Frontend | `frontend/src/` | React UI, Zustand state, API and event clients, i18n |
| API | `backend/api/` | REST routes, schemas, auth boundaries, streaming |
| Core services | `backend/services/` | Generation, dubbing, audio processing, persistence |
| Engines | `backend/engines/` | Isolated and optional engine adapters |
| Worker system | `backend/worker/` | Authenticated remote compute and job transport |
| Data | `omnivoice_data/` | Projects, voices, settings, logs, and SQLite state |
| Delivery | `scripts/`, `deploy/`, `.github/workflows/` | Development, packaging, containers, releases, CI |
### Network boundary
- The desktop talks to a loopback-only backend on `localhost:3900`.
- Loopback API calls need no server key. Remote access requires a share PIN or API key.
- Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
- Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.
<a id="api"></a>
## Local speech platform and OpenAI-compatible API
Point an OpenAI-compatible audio client at the local backend:
```diff
- base_url="https://api.openai.com/v1"
+ base_url="http://localhost:3900/v1"
```
| Endpoint | Purpose |
|---|---|
| `POST /v1/audio/speech` | TTS to `mp3`, `opus`, `aac`, `flac`, `wav`, or `pcm`; select a profile with `voice` and an engine with `model` |
| `POST /v1/audio/transcriptions` | STT to `json`, `text`, `verbose_json`, `srt`, or `vtt` |
| `WS /v1/audio/transcriptions/stream` | Live PCM/WebM transcription with partial, utterance, and session-final events |
| `GET /.well-known/voicestudio-speech` | Discover HTTP, WebSocket, MCP, and native dictation-control transports |
| `GET /v1/audio/voices` | List local voice profiles and engines |
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:3900/v1", api_key="local")
with client.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) as response:
response.stream_to_file("speech.wav")
```
```bash
# Quick test via cURL
curl http://localhost:3900/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"model": "tts-1", "input": "Made on my own hardware.", "voice": "default", "response_format": "wav"}' \
--output speech.wav
```
The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps,
and TUIs trigger the system-wide dictation flow or reuse its native text
insertion. See the [speech platform guide](docs/speech-platform.md). The full API
reference is in **Settings → OpenAPI Reference**. For LAN, Tailscale, or proxy
access, read [API authentication](docs/api-auth.md) before exposing the backend.
### Agent skills
Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other [skills.sh](https://skills.sh)-compatible agents:
```bash
npx skills add debpalash/VoiceStudio
```
- `omnivoice`: synthesize speech and transcribe audio through local VoiceStudio.
- `oss-maintainer`: the repository's open-source maintenance workflow.
### Model Context Protocol (MCP)
VoiceStudio mounts an MCP server at `http://localhost:3900/mcp` for Claude Desktop, Cursor, and AI agents:
```json
{
"mcpServers": {
"voicestudio": {
"url": "http://localhost:3900/mcp"
}
}
}
```
For clients requiring stdio transport, use the bundled local shim (`docs/mcp.json`):
```json
{
"mcpServers": {
"voicestudio": {
"command": "python",
"args": ["-m", "backend.mcp_shim"],
"cwd": "/path/to/VoiceStudio"
}
}
}
```
See the [MCP guide](docs/mcp.md) for tools (`generate_speech`, `clone_voice`, `transcribe`), file streaming modes, and client bindings.
### Google Colab
[![Open in Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb)
The [notebook](notebooks/OmniVoice_Studio_Colab.ipynb) runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.
<a id="documentation"></a>
> **Electron is the primary desktop app.** The next desktop release ships Electron, with one final Tauri sunset update. Bug reports and contributions remain welcome; include the app version and whether you use Electron or Tauri.
## Documentation
| Need | Read |
| Need | Start here |
|---|---|
| Install | [macOS](docs/install/macos.md) · [Windows](docs/install/windows.md) · [Linux](docs/install/linux.md) · [Docker](docs/install/docker.md) |
| Fix setup | [Troubleshooting](docs/install/troubleshooting.md) · [model downloads](docs/downloading-models.md) · [Hugging Face token](docs/setup/huggingface-token.md) |
| Choose an engine | [Engine guides](docs/engines/README.md) · [benchmarks](docs/benchmarks.md) · [expressive speech](docs/expressive-speech.md) |
| Tune hardware | [Performance](docs/performance.md) · [remote workers](docs/remote-workers.md) |
| Build integrations | [Speech platform](docs/speech-platform.md) · [Private production API](docs/production-private-api.md) · [API auth](docs/api-auth.md) · [MCP](docs/mcp.md) · [examples](examples/README.md) |
| Build VoiceStudio | [Contributing](.github/CONTRIBUTING.md) · [engine acceptance](docs/engine-acceptance.md) |
| Track changes | [Changelog](CHANGELOG.md) · [roadmap](docs/ROADMAP.md) · [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) |
| Remove everything | [Uninstall guide](docs/install/uninstall.md) |
| Setup help | [Troubleshooting](docs/install/troubleshooting.md) · [Model downloads](docs/downloading-models.md) |
| Models & audio quality | [Engine guides](docs/engines/README.md) · [Benchmarks](docs/benchmarks.md) |
| Integrations | [Local API](docs/speech-platform.md) · [MCP](docs/mcp.md) · [Examples](examples/README.md) |
| Development | [Contributing](.github/CONTRIBUTING.md) · [Electron](electron/README.md) · [Changelog](CHANGELOG.md) |
<a id="faq"></a>
Agent skills: `npx skills add debpalash/VoiceStudio` — choose **voicestudio** for audio workflows or **voicestudio-maintainer** for repository maintenance.
## FAQ
## Sponsors
<details>
<summary><strong>Does it work on Apple Silicon and Intel Macs?</strong></summary>
<a href="https://forms.gle/2PYCvd39hbwijzX37"><img src="docs/media/sponsor-slot.svg" alt="Your brand — apply for a featured VoiceStudio sponsor slot" width="640" /></a>
Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See [macOS installation](docs/install/macos.md).
</details>
**Become a featured partner.** [Apply for a paid placement](https://forms.gle/2PYCvd39hbwijzX37) · [Email us](mailto:partner@voicestudio.sh)
<details>
<summary><strong>How much VRAM do I need?</strong></summary>
Support development: [Ko-fi](https://ko-fi.com/debpalash) · [PayPal](https://paypal.me/palashCoder) · [Sponsorship details](SPONSORS.md)
A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the [benchmarks](docs/benchmarks.md) and engine guide.
</details>
## License & responsible use
<details>
<summary><strong>Why does a longer reference clip not always improve the clone?</strong></summary>
Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see [data preparation](docs/data_preparation.md) and [training](docs/training.md).
</details>
<details>
<summary><strong>Can I use generated audio commercially?</strong></summary>
VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.
</details>
<details>
<summary><strong>Does VoiceStudio collect data?</strong></summary>
Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at **Settings → Privacy**.
</details>
<details>
<summary><strong>How do I remove VoiceStudio and its data?</strong></summary>
Use `scripts/uninstall.sh` on macOS/Linux or `scripts\uninstall.ps1` on Windows. Both show a dry run before deletion. See the [uninstall guide](docs/install/uninstall.md) for every path.
</details>
## Community and contributing
- [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues) for reproducible bugs and feature requests.
- [Discord](https://discord.gg/bzQavDfVV9) for setup help and project discussion.
- [Good first issues](https://github.com/debpalash/VoiceStudio/labels/good%20first%20issue) for a scoped starting point.
- [Contributing guide](.github/CONTRIBUTING.md) for setup, tests, and pull requests.
<p align="center">
<a href="https://star-history.com/#debpalash/VoiceStudio&Date">
<img src="https://api.star-history.com/svg?repos=debpalash/VoiceStudio&type=Date" alt="Star History Chart" width="100%" />
</a>
</p>
## Support development
VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.
[Ko-fi](https://ko-fi.com/debpalash) · [PayPal](https://paypal.me/palashCoder) · [Sponsorship details](SPONSORS.md)
## Responsible use and safety
VoiceStudio enables zero-shot voice cloning and speech generation on personal hardware. Please use it responsibly:
- **Consent:** Only clone or synthesize voices with explicit permission from the speaker.
- **Audio provenance:** VoiceStudio integrates [AudioSeal](https://github.com/facebookresearch/audioseal) imperceptible watermarking by default to detect and identify synthetic speech without altering sound quality.
- **Local privacy:** For the default local workflow, audio recordings, transcripts, voices, and projects remain strictly on your local disk; data leaves your device only when you explicitly configure remote workers or external ASR endpoints.
## License
VoiceStudio is licensed under [AGPL-3.0](LICENSE). You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact **VoiceStudio@palash.dev**. See [LICENSE-NOTICE.md](LICENSE-NOTICE.md) for the plain-language scope.
Optional engines and downloaded models retain their own licenses. The bundled `omnivoice/` Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.
## Acknowledgments
VoiceStudio builds on [OmniVoice](https://github.com/k2-fsa/OmniVoice), [WhisperX](https://github.com/m-bain/whisperX), [Demucs](https://github.com/facebookresearch/demucs), [Pyannote](https://github.com/pyannote/pyannote-audio), [CTranslate2](https://github.com/OpenNMT/CTranslate2), [AudioSeal](https://github.com/facebookresearch/audioseal), [Tauri](https://tauri.app), [Supertonic](https://huggingface.co/Supertone/supertonic-3), [Sherpa-ONNX](https://github.com/k2-fsa/sherpa-onnx), [GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS), and [PocketTTS](https://kyutai.org).
<div align="center">
<strong><a href="https://github.com/debpalash/VoiceStudio/releases/latest">Download VoiceStudio</a></strong> ·
<a href="https://github.com/debpalash/VoiceStudio">Star the project</a> ·
<a href="https://discord.gg/bzQavDfVV9">Join Discord</a>
</div>
[AGPL-3.0](LICENSE). Models have their own licenses; review them before commercial use. Clone voices only with permission. See [license details](LICENSE-NOTICE.md).
+53 -669
View File
@@ -1,697 +1,81 @@
*本文档是 [README.md](README.md) 的简体中文翻译;若与英文版有出入,以英文版为准。*
<div align="center">
<img src="docs/logo.png" alt="VoiceStudio 徽标" width="120" height="120" />
<img src="docs/logo.png" alt="VoiceStudio" width="88" />
<h1>VoiceStudio</h1>
<p><sub><em>原名 OmniVoice-Studio</em></sub></p>
<h3>创造声音,讲述故事,文件始终属于你。♡</h3>
<p>在一个开源桌面工作室里完成克隆、设计、配音、听写和有声书制作。<br/><b>默认本地优先。</b>没有订阅,也没有用量计费;联网服务始终由你主动选择。</p>
<p><strong>开源声音克隆与工作流引擎。在本地构建。</strong></p>
<p>使用本地 AI 克隆声音、翻译配音、语音听写和制作有声书。</p>
<p>
<a href="#quickstart">快速开始</a> ·
<a href="#features">功能</a> ·
<a href="#why-voicestudio">为什么选择 VoiceStudio</a> ·
<a href="#tts-engines">引擎</a> ·
<a href="#openai-api">API</a> ·
<a href="#sponsor--donate">捐赠</a> ·
<a href="#contributing">参与贡献</a> ·
<a href="https://github.com/debpalash/VoiceStudio/releases/latest">下载</a> ·
<a href="#开始使用">开始使用</a> ·
<a href="#文档">文档</a> ·
<a href="https://discord.gg/bzQavDfVV9">Discord</a> ·
<a href="README.md"><strong>English</strong></a>
</p>
<p>
<a href="https://github.com/debpalash/VoiceStudio/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/debpalash/VoiceStudio/ci.yml?branch=main&style=flat-square&label=CI" alt="CI 状态" /></a>
<a href="https://github.com/debpalash/VoiceStudio/stargazers"><img src="https://img.shields.io/github/stars/debpalash/VoiceStudio?style=flat-square&color=f59e0b" alt="Star 数" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/github/v/release/debpalash/VoiceStudio?style=flat-square&color=10b981" alt="版本" /></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/license-AGPL--3.0-blue?style=flat-square" alt="许可证" /></a>
<a href="https://github.com/debpalash/VoiceStudio/issues"><img src="https://img.shields.io/github/issues/debpalash/VoiceStudio?style=flat-square&color=ef4444" alt="Issues" /></a>
<a href="https://discord.gg/bzQavDfVV9"><img src="https://img.shields.io/badge/Discord-Join_Community-5865F2?style=flat-square&logo=discord&logoColor=white" alt="Discord" /></a>
<a href="https://ko-fi.com/debpalash"><img src="https://img.shields.io/badge/Ko--fi-Support_Us-FF5E5B?style=flat-square&logo=ko-fi&logoColor=white" alt="Ko-fi" /></a>
<a href="https://paypal.me/palashCoder"><img src="https://img.shields.io/badge/PayPal-Donate-00457C?style=flat-square&logo=paypal&logoColor=white" alt="PayPal" /></a>
</p>
<p>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/⬇_Download-macOS_·_Windows_·_Linux-10b981?style=for-the-badge" alt="下载最新版本" /></a>
<a href="README.md">English</a>
</p>
</div>
<br/>
![Electron 应用演示:声音克隆、声音设计、视频配音和模型管理](docs/media/electron/voicestudio.gif)
<div align="center">
<img src="docs/media/0.5.0/quick-switch.gif" alt="VoiceStudio — 从状态栏快速切换 TTS 引擎" width="100%"/>
</div>
<p align="center"><sub>新 Electron 桌面界面,使用此分支及内置演示声音录制。正式发布版本的界面可能有所不同。</sub></p>
> **声音很私人,创作空间也应该真正属于你。** VoiceStudio 的核心流程运行在你的硬件上:克隆、设计、配音、听写,并以 646 种语言创作,不需要订阅,也没有用量计费。联网引擎和服务始终是清晰可见的可选项,而不是隐藏依赖。
## 用 VoiceStudio 创作
> [!WARNING]
> **活跃 Beta 阶段。** 各版本之间可能出现故障——如需最新修复,请从源码运行。非常欢迎 Bug 报告和 PR[提交 Issue](https://github.com/debpalash/VoiceStudio/issues) 或 [加入 Discord](https://discord.gg/bzQavDfVV9)
- **声音克隆与设计**:上传参考录音,或用文字描述你想要的声音。
- **视频配音**:转录、翻译、分配说话人,并编辑语音时间轴
- **语音听写**:通过悬浮录音组件录制、转录和复制文字。
- **长篇创作**:制作多角色脚本、有声书和批量任务。
- **模型管理**:选择语音合成与转录引擎、语言及计算设备。
<a id="quickstart"></a>
本地工作流在你的硬件上运行。远程服务为可选功能;使用情况分析须经同意才会启用。
## ⚡ 快速开始
<table>
<tr>
<td><img src="docs/media/electron/voice-cloning.png" alt="Electron 声音克隆工作区与内置演示声音" width="100%" /></td>
<td><img src="docs/media/electron/dubbing.png" alt="Electron 视频配音工作区" width="100%" /></td>
</tr>
<tr><td align="center">声音克隆</td><td align="center">视频配音</td></tr>
<tr>
<td><img src="docs/media/electron/voice-design.png" alt="Electron 声音设计工作区" width="100%" /></td>
<td><img src="docs/media/electron/models.png" alt="本地语音模型管理" width="100%" /></td>
</tr>
<tr><td align="center">声音设计</td><td align="center">本地模型</td></tr>
</table>
<div align="center">
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/macOS-DMG_(Apple_Silicon)-000?style=for-the-badge&logo=apple&logoColor=white" alt="下载 macOS DMG" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Windows-MSI_(x64)-0078D4?style=for-the-badge&logo=windows&logoColor=white" alt="下载 Windows MSI" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Linux-AppImage_(x64)-FCC624?style=for-the-badge&logo=linux&logoColor=black" alt="下载 Linux AppImage" /></a>
<br/>
<sub>三个按钮都会打开最新发布页——在资源列表中下载对应你系统的安装包。</sub><br/>
<sub><b>macOS</b>首次启动需要一次性批准——右键点击 → <b>打开</b>macOS 15 上为 系统设置 → 隐私与安全性 → <b>“仍要打开”</b>)。无需终端。<a href="docs/install/macos.md#gatekeeper-quarantine">为什么?</a> · <b>Intel Mac</b>不支持本地后端(<a href="https://github.com/debpalash/VoiceStudio/issues/889">#889</a>)——<a href="docs/install/macos.md">详情</a>。</sub>
</div>
## 开始使用
选择你的操作系统,按指南从头到尾操作
从 [Releases](https://github.com/debpalash/VoiceStudio/releases/latest) 下载,然后阅读对应平台的安装指南
- 🍎 **macOS** — [docs/install/macos.md](docs/install/macos.md)
- 🪟 **Windows** — [docs/install/windows.md](docs/install/windows.md)
- 🐧 **Linux** — [docs/install/linux.md](docs/install/linux.md)
- 🐳 **Docker** — [docs/install/docker.md](docs/install/docker.md) · [Docker Hub: `palashdeb/omnivoice-studio`](https://hub.docker.com/r/palashdeb/omnivoice-studio)
**[macOS](docs/install/macos.md) · [Windows](docs/install/windows.md) · [Linux](docs/install/linux.md) · [Docker](docs/install/docker.md)**
打开声音克隆页面,选择已有声音或添加清晰的参考录音,输入文字并生成。按提示安装所需模型。硬件要求因引擎而异,详见[性能指南](docs/performance.md)
**从源码运行 Electron 预览版:**
```bash
# Docker 快速运行 (CPU / 本地环回模式)
docker run -d -p 127.0.0.1:3900:3900 -v omnivoice-data:/app/omnivoice_data --name voicestudio palashdeb/omnivoice-studio:stable
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
cd electron
bun run dev
```
**三步克隆出你的第一个声音:**
环境要求和后端配置见 [Electron 开发指南](electron/README.md)。项目仍在积极开发中,可通过 [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues) 反馈问题。
1. **安装并启动。** 首次启动会自动搭建 Python 运行环境并下载模型权重——启动画面会逐步显示进度(仅首次,需要几分钟;之后即开即用)。
2. 从启动台打开**语音克隆**,拖入任意声音的 **3 秒音频**
3. **输入一句话,点击生成。** 音频在你的设备上生成并保存,支持 646 种语言(商业使用前请审阅所选模型与分词器的许可条款)。
## 文档
### 🎧 音频示例
在线试听 VoiceStudio 本地生成的实际音频样例:
| 工作流 | 提示词 / 参考音频 | 生成音频 |
|---|---|---|
| **声音克隆** | [demo_voice.wav](backend/assets/samples/demo_voice.wav) | [demo_clone_output.wav](backend/assets/samples/demo_clone_output.wav) |
| **声音设计** (美语新闻主播) | *"清晰、权威的美国广播级音色"* | [demo_voice_design_us_news_anchor.wav](backend/assets/samples/voice_design/demo_voice_design_us_news_anchor.wav) |
| **声音设计** (英式有声书) | *"温暖生动的英式故事讲述音色"* | [demo_voice_design_audiobook_uk_narrator.wav](backend/assets/samples/voice_design/demo_voice_design_audiobook_uk_narrator.wav) |
| **视频配音** (多语种) | [source.src.wav](backend/assets/samples/demo/dubbing/source.src.wav) | [西班牙语](backend/assets/samples/demo/dubbing/dubbed_es.src.wav) · [法语](backend/assets/samples/demo/dubbing/dubbed_fr.src.wav) · [日语](backend/assets/samples/demo/dubbing/dubbed_ja.src.wav) · [中文](backend/assets/samples/demo/dubbing/dubbed_zh.src.wav) |
觉得慢?[docs/performance.md](docs/performance.md) 讲清了生成时间到底花在哪里、有哪些调优开关,以及“它变慢了”的三个经典原因。各引擎/设备的实测数据见 [docs/benchmarks.md](docs/benchmarks.md)。
> 正在从 **[CorentinJ/Real-Time-Voice-Cloning](https://github.com/CorentinJ/Real-Time-Voice-Cloning)**(现已归档)迁移过来?我们有专门的迁移指南:[docs/migration/real-time-voice-cloning.md](docs/migration/real-time-voice-cloning.md)。
<details>
<summary><b>🧰 卡住了?自检、Token 与受限网络</b></summary>
<br/>
先运行内置自检——在应用中打开 **设置 → 关于 → “运行自检”**,或在源码检出目录中执行
`uv run python backend/main.py --diagnose`(加 `--deep` 还会实际加载当前引擎进行测试)。然后查看
[docs/install/troubleshooting.md](docs/install/troubleshooting.md) 中排名前
10 的安装错误。运行时出错时,应用内的错误界面会直接深链到对应条目;**设置 → 关于 →
“保存诊断包”** 会把脱敏日志与自检报告打包,方便附在 Bug 报告里。
Hugging Face Token 的配置见
[docs/setup/huggingface-token.md](docs/setup/huggingface-token.md)。说话人分离相关的模型访问门槛见
[docs/features/diarization.md](docs/features/diarization.md)。下载速度、⚡ 快速下载(Xet)状态,以及受限网络 / 镜像选项见
[docs/downloading-models.md](docs/downloading-models.md)。
</details>
---
<a id="features"></a>
## ✨ 功能
八大主打功能——折叠区里还有十二项等你展开。
<table>
<tr>
<td align="center" width="25%">
<h3>🎙️ 语音克隆</h3>
<p>3 秒音频 → 复刻任何声音。<br/><b>646 种语言</b>,零样本。</p>
</td>
<td align="center" width="25%">
<h3>🎨 声音设计</h3>
<p>性别、年龄、口音、音高、语速、<br/>情感、方言——<b>随心调节</b>。</p>
</td>
<td align="center" width="25%">
<h3>🎬 视频配音</h3>
<p>YouTube 链接或文件 → 转录 →<br/>翻译 → 重新配音 → <b>MP4</b>。</p>
</td>
<td align="center" width="25%">
<h3>📖 有声书编辑器</h3>
<p>导入文本、EPUB 或 PDF。自动分章、<br/>响度归一、元数据。导出 <b>.m4b</b>。</p>
</td>
</tr>
<tr>
<td align="center" valign="top">
<h3>🎭 故事模式</h3>
<p>多声音编辑器。逐行分配声音、<br/>预览、<b>导出完整配音阵容</b>。</p>
</td>
<td align="center" valign="top">
<h3>⌨️ 听写工具</h3>
<p>在<b>任何应用</b>中按 <kbd>⌘</kbd>+<kbd>⇧</kbd>+<kbd>Space</kbd>。<br/>转录、自动粘贴、随即消失。</p>
</td>
<td align="center" valign="top">
<h3>🔐 本地优先</h3>
<p>核心创作流程<br/><b>留在你的设备上</b>。</p>
</td>
<td align="center" valign="top">
<h3>🤖 MCP 服务器</h3>
<p>从 <b>Claude</b>、Cursor 或<br/>任何 MCP 客户端使用 VoiceStudio。</p>
</td>
</tr>
</table>
<details>
<summary><b>……还有 12 项</b>——人声分离、说话人分离、批量处理、水印、诊断等等</summary>
<br/>
- 🔊 **人声分离** — 基于 Demucs:把语音从音乐中分离出来,同时保留背景音床。
- 👥 **说话人分离** — Pyannote + WhisperX 自动识别谁说了什么。
- 📦 **批量队列** — 拖入 50 个视频就可以走开;每个任务都有独立进度条。
- 🛡️ **AI 水印** — AudioSeal(Meta):不可见,且能在压缩后留存。
- 🔬 **诊断** — 自检套件、错误日志、脱敏诊断包。
-**GPU 自动检测** — CUDA · MPS · ROCmLinux,需手动开启)· CPU;显存 ≤8 GB 时自动卸载。
- 🧭 **引擎路由** — 逐引擎 GPU 预检;绝不静默回退到 CPU。
- 🧩 **可扩展** — 继承 `TTSBackend`,约 50 行代码即可接入任意引擎。
- 🎒 **便携声音角色** — 将声音导出为 `.ovsvoice` 包:身份 + 水印。
- ♾️ **无限长 TTS** — 按句分块生成,没有长度上限,可经 WebSocket 流式输出。
- 🌐 **远程后端** — 让 UI 指向远程服务器;对 Tailscale 友好,支持 Bearer 认证。
- 🧠 **听写 + LLM** — 用本地 LLM 润色转录文本,可选回声消除。
</details>
---
<a id="why-voicestudio"></a>
## 💡 为什么选择 VoiceStudio
云端语音工具很方便,但工作流会依赖账号、用量计费和他人的基础设施。VoiceStudio 在你的硬件上提供完整工作室;只有你主动选择时,才会使用联网集成。
| | **ElevenLabs** | **VoiceStudio** |
|---|---|---|
| **价格** | 订阅与用量限制 | 免费且开源(AGPL-3.0)· 专有用途可选 [商业许可证](#license) |
| **语音克隆** | ✅ 3 秒音频 | ✅ 3 秒音频,零样本 |
| **声音设计** | ✅ 性别、年龄 | ✅ 性别、年龄、口音、音高、风格、方言 |
| **有声书 / 故事** | ❌ | ✅ 完整有声书编辑器 + 多声音故事(EPUB/PDF 导入,.m4b 导出) |
| **语言** | 取决于套餐和模型 | **646** |
| **视频配音** | ✅ 仅云端 | ✅ 完全本地 |
| **数据隐私** | 音频在远端处理 | 核心流程在本地运行;联网服务必须主动选择 |
| **API 密钥** | 需要账号 | 本地流程不需要 |
| **GPU 支持** | 不适用(云端) | CUDA · Apple Silicon · ROCmLinux)· CPU |
| **桌面应用** | ❌ | ✅ macOS · Windows · Linux |
| **TTS 引擎** | 1 | **16** — [完整矩阵](#tts-engines) |
| **ASR 引擎** | 1 | **11** — [完整阵容](#asr-engines) |
| **MCP 服务器** | ❌ | ✅ 可从 Claude、Cursor 及任何 MCP 客户端使用 |
| **自检** | ❌ | ✅ 诊断套件、错误日志、脱敏调试包 |
| **可定制** | ❌ 闭源 | ✅ 随你 Fork、扩展、发布 |
专业级语音 AI,去掉订阅,也去掉云端。
<div align="center">
<br/>
<b>心动了?来和我们一起构建吧。</b><br/>
<a href="https://discord.gg/bzQavDfVV9"><img src="https://img.shields.io/badge/Join_Discord-5865F2?style=for-the-badge&logo=discord&logoColor=white" alt="加入 Discord" /></a>
<br/><br/>
</div>
---
## 🖥️ 系统要求
| | **最低配置** | **推荐配置** |
|---|---|---|
| **操作系统** | Windows 10、macOS 12+Apple Silicon)、Ubuntu 24.04+glibc 2.39+ | 任意现代 64 位操作系统 |
| **内存** | 8 GB | 16 GB+ |
| **显存(GPU** | 4 GB(自动将 TTS 卸载到 CPU | 8 GB+NVIDIA RTX 3060+ |
| **硬盘** | 10 GB 可用空间(模型 + 缓存) | 20 GB+ SSD |
| **Python** | 3.10+(由 `uv` 管理) | 3.113.12 |
| **GPU** | 可选——CPU 也能跑 | NVIDIA CUDA · Apple Silicon MPS · AMD ROCm(仅 Linux |
> [!TIP]
> 对于显存 **≤8 GB** 的 GPUVoiceStudio 会在转录期间自动将 TTS 卸载到 CPU——无需配置。不需要专用 GPU;整条流水线都可以在 CPU 上运行(只是慢一些)。
> [!NOTE]
> **AMD GPU** ROCm 加速**仅限 Linux 且需手动开启**——在首次运行的设置界面选择 **“AMD GPU (ROCm)”**,或设置 `OMNIVOICE_TORCH_VARIANT=rocm`[docs/install/linux.md](docs/install/linux.md#amd-gpu-rocm))。在 **Docker/Podman** 中请改用专门的 ROCm 镜像:`ghcr.io/debpalash/omnivoice-studio:rocm`[docs/install/docker.md](docs/install/docker.md#pull-and-run-amd-gpu--rocm))。**在 Windows 上,AMD GPU(含 Ryzen AI 核显)只能以 CPU 运行**PyTorch 没有 Windows 版 ROCm 轮子,因此 Windows 上的 GPU 加速仅限 NVIDIA/CUDA[docs/install/windows.md](docs/install/windows.md#gpu-support))。
> [!IMPORTANT]
> **macOS Intelx86_64)不支持本地后端:** 应用 UI 可以安装,但 Python 后端无法运行,因为 PyTorch 已不再发布 Intel Mac 轮子([#889](https://github.com/debpalash/VoiceStudio/issues/889))。Intel Mac 用户仍可让 UI 指向另一台机器上的远程后端——参见 [docs/install/macos.md](docs/install/macos.md)。
<a id="hardware-recommendations"></a>
### 💡 按硬件推荐引擎配置
| 硬件配置 | 推荐 TTS 引擎 | 推荐 ASR 语音识别 | 优势 |
|---|---|---|---|
| **Apple Silicon (M1M4)** | [MLX-Audio](docs/engines/mlx-audio.md) · [OmniVoice](docs/engines/omnivoice.md) (MPS) | [MLX Whisper](docs/engines/mlx-whisper.md) · [Parakeet MLX](docs/engines/parakeet-mlx.md) | 原生统一内存,macOS 上延迟最低、性能最强 |
| **NVIDIA 显卡 (8 GB+ 显存)** | [OmniVoice](docs/engines/omnivoice.md) · [CosyVoice 3](docs/engines/cosyvoice.md) | [WhisperX](docs/engines/whisperx.md) | 极致零样本克隆品质、字级时间戳对齐与说话人分离 |
| **低显存 / 仅 CPU 设备** | [PocketTTS](docs/engines/pockettts.md) · [Sherpa-ONNX](docs/engines/sherpa-onnx.md) · [KittenTTS](docs/engines/kittentts.md) | [Moonshine](docs/engines/moonshine.md) · [Faster-Whisper](docs/engines/faster-whisper.md) (`int8`) | 超低内存占用,针对 CPU 指令集深度优化 |
<a id="tts-engines"></a>
### 🗣️ TTS 引擎
**16 个引擎,一个选择器。** VoiceStudio(默认,支持 600+ 语言)始终可用;另有七个引擎可选装并自动检测(CosyVoice 3、GPT-SoVITS、VoxCPM2、MOSS-TTS-Nano、KittenTTS、MLX-Audio、Sherpa-ONNX),外加八个按需延迟安装的引擎(IndexTTS 2.5、OmniVoice GGUF、OmniVoice 子进程版、PocketTTS、Supertonic 3、MOSS-TTS-v1.5、dots.tts、Confucius4-TTS)。在 **设置 → TTS 引擎** 中切换;所选引擎将应用于所有语音合成场景。**每个引擎都有独立指南:[docs/engines](docs/engines/README.md)(英文)。**
<details>
<summary><b>📊 完整矩阵</b>——16 个引擎 × 平台 × 克隆/指令 × 许可证</summary>
<br/>
| 引擎 | 语言 | 克隆 | 指令 | Linux | macOS ARM | Windows | 许可证 |
|--------|:---------:|:-----:|:--------:|:-----:|:---------:|:-------:|:-------:|
| **VoiceStudio**(默认,由 k2-fsa/OmniVoice 驱动) | 600+ | ✅ | ✅ | ✅ CUDA/CPU | ✅ MPS | ✅ CUDA/CPU | 内置 |
| **CosyVoice 3** | 9 + 18 种方言 | ✅ | ✅ | ✅ CUDA/CPU | ✅ MPS | ✅ CUDA/CPU | Apache-2.0 |
| **GPT-SoVITS** | 5 | ✅ | — | ✅ CUDA/CPU | — | ✅ CUDA/CPU | MIT |
| **VoxCPM2** | 30 | ✅ | ✅ | ✅ CUDA/CPU | ✅ MPS | ✅ CUDA/CPU | Apache-2.0 |
| **MOSS-TTS-Nano** | 20 | ✅ | — | ✅ CUDA/CPU | ✅ CPU | ✅ CUDA/CPU | Apache-2.0 |
| **KittenTTS** | 英语 | — | — | ✅ CPU | ✅ CPU | ✅ CPU | MIT |
| **MLX-Audio**Kokoro、Qwen3-TTS、CSM、Dia 等) | 多语言 | 因模型而异 | 因模型而异 | ❌ | ✅ 原生 | ❌ | 因模型而异 |
| **Sherpa-ONNX** | 20+ | — | — | ✅ CUDA/CPU | ✅ CPU | ✅ CUDA/CPU | Apache-2.0 |
| **IndexTTS 2.5** ⚡ | 中文 · 英语 · 日语 · 西班牙语 · 阿拉伯语 | ✅ | — | ✅ CUDA | — | ✅ CUDA | Bilibili 模型许可¹ |
| **OmniVoice GGUF** ⚡ | 600+ | ✅ | ✅ | ✅ CPU | ✅ CPU | ✅ CPU | 内置 |
| **Supertonic 3** ⚡ | 31 | — | — | ✅ CPU | ✅ CPU | ✅ CPU | OpenRAIL-M |
| **MOSS-TTS-v1.5** ⚡(8B | 31 | ✅ | — | ✅ CUDA/CPU | ✅ CPU | ✅ CUDA/CPU | Apache-2.0 |
| **dots.tts** ⚡(2B | 24 | ✅ | — | ✅ CUDA/CPU | ✅ CPU | ❌ | Apache-2.0 |
| **Confucius4-TTS** ⚡ | 14 | ✅ | — | ✅ CUDA/CPU | ✅ CPU | ✅ CUDA/CPU | Apache-2.0 |
¹ 若月活跃用户超过 1 亿,或年收入超过人民币 10 亿元,使用 IndexTTS 2.5
前必须另行取得 Bilibili 的书面许可。启用可选边车前,请审阅其
[模型许可](https://huggingface.co/IndexTeam/IndexTTS-2.5/blob/main/LICENSE)。
> **CUDA** = GPU 加速 · **MPS** = Apple Silicon Metal · **CPU** = 随处可运行,大模型较慢 · KittenTTS 和 MOSS-TTS-Nano 可在 CPU 上实时运行 · MLX-Audio 仅限 Apple Silicon · ⚡ = 延迟注册(首次使用时安装)
>
> **克隆**能力的意义不止于单段生成:视频配音(以及任何固定了声音的批量任务)需要参考音频克隆来保持说话人身份,因此把不支持克隆的引擎(KittenTTS、Sherpa-ONNX、Supertonic 3)设为当前引擎时,这些任务会在开始前就给出可操作的失败提示,而不是静默回退到 VoiceStudio。
>
> **MOSS-TTS-v1.5**8B,约 16 GB)、**dots.tts**2B,约 9 GB)和 **Confucius4-TTS** 是重量级可选引擎,从本地克隆在各自独立的 venv 中运行。三者均不支持 Apple Silicon MPS(在 Mac 上以 CPU 运行);dots.tts 没有 Windows 路径;Confucius4 建议使用 CUDACPU 可用,约为实时时长的 17 倍)。详情:[MOSS-TTS-v1.5](docs/engines/moss-tts-v15.md) · [dots.tts](docs/engines/dots-tts.md) · [Confucius4-TTS](docs/engines/confucius4-tts.md)。
</details>
<a id="asr-engines"></a>
### 🎧 ASR 引擎
**11 个引擎**——它们驱动听写、视频配音和字幕。**WhisperX** 是跨平台的默认引擎(约 100 种语言,词级时间对齐);其余引擎均为可选装并自动检测。在 **设置 → 引擎** 中切换。十个完全在本地设备上运行;第十一个(OpenAI 兼容)是可选的远程客户端,可用于 Qwen3-ASR 或任何兼容的服务器。
<details>
<summary><b>📊 完整阵容</b>——11 个引擎、各自的强项与计算类型说明</summary>
<br/>
| 引擎 | `OMNIVOICE_ASR_BACKEND` | 语言 | 最适合 |
|--------|-------------------------|:---------:|----------|
| **WhisperX**(默认) | `whisperx` | ~100 | 配音与字幕——通过 wav2vec2 强制对齐实现词级时间对齐 |
| **Faster-Whisper** | `faster-whisper` | ~100 | Linux / macOS / Windows 上的快速转录(CTranslate2 |
| **Faster-Whisper(隔离)** | `faster-whisper-isolated` | ~100 | 与 Faster-Whisper 相同,但在子进程中崩溃隔离——ASR 崩溃不会拖垮整个应用 |
| **MLX Whisper** | `mlx-whisper` | ~100 | Apple Silicon 原生速度(Apple MLX / Metal |
| **PyTorch Whisper** | `pytorch-whisper` | ~100 | 经 🤗 Transformers 的 CUDA / CPU 兜底方案(无需 cuDNN 8 |
| **Parakeet TDT** | `nemo-parakeet` | 英语 + 25 种欧洲语言 | 即使在 CPU 上也能以约 10 倍实时速度达到 SOTA 精度,自动语言检测(NVIDIA NeMoCUDA/CPU |
| **Moonshine** | `moonshine` | 英语 | 边缘设备 / 低延迟,ONNX |
| **FunASR** | `funasr` | 50+ | 多语言一体化——内置 VAD + 行内说话人分离(SenseVoice |
| **sherpa-onnx**(实时听写) | `sherpa-onnx-asr` | 25 种欧洲语言 + 90+ | 实时、快于实时的听写——小体积流式/离线 ONNX 模型(Parakeet TDT v3/v2、流式 Zipformer 与 Paraformer、Whisper Tiny),CPU 运行,macOS / Windows / Linux 表现完全一致。在 **设置 → 语音** 中按模型选择。 |
| **OpenAI 兼容** ⚠️ 远程 | `openai-compat-asr` | 取决于服务器 | 当下通往 **Qwen3-ASR** 的路径(自托管服务器,无需等 transformers 支持)、任何 OpenAI 兼容的转录端点,或 OpenAI 官方 API——无需安装,在 **设置 → 引擎**(ASR 标签页)中配置并测试连接。音频会离开你的设备,发送到你指定的任何服务器;参见 [docs/engines/openai-compatible-asr.md](docs/engines/openai-compatible-asr.md)。 |
> Whisper 系列引擎覆盖约 100 种语言;**FunASR / SenseVoice** 额外提供一条多语言一体化路径,内置语音活动检测与行内说话人分离。**sherpa-onnx** 驱动实时听写的模型选择器——你边说,文字边出现。除可选的 OpenAI 兼容远程客户端外,所有引擎都在本地设备上运行——无需 API 密钥,无需云端。
> **GPU 不支持高效 float16** 在较老的 NVIDIA GPUMaxwell/Pascal、GTX 16xx)上,或在 CTranslate2/cuDNN 版本不匹配之后,CTranslate2 系 ASR 引擎(WhisperX、Faster-Whisper)无法运行 `float16`VoiceStudio 会自动改用 `int8` 重试——无需配置。如果转录仍然失败,可用 `ASR_COMPUTE_TYPE` 环境变量固定计算类型(逃生舱口):`ASR_COMPUTE_TYPE=int8`CPU 用 `float32`)。将其设为 `int8` 并重启后端。
</details>
---
## 🏗️ 架构
```
┌─────────────────────────────────────────────────────────────┐
│ Frontend (React) │
│ DubTab · VoiceConsole · Stories · Audiobook · Gallery │
│ Dictation · BatchQueue · Diagnostics · MCP Client │
├─────────────────────────────────────────────────────────────┤
│ Backend (FastAPI) │
│ 100+ API endpoints · SSE+WSS streaming · SQLite │
├──────────┬──────────┬──────────┬──────────┬────────────────┤
│ WhisperX │ Demucs │VoiceStudio │ Pyannote │ Engine Routing │
│ (+7 ASR │ Source │ (+10 │ Diariz- │ ↳ GPU preflight │
│ engines) │ Sep. │ TTS) │ ation │ ↳ No silent CPU │
└──────────┴──────────┴──────────┴──────────┴────────────────┘
CUDA / MPS / ROCm / CPU (auto-detected + routed)
```
<a id="openai-api"></a>
## 🔌 OpenAI 兼容 API
已经有会说 OpenAI 音频 API 的脚本、智能体或工具?把它指向 `http://localhost:3900/v1` 即可——不需要密钥,也不用改代码。后端为音频端点内置了即插即用的兼容接口,直接接到你当前启用的 TTS/ASR 引擎(没错,`voice` 参数接受你克隆的声音配置 ID)。
| 端点 | 作用 |
| 需求 | 链接 |
|---|---|
| `POST /v1/audio/speech` | TTS——输入文本;输出 `mp3` / `wav` / `flac` / `opus` / `pcm``tts-1` / `tts-1-hd` 映射到你当前启用的引擎;也接受 OpenAI 的声音名称(`alloy` 等)。 |
| `POST /v1/audio/transcriptions` | STT——输入音频文件;输出 `json``text``verbose_json``srt``vtt``whisper-1` 映射到你当前启用的 ASR 引擎。 |
| `GET /v1/audio/voices` | VoiceStudio 扩展——列出所有声音配置和引擎,客户端可据此发现你的克隆声音。 |
| 安装帮助 | [故障排查](docs/install/troubleshooting.md) · [模型下载](docs/downloading-models.md) |
| 模型与音质 | [引擎指南](docs/engines/README.md) · [基准测试](docs/benchmarks.md) |
| 集成 | [本地 API](docs/speech-platform.md) · [MCP](docs/mcp.md) · [示例](examples/README.md) |
| 参与开发 | [贡献指南](.github/CONTRIBUTING.md) · [Electron](electron/README.md) · [更新日志](CHANGELOG.md) |
```sh
curl http://localhost:3900/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"model": "tts-1", "voice": "alloy", "input": "Generated on my own hardware.", "response_format": "wav"}' \
--output speech.wav
```
安装智能体技能:`npx skills add debpalash/VoiceStudio`
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:3900/v1", api_key="none") # any string works — nothing checks it
## 支持 VoiceStudio
result = client.audio.transcriptions.create(model="whisper-1", file=open("clip.wav", "rb"))
print(result.text)
```
[Ko-fi](https://ko-fi.com/debpalash) · [PayPal](https://paypal.me/palashCoder) · [赞助项目](SPONSORS.md) · [商务合作](mailto:partner@voicestudio.sh)
想要完整的接口(100+ 端点)?完整的 REST API 参考已内嵌在应用中——**设置 → OpenAPI 参考**(由 Scalar 驱动),或点击页脚的 `{}` 按钮
**让语音应用开发者看到你的品牌。** 了解应用底部栏、集成目录、文档和 README 的付费展示合作。[申请合作](https://forms.gle/2PYCvd39hbwijzX37)或[发送邮件](mailto:partner@voicestudio.sh)
### 📓 在 Google Colab 上运行
## 许可与负责任使用
[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb)
没有本地 GPU?官方笔记本([notebooks/OmniVoice_Studio_Colab.ipynb](notebooks/OmniVoice_Studio_Colab.ipynb))可在免费的 Colab T4 上启动完整应用(包含 Web 界面):在笔记本内直接构建前端,用 uv 安装后端(复用 Colab 预装的 CUDA PyTorch),并通过 Colab 内置端口代理打开界面。无需第三方隧道,也无需任何 API 密钥。随后还有一套覆盖全部主要功能的 API 导览,全部可在笔记本内直接播放:多语言 TTS、声音克隆与声音设计、已保存的声音档案、语音转写、AI 水印检测、OpenAI 兼容 API、多角色故事、带章节的 m4b 有声书,以及一个附带人声分离音轨的迷你视频配音。
### 🤝 智能体技能(Agent Skills
用一条命令教会你的 AI 智能体(Claude Code、Cursor、Codex 等)使用 VoiceStudio
```sh
npx skills add debpalash/omnivoice-studio
```
内含两个 [skills](https://skills.sh)**`omnivoice`**——让任何智能体通过你的本地安装进行语音合成与转录(包括你克隆的声音),免费且离线;以及 **`oss-maintainer`**——本项目所遵循的维护者方法论,适合任何用智能体运营自己开源项目的人。
### 🔌 模型上下文协议(MCP 服务器)
VoiceStudio 在 `http://localhost:3900/mcp` 挂载了 MCP 服务,可供 Claude Desktop、Cursor 与自主智能体调用:
```json
{
"mcpServers": {
"voicestudio": {
"url": "http://localhost:3900/mcp"
}
}
}
```
对于需要 stdio 管道传输的客户端,请使用内置的本地桥接脚本(`docs/mcp.json`):
```json
{
"mcpServers": {
"voicestudio": {
"command": "python",
"args": ["-m", "backend.mcp_shim"],
"cwd": "/path/to/VoiceStudio"
}
}
}
```
支持 `generate_speech``clone_voice``transcribe` 等工具与流式文件输出模式,详见 [docs/mcp.md](docs/mcp.md)。
---
## 🗺️ 路线图
### 🔜 即将推出
- 🎬 **唇形同步 v2** — 使用 wav2lip 进行视觉语音时间对齐
- 🌐 **在线演示** — 无需安装即可体验 VoiceStudio
- 🔌 **插件市场** — 社区贡献的 TTS 引擎与特效
- 🎵 **实时变声器** — 通话中的麦克风实时变声
<details>
<summary><b>✅ 已经发布的一切</b>——按类别列出的“成绩单”</summary>
<br/>
| 分类 | 功能 |
|----------|----------|
| **长内容** | 有声书编辑器(文本/EPUB/PDF → 分章 .m4b)、Stories 多声音编辑器、两遍响度归一母带处理、渲染中断后的崩溃续渲、发音控制 + SSML-lite 韵律 |
| **配音** | 完整流水线(转录→翻译→合成→封装)、场景感知分割、唇形同步评分、流式 TTS、逐说话人声音分配、Smart Fit 时长匹配 + 二次 QC、独立的配音主页 |
| **声音** | 零样本克隆、声音设计、A/B 对比、声音预览控件、支持收藏/标签的声音库、便携声音角色包(`.ovsvoice`)、声音控制台工作区 |
| **音频** | Demucs 人声分离、逐段增益、选择性音轨导出、分轨/SRT/VTT/MP3 导出、按句分块实现的无限长 TTS |
| **多语言** | 多语言批量选择器、顺序 GPU 执行的批量配音队列 |
| **说话人分离** | Pyannote 机器学习分离、自动说话人克隆提取、逐说话人声音分配 |
| **ASR** | 9 个引擎(WhisperX、Faster-Whisper、隔离版 Faster-Whisper、MLX Whisper、PyTorch Whisper、Parakeet TDT、Moonshine、FunASR/SenseVoice、sherpa-onnx 实时听写)、崩溃隔离的子进程后端 |
| **TTS** | 14 个引擎(VoiceStudio、CosyVoice 3、GPT-SoVITS、VoxCPM2、MOSS-TTS-Nano、KittenTTS、MLX-Audio、Sherpa-ONNX+ 延迟安装:IndexTTS 2.5、OmniVoice GGUF、Supertonic 3、MOSS-TTS-v1.5、dots.tts、Confucius4-TTS)、带 GPU 预检的引擎路由 |
| **基础设施** | Docker 部署、CUDA/MPS/ROCm 自动检测、cuDNN 8 兼容、显存感知模型卸载、引擎路由(绝不静默回退 CPU)、诊断套件与错误日志、受限网络镜像支持 |
| **AI 溯源** | AudioSeal 不可见水印(类似 SynthID)、视频徽标叠加、水印检测 API |
| **用户体验** | 撤销/重做、键盘快捷键、拖放、会话持久化、首次启动按屏幕推荐界面缩放,以及原生 WebKitGTK 缩放 |
| **实时事件** | WebSocket 事件总线——数据变更时即时刷新侧边栏、指数退避重连 |
| **状态管理** | Zustand 状态迁移——`uiSlice``pillSlice``dubSlice``generateSlice``prefsSlice``glossarySlice` |
| **桌面** | 跨平台 Tauri 安装程序(macOS DMG——Apple SiliconIntel 不支持本地后端,#889——Windows MSI、Linux deb/AppImage)、自动更新基础设施、单实例约束、关闭最小化到托盘、macOS Gatekeeper 修复 |
| **听写** | 全局系统级热键(`⌘+⇧+Space`)、无边框浮动控件、WebSocket 流式 ASR、自动粘贴、可自定义热键、本地 LLM 转录润色 |
| **批量流水线** | 完整批量 TTS:提取 → 转录 → 翻译 → 生成 → 混音 → 导出,带实时进度追踪 |
| **MCP 服务器** | 让 VoiceStudio 成为 Claude、Cursor 及任何 MCP 客户端的本地 TTS/STT 提供方 |
| **远程后端** | 让桌面 UI 指向远程后端 URL,支持 Bearer 认证(附 Tailscale 文档) |
| **可靠性** | 启动开屏的卡死看门狗、逐引擎 GPU 兼容矩阵、引擎二进制不可执行时的可操作报错、setuptools 自动修复 |
</details>
---
<a id="sponsor--donate"></a>
## 💜 赞助 / 捐赠
VoiceStudio 由一位开发者使用 Claude Code 和 AI 智能体独立打造——而智能体账单是实打实的(过去三个月花了数千美元)。如果 VoiceStudio 为你创造了价值,帮忙分担一小部分账单,就能让开发保持全职推进。
<div align="center">
**本月智能体账单基金**
<img src="https://img.shields.io/badge/raised_%2410_of_%24200-5%25-EAB308?style=for-the-badge" alt="已筹 $10 / $200" />
<br/><br/>
<a href="https://ko-fi.com/debpalash"><img src="https://img.shields.io/badge/Ko--fi-Support_❤️-FF5E5B?style=for-the-badge&logo=ko-fi&logoColor=white" alt="Ko-fi" /></a>
&nbsp;&nbsp;
<a href="https://paypal.me/palashCoder"><img src="https://img.shields.io/badge/PayPal-Donate-00457C?style=for-the-badge&logo=paypal&logoColor=white" alt="PayPal" /></a>
<br/>
<sub>每一美元都直接用于支付智能体账单——让 VoiceStudio 的开发持续不断。</sub>
<br/><br/>
<sub><b>来自 VoiceStudio 作者的更多应用</b>——同样的本地优先理念:
<a href="https://github.com/debpalash/Opal"><b>Opal</b> 💠</a>(播放一切——AI 时代的媒体播放器)·
<a href="https://github.com/debpalash/memxt"><b>memxt</b> 🧠</a>Claude Code 与编码智能体的本地记忆)。
给它们点个 ⭐ 也是一种支持 → <a href="#more-from-the-maker">详见下文</a>。</sub>
</div>
<a id="sponsors"></a>
### 🌟 赞助商
VoiceStudio **免费**且采用 **AGPL-3.0** 许可——没有付费版,没有 SaaS 收入。赞助商让开发得以持续,作为回报,可以在这里、在应用内(顶级档位还包括项目官网)获得一个徽标位。这是一份感谢,绝不是付费墙。**[查看档位并成为赞助商 →](SPONSORS.md)**
<div align="center">
<!-- SPONSORS:START — logo slots are filled here as sponsors come aboard; see SPONSORS.md -->
**这里可以是你的徽标** — [成为赞助商](SPONSORS.md)
<!-- SPONSORS:END -->
</div>
<sub>💡 GitHub 也会在本仓库顶部显示一个 **Sponsor** 按钮,经由 <a href=".github/FUNDING.yml"><code>.github/FUNDING.yml</code></a> 指向相同的链接。</sub>
---
## 💬 社区
<div align="center">
<a href="https://discord.gg/bzQavDfVV9"><img src="https://img.shields.io/badge/💬_Discord-Join_Community-5865F2?style=for-the-badge&logo=discord&logoColor=white" alt="加入 Discord" /></a>
<br/>
<sub>设置类问题我们几小时内就会回复,而不是几天。</sub>
</div>
<details>
<summary><b>里面都在聊什么</b></summary>
<br/>
| 频道 | 那里发生什么 |
|---------|--------------------|
| `#announcements` | 发布消息与重大时刻——新版本最先在这里公布 |
| `#releases` + `#changelog` | 每一个构建,以及里面究竟有什么 |
| `#issues` | 以论坛帖子形式提交的 Bug 报告——直接分诊进 GitHub Issues |
| `#ideas` | 功能请求,供讨论与投票 |
| `#discuss-ideas` | 动手之前的设计讨论 |
| `#general` | 安装帮助、GPU 疑难排查,以及晒你的配音成果 |
</details>
---
<a id="contributing"></a>
## 🤝 参与贡献
非常欢迎——Bug 修复、新的 TTS 引擎适配器、UI 改进、文档、翻译。统统欢迎。
- 📖 阅读 **[贡献指南](.github/CONTRIBUTING.md)** 了解环境搭建、代码风格和 PR 工作流
- 🐛 浏览 [good first issues](https://github.com/debpalash/VoiceStudio/labels/good%20first%20issue)
- 💬 加入我们的 [Discord](https://discord.gg/bzQavDfVV9) 讨论想法或寻求帮助
---
## ❓ 常见问题
<details>
<summary><b>真的能和 ElevenLabs 一样好吗?</b></summary>
<br/>
诚实的回答:<b>取决于你要做什么。</b>
<b>VoiceStudio 真正有竞争力的地方:</b>从干净的参考音频进行语音克隆(最先进的开源扩散 TTS)、语言覆盖(646 种语言对他们的 32 种),以及所有结构性优势——没有按字符计费、没有用量上限、音频不离开你的设备、完整的流水线可定制性(14 个 TTS 引擎、10 个 ASR 引擎、翻译方案随你选)。
<b>ElevenLabs 仍然领先的地方:</b>开箱即用的稳定性与打磨程度,尤其是英语 TTS。他们的单一模型经过深度调优;我们的质量取决于你选择的引擎、你的硬件,以及(对克隆而言)参考音频——干燥、近麦的音频比嘈杂或有回声的音频克隆效果好得多。
<b>具体到配音:</b>配音是一条链——转录 → 翻译 → 克隆 → 合成——在<i>你的</i>素材上,它只取决于最薄弱的一环。如果部分输出语无伦次,先检查片段表里的<i>原文</i>:当转录本身就错了,换一个 ASR 引擎或使用更干净的源音频——修复点通常在这里,而不是声音。
拿你的真实素材试试——免费,下载一次即可。许多用户直接用它替换了 ElevenLabs;也有人两个都留着。这两种结果我们都乐见。
</details>
<details>
<summary><b>能在 Apple SiliconM1/M2/M3/M4)上运行吗?</b></summary>
<br/>
可以。MPS 加速会被自动检测。在 Apple 硬件上,MLX 优化的 Whisper 模型可提供更快的转录速度。<b>不支持 Intel Mac</b>:应用 UI 可以安装,但本地 Python 后端无法运行,因为 PyTorch 已不再发布 Intel Mac 轮子(<a href="https://github.com/debpalash/VoiceStudio/issues/889">#889</a>)——Intel Mac 只能配合远程后端使用。
</details>
<details>
<summary><b>需要多少显存?</b></summary>
<br/>
<b>最低 4 GB。</b> 显存 ≤8 GB 时,TTS 模型会在转录期间自动卸载到 CPU。8 GB 以上时,所有组件同时在 GPU 上运行。完全没有 GPU?CPU 模式也能用——只是慢一些(TTS 约慢 3 倍)。
</details>
<details>
<summary><b>可以用于商业用途吗?</b></summary>
<br/>
<b>可以——商业使用免费</b>,基于 <a href="https://www.gnu.org/licenses/agpl-3.0.html">AGPL-3.0</a>:运行它、出售用它生成的音频、为客户的视频配音、在团队中部署。只有一项义务:如果你<b>修改</b>了 VoiceStudio 并通过网络向他人提供该修改版本,你必须依据相同条款分享修改后的源代码。想把它嵌入闭源产品?可获取商业许可证——参见<a href="#license">许可证</a>。
</details>
<details>
<summary><b>支持哪些语言?</b></summary>
<br/>
通过 VoiceStudio 模型的 TTS 支持 646 种语言。转录(WhisperX)支持 99 种语言。翻译覆盖范围取决于目标语言对。
</details>
<details>
<summary><b>可以添加自己的 TTS 引擎吗?</b></summary>
<br/>
可以。在 <code>backend/services/tts_backend.py</code> 中继承 <code>TTSBackend</code>,并将其添加到 <code>_REGISTRY</code> 字典中——约 50 行代码。十四个内置引擎均以此方式实现;参见 <a href="#tts-engines">TTS 引擎</a>。
</details>
<details>
<summary><b>VoiceStudio 会收集我的任何数据吗?</b></summary>
<br/>
<b>除非你明确同意,否则不会。</b>首次运行时应用会<i>询问</i>你——一个页面、两个同等分量的按钮,没有预先勾选。在你回答“是”之前,VoiceStudio 什么都不发送:没有分析、没有遥测、没有账号、没有“回传”。跳过提问就等于“否”。无论如何,你的文本、音频、声音和项目永远不会离开你的设备。
如果你选择同意(也可随时在 <b>设置 → 隐私 → “帮助改进 VoiceStudio”</b> 中开关),发送的只是匿名、不含内容的使用统计:生成信息(引擎、语言、生成耗时、字符<i>数量</i>、错误<i>类型</i>),以及应用生命周期——一次安装信号、版本更新(版本号之间)、崩溃(错误类别和<i>分桶后的</i>运行时长,绝不含日志)、错误<i>类型</i>(有上限、去重),以及卸载时的一次告别信号。绝不包含你的文本、音频、文件名或任何可识别信息——这由代码中的属性白名单强制保证(<code>backend/core/analytics.py</code>),而不只是一句承诺。源码构建根本没有分析数据的接收端,因此根本不会询问。你自己的统计数字在 <b>设置 → 用量</b> 中查看,本地计算,不发送到任何地方。
</details>
<details>
<summary><b>如何卸载它 / 删除它的所有数据?</b></summary>
<br/>
VoiceStudio 完全本地运行——卸载就是删除应用及其写入的文件夹(模型缓存、Python 环境、你的声音/项目、配置)。运行 <code>scripts/uninstall.sh</code>macOS/Linux)或 <code>scripts\uninstall.ps1</code>Windows)——它会先以干跑方式列出每个文件夹及其大小,加 <code>--yes</code> 才会真正删除。完整的各平台路径列表和应用移除步骤见 <a href="docs/install/uninstall.md"><b>docs/install/uninstall.md</b></a>。
</details>
## 🛡️ 负责任使用与安全
VoiceStudio 在个人硬件上提供零样本语音克隆与语音创作能力。我们提倡负责任的技术使用:
- **明确授权:** 严禁在未经说话人本人知情并明确授权的情况下克隆其声音。
- **AI 溯源:** VoiceStudio 默认集成 [AudioSeal](https://github.com/facebookresearch/audioseal) 不可见神经音频水印,在完全不影响听感音质的前提下精准标记合成语音。
- **本地隐私:** 默认本地工作流下,所有音频、声音档案、项目与转录文本始终保存在你的本地设备上;仅当你主动配置远程工作节点或第三方 ASR 端点时,相应数据才会传输到对应服务。
---
<a id="license"></a>
## 📜 许可证
VoiceStudio 是基于 [**GNU Affero 通用公共许可证 v3.0AGPL-3.0**](https://www.gnu.org/licenses/agpl-3.0.html) 的自由开源软件。
**可免费用于任何用途——包括商业和企业内部用途。** 运行它、出售用它生成的音频、为自己或客户的视频配音、在团队中推广——全部免费,无需许可证。作为一份**网络著佐权(copyleft)**许可证,AGPL 增加了一项义务:如果你**修改**了 VoiceStudio 并通过网络向他人提供该修改版本,你必须依据相同的 AGPL-3.0 条款向他们提供该修改版本的完整对应源代码。
希望将 VoiceStudio 嵌入**闭源或专有**产品或服务、又不受 AGPL-3.0 著佐权义务约束的组织,可获取**商业许可证**。**定价方案即将推出。** 咨询:**VoiceStudio@palash.dev**。
捆绑的 `omnivoice/` TTS 模型(作者 Han Zhu)在上游仍为 Apache-2.0 许可。完整且具约束力的条款请参见 [`LICENSE`](LICENSE)。
---
## 🙏 致谢
VoiceStudio 站在这些杰出开源工作的肩膀上:
| 项目 | 作用 |
|---------|------|
| [**VoiceStudio (k2-fsa)**](https://github.com/k2-fsa/OmniVoice) | 零样本扩散 TTS 引擎——核心语音合成模型 |
| [**WhisperX**](https://github.com/m-bain/whisperX) | 词级别语音识别与时间对齐 |
| [**Demucs (Meta)**](https://github.com/facebookresearch/demucs) | 音乐源分离,用于人声分离 |
| [**Pyannote**](https://github.com/pyannote/pyannote-audio) | 说话人分离——谁说了什么 |
| [**CTranslate2**](https://github.com/OpenNMT/CTranslate2) | CPU 和 GPU 上的优化 Transformer 推理 |
| [**AudioSeal (Meta)**](https://github.com/facebookresearch/audioseal) | 用于 AI 溯源的不可见神经音频水印 |
| [**Tauri**](https://tauri.app) | 原生桌面应用框架 |
| [**Supertone / Supertonic 3**](https://huggingface.co/Supertone/supertonic-3) | ONNX TTS 引擎——31 种语言,CPU 高效 |
| [**Sherpa-ONNX**](https://github.com/k2-fsa/sherpa-onnx) | 支持 WASM 的通用 TTS/ASR 运行时 |
| [**GPT-SoVITS**](https://github.com/RVC-Boss/GPT-SoVITS) | 零样本 TTS 引擎——5 种语言,RTF 0.014 |
---
<a id="more-from-the-maker"></a>
## 🧰 来自同一作者的更多本地开源项目
喜欢这种本地优先的理念?它是一脉相承的——同一位作者,同一条准则:**你的数据只留在你的设备上。** 全部项目见 [palash.dev](https://palash.dev)。
<table>
<tr>
<td align="center" width="50%" valign="top">
<br/>
<a href="https://github.com/debpalash/Opal"><img src="https://raw.githubusercontent.com/debpalash/Opal/main/assets/opal_logo.png" width="96" alt="Opal 徽标"/></a>
<h3><a href="https://github.com/debpalash/Opal">Opal 💠</a></h3>
<p><b>播放一切。</b>AI 时代的媒体播放器。</p>
<p><sub>视频、动漫、漫画、种子、Jellyfin 和 Plex——一个播放器全部搞定,并内置本地 AI 记忆与上下文。使用 Zig 编写,支持 macOS 和 Windows。</sub></p>
<p>
<a href="https://github.com/debpalash/Opal/stargazers"><img src="https://img.shields.io/github/stars/debpalash/Opal?style=flat-square&color=f59e0b" alt="Opal Star 数"/></a>
<a href="https://palash.dev/opal"><img src="https://img.shields.io/badge/site-palash.dev%2Fopal-8b5cf6?style=flat-square" alt="Opal 官网"/></a>
</p>
</td>
<td align="center" width="50%" valign="top">
<br/>
<a href="https://github.com/debpalash/memxt"><img src="https://raw.githubusercontent.com/debpalash/memxt/main/assets/logo-mark.svg" width="96" alt="memxt 徽标"/></a>
<h3><a href="https://github.com/debpalash/memxt">memxt 🧠</a></h3>
<p><b>经基准测试验证的最快开源 AI 记忆系统。</b></p>
<p><sub>为 Claude Code 和编码智能体提供本地长期记忆——基于 SQLite + 嵌入向量的 MCP 服务器,100% 在你的设备上运行。你的智能体终于能记住昨天了。</sub></p>
<p>
<a href="https://github.com/debpalash/memxt/stargazers"><img src="https://img.shields.io/github/stars/debpalash/memxt?style=flat-square&color=f59e0b" alt="memxt Star 数"/></a>
<a href="https://github.com/debpalash/memxt#readme"><img src="https://img.shields.io/badge/docs-README-10b981?style=flat-square" alt="memxt 文档"/></a>
</p>
</td>
</tr>
</table>
---
<div align="center">
<br/>
如果你读到了这里,你就是我们的同路人。<br/>
**[⭐ 给这个仓库点个 Star](https://github.com/debpalash/VoiceStudio)**,让更多人能找到它。<br/>
**[💬 加入 Discord](https://discord.gg/bzQavDfVV9)**,分享你的作品。<br/>
**[❤️ 支持开发](https://ko-fi.com/debpalash)**——资助让 VoiceStudio 持续发布的 AI 智能体账单。
<br/>
<a href="https://star-history.com/#debpalash/VoiceStudio&Date">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/svg?repos=debpalash/VoiceStudio&type=Date&theme=dark" />
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/svg?repos=debpalash/VoiceStudio&type=Date" />
<img alt="Star 历史" src="https://api.star-history.com/svg?repos=debpalash/VoiceStudio&type=Date&theme=dark" width="600" />
</picture>
</a>
</div>
应用采用 [AGPL-3.0](LICENSE) 许可。模型遵循各自的许可,商用前请确认其条款。克隆声音前须取得本人许可。详见[许可说明](LICENSE-NOTICE.md)。
+114 -2
View File
@@ -32,17 +32,129 @@ def _public_routing_reason(status: object, diagnostic: object) -> str:
return _ROUTING_BY_STATUS.get(status, _ROUTING_UNAVAILABLE)
# Categories for WHY an engine is unavailable. The probe's own sentence cannot
# cross the boundary — it carries exception text, local paths and sometimes
# credentials — but "Engine unavailable. Check installation and configuration."
# told the user nothing at all, and "Last error: A previous engine check
# failed." reads like a crash rather than "you have not installed this yet"
# (#1866). Classifying the private diagnostic into an owned sentence keeps the
# boundary intact and still names the kind of problem and the place to fix it.
_UNAVAILABLE_NOT_INSTALLED = (
"This engine's package isn't installed yet. Install it from "
"Model Catalogue."
)
# An engine gated behind an in-app license review (Supertonic-3, PocketTTS).
# The Model Catalogue shows its Accept button only when the reason matches
# /license not accepted/i (EngineCompatibilityMatrix.reasonMentionsLicense), so
# this sentence must keep those words: collapsing it into the generic line hid
# the only way to enable those engines.
_UNAVAILABLE_LICENSE = (
"License not accepted yet. Review and accept it in "
"Model Catalogue to enable this engine."
)
# An engine that cannot run on this machine at all: Apple-Silicon-only MLX,
# PyTorch with no Intel Mac build. "Isn't installed yet" or "check
# installation" sent people after an install that could never work.
_UNAVAILABLE_PLATFORM = (
"This engine doesn't run on this computer's platform. Its guide lists "
"the platforms it supports."
)
# Apple Silicon whose PyTorch cannot use the GPU (MPS): the platform is
# right, the installation is not. MLX-Audio / MLX-Whisper need MPS (#390).
_UNAVAILABLE_NO_MPS = (
"This engine needs Apple's GPU (MPS), and this installation's PyTorch "
"can't use it. Updating macOS or reinstalling VoiceStudio usually "
"restores it."
)
_UNAVAILABLE_NEEDS_CONFIG = (
"This engine needs to be configured before it can run. Open "
"Model Catalogue to finish setting it up."
)
_UNAVAILABLE_FILE_MISSING = (
"A file this engine needs is missing or unreadable. Reinstall it from "
"Model Catalogue."
)
# The same two cases for an engine the app cannot install for you. "Install it
# from Model Catalogue" sent people to a page with no Install button
# for that engine — most of the catalogue — which reads as the app being
# broken. The row's own guide link (``docs_url``) is the real next step.
_UNAVAILABLE_NOT_INSTALLED_MANUAL = (
"This engine isn't installed yet, and it has no one-click install. "
"Its guide lists the install steps."
)
_UNAVAILABLE_FILE_MISSING_MANUAL = (
"A file this engine needs is missing or unreadable. Its guide lists the "
"install steps."
)
_MANUAL_INSTALL_VARIANT = {
_UNAVAILABLE_NOT_INSTALLED: _UNAVAILABLE_NOT_INSTALLED_MANUAL,
_UNAVAILABLE_FILE_MISSING: _UNAVAILABLE_FILE_MISSING_MANUAL,
}
# Matched against the lowered probe text. Ordered most specific first: a
# missing file often also says "not installed", and the file case has the more
# useful remedy of the two.
_UNAVAILABLE_SIGNATURES = (
# First: its probe text also says "Open Model Catalogue", and the
# license is the one gap only the user can close.
(_UNAVAILABLE_LICENSE, ("license not accepted",)),
# Before the install and file checks: a platform reason often also says
# "unavailable" or names a missing wheel, and no install can fix it. Not
# "apple silicon only": mlx-audio says that on an M-series Mac too, when
# the package is merely missing and installing does help.
(_UNAVAILABLE_PLATFORM, (
"requires apple silicon", "not supported on this platform",
"unavailable on intel macs", "no macos x86_64 wheel",
"no windows install", "not supported on windows",
)),
(_UNAVAILABLE_NO_MPS, ("torch mps unavailable",)),
(_UNAVAILABLE_FILE_MISSING, (
"file is missing", "file is empty", "file is unreadable",
"script missing", "binary", "not found at",
)),
(_UNAVAILABLE_NEEDS_CONFIG, (
"environment variable", "configure a server endpoint", "api key",
"unconfigured", "set the", "base url",
)),
(_UNAVAILABLE_NOT_INSTALLED, (
"not installed", "package missing", "not available", "no module named",
"import ", "unavailable:", "failed to load",
)),
)
def _public_unavailable_reason(diagnostic: object) -> str:
"""Map a private availability probe to an accurate stable category."""
private = diagnostic.lower() if isinstance(diagnostic, str) else ""
for public, markers in _UNAVAILABLE_SIGNATURES:
if any(marker in private for marker in markers):
return public
return _UNAVAILABLE
def public_backends(entries: list[dict]) -> list[dict]:
"""Copy registry entries while replacing service diagnostics.
Availability probes may contain exception text, local paths, tracebacks, or
credentials. Installation hints are registry-authored and remain intact.
credentials. Registry-authored fields are not probe output and remain
intact: ``install_hint``, ``setup_snippet`` and ``docs_url`` are all
VoiceStudio-owned constants keyed on the engine id, so an unavailable row
still has something actionable to show and somewhere to send the user
(#1866) even though ``reason``/``last_error`` are replaced here.
"""
safe: list[dict] = []
for entry in entries:
item = dict(entry)
if item.get("reason") is not None:
item["reason"] = _UNAVAILABLE
reason = _public_unavailable_reason(item["reason"])
# Only a row that explicitly says it has NO one-click install gets
# the manual wording. Rows without the field (ASR, LLM,
# translation — some of which have installers of their own) keep
# the line that points at Model Catalogue.
if item.get("one_click_install") is False:
reason = _MANUAL_INSTALL_VARIANT.get(reason, reason)
item["reason"] = reason
if item.get("last_error") is not None:
item["last_error"] = _PREVIOUS_FAILURE
if item.get("routing_reason") is not None:
+42 -19
View File
@@ -57,11 +57,12 @@ _PREVIEW_SEED = 42
# 32 reliably converges to speech across the gallery's instruct/script space
# at a one-time (cached) render cost.
_PREVIEW_NUM_STEP = 32
# Spectral-flatness floor below which a render is a degenerate tonal artifact
# rather than speech. Real, mastered speech sits ~0.040.07; a tonal buzz
# collapses to <0.005. 0.015 separates the two with wide margin and sits well
# below even breathy/whisper voices (which are broadband → high flatness).
_DEGENERATE_FLATNESS = 0.015
# Reject near-pure tonal artifacts using mean framed spectral flatness.
# Calibrated against the tracked speech demos exercised by
# test_archetype_preview_quality.py: the quietest (Mandarin dubbing, 44.1 kHz)
# measures ~7.7e-6, while the worst tested tonal buzz measures ~3.3e-9.
# 1e-7 leaves >10x margin on both sides without rejecting low-flatness speech.
_DEGENERATE_FLATNESS = 1e-7
def _preview_key(a: dict) -> str:
@@ -248,24 +249,46 @@ def _is_blank_audio(audio_tensor) -> bool:
return False
_FLATNESS_FRAME = 1024
_FLATNESS_HOP = 512
#: Frames quieter than this fraction of the loudest frame's energy are the gaps
#: between words, not speech; their spectrum is the noise floor and averaging it
#: in drags the measurement toward the value of whatever silence sounds like.
_FLATNESS_FRAME_FLOOR = 1e-4
def _spectral_flatness(audio_tensor) -> Optional[float]:
"""Geometric-mean / arithmetic-mean of the power spectrum.
"""Mean per-frame geometric-mean / arithmetic-mean of the power spectrum.
~1.0 for broadband noise, →0 for a pure tone. The degenerate diffusion
renders this guards against are near-pure tonal buzzes (flatness <0.005),
distinct from both silence (caught by ``_is_blank_audio``) and real speech
(~0.04+). Returns ``None`` if it can't be computed so callers don't act on
a bad measurement.
renders this guards against are near-pure tonal buzzes, distinct from both
silence (caught by ``_is_blank_audio``) and real speech. Returns ``None``
if it can't be computed so callers don't act on a bad measurement.
Measured over short frames and averaged — the standard definition. A single
FFT of the whole clip (what this used to do) is not the same quantity: its
frequency resolution grows with clip length, so speech harmonics carve
ever-deeper nulls into the spectrum and the geometric mean collapses. That
made the result depend on how long the clip was rather than on what it
sounded like, and put real speech below the rejection threshold.
"""
try:
import torch
t = audio_tensor if isinstance(audio_tensor, torch.Tensor) else torch.as_tensor(audio_tensor)
t = t.detach().to("cpu", dtype=torch.float32).flatten()
if t.numel() < 1024 or not torch.isfinite(t).all():
t = t.detach().to("cpu", dtype=torch.float32)
if t.ndim > 1:
t = t.mean(dim=0)
t = t.flatten()
if t.numel() < _FLATNESS_FRAME or not torch.isfinite(t).all():
return None
spec = torch.fft.rfft(t * torch.hann_window(t.numel())).abs().pow(2) + 1e-12
return float(torch.exp(torch.mean(torch.log(spec))) / torch.mean(spec))
frames = t.unfold(0, _FLATNESS_FRAME, _FLATNESS_HOP)
spec = torch.fft.rfft(frames * torch.hann_window(_FLATNESS_FRAME)).abs().pow(2) + 1e-12
energy = spec.sum(dim=1)
spec = spec[energy > energy.max() * _FLATNESS_FRAME_FLOOR]
if spec.shape[0] == 0:
return None
return float((torch.exp(spec.log().mean(dim=1)) / spec.mean(dim=1)).mean())
except Exception: # never let the checker itself block a render
return None
@@ -405,8 +428,8 @@ def _preview_source(a: dict) -> tuple[str, str]:
return "cached", ""
if _no_voice_model_downloaded():
return "no_model", (
"You're offline and no voice model is downloaded yet — "
"Model Catalogue → Models → Download."
"You're offline and no voice model is downloaded yet — download "
"one from the engine's Weights list in Model Catalogue."
)
return "rendering", "Rendering this preview on your machine — it may take a moment."
@@ -544,8 +567,8 @@ async def preview_archetype(
if _no_voice_model_downloaded():
detail = (
"You're offline and no voice model is downloaded yet — "
"Model Catalogue → Models → Download. (Or turn on pre-rendered "
"voice previews in Model Catalogue → Models.)"
"download one from the engine's Weights list in Model Catalogue. (Or turn "
"on pre-rendered voice previews in Settings → Storage.)"
)
else:
detail = (
@@ -609,7 +632,7 @@ async def use_archetype(archetype_id: str, name: Optional[str] = Query(None)):
if _no_voice_model_downloaded():
detail = (
"Creating a voice needs the voice model — no voice model is "
"downloaded yet. Model Catalogue → Models → Download."
"downloaded yet. Download one from the engine's Weights list in Model Catalogue."
)
else:
detail = (
+23 -2
View File
@@ -420,8 +420,11 @@ def _omnivoice_sampling_kwargs(opts: ExpressiveOptions) -> dict:
today exactly: num_step 32, guidance 2.0, and NO temperature/postprocess
kwargs (the model keeps its own defaults). Emotion is never forwarded —
the VoiceStudio config rejects unknown kwargs."""
from services.performance_profiles import tts_defaults
defaults = tts_defaults()
kw = {
"num_step": opts.num_step if opts.num_step is not None else LONGFORM_NUM_STEP,
"num_step": opts.num_step if opts.num_step is not None else defaults.get("num_step", LONGFORM_NUM_STEP),
"guidance_scale": (
opts.guidance_scale if opts.guidance_scale is not None else LONGFORM_GUIDANCE_SCALE
),
@@ -432,6 +435,8 @@ def _omnivoice_sampling_kwargs(opts: ExpressiveOptions) -> dict:
kw["class_temperature"] = opts.class_temperature
if opts.postprocess_output is not None:
kw["postprocess_output"] = opts.postprocess_output
elif "postprocess_output" in defaults:
kw["postprocess_output"] = defaults["postprocess_output"]
return kw
@@ -718,6 +723,10 @@ def _remote_chapter_call(chapter, *, engine_id, default_voice, voice_map,
"expressive": opts.to_manifest(), "watermark": bool(watermark_enabled()),
}
signature = hashlib.sha256(json.dumps(params, sort_keys=True, default=str).encode()).hexdigest()
# The worker synthesizes from ``spans``, but the gateway and scheduler read
# top-level ``text`` to scale the remote execution deadline. Add this after
# the signature so existing content-addressed remote cache keys still hit.
params["text"] = "\n".join(row["text"] for row in rows)
wav_path = os.path.join(cache_dir, f"remote-{signature}.wav")
def decode(result):
@@ -739,7 +748,7 @@ async def _run_chapter(chapter, *, operation="audiobook", decision, job, default
voice_map, lexicon, cache_dir):
"""Run one chapter through the gateway; local preparation stays lazy."""
from services import gpu_gateway
from services.tts_backend import active_backend_id
from services.tts_backend import active_backend_id, get_backend_class
engine_id = active_backend_id()
remote, remote_cache = _remote_chapter_call(
@@ -753,15 +762,27 @@ async def _run_chapter(chapter, *, operation="audiobook", decision, job, default
return remote_cache, float(info.duration), True, None
async def prepare_local():
from services.model_manager import generate_timeout_s
synth, sr, resolve, local_engine = await _prepare_synth(
default_voice, language=language, opts=opts, voice_map=voice_map
)
try:
timeout_engine = get_backend_class(local_engine)
except ValueError:
# Tests and third-party integrations may inject a synth under a
# non-catalogue id. Keep the canonical host/text policy available;
# registered production engines still add their routing metadata.
timeout_engine = None
return gpu_gateway.LocalCall(
fn=lambda: _render_chapter_cached(
chapter, synth, sr, local_engine, resolve, cache_dir, lexicon,
language, opts, voice_map,
),
what="Audiobook chapter",
timeout=generate_timeout_s(
remote.params["text"], engine=timeout_engine
),
)
return await gpu_gateway.run(
+463 -146
View File
@@ -9,6 +9,8 @@ the SQLite `jobs` table for history, but the queue itself restarts empty
on backend restart — intentional, since GPU jobs can't be safely resumed.
"""
import os
import json
import shutil
import uuid
import time
import asyncio
@@ -16,20 +18,42 @@ import logging
from typing import Optional, List
from fastapi import APIRouter, File, UploadFile, HTTPException, Form
from fastapi.responses import JSONResponse
from pydantic import BaseModel
from core.config import DATA_DIR
from core import failure
from core.logging_utils import log_safe
from core.file_cleanup import FileCleanupError, unlink_if_present
from services.dub_batching import (
BATCH_WIDTH_ENV,
batch_timeout_s as _batch_timeout_s,
native_batch_width as _native_batch_width,
)
from services import gpu_gateway
from services.segment_bundle import extract_segment_wavs, remove_segment_wavs
from services.tts_backend import active_backend_id, resolve_generation_backend
router = APIRouter()
logger = logging.getLogger("omnivoice.batch")
# Compatibility values emitted by the established Tauri Batch picker. They
# are taxonomy tokens, not arbitrary prose, and are resolved server-side so
# native watch-folder uploads and both desktop clients use the same voice.
_BATCH_PRESET_INSTRUCT = {
"narrator": "male, middle-aged, low pitch, british accent",
"excited_child": "child, high pitch",
"anxious_whisper": "young adult, whisper",
"surprised_woman": "female, young adult, high pitch",
"elderly_story": "male, elderly, very low pitch",
"sichuan": "female, young adult, moderate pitch, \u56db\u5ddd\u8bdd",
}
# ── In-memory queue ─────────────────────────────────────────────────────
_queue: asyncio.Queue = None # Lazily initialised
_worker_task: asyncio.Task = None # Background consumer
_processing_job_ids: set[str] = set()
_jobs: dict = {} # job_id → status dict
@@ -40,11 +64,15 @@ class BatchJobStatus(BaseModel):
langs: List[str]
voice_id: Optional[str] = None
preserve_bg: bool = True
translation_provider: Optional[str] = None
created_at: float
started_at: Optional[float] = None
finished_at: Optional[float] = None
error: Optional[str] = None
progress: Optional[dict] = None
attempts: int = 1
retry_ready: bool = True
setup_required: Optional[dict] = None
def _ensure_queue():
@@ -66,6 +94,7 @@ async def _worker():
job["status"] = "running"
job["started_at"] = time.time()
_processing_job_ids.add(job_id)
logger.info("Batch job %s starting: %s", job_id, job["filename"])
try:
@@ -95,6 +124,9 @@ async def _worker():
job["finished_at"] = time.time()
logger.error("Batch job %s failed: %s", job_id, e, exc_info=True)
finally:
_processing_job_ids.discard(job_id)
if job["status"] == "cancelled":
job["retry_ready"] = True
_queue.task_done()
@@ -104,18 +136,62 @@ def _set_progress(job, stage, percent=0, **extra):
#: Override for the native dub batch width. Set to 1 to disable batching.
BATCH_WIDTH_ENV = "OMNIVOICE_DUB_BATCH_WIDTH"
#: Hard ceiling on the override — a batch this wide is already amortizing
#: almost all of the per-call setup, and beyond it the failure mode is an OOM
#: that costs more than the saving.
_MAX_BATCH_WIDTH = 16
# Bound each allocation while persisting multipart uploads. Video inputs can
# be many gigabytes; `await UploadFile.read()` with no size used to mirror the
# entire file in process memory before writing it back out.
_UPLOAD_CHUNK_BYTES = 1024 * 1024
_REMOTE_BATCH_OPERATION = "batch_segments"
async def _resolve_batch_execution(voice: dict):
"""Resolve Batch's TTS target without loading local weights remotely."""
engine_id = active_backend_id()
decision = gpu_gateway.decide("batch")
if decision.remote:
await gpu_gateway.preflight(
engine_id,
decision,
operation=_REMOTE_BATCH_OPERATION,
)
return engine_id, decision, None
backend = await resolve_generation_backend(
require_cloning=voice["requires_cloning"],
cloning_purpose="this batch job's pinned voice",
)
return engine_id, decision, backend
def _decode_remote_batch(
result: gpu_gateway.RemoteResult,
batch_dir: str,
expected: set[int],
) -> tuple[dict[int, str], int]:
"""Validate and unpack one worker result before accepting remote success."""
import soundfile as sf
target = os.path.join(batch_dir, ".remote", result.task_id)
paths = extract_segment_wavs(result.path or "", target)
try:
if set(paths) != expected:
missing = sorted(expected - set(paths))
extra = sorted(set(paths) - expected)
raise ValueError(
f"segment bundle mismatch (missing={missing}, extra={extra})"
)
rates = {int(sf.info(path).samplerate) for path in paths.values()}
if len(rates) != 1 or next(iter(rates), 0) <= 0:
raise ValueError("segment bundle has inconsistent sample rates")
return paths, rates.pop()
except BaseException:
remove_segment_wavs(paths)
raise
async def _save_upload(upload: UploadFile, destination: str) -> None:
try:
@@ -130,64 +206,63 @@ async def _save_upload(upload: UploadFile, destination: str) -> None:
raise
def _native_batch_width(backend) -> int:
"""How many segments to render in one native batch on THIS host.
def _batch_voice(voice_id: str | None) -> dict:
"""Resolve one queue-wide voice into concrete generation inputs.
A native batch widens the forward pass, so the width cannot be a constant.
The default engine declares ``min_vram_gb = 6.0`` for a SINGLE job; an
unconditional 8-wide batch would OOM the 4-8 GB CUDA cards and the MPS
Macs where the per-segment path succeeds today — turning a throughput
optimization into a regression on exactly the hardware that already
struggles (#1616 is a 4 GB card reporting capacity failures). Default
behaviour must not get riskier on a host, so the width is derived from
measured headroom and falls back to 1 (no batching) when unknown.
CPU hosts get 1: batching there buys no kernel amortization and only
multiplies peak RAM.
Clone profiles contribute their reference; designed profiles contribute
their healed instruction and seed. Legacy ``preset:`` selections become
the same instruction used by Dubbing instead of falling through to the
engine default.
"""
override = os.environ.get(BATCH_WIDTH_ENV, "").strip()
if override:
try:
return max(1, min(_MAX_BATCH_WIDTH, int(override)))
except (TypeError, ValueError):
logger.warning(
"%s=%r is not an integer — deriving the batch width from the host instead.",
BATCH_WIDTH_ENV, override,
)
try:
from core.device_caps import detect_host_caps
caps = detect_host_caps()
except Exception: # noqa: BLE001 — an unprobeable host takes the safe path
return 1
if caps.family == "cpu" or not caps.vram_gb:
return 1
headroom = caps.vram_gb - float(getattr(backend, "min_vram_gb", 0.0) or 0.0)
if headroom < 2.0:
return 1
if headroom < 6.0:
return 2
if headroom < 12.0:
return 4
return 8
resolved = {
"ref_audio": None,
"ref_text": None,
"instruct": "",
"seed": None,
"requires_cloning": False,
}
if not voice_id:
return resolved
if voice_id.startswith("preset:"):
preset_id = voice_id.removeprefix("preset:")
instruct = _BATCH_PRESET_INSTRUCT.get(preset_id)
if instruct is None:
raise ValueError("That built-in voice preset no longer exists")
from omnivoice.utils.voice_design import sanitize_instruct
resolved["instruct"] = sanitize_instruct(instruct)
return resolved
def _batch_timeout_s(texts: list[str], backend) -> float:
"""Execution budget for one native batch.
from core.config import VOICES_DIR
from core.db import db_conn
Not the sum of the per-item budgets: ``generate_timeout_s`` returns a
floor (300s GPU / 600s CPU) plus per-length overage, so summing it across
eight items yields a ~2400s budget — and a wedged batch would hold a
GPU-pool worker for forty minutes before the reset this file depends on
(#730). One floor covers wedge detection for the whole call; only the
length-driven overage is genuinely additive.
"""
from services.model_manager import generate_timeout_s
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?",
(voice_id,),
).fetchone()
if row is None:
raise ValueError("That saved voice no longer exists")
floor = generate_timeout_s("", engine=backend)
overage = sum(
max(0.0, generate_timeout_s(text, engine=backend) - floor) for text in texts
)
return floor + overage
if row["kind"] == "design":
from omnivoice.utils.voice_design import heal_design_instruct
resolved["instruct"] = heal_design_instruct(row["instruct"], row["vd_states"])
resolved["seed"] = int(row["seed"]) if row["seed"] is not None else None
return resolved
relative = row["locked_audio_path"] if row["is_locked"] else row["ref_audio_path"]
if not relative:
raise ValueError("That saved voice has no reference audio")
ref_audio = os.path.join(VOICES_DIR, relative)
if not os.path.isfile(ref_audio):
raise ValueError("That saved voice's reference audio is missing")
resolved.update({
"ref_audio": ref_audio,
"ref_text": row["ref_text"],
"requires_cloning": True,
})
return resolved
async def _run_batch_pipeline(job_id: str, job: dict):
@@ -280,19 +355,30 @@ async def _run_batch_pipeline(job_id: str, job: dict):
return
# ── Engine resolution (issue #312 class) ────────────────────────────
# Batch used to hardcode VoiceStudio via get_model() regardless of the
# engine selected in Model Catalogue → Engines. require_cloning only when a
# specific voice is pinned (job["voice_id"]) — an unpinned job is fine on
# any active engine. Resolved ONCE for the whole job (every language
# Batch used to hardcode VoiceStudio regardless of the engine selected in
# Model Catalogue. Clone profiles require a cloning-capable engine; presets
# and designed voices use instruction mode. Resolve once for the whole job.
# below shares the same active engine); an uncaught ValueError here
# propagates to _worker()'s existing except-Exception handling, which
# already records a structured job failure via core.failure.build_failure.
from services.tts_backend import resolve_generation_backend
backend = await resolve_generation_backend(
require_cloning=bool(job.get("voice_id")),
cloning_purpose="this batch job's pinned voice",
)
sr = backend.sample_rate
voice = _batch_voice(job.get("voice_id"))
engine_id, execution_target, backend = await _resolve_batch_execution(voice)
sr = backend.sample_rate if backend is not None else 0
from services.performance_profiles import tts_defaults
_profile_defaults = tts_defaults(engine_id)
_batch_num_step = _profile_defaults.get("num_step", 16)
_batch_postprocess = _profile_defaults.get("postprocess_output", True)
batch_run = gpu_gateway.JobRun("batch")
async def _prepare_local_batch() -> gpu_gateway.LocalCall:
nonlocal backend, sr
if backend is None:
backend = await resolve_generation_backend(
require_cloning=voice["requires_cloning"],
cloning_purpose="this batch job's pinned voice",
)
sr = backend.sample_rate
return gpu_gateway.LocalCall(fn=lambda: None, what="Batch TTS fallback")
# ── 3. Translate + Generate per language ───────────────────────────
total_langs = len(langs)
@@ -311,40 +397,65 @@ async def _run_batch_pipeline(job_id: str, job: dict):
translated_segments = list(segments) # copy
if target_lang != source_lang:
try:
def _translate_batch(segs, src, tgt):
"""Translate segment texts via Google Translate."""
from deep_translator import GoogleTranslator
TRANSLATE_CODES = {
"en": "en", "es": "es", "fr": "fr", "de": "de",
"it": "it", "pt": "pt", "ru": "ru", "ja": "ja",
"ko": "ko", "zh": "zh-CN", "ar": "ar", "hi": "hi",
"tr": "tr", "pl": "pl", "nl": "nl", "sv": "sv",
}
src_code = TRANSLATE_CODES.get(src, src) or "auto"
tgt_code = TRANSLATE_CODES.get(tgt, tgt)
translator = GoogleTranslator(source=src_code, target=tgt_code)
out = []
for s in segs:
s_copy = dict(s)
text = s.get("text", "").strip()
if text:
try:
s_copy["text"] = translator.translate(text) or text
except Exception as e:
logger.warning("Translate seg failed: %s", e)
out.append(s_copy)
return out
# Use the same provider dispatch as interactive Dubbing. The old
# batch-only implementation hardcoded Google and silently kept the
# source text on failure, which could make an English track labelled
# "es" while also sending text online despite an offline selection.
from api.routers.dub_translate import dub_translate
from schemas.requests import TranslateRequest
translated_segments = await loop.run_in_executor(
_cpu_pool, _translate_batch,
segments, source_lang, target_lang,
from core import prefs
provider = job.get("translation_provider") or prefs.get("translation_backend", "argos")
translation = await dub_translate(TranslateRequest(
segments=[
{
"id": str(segment["id"]),
"text": segment.get("text", ""),
"start": segment.get("start"),
"end": segment.get("end"),
}
for segment in segments
],
source_lang=source_lang,
target_lang=target_lang,
provider=provider,
quality="fast",
))
if isinstance(translation, JSONResponse):
try:
payload = json.loads(translation.body)
detail = payload.get("error") or payload.get("detail")
if payload.get("code") == "argos_pack_missing":
job["setup_required"] = {
"kind": "argos_packs",
"source_lang": source_lang,
"target_langs": [
pair["target_lang"]
for pair in payload.get("pairs", [])
if isinstance(pair, dict) and pair.get("target_lang")
],
}
except Exception: # noqa: BLE001 — retain the stable fallback
detail = None
raise RuntimeError(
detail or f"{provider} could not translate this batch"
)
except ImportError:
logger.warning("deep_translator not installed, skipping translation for %s", target_lang)
except Exception as e:
logger.warning("Translation failed for %s: %s, using original", target_lang, e)
translated_segments = segments
rows = {
str(row.get("id")): row
for row in translation.get("translated", [])
if isinstance(row, dict)
}
failed = [row for row in rows.values() if row.get("error")]
if failed or len(rows) != len(segments):
raise RuntimeError(
f"{provider} translation failed for "
f"{len(failed) or len(segments) - len(rows)} segment(s)"
)
translated_segments = [
{**segment, "text": rows[str(segment["id"])]["text"]}
for segment in segments
]
if job["status"] == "cancelled":
return
@@ -362,6 +473,90 @@ async def _run_batch_pipeline(job_id: str, job: dict):
from services.audio_io import atomic_save_wav
import torch
remote_segments: dict[int, str] = {}
valid_rows = [
(i, segment)
for i, segment in enumerate(translated_segments)
if segment.get("end", 0) - segment.get("start", 0) > 0.05
and segment.get("text", "").strip()
]
if execution_target.remote and valid_rows:
remote_rows = [
{
"index": i,
"text": segment.get("text", "").strip(),
"language": target_lang,
"ref_text": voice["ref_text"],
"instruct": voice["instruct"] or None,
"duration": segment.get("end", 0) - segment.get("start", 0),
"num_step": _batch_num_step,
"postprocess_output": _batch_postprocess,
"guidance_scale": 2.0,
"speed": 1.0,
"effect_preset": "batch",
"seed": (
voice["seed"] + i if voice["seed"] is not None else None
),
# The assembled track receives one watermark below. Marking
# each line here would double-process remote output.
"watermark": False,
}
for i, segment in valid_rows
]
expected = {row["index"] for row in remote_rows}
def _remote_state(state: dict) -> None:
fraction = max(0.0, min(1.0, float(state.get("progress") or 0.0)))
_set_progress(
job,
"generate",
percent=int(((lang_idx + fraction) / total_langs) * 100),
current_lang=target_lang,
current_segment=min(len(remote_rows), round(fraction * len(remote_rows))),
total_segments=len(remote_rows),
execution_target=execution_target.label,
execution_phase=state.get("phase"),
)
route_task = asyncio.create_task(
gpu_gateway.run(
"batch",
local=gpu_gateway.LocalCall(prepare=_prepare_local_batch),
remote=gpu_gateway.RemoteCall(
engine=engine_id,
operation=_REMOTE_BATCH_OPERATION,
params={
"segments": remote_rows,
"ref_audio": [voice["ref_audio"] for _ in remote_rows],
"input_seconds": sum(
float(row.get("duration") or 0.0) for row in remote_rows
),
},
idempotency_key=f"batch:{job_id}:{target_lang}",
decode=lambda result: _decode_remote_batch(
result, batch_dir, expected
),
),
decision=execution_target,
job=batch_run,
on_state=_remote_state,
)
)
while not route_task.done():
await asyncio.wait({route_task}, timeout=0.25)
if job["status"] == "cancelled":
route_task.cancel()
try:
await route_task
except asyncio.CancelledError:
pass
return
routed = route_task.result()
if routed is not None:
remote_segments, sr = routed
# A remote-only empty transcript still needs a valid silent-track rate.
sr = sr or 24_000
total_samples = int(duration * sr)
full_audio = torch.zeros(1, total_samples)
total_segs = len(translated_segments)
@@ -372,26 +567,15 @@ async def _run_batch_pipeline(job_id: str, job: dict):
# the established one-segment behavior below.
from services.tts_backend import TTSBackend
batched_audio: dict[int, torch.Tensor] = {}
has_native_batch = type(backend).generate_batch is not TTSBackend.generate_batch
has_native_batch = (
backend is not None
and type(backend).generate_batch is not TTSBackend.generate_batch
)
if has_native_batch:
from services.text_normalization import normalize_for_tts
batch_ref_audio = None
batch_ref_text = None
if job.get("voice_id"):
from core.db import db_conn
from core.config import VOICES_DIR as _VD
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?",
(job["voice_id"],),
).fetchone()
if row:
if row["is_locked"] and row["locked_audio_path"]:
batch_ref_audio = os.path.join(_VD, row["locked_audio_path"])
elif row["ref_audio_path"]:
batch_ref_audio = os.path.join(_VD, row["ref_audio_path"])
batch_ref_text = row["ref_text"]
batch_ref_audio = voice["ref_audio"]
batch_ref_text = voice["ref_text"]
batch_width = _native_batch_width(backend)
@@ -428,17 +612,20 @@ async def _run_batch_pipeline(job_id: str, job: dict):
]
def _render_native_batch():
if voice["seed"] is not None:
torch.manual_seed(voice["seed"])
generated = backend.generate_batch(
batch_texts,
language=target_lang,
ref_audio=batch_ref_audio,
ref_text=batch_ref_text,
instruct=voice["instruct"] or None,
duration=batch_durations,
num_step=16,
num_step=_batch_num_step,
guidance_scale=2.0,
speed=1.0,
denoise=True,
postprocess_output=True,
postprocess_output=_batch_postprocess,
)
if len(generated) != len(batch_indices):
raise RuntimeError(
@@ -473,6 +660,7 @@ async def _run_batch_pipeline(job_id: str, job: dict):
for i, seg in enumerate(translated_segments):
if job["status"] == "cancelled":
remove_segment_wavs(remote_segments)
return
_set_progress(
@@ -499,32 +687,18 @@ async def _run_batch_pipeline(job_id: str, job: dict):
from services.text_normalization import normalize_for_tts
text = normalize_for_tts(text, lang)
ref_audio = None
ref_text = None
# Use voice_id if provided
if job.get("voice_id"):
from core.db import db_conn
from core.config import VOICES_DIR as _VD
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?",
(job["voice_id"],),
).fetchone()
if row:
if row["is_locked"] and row["locked_audio_path"]:
ref_audio = os.path.join(_VD, row["locked_audio_path"])
elif row["ref_audio_path"]:
ref_audio = os.path.join(_VD, row["ref_audio_path"])
ref_text = row.get("ref_text")
try:
if backend is None:
raise RuntimeError("the local TTS fallback was not prepared")
if voice["seed"] is not None:
torch.manual_seed(voice["seed"] + i)
audio_out = backend.generate(
text=text, language=lang,
ref_audio=ref_audio, ref_text=ref_text,
duration=dur, num_step=16,
ref_audio=voice["ref_audio"], ref_text=voice["ref_text"],
instruct=voice["instruct"] or None,
duration=dur, num_step=_batch_num_step,
guidance_scale=2.0, speed=1.0,
denoise=True, postprocess_output=True,
denoise=True, postprocess_output=_batch_postprocess,
)
if not getattr(backend, "applies_own_mastering", False):
audio_out = apply_mastering(audio_out, sample_rate=sr)
@@ -548,15 +722,43 @@ async def _run_batch_pipeline(job_id: str, job: dict):
# Budget is the shared length-scaled one (#1190): a long segment
# on CPU-class hardware no longer dies on the flat 300s.
from services.model_manager import generate_timeout_s
if has_native_batch and i not in batched_audio:
await _prefetch_batch(i)
if i in batched_audio:
audio_tensor = batched_audio.pop(i)
remote_path = remote_segments.pop(i, None)
if remote_path is not None:
import soundfile as sf
try:
audio_array, remote_sr = sf.read(
remote_path,
dtype="float32",
always_2d=True,
)
if int(remote_sr) != sr:
raise ValueError(
f"remote segment sample rate changed from {sr} to {remote_sr}"
)
audio_tensor = torch.from_numpy(audio_array.T).mean(
dim=0,
keepdim=True,
)
finally:
remove_segment_wavs({i: remote_path})
else:
audio_tensor = await run_on_gpu_pool_guarded(
_gen, what="Batch generate",
timeout=generate_timeout_s(seg_text, engine=backend),
)
if backend is None:
await _prepare_local_batch()
# This path means a validated remote bundle lost a row
# after dispatch. Recover only that row; native batches
# were not planned for this language.
has_native_batch = False
if has_native_batch and i not in batched_audio:
await _prefetch_batch(i)
if i in batched_audio:
audio_tensor = batched_audio.pop(i)
else:
audio_tensor = await run_on_gpu_pool_guarded(
_gen,
what="Batch generate",
timeout=generate_timeout_s(seg_text, engine=backend),
)
# Fit to slot
target_samples_seg = int(seg_duration * sr)
@@ -604,6 +806,8 @@ async def _run_batch_pipeline(job_id: str, job: dict):
f"left silent: {e}"
)
remove_segment_wavs(remote_segments)
# ── 3c. Save dubbed audio track ───────────────────────────────
# Invisible provenance mark on the assembled track (#1169), tensor
# stage, before the WAV write / aac mux — batch dubs used to ship
@@ -670,6 +874,7 @@ async def _run_batch_pipeline(job_id: str, job: dict):
outputs[target_lang] = output_path
job["outputs"] = outputs
job.pop("setup_required", None)
_set_progress(job, "done", 100)
@@ -681,6 +886,7 @@ async def enqueue_batch_job(
langs: str = Form("es"), # comma-separated lang codes
voice_id: Optional[str] = Form(None),
preserve_bg: bool = Form(True),
translation_provider: Optional[str] = Form(None),
):
"""Enqueue a video for batch dubbing.
@@ -694,6 +900,14 @@ async def enqueue_batch_job(
if not lang_list:
raise HTTPException(400, "At least one target language is required")
# Validate the snapshot before persisting a potentially large upload.
# Resolve it again in the worker so deleting or editing a queued profile
# cannot silently fall back to the engine's default voice.
try:
await asyncio.to_thread(_batch_voice, voice_id)
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
# TTS-only install: no ASR model on disk → typed 409 with a download CTA
# now, instead of accepting the job and having the transcribe stage
# silently auto-download multi-GB whisper weights (or fail) in the worker.
@@ -702,6 +916,19 @@ async def enqueue_batch_job(
if missing is not None:
raise HTTPException(409, {**missing, "message": asr_model_missing_detail(missing)})
# Snapshot the selected translation engine when the user enqueues the job,
# so a later Settings change cannot alter work already waiting in the queue.
from core import prefs
from services import translation_engines
provider = translation_provider or prefs.get("translation_backend", "argos")
if not translation_engines.get_engine(provider):
raise HTTPException(400, "Unknown translation engine")
if not translation_engines.is_installed(provider):
raise HTTPException(409, "Install the selected translation engine before adding this batch")
if not translation_engines.is_ready(provider):
raise HTTPException(409, "Configure the selected translation provider before adding this batch")
# Save the uploaded video
batch_dir = os.path.join(DATA_DIR, "batch")
os.makedirs(batch_dir, exist_ok=True)
@@ -718,7 +945,9 @@ async def enqueue_batch_job(
"langs": lang_list,
"voice_id": voice_id,
"preserve_bg": preserve_bg,
"translation_provider": provider,
"created_at": time.time(),
"attempts": 1,
"started_at": None,
"finished_at": None,
"error": None,
@@ -741,6 +970,8 @@ def list_batch_jobs(status: Optional[str] = None, limit: int = 50):
if status:
if status == "active":
jobs = [j for j in jobs if j["status"] in ("queued", "running")]
elif status == "retryable":
jobs = [j for j in jobs if j["status"] in ("failed", "cancelled")]
else:
jobs = [j for j in jobs if j["status"] == status]
jobs.sort(key=lambda j: j["created_at"], reverse=True)
@@ -764,14 +995,88 @@ def cancel_batch_job(job_id: str):
raise HTTPException(404, "Job not found")
if job["status"] in ("done", "failed", "cancelled"):
return {"already": job["status"]}
was_running = job["status"] == "running" or job_id in _processing_job_ids
job["status"] = "cancelled"
job["retry_ready"] = not was_running
job["finished_at"] = time.time()
return {"cancelled": True}
@router.post("/batch/jobs/{job_id}/retry")
async def retry_batch_job(job_id: str):
"""Retry a terminal job using its original app-owned upload and settings."""
job = _jobs.get(job_id)
if not job:
raise HTTPException(404, "Job not found")
if job["status"] not in ("failed", "cancelled"):
raise HTTPException(409, f"Job is {job['status']}, not retryable")
if job_id in _processing_job_ids or not job.get("retry_ready", True):
raise HTTPException(409, "The cancelled job is still stopping")
if not os.path.isfile(job.get("video_path") or ""):
raise HTTPException(409, "The original batch input is no longer available")
try:
await asyncio.to_thread(_batch_voice, job.get("voice_id"))
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
from services.asr_backend import asr_model_missing_detail, asr_model_missing_error
missing = await asyncio.to_thread(asr_model_missing_error)
if missing is not None:
raise HTTPException(409, {**missing, "message": asr_model_missing_detail(missing)})
from services import translation_engines
provider = job.get("translation_provider") or "argos"
if not translation_engines.is_ready(provider):
raise HTTPException(409, "Configure the selected translation provider before retrying")
if provider == "argos" and job.get("source_lang"):
status = await asyncio.to_thread(
translation_engines.argos_pack_status,
job["source_lang"],
job["langs"],
)
if any(not pair["installed"] for pair in status["pairs"]):
raise HTTPException(409, "Install the required Argos language packs before retrying")
batch_root = os.path.realpath(os.path.join(DATA_DIR, "batch"))
output_dir = os.path.realpath(os.path.join(batch_root, job_id))
if os.path.dirname(output_dir) != batch_root:
raise HTTPException(status_code=400, detail="Invalid batch job path")
try:
if os.path.isdir(output_dir):
await asyncio.to_thread(shutil.rmtree, output_dir)
except OSError as exc:
raise HTTPException(
status_code=500,
detail="Could not reset the batch output files. Close any app using them and retry.",
) from exc
for key in (
"duration",
"segments",
"source_lang",
"outputs",
"warnings",
"setup_required",
"retry_ready",
):
job.pop(key, None)
job.update({
"status": "queued",
"started_at": None,
"finished_at": None,
"error": None,
"progress": None,
"attempts": int(job.get("attempts", 1)) + 1,
})
_ensure_queue()
await _queue.put(job_id)
return {"job_id": job_id, "status": "queued", "queue_position": _queue.qsize()}
@router.delete("/batch/jobs/{job_id}")
def delete_batch_job(job_id: str):
"""Delete a batch job record and its video file."""
"""Delete a batch job record and every app-owned input/output file."""
job = _jobs.get(job_id)
if not job:
raise HTTPException(404, "Job not found")
@@ -783,6 +1088,18 @@ def delete_batch_job(job_id: str):
status_code=500,
detail="Could not delete the batch video file. Close any app using it and retry.",
) from exc
batch_root = os.path.realpath(os.path.join(DATA_DIR, "batch"))
output_dir = os.path.realpath(os.path.join(batch_root, job_id))
if os.path.dirname(output_dir) != batch_root:
raise HTTPException(status_code=400, detail="Invalid batch job path")
try:
if os.path.isdir(output_dir):
shutil.rmtree(output_dir)
except OSError as exc:
raise HTTPException(
status_code=500,
detail="Could not delete the batch output files. Close any app using them and retry.",
) from exc
_jobs.pop(job_id, None)
return {"deleted": True}
+99 -13
View File
@@ -28,6 +28,17 @@ router = APIRouter()
logger = logging.getLogger("omnivoice.capture")
def _timing(value):
"""A segment timing, or ``None`` when the engine could not determine one.
``dict.get(key, 0)`` hands back a stored ``None`` rather than the default,
because the key is present — so rounding it raised and took a transcript
that was otherwise fine down with it (#1904). Pass the null through instead:
the segment list renders whichever half of the range is known.
"""
return round(value, 2) if isinstance(value, (int, float)) else None
def _truthy(value: Optional[str]) -> bool:
"""Parse a multipart form flag. Treats '1'/'true'/'yes'/'on'/'auto'
(any case) as on; everything else — including None — as off."""
@@ -49,7 +60,8 @@ async def transcribe_audio(
language: Optional language hint (not currently used; auto-detected).
model: Whisper model size (legacy; ignored in dual-mode architecture).
mode: 'fast' (default) uses MLX Turbo for speed; 'accurate' uses
WhisperX with forced alignment for word-level timing.
the selected ASR engine with word-level timing. 'reference' uses
the selected ASR engine without word-level timing.
refine: Opt-in local-LLM cleanup of the final text (disfluencies,
self-corrections, punctuation) — same pipeline the live
dictation socket uses. Off by default so MCP/CLI callers don't
@@ -80,7 +92,9 @@ async def transcribe_audio(
tmp.write(content)
tmp.close()
use_accurate = (mode or "").strip().lower() == "accurate"
requested_mode = (mode or "").strip().lower()
use_accurate = requested_mode == "accurate"
use_active_asr = requested_mode in {"accurate", "reference"}
# TTS-only install: no ASR model on disk → typed 409 with a download
# CTA, BEFORE any backend is constructed (the whisper backends
@@ -88,7 +102,8 @@ async def transcribe_audio(
from services.asr_backend import asr_model_missing_detail, asr_model_missing_error
missing = await asyncio.to_thread(
asr_model_missing_error,
purpose="transcribe" if use_accurate else "dictation",
purpose="transcribe" if use_active_asr else "dictation",
require_installed=requested_mode == "reference",
)
if missing is not None:
raise HTTPException(
@@ -97,7 +112,7 @@ async def transcribe_audio(
)
def _run():
if use_accurate:
if use_active_asr:
# Accurate mode: full WhisperX with forced alignment —
# for when the user explicitly wants word-level timing.
# `load_*`, not `get_*`: the selector alone hands back an
@@ -105,8 +120,8 @@ async def transcribe_audio(
# chain is broken, which then 500s at `.transcribe()`. The
# loader degrades to the next healthy engine (#1185).
from services.asr_backend import load_active_asr_backend
backend = load_active_asr_backend()
result = backend.transcribe(tmp.name, word_timestamps=True)
backend = load_active_asr_backend(require_installed=True) if requested_mode == "reference" else load_active_asr_backend()
result = backend.transcribe(tmp.name, word_timestamps=use_accurate)
else:
# Fast mode (default): use the fastest available engine
# (MLX Turbo on Apple Silicon). Skip word_timestamps for
@@ -114,7 +129,8 @@ async def transcribe_audio(
from services.asr_backend import get_capture_asr_backend
backend = get_capture_asr_backend()
result = backend.transcribe(tmp.name, word_timestamps=False)
return result, backend.id
sherpa_model_id = getattr(getattr(backend, "spec", None), "id", None)
return result, backend.id, sherpa_model_id
from services.model_manager import _gpu_pool
from services.asr_backend import (
@@ -124,7 +140,7 @@ async def transcribe_audio(
)
t0 = time.perf_counter()
try:
result, engine_id = await run_transcribe_guarded(
result, engine_id, sherpa_model_id = await run_transcribe_guarded(
_gpu_pool, _run, what="Dictation",
)
except ASRTimeoutError as e:
@@ -140,6 +156,69 @@ async def transcribe_audio(
status_code=409,
detail={**e.payload, "message": asr_model_missing_detail(e.payload)},
)
# Some sherpa-onnx NeMo-TDT builds load successfully but decode an
# entire spoken clip to no tokens. Live dictation already recovers
# from that failure; the shared file endpoint must do the same because
# it also powers uploaded transcription and automatic profile text.
# Retry only through an already-installed fallback, and demote the
# silent model only when the second recognizer actually heard words.
initial_text = str(result.get("text") or "").strip()
if not initial_text and result.get("segments"):
initial_text = " ".join(
str(segment.get("text") or "")
for segment in result["segments"]
if isinstance(segment, dict)
).strip()
recovered_from = None
if not use_active_asr and sherpa_model_id and not initial_text:
fallback_missing = await asyncio.to_thread(
asr_model_missing_error,
purpose="dictation",
skip_sherpa=True,
require_installed=True,
)
if fallback_missing is None:
def _run_fallback():
from services.asr_backend import get_capture_asr_backend
fallback = get_capture_asr_backend(skip_sherpa=True)
return (
fallback.transcribe(tmp.name, word_timestamps=False),
fallback.id,
)
try:
fallback_result, fallback_engine_id = await run_transcribe_guarded(
_gpu_pool,
_run_fallback,
what="Dictation fallback",
)
fallback_text = str(fallback_result.get("text") or "").strip()
if not fallback_text and fallback_result.get("segments"):
fallback_text = " ".join(
str(segment.get("text") or "")
for segment in fallback_result["segments"]
if isinstance(segment, dict)
).strip()
if fallback_text:
from services.sherpa_dictation import demote_model
await asyncio.to_thread(demote_model, sherpa_model_id)
result = fallback_result
engine_id = fallback_engine_id
recovered_from = sherpa_model_id
logger.warning(
"File transcription recovered from silent dictation model %s "
"through installed engine %s",
sherpa_model_id,
fallback_engine_id,
)
except Exception:
logger.exception(
"Installed fallback failed after dictation model %s returned no text",
sherpa_model_id,
)
elapsed = round(time.perf_counter() - t0, 2)
# Normalize result shape
@@ -162,10 +241,15 @@ async def transcribe_audio(
from services.text_polish import polish_text
full_text = polish_text(full_text)
# Calculate audio duration from segments if available
# Calculate audio duration from segments if available. A segment whose
# timing the engine could not determine carries end=None (sherpa's
# _sherpa_result when the sample rate yields no duration, and every
# plain-text OpenAI-compatible response), so measure only the ones that
# have a number and keep 0.0 when none do.
duration = 0.0
if segments:
duration = max(s.get("end", 0) for s in segments)
ends = [e for e in (s.get("end") for s in segments) if isinstance(e, (int, float))]
duration = max(ends) if ends else 0.0
detected_lang = result.get("language", language or "unknown")
@@ -186,7 +270,7 @@ async def transcribe_audio(
logger.info(
"Capture transcription done: engine=%s, elapsed=%.2fs, duration=%.1fs, mode=%s, refined=%s",
engine_id, elapsed, duration, "accurate" if use_accurate else "fast",
engine_id, elapsed, duration, requested_mode if use_active_asr else "fast",
refined_text is not None,
)
@@ -194,8 +278,8 @@ async def transcribe_audio(
"text": full_text,
"segments": [
{
"start": round(s.get("start", 0), 2),
"end": round(s.get("end", 0), 2),
"start": _timing(s.get("start", 0)),
"end": _timing(s.get("end", 0)),
"text": s.get("text", "").strip(),
}
for s in segments
@@ -207,6 +291,8 @@ async def transcribe_audio(
}
if refined_text is not None:
response["refined_text"] = refined_text
if recovered_from is not None:
response["model_silent"] = recovered_from
return response
finally:
try:
+29 -4
View File
@@ -55,6 +55,17 @@ from services.text_polish import polish_text
router = APIRouter()
logger = logging.getLogger("omnivoice.capture_ws")
def _timing(value):
"""A segment timing, or ``None`` when the engine could not determine one.
``dict.get(key, 0)`` returns a stored ``None`` rather than the default, so
rounding it raised (#1904). The null is the honest answer here — this module
emits it deliberately for un-endpointed utterances and the segment list
renders whichever half of the range is known.
"""
return round(value, 2) if isinstance(value, (int, float)) else None
SPEECH_PROTOCOL = "voicestudio.speech.v1"
PLATFORM_STREAM_PATH = "/v1/audio/transcriptions/stream"
@@ -349,6 +360,7 @@ async def ws_transcribe(websocket: WebSocket):
audio_chunks: list[bytes] = []
total_bytes = 0
last_audio_time = time.monotonic()
paused = False
running = True
partial_text = ""
# Track whether the client initiated the disconnect. When True the
@@ -366,7 +378,7 @@ async def ws_transcribe(websocket: WebSocket):
message as the authoritative result and skip the duplicate HTTP
POST that used to run on every dictation.
"""
nonlocal total_bytes, last_audio_time, running, client_disconnected
nonlocal total_bytes, last_audio_time, running, client_disconnected, paused
try:
while running:
msg = await websocket.receive()
@@ -397,6 +409,10 @@ async def ws_transcribe(websocket: WebSocket):
total_bytes += len(data)
last_audio_time = time.monotonic()
continue
if msg.get("text") in ("PAUSE", "RESUME"):
paused = msg["text"] == "PAUSE"
last_audio_time = time.monotonic()
continue
if _is_end_control(msg.get("text")):
# Client signals end-of-audio but stays connected for `final`.
running = False
@@ -429,6 +445,9 @@ async def ws_transcribe(websocket: WebSocket):
if not running:
break
if paused:
continue
# Check silence timeout
if time.monotonic() - last_audio_time > SILENCE_TIMEOUT_S and total_bytes > MIN_BUFFER_BYTES:
running = False
@@ -1187,13 +1206,19 @@ async def _transcribe_buffer_full(
from services.refinement import collapse_repetitive_artifacts
full_text = collapse_repetitive_artifacts(full_text)
duration = max((s.get("end", 0) for s in segments), default=0.0)
# end=None means the engine could not determine the timing — this
# module writes exactly that in its own streaming payloads, and
# sherpa's _sherpa_result does too when the sample rate yields no
# duration. Measure only real numbers, and pass the nulls through
# rather than rounding them (#1904).
ends = [e for e in (s.get("end") for s in segments) if isinstance(e, (int, float))]
duration = max(ends) if ends else 0.0
return {
"text": full_text,
"segments": [
{"start": round(s.get("start", 0), 2),
"end": round(s.get("end", 0), 2),
{"start": _timing(s.get("start", 0)),
"end": _timing(s.get("end", 0)),
"text": s.get("text", "").strip()}
for s in segments
],
+18 -1
View File
@@ -21,7 +21,7 @@ import logging
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel
from typing import Optional
from typing import Literal, Optional
from api.dependencies import require_local
from api.public_engine_metadata import public_unavailability
@@ -86,6 +86,23 @@ def list_dictation_models():
}
@router.get("/dictation/readiness", dependencies=[Depends(require_local)])
def dictation_readiness(
model_id: str | None = None,
purpose: Literal["dictation", "transcribe"] = "dictation",
) -> dict:
"""Check a selected ASR path without loading or downloading weights."""
from services.asr_backend import asr_model_missing_error
missing = asr_model_missing_error(
purpose=purpose,
sherpa_model_id=(model_id or _read_prefs()["model_id"])
if purpose == "dictation"
else None,
)
return {"ready": missing is None, "missing": missing}
@router.get("/dictation/prefs", dependencies=[Depends(require_local)])
def get_dictation_prefs():
return _read_prefs()
+231 -21
View File
@@ -355,6 +355,138 @@ async def dub_import_srt(job_id: str, file: UploadFile = File(...)):
}
def _select_downloaded_caption_track(
tracks: dict[str, list[dict]], preferred: str | None,
) -> str | None:
"""Choose the closest original-language caption track deterministically."""
available = [key for key, cues in tracks.items() if isinstance(cues, list) and cues]
if not available:
return None
preferred_tag = (preferred or "").strip().lower().replace("_", "-")
preferred_base = preferred_tag.split("-", 1)[0]
def rank(key: str) -> tuple[int, int, int, str]:
tag = key.strip().lower().replace("_", "-")
base = tag.split("-", 1)[0]
if preferred_tag:
language_rank = 0 if tag == preferred_tag else 1 if base == preferred_base else 2
else:
language_rank = 0
return (
language_rank,
0 if tag.endswith("-orig") else 1,
0 if "-" not in tag else 1,
tag,
)
return min(available, key=rank)
def _prepare_downloaded_caption_segments(cues: list[dict], duration: float) -> list[dict]:
"""Normalize downloaded VTT cues into safe, sequential Dub segments."""
def cue_start(cue: dict) -> float:
try:
return float(cue.get("start") or 0.0)
except (TypeError, ValueError):
return 0.0
def remove_repeated_prefix(previous: str, current: str) -> str:
previous_words = previous.split()
current_words = current.split()
folded_previous = [word.casefold() for word in previous_words]
folded_current = [word.casefold() for word in current_words]
for count in range(min(len(previous_words), len(current_words)), 0, -1):
if folded_previous[-count:] == folded_current[:count]:
return " ".join(current_words[count:])
return current
prepared: list[dict] = []
previous_end = 0.0
ordered = sorted((cue for cue in cues if isinstance(cue, dict)), key=cue_start)
for index, cue in enumerate(ordered):
try:
raw_start = max(0.0, float(cue.get("start") or 0.0))
end = float(cue.get("end") or raw_start)
except (TypeError, ValueError):
continue
text = " ".join(str(cue.get("text") or "").split())
if duration > 0:
if raw_start >= duration:
continue
end = min(end, duration)
if prepared and raw_start < previous_end:
text = remove_repeated_prefix(prepared[-1]["text"], text)
if not text:
prepared[-1]["end"] = round(max(previous_end, end), 3)
previous_end = max(previous_end, end)
continue
# Caption hosts commonly emit slightly overlapping cues. Dubbing needs
# a monotonic timeline, so trim the later cue rather than manufacture
# overlapping speech slots.
start = max(raw_start, previous_end)
if not text or end <= start:
continue
prepared.append({
"id": str(index),
"start": round(start, 3),
"end": round(end, 3),
"text": text,
"speaker_id": "Speaker 1",
})
previous_end = end
cleaned = clean_up_segments(prepared)
return [
{
**segment,
"id": index,
"text_original": segment.get("text", ""),
}
for index, segment in enumerate(cleaned)
]
@router.post("/dub/use-downloaded-captions/{job_id}")
def dub_use_downloaded_captions(job_id: str):
"""Seed a prepared Dub job from its downloaded caption track."""
job = _get_job(job_id)
if not job:
raise HTTPException(status_code=404, detail="Job not found")
tracks = job.get("youtube_subs")
if not isinstance(tracks, dict):
raise HTTPException(status_code=404, detail="No downloaded captions are available")
caption_lang = _select_downloaded_caption_track(
tracks,
job.get("source_lang_override") or job.get("source_lang"),
)
if caption_lang is None:
raise HTTPException(status_code=404, detail="No downloaded captions are available")
segments = _prepare_downloaded_caption_segments(
tracks[caption_lang],
float(job.get("duration") or 0.0),
)
if not segments:
raise HTTPException(status_code=422, detail="Downloaded captions contain no usable cues")
source_lang = job.get("source_lang_override") or _detected_source_lang(caption_lang)
job["segments"] = segments
job["source_lang"] = source_lang
job["full_transcript"] = " ".join(segment["text"] for segment in segments)
# Caption files contain timing and text, but no trustworthy speaker or
# reference-audio attribution. Never retain stale clone maps from a prior
# transcript on the same job.
job["segment_clones"] = {}
job["speaker_clones"] = {}
job.pop("cast_sources", None)
_save_job(job_id, job)
return {
"segments": segments,
"source_lang": source_lang,
"caption_lang": caption_lang,
"available": sorted(tracks.keys()),
}
@router.post("/dub/cleanup-segments/{job_id}")
def dub_cleanup_segments(job_id: str):
"""Re-run merge/stitch passes on a job's existing segments to drop fragments."""
@@ -375,8 +507,7 @@ def dub_abort(job_id: str):
had_procs = bool(_active_procs.get(job_id))
_kill_job_procs(job_id)
try:
if task_manager.cancel_task(job_id) is False:
raise RuntimeError("task cancellation was declined")
had_task = task_manager.cancel_task(job_id)
except Exception as exc:
logger.warning("Dub task cancellation failed")
raise HTTPException(
@@ -386,7 +517,13 @@ def dub_abort(job_id: str):
job = _dub_jobs.get(job_id)
if job is not None:
job["aborted"] = True
return {"aborted": True, "had_active_procs": had_procs}
# Cancellation is idempotent: a missing active task means it already
# stopped between the renderer aborting its stream and this request.
return {
"aborted": True,
"had_active_procs": had_procs,
"had_active_task": had_task,
}
@router.get("/dub/history")
@@ -540,12 +677,30 @@ _DUB_SOURCE_LANG_CODES = frozenset({
def _source_lang_override(value: str | None) -> str | None:
"""Normalize a user-selected source language; auto/und means detect."""
"""Normalize a user-selected source language; auto/und means detect.
A rejection NAMES the code it rejected. "Invalid source language code" on
its own cannot be acted on or reported usefully: it does not say which of
the ninety-odd codes was wrong, so neither the user nor a maintainer
reading the auto-filed issue can tell whether the picker offered something
the backend does not accept, or a stale preference from an older build is
still being sent (#1960).
The value is a language code the user chose from a menu not private
data and the neighbouring engine validator already echoes its input the
same way.
"""
code = (value or "").strip().lower()
if code in {"", "auto", "und"}:
return None
if code not in _DUB_SOURCE_LANG_CODES:
raise HTTPException(status_code=400, detail="Invalid source language code")
raise HTTPException(
status_code=400,
detail=(
f"Invalid source language code: {code!r}. Pick a language from "
"the Dubbing source-language menu, or leave it on auto-detect."
),
)
return code
@@ -597,8 +752,19 @@ async def dub_upload(
os.makedirs(job_dir, exist_ok=True)
video_path = os.path.join(job_dir, f"original{ext}")
with open(video_path, "wb") as f:
f.write(await video.read())
def _stream_upload_to_disk() -> None:
# UploadFile is already a spooled file. Copy it in bounded chunks on a
# worker thread instead of materialising a multi-GB video in RAM and
# blocking every API request while the event loop writes it.
video.file.seek(0)
with open(video_path, "wb") as output:
shutil.copyfileobj(video.file, output, length=1024 * 1024)
try:
await asyncio.to_thread(_stream_upload_to_disk)
finally:
await video.close()
filename = video.filename or f"video{ext}"
task_id = f"prep_{job_id}"
@@ -714,7 +880,7 @@ _prep_event_helper = dub_pipeline.prep_event # alias; we keep the module-local
#: into one reference, which is how "made up" clone voices happen).
CLONE_SKIP_HEURISTIC_MSG = (
"auto voice cloning skipped: speaker labels are gap-based estimates — "
"set up diarization (Model Catalogue → Models → pyannote) for per-speaker clones"
"set up diarization (Model Catalogue → Other weights → pyannote) for per-speaker clones"
)
@@ -941,6 +1107,23 @@ async def dub_transcribe_stream(
job = _get_job(job_id)
# The durable job is written before the terminal SSE events below. If
# the renderer, proxy, or backend connection drops in that narrow
# window, reconnecting must replay the completed result instead of
# running a second whole-file ASR pass. This is deliberately gated by
# an explicit completion marker so partial work and imported subtitle
# rows still take their established paths.
if job and job.get("transcription_complete") and isinstance(job.get("segments"), list):
yield _sse_event("final", {
"segments": job["segments"],
"source_lang": job.get("source_lang") or "en",
"full_transcript": job.get("full_transcript") or "",
"speaker_clones": job.get("cast_sources", {}),
"cast_sources": job.get("cast_sources", {}),
})
yield _sse_event("done", {})
return
preflight_error: Optional[str] = None
# Extra machine-readable fields merged into the preflight `error` SSE event
# (e.g. the typed asr_model_missing payload → download-CTA in the UI).
@@ -1435,6 +1618,7 @@ async def dub_transcribe_stream(
from services.model_manager import (
DIARIZATION_ERR_LICENSE,
DIARIZATION_ERR_NO_TOKEN,
DIARIZATION_ERR_MISSING,
)
from core import error_docs_map
@@ -1486,7 +1670,7 @@ async def dub_transcribe_stream(
f"unavailable, so the ASR engine's built-in speaker "
f"turns were used and the detected count may differ "
f"from the {num_speakers} you set. Set up diarization "
f"(Model Catalogue → Models → pyannote) to enforce an exact "
f"(Model Catalogue → Other weights → pyannote) to enforce an exact "
f"speaker count."
)
return resplit, {
@@ -1527,7 +1711,23 @@ async def dub_transcribe_stream(
from services import token_resolver
resolved = token_resolver.resolve()
if err_sentinel == DIARIZATION_ERR_NO_TOKEN or not resolved:
if err_sentinel == DIARIZATION_ERR_MISSING:
from services.diarization_runtime import SORTFORMER, selected_backend
native_selected = selected_backend() == SORTFORMER
detail = (
"Native Sortformer files are missing. Install audiocpp_cli beside "
"the audio.cpp native bundle in Settings > Models > "
"Diarisation, then retry transcription. "
"Using silence gaps for now; rapid speaker turns may be merged."
) if native_selected else (
"Speaker diarization files are missing or incomplete. "
"Install or repair pyannote in Settings > Models > Diarisation, "
"then retry transcription. No models were downloaded during "
"this job. Using silence gaps for now; rapid speaker turns "
"may be merged."
)
error_class = "DIARIZATION_MODEL_MISSING"
elif err_sentinel == DIARIZATION_ERR_NO_TOKEN:
detail = (
"Speaker diarization is disabled because no HuggingFace token "
"was found in any source (Settings → API Keys, the HF_TOKEN "
@@ -1540,12 +1740,12 @@ async def dub_transcribe_stream(
)
error_class = "HF_AUTH_FAILED"
elif err_sentinel == DIARIZATION_ERR_LICENSE:
who = resolved.username or "(whoami suppressed)"
who = resolved.username if resolved else "(not signed in)"
detail = (
f"Speaker diarization model is gated — the "
f"pyannote/speaker-diarization-3.1 license has not been "
f"accepted on HuggingFace by this account "
f"(source={resolved.source}, user={who}). Visit "
f"(user={who}). Visit "
f"huggingface.co/pyannote/speaker-diarization-3.1 AND "
f"huggingface.co/pyannote/segmentation-3.0 while signed "
f"in and click 'Agree and access repository' on both, "
@@ -1557,17 +1757,13 @@ async def dub_transcribe_stream(
else:
# err_sentinel == DIARIZATION_ERR_LOAD (or unexpected None
# with a resolved token — historical safety net).
who = resolved.username or "(whoami suppressed)"
detail = (
f"Speaker diarization model failed to load even though an HF "
f"token was found (source={resolved.source}, user={who}). "
f"Most common causes: the pyannote/speaker-diarization-3.1 "
f"license has not been accepted on HuggingFace, or there is "
f"a pyannote/torch version mismatch. See backend logs for "
f"The installed speaker diarization model failed to load. "
f"See Settings > Logs > Backend for "
f"the underlying error. Falling back to a silence-gap "
f"heuristic; rapid speaker turns may be merged."
)
error_class = "PYANNOTE_LICENSE_REQUIRED"
error_class = "DIARIZATION_LOAD_FAILED"
warning = {
"detail": detail + _hint_suffix(),
"error_class": error_class,
@@ -1588,7 +1784,13 @@ async def dub_transcribe_stream(
# provided (#274). pyannote's apply() accepts num_speakers;
# omit it entirely when None so we don't depend on the kwarg
# existing in every pyannote build.
if num_speakers:
from services.diarization_native import NativeSortformer
if isinstance(diar_pipe, NativeSortformer):
diar = diar_pipe(
asr_audio_target, num_speakers=num_speakers, job_id=job_id,
cancel_check=lambda: bool(job.get("aborted")) or task_manager.is_cancelled(job_id),
)
elif num_speakers:
logger.info("Diarizing with num_speakers=%d (user hint)", num_speakers)
diar = diar_pipe(asr_audio_target, num_speakers=num_speakers)
else:
@@ -1614,7 +1816,7 @@ async def dub_transcribe_stream(
len(asr_phrase_segments), separation,
)
return recovered_segments, None, "phrase_embeddings"
return resplit, None, "pyannote"
return resplit, None, "audiocpp-sortformer" if isinstance(diar_pipe, NativeSortformer) else "pyannote"
except Exception as e:
logger.exception("Diarization failed")
# Inline ASR turns beat the silence-gap heuristic as a crash
@@ -1663,6 +1865,9 @@ async def dub_transcribe_stream(
final_segs, diar_warning, labels_source = done.pop().result()
break
yield _sse_event("ping", {})
if job.get("aborted") or task_manager.is_cancelled(job_id):
yield _sse_event("aborted", {})
return
if diar_warning:
logger.warning("diarization fallback: %s", diar_warning.get("detail"))
payload = {
@@ -1677,6 +1882,8 @@ async def dub_transcribe_stream(
payload["speaker_hint"] = diar_warning["speaker_hint"]
yield _sse_event("warning", payload)
from services.segmentation import deduplicate_chunk_segments
final_segs = deduplicate_chunk_segments(final_segs)
job["segments"] = final_segs
# Auto-speaker-clone: sample each detected speaker's voice from the
@@ -1827,6 +2034,7 @@ async def dub_transcribe_stream(
detected_lang
)
job["full_transcript"] = " ".join(s.get("text", "") for s in final_segs)
job["transcription_complete"] = True
_save_job(job_id, job)
# Restore TTS model to GPU now that ASR is done. unload() blocks
@@ -2088,6 +2296,8 @@ async def dub_transcribe(job_id: str, num_speakers: Optional[int] = None):
raise
if job.get("aborted"):
raise HTTPException(status_code=499, detail="Transcription aborted")
from services.segmentation import deduplicate_chunk_segments
segments_result = deduplicate_chunk_segments(segments_result)
job["segments"] = segments_result
source_lang = job.get("source_lang")
_save_job(job_id, job)
+103 -35
View File
@@ -15,7 +15,7 @@ from core.http_headers import content_disposition
from core.logging_utils import log_safe
from core.path_security import UnsafePath, resolve_within
from core.tasks import task_manager
from fastapi import APIRouter, Header, HTTPException, Query, Response
from fastapi import APIRouter, Header, HTTPException, Query, Request, Response
from fastapi.responses import FileResponse, StreamingResponse
from services.ffmpeg_utils import (
bed_mix_filter,
@@ -38,6 +38,31 @@ router = APIRouter()
logger = logging.getLogger("omnivoice.api")
async def _preserved_background(job: dict, job_id: str, lang: str, *, prepare: bool = True) -> str:
"""All mixed preview/download paths share the same dialogue-only bed."""
from services.dub_background import surgical_background
bed = _optional_dub_artifact(job.get("no_vocals_path"), job_id)
source = _optional_dub_artifact(job.get("video_path"), job_id) or _optional_dub_artifact(job.get("audio_path"), job_id)
if not bed or not source:
raise HTTPException(status_code=409, detail={"code": "dub_background_unavailable", "message": "Original audio and background separation are required"})
track = (job.get("dubbed_tracks") or {}).get(lang) or {}
segments = track.get("source_segments") or job.get("segments") or []
if not segments:
raise HTTPException(status_code=409, detail={"code": "dub_background_unavailable", "message": "Dialogue timing is required"})
if not prepare:
return bed
strategy = track.get("timing_strategy") or job.get("timing_strategy")
plans = job.get("fit_plans" if strategy == "smart_fit" else "video_stretch_plans") or {}
entry = (plans.get(lang) or {}) if strategy in {"smart_fit", "stretch_video"} else {}
directory = os.path.join(_existing_job_dir_or_404(job_id), "exports")
os.makedirs(directory, exist_ok=True)
try:
return await surgical_background(source, bed, directory, segments, entry.get("plan") or [], float(entry.get("orig_duration") or job.get("duration") or 0))
except (ValueError, RuntimeError) as exc:
raise HTTPException(status_code=409, detail={"code": "dub_background_unavailable", "message": str(exc)}) from exc
def _unique_stamp() -> str:
"""Return a short unique suffix like '20260415T142301-ab12cd34' for export files."""
return f"{time.strftime('%Y%m%dT%H%M%S')}-{uuid.uuid4().hex[:8]}"
@@ -596,7 +621,7 @@ def _build_audio_export_cmd(
# Mix the dubbed voice over the original background bed (same weights
# as the video mux path) so ambience/music is preserved.
cmd += ["-i", bg_path, "-filter_complex",
bed_mix_filter("1:a", "0:a"),
bed_mix_filter("1:a", "0:a", bed_gain=1.0),
"-map", "[aout]"]
cmd += codec
cmd.append(out_path)
@@ -691,7 +716,7 @@ async def dub_download(
else:
output_name = f"dubbed_audio_{stamp}.m4a"
out_path = os.path.join(exports_dir, output_name)
bg = _optional_dub_artifact(job.get("no_vocals_path"), job_id) if preserve_bg else None
bg = await _preserved_background(job, job_id, lang_code) if preserve_bg else None
cmd = _build_audio_export_cmd(ffmpeg, track_info["path"], bg, out_path, fmt)
try:
rc, _, stderr = await run_ffmpeg(cmd, timeout=1800.0)
@@ -839,17 +864,16 @@ async def dub_download(
retimed_idx = input_idx
input_idx += 1
bg_audio = _optional_dub_artifact(job.get("no_vocals_path"), job_id) if preserve_bg else None
bg_idx = None
if bg_audio and filtered_tracks:
cmd += ["-i", bg_audio]
bg_idx = input_idx
input_idx += 1
tracks_to_process = []
for lang_code, track_info in filtered_tracks.items():
if preserve_bg:
bg_audio = await _preserved_background(job, job_id, lang_code)
cmd += ["-i", bg_audio]
bg_idx = input_idx
input_idx += 1
cmd += ["-i", track_info["path"]]
tracks_to_process.append({"lang_code": lang_code, "idx": input_idx, "info": track_info})
tracks_to_process.append({"lang_code": lang_code, "idx": input_idx, "bg_idx": bg_idx, "info": track_info})
input_idx += 1
filter_parts: list[str] = []
@@ -918,7 +942,7 @@ async def dub_download(
for i, t in enumerate(tracks_to_process):
tail = f",apad=whole_dur={apad_dur:.4f}" if apad_dur else ""
filter_parts.append(bed_mix_filter(
f"{bg_idx}:a", f"{t['idx']}:a", out=f"aout{i}", tail=tail, uniq=str(i),
f"{t['bg_idx']}:a", f"{t['idx']}:a", out=f"aout{i}", tail=tail, uniq=str(i), bed_gain=1.0,
))
t["out_label"] = f"[aout{i}]"
for t in tracks_to_process:
@@ -1050,8 +1074,8 @@ _MEDIA_TYPES = {
}
@router.get("/dub/media/{job_id}")
async def dub_get_media(job_id: str):
@router.api_route("/dub/media/{job_id}", methods=["GET", "HEAD"])
async def dub_get_media(job_id: str, request: Request):
_job_dir_or_400(job_id)
job = _get_job(job_id)
if not job:
@@ -1064,7 +1088,15 @@ async def dub_get_media(job_id: str):
# silent black box. Default to video/mp4 because the ingest pipeline
# remuxes URL downloads to mp4 (dub_pipeline.yt_download_sync).
ext = os.path.splitext(video_path)[1].lower()
return FileResponse(video_path, media_type=_MEDIA_TYPES.get(ext, "video/mp4"))
media_type = _MEDIA_TYPES.get(ext, "video/mp4")
headers = {
"Cache-Control": "private, max-age=31536000, immutable",
"Accept-Ranges": "bytes",
}
if request.method == "HEAD":
headers["Content-Length"] = str(os.path.getsize(video_path))
return Response(media_type=media_type, headers=headers)
return FileResponse(video_path, media_type=media_type, headers=headers)
# One mux at a time per preview file. Without this, two overlapping requests
# (e.g. the <video> element remounting right after a re-dub) both ran ffmpeg
@@ -1081,8 +1113,9 @@ def _preview_lock(path: str) -> asyncio.Lock:
return lock
@router.get("/dub/preview-video/{job_id}")
@router.api_route("/dub/preview-video/{job_id}", methods=["GET", "HEAD"])
async def dub_preview_video(
request: Request,
job_id: str,
lang: str = Query(..., description="Language code of the dubbed track to mux in"),
preserve_bg: bool = Query(True),
@@ -1110,7 +1143,7 @@ async def dub_preview_video(
video_path = _dub_artifact(job.get("video_path"), job_id, missing_detail="Source video missing")
bg_audio = _optional_dub_artifact(job.get("no_vocals_path"), job_id) if preserve_bg else None
bg_audio = await _preserved_background(job, job_id, lang, prepare=request.method != "HEAD") if preserve_bg else None
has_bg = bool(bg_audio)
# realpath-normalised + containment-checked inline BEFORE any filesystem
@@ -1122,9 +1155,9 @@ async def dub_preview_video(
if not exports_dir.startswith(_base + os.sep):
raise HTTPException(status_code=400, detail="Invalid job id")
os.makedirs(exports_dir, exist_ok=True)
bg_suffix = "bg" if (preserve_bg and has_bg) else "nobg"
bg_suffix = "surgical_v2_" + Path(bg_audio).stem if (preserve_bg and has_bg) else "nobg"
preview_path = os.path.realpath(
os.path.join(exports_dir, f"preview_{lang}_{bg_suffix}.mp4")
os.path.join(exports_dir, f"preview_v2_{lang}_{bg_suffix}.mp4")
)
if not preview_path.startswith(_base + os.sep):
raise HTTPException(status_code=400, detail="Invalid path")
@@ -1138,6 +1171,18 @@ async def dub_preview_video(
and os.path.getmtime(preview_path) >= track_mtime
)
# Vidstack probes extensionless routes with HEAD before choosing a native
# provider. Confirm that this preview is valid without starting an ffmpeg
# mux; the following GET builds it lazily when needed.
if request.method == "HEAD":
headers = {
"Cache-Control": "private, max-age=31536000, immutable",
"Accept-Ranges": "bytes",
}
if _cache_ok():
headers["Content-Length"] = str(os.path.getsize(preview_path))
return Response(media_type="video/mp4", headers=headers)
async def _mux_preview():
# Mux into a temp file and os.replace() into place so a concurrent
# reader never sees a partially-written preview (#281: video stuck
@@ -1249,7 +1294,7 @@ async def dub_preview_video(
audio_map = f"{track_idx}:a:0"
if bg_idx is not None:
tail = f",apad=whole_dur={apad_dur:.4f}" if apad_dur else ""
filter_parts.append(bed_mix_filter(f"{bg_idx}:a", f"{track_idx}:a", tail=tail))
filter_parts.append(bed_mix_filter(f"{bg_idx}:a", f"{track_idx}:a", tail=tail, bed_gain=1.0))
audio_map = "[aout]"
elif apad_dur:
filter_parts.append(f"[{track_idx}:a]apad=whole_dur={apad_dur:.4f}[aout]")
@@ -1266,7 +1311,7 @@ async def dub_preview_video(
cmd += ["-c:v", "libx264", "-preset", "medium", "-crf", "20", "-pix_fmt", "yuv420p"]
else:
cmd += ["-c:v", "copy"]
cmd += ["-c:a", "aac", "-b:a", "192k"]
cmd += ["-c:a", "aac", "-b:a", "192k", "-movflags", "+faststart"]
# `-shortest` would cut the retimed video at the (slightly different)
# audio length and lose the trailing frame; only use it on the copy path.
if not stretch_entry and retime_decision is None:
@@ -1311,22 +1356,37 @@ async def dub_preview_video(
if not _cache_ok():
await _mux_preview()
# no-store: the URL is stable across re-dubs, so any HTTP-level caching
# in the WebView would keep showing the previous dub after a re-generate
# (#281: "edits don't change the result").
# The renderer includes the segment-fingerprint revision in the URL, so a
# regenerated track gets a fresh cache key. Keep each completed preview:
# switching Original/Dub then reuses local ranges instead of re-reading a
# multi-hundred-megabyte MP4 from the backend.
return FileResponse(
preview_path,
media_type="video/mp4",
headers={"Cache-Control": "no-store"},
headers={"Cache-Control": "private, max-age=31536000, immutable", "Accept-Ranges": "bytes"},
)
def _compute_onsets_sync(src_path: str) -> list[float]:
def _compute_timeline_sync(src_path: str) -> tuple[list[float], list[float]]:
"""Blocking part of onset analysis — runs in a worker thread."""
import numpy as np
import soundfile as sf
from services.onset_align import detect_speech_onsets
audio, sr = sf.read(src_path, dtype="float32")
return detect_speech_onsets(audio, sr)
onsets = detect_speech_onsets(audio, sr)
mono = np.asarray(audio, dtype=np.float32)
if mono.ndim > 1:
mono = mono.mean(axis=1)
mono = mono.reshape(-1)
if mono.size == 0:
return onsets, []
bucket_count = min(2048, int(mono.size))
bucket_width = max(1, (int(mono.size) + bucket_count - 1) // bucket_count)
padded_size = bucket_count * bucket_width
if padded_size != mono.size:
mono = np.pad(mono, (0, padded_size - int(mono.size)))
peaks = np.max(np.abs(mono.reshape(bucket_count, bucket_width)), axis=1)
return onsets, [round(float(value), 5) for value in peaks]
@router.get("/dub/onsets/{job_id}")
@@ -1365,20 +1425,24 @@ async def dub_get_onsets(job_id: str):
):
with open(cache_path, "r", encoding="utf-8") as f:
cached = json.load(f)
if isinstance(cached, dict) and isinstance(cached.get("onsets"), list):
if (
isinstance(cached, dict)
and isinstance(cached.get("onsets"), list)
and isinstance(cached.get("peaks"), list)
):
return cached
except (OSError, ValueError):
pass # unreadable/corrupt cache → recompute below
try:
onsets = await asyncio.to_thread(_compute_onsets_sync, src_path)
onsets, peaks = await asyncio.to_thread(_compute_timeline_sync, src_path)
except Exception as e:
raise HTTPException(
status_code=500,
detail=f"Onset analysis failed: {str(e)[:200]}",
)
payload = {"onsets": onsets, "source": source}
payload = {"onsets": onsets, "peaks": peaks, "source": source}
try:
os.makedirs(os.path.dirname(cache_path), exist_ok=True)
tmp_path = cache_path + ".tmp"
@@ -1621,13 +1685,13 @@ async def dub_download_audio(
exports_dir = os.path.join(job_dir, "exports")
os.makedirs(exports_dir, exist_ok=True)
bg_audio = _optional_dub_artifact(job.get("no_vocals_path"), job_id) if preserve_bg else None
bg_audio = await _preserved_background(job, job_id, lang_label) if preserve_bg else None
if bg_audio:
ffmpeg = find_ffmpeg()
final_audio_path = os.path.join(exports_dir, f"mixed_dub_{stamp}.wav")
cmd = [
ffmpeg, "-i", bg_audio, "-i", wav_path,
"-filter_complex", bed_mix_filter("0:a", "1:a"),
"-filter_complex", bed_mix_filter("0:a", "1:a", bed_gain=1.0),
"-map", "[aout]", "-c:a", "pcm_s16le", "-y", final_audio_path
]
try:
@@ -1638,8 +1702,9 @@ async def dub_download_audio(
raise Exception("ffmpeg mix produced no output file")
wav_path = final_audio_path
logger.info("Dub audio mix completed")
except Exception:
except Exception as exc:
logger.exception("Failed to mix audio")
raise HTTPException(status_code=500, detail={"code": "dub_background_unavailable", "message": "Could not preserve background audio"}) from exc
base_name = os.path.splitext(job.get('filename', 'audio'))[0]
safe_name = ''.join(c for c in base_name if c.isalnum() or c in '-_ ').strip() or 'audio'
@@ -1907,20 +1972,23 @@ async def dub_download_mp3(
os.makedirs(exports_dir, exist_ok=True)
source_path = wav_path
bg_audio = _optional_dub_artifact(job.get("no_vocals_path"), job_id) if preserve_bg else None
bg_audio = await _preserved_background(job, job_id, lang_label) if preserve_bg else None
if bg_audio:
mixed_path = os.path.join(exports_dir, f"mixed_mp3_{stamp}.wav")
cmd_mix = [
ffmpeg, "-i", bg_audio, "-i", wav_path,
"-filter_complex", bed_mix_filter("0:a", "1:a"),
"-filter_complex", bed_mix_filter("0:a", "1:a", bed_gain=1.0),
"-map", "[aout]", "-c:a", "pcm_s16le", "-y", mixed_path
]
try:
rc, _, _ = await run_ffmpeg(cmd_mix, timeout=900.0)
if rc == 0 and os.path.exists(mixed_path) and os.path.getsize(mixed_path) > 0:
source_path = mixed_path
except Exception:
else:
raise RuntimeError("Background mixing failed")
except Exception as exc:
logger.exception("Failed to mix audio for MP3")
raise HTTPException(status_code=500, detail={"code": "dub_background_unavailable", "message": "Could not preserve background audio"}) from exc
mp3_path = os.path.join(exports_dir, f"dubbed_{stamp}.mp3")
# Accept '128', '192k' etc. — normalize to ffmpeg's 'Nk' form and clamp
+409 -180
View File
@@ -5,8 +5,6 @@ import struct
import logging
import time
import asyncio
import shutil
import zipfile
import torch
import torchaudio
from fastapi import APIRouter, HTTPException
@@ -16,7 +14,8 @@ from core.config import DUB_DIR, VOICES_DIR, dub_seg_path
from core.tasks import task_manager
from schemas.requests import DubRequest
from services.model_manager import _gpu_pool, run_on_gpu_pool_guarded
from services.tts_backend import resolve_generation_backend, active_backend_id
from services.tts_backend import TTSBackend, resolve_generation_backend, active_backend_id
from services.dub_batching import batch_timeout_s, native_batch_width
from services import gpu_gateway
from services.audio_dsp import apply_mastering, normalize_audio, apply_effects_chain, get_effect_chain
from services.audio_io import atomic_save_wav, _safe_torchaudio_save
@@ -31,22 +30,29 @@ from services.ffmpeg_utils import (
)
from services.rvc import apply_rvc, is_enabled as rvc_is_enabled
from services.incremental import segment_fingerprint, fit_fingerprint
from services.fit_planner import UNDERRUN_TOLERANCE, FitParams, plan_fit
from services.fit_planner import FitParams, plan_fit
from services.watermark import mark_synthetic
from services.speaker_clone import auto_profile_id
from services.segment_bundle import extract_segment_wavs
from api.routers.dub_core import _get_job, _save_job
from omnivoice.utils.voice_design import heal_design_instruct
logger = logging.getLogger("omnivoice.dub")
# Maximum compression ratio we'll attempt with pitch-preserving stretch
# before declaring "no way to fit cleanly" and falling back. atempo
# remains intelligible up to ~1.5× then introduces audible WSOLA
# artefacts; above ~1.8× speech becomes a fast garbled stream that no
# DSP can rescue. The contributing-factor pipeline (CPS-aware slot-fit
# in services/speech_rate.py, gap absorption below) keeps us under this
# in practice — this is only a guard rail.
MAX_STRETCH_RATIO = 1.8
class _RemoteDubBackend:
"""Sample-rate carrier while Dubbing runs without local TTS weights."""
sample_rate = 24_000
async def _resolve_dub_execution():
"""Resolve routing without loading local weights for a remote dub."""
engine_id = active_backend_id()
decision = gpu_gateway.decide("dub_segments")
if decision.remote:
await gpu_gateway.preflight(engine_id, decision, operation="dub_segments")
return engine_id, decision, _RemoteDubBackend()
return engine_id, decision, await resolve_generation_backend(require_cloning=True)
def _prepare_oom_retry(error: Exception, *, execution_target: str) -> bool:
@@ -362,6 +368,12 @@ def forget_missing_ref_warnings(job_id: str) -> None:
_MISSING_REF_WARNED.pop(str(job_id), None)
def _ref_within_limit(info) -> bool:
from services.speaker_clone import MAX_REF_DURATION_S
return bool(info) and float(info.get("duration") or 0.0) <= MAX_REF_DURATION_S
def resolve_consistent_ref(job: dict, speaker_key: str, memo: dict | None = None):
"""ONE clone reference for every segment of `speaker_key`.
@@ -370,8 +382,9 @@ def resolve_consistent_ref(job: dict, speaker_key: str, memo: dict | None = None
the per-line path uses as its fallback;
2. no speaker clone (heuristic diarization skips extraction entirely
the key case): a deterministic pick among that speaker's per-segment
clips: longest clip 3 s, tie-break lowest segment id. Clips all
shorter than 3 s degrade to "longest overall", same tie-break.
clips: longest usable clip 3 s within the shared reference limit,
tie-break lowest segment id. Short clips fall back to longest usable.
Oversized references must never strand every short line for a speaker.
Returns the clone info dict ({"ref_audio", "ref_text", ...}) or None.
Pure function of the job dict; `memo` (keyed by speaker_key) just avoids
@@ -380,7 +393,11 @@ def resolve_consistent_ref(job: dict, speaker_key: str, memo: dict | None = None
if memo is not None and speaker_key in memo:
return memo[speaker_key]
from services.speaker_clone import MAX_REF_DURATION_S
ref = _find_speaker_clone(job.get("speaker_clones") or {}, speaker_key)
if ref and float(ref.get("duration") or 0.0) > MAX_REF_DURATION_S:
ref = None
if ref is None:
seg_clones = job.get("segment_clones") or {}
candidates = []
@@ -392,7 +409,8 @@ def resolve_consistent_ref(job: dict, speaker_key: str, memo: dict | None = None
continue
sid = str(row.get("id", ""))
info = seg_clones.get(sid)
if info and info.get("ref_audio"):
if (info and info.get("ref_audio")
and float(info.get("duration") or 0.0) <= MAX_REF_DURATION_S):
candidates.append((sid, info))
if candidates:
usable = [
@@ -436,6 +454,9 @@ def _remote_voice(job: dict, profile_id: str | None, seg_id, voice_match: str,
info = ((job.get("segment_clones") or {}).get(str(seg_id))
or _find_speaker_clone(job.get("speaker_clones") or {}, key))
single_use = str(seg_id) in (job.get("segment_clones") or {})
if not _ref_within_limit(info):
info = resolve_consistent_ref(job, key, memo)
single_use = False
if info:
ref_audio, ref_text = info.get("ref_audio"), info.get("ref_text")
elif profile_id:
@@ -461,21 +482,12 @@ def _remote_voice(job: dict, profile_id: str | None, seg_id, voice_match: str,
def _decode_remote_dub(result: gpu_gateway.RemoteResult) -> dict[int, str]:
"""Extract the worker bundle into a task-scoped directory, path-safely."""
target = os.path.join(DUB_DIR, ".remote", result.task_id)
os.makedirs(target, exist_ok=True)
paths: dict[int, str] = {}
with zipfile.ZipFile(result.path) as archive:
for member in archive.infolist():
match = re.fullmatch(r"segments/(\d+)\.wav", member.filename)
if not match:
raise ValueError(f"unexpected dub artifact member: {member.filename}")
index = int(match.group(1))
destination = os.path.join(target, f"{index}.wav")
partial = f"{destination}.part"
with archive.open(member) as source, open(partial, "wb") as output:
shutil.copyfileobj(source, output)
os.replace(partial, destination)
paths[index] = destination
return paths
try:
return extract_segment_wavs(result.path or "", target)
except ValueError as exc:
# Preserve the established route-specific error wording consumed by
# diagnostics and regression tests.
raise ValueError(str(exc).replace("segment artifact", "dub artifact")) from exc
router = APIRouter()
@@ -491,16 +503,25 @@ async def dub_generate(job_id: str, req: DubRequest):
)
# ── Engine resolution (issue #312 class) ────────────────────────────────
# Dub used to hardcode VoiceStudio via get_model() regardless of the engine
# selected in Model Catalogue → Engines — a SILENT fallback. Every real dub
# segment's ref_audio resolves to either an auto:<speaker>/auto-seg:<id>
# clone cut from the source video or a saved voice-profile row (see
# `_gen` below), so require_cloning=True: an engine that can't clone
# would either mis-clone per segment or fail deep into the job. Checked
# ONCE here, before the streaming task starts, so a doomed job fails fast
# with one clear message instead of N per-segment ones.
# Every rendered segment clones either source speech or a saved profile, so
# local execution still requires a cloning-capable engine. Remote execution
# validates the selected worker here without loading duplicate local weights;
# its local backend is prepared only if gateway fallback actually selects it.
try:
backend = await resolve_generation_backend(require_cloning=True)
engine_id, decision, backend = await _resolve_dub_execution()
except gpu_gateway.ModelNotDownloaded as e:
raise HTTPException(
status_code=409,
detail={
"error": "model_not_downloaded",
"message": str(e),
"engine": e.engine,
"repo_ids": e.repo_ids,
"target": e.target,
"target_label": e.target_label,
"downloadable": e.downloadable,
},
) from e
except ValueError as e:
raise HTTPException(status_code=400, detail=str(e))
except Exception as e:
@@ -516,6 +537,14 @@ async def dub_generate(job_id: str, req: DubRequest):
)
raise HTTPException(status_code=503, detail=payload["detail"]) from e
# Resolve the global profile once for this job. Explicit Production
# overrides remain authoritative, while ordinary Dubbing now follows the
# same Fast/Balanced/Quality/Max contract as Clone and long-form work.
from services.performance_profiles import tts_defaults
_profile_defaults = tts_defaults(engine_id)
_job_num_step = req.num_step if req.num_step is not None else _profile_defaults.get("num_step", 16)
_job_postprocess = _profile_defaults.get("postprocess_output", True)
async def _stream(task_id):
total = len(req.segments)
all_segment_wavs = []
@@ -677,8 +706,29 @@ async def dub_generate(job_id: str, req: DubRequest):
_wav_kind = (
_kind_map.get(lang_code) if isinstance(_kind_map, dict) else job.get("seg_wav_kind")
)
if strategy != "strict_slot" and regen_only is not None and _wav_kind != "natural":
if regen_only is not None and _wav_kind != "natural":
regen_only = None
# A partial rerun must repair absent or corrupt caches, including before
# remote batching decides which lines need synthesis.
if regen_only is not None:
for index, segment in enumerate(req.segments):
sid = seg_ids[index] if index < len(seg_ids) else f"seg_{index}"
if sid in regen_only or not segment.text.strip():
continue
cache = _seg_lang_path(sid)
if not os.path.exists(cache) and _legacy_seg_cache_ok(job, lang_code):
for key in (sid, index):
legacy = dub_seg_path(job_id, key)
if os.path.exists(legacy):
cache = legacy
break
try:
info = torchaudio.info(cache)
intact = info.num_frames > 0 and _cached_payload_intact(cache, info)
except Exception:
intact = False
if not intact:
regen_only.add(sid)
# Manifest: stable segment id per current index. Per-segment WAVs are
# named by stable id (dub_seg_path) so regen reuses the right audio after
# reorder; index-keyed readers (preview/export) resolve via this manifest.
@@ -701,11 +751,70 @@ async def dub_generate(job_id: str, req: DubRequest):
_t_start = time.perf_counter()
_t_cache = 0.0
_t_tts = 0.0
_batched_audio: dict[int, torch.Tensor] = {}
_profile_row_cache: dict[str, object | None] = {}
_has_native_batch = (
getattr(type(backend), "generate_batch", TTSBackend.generate_batch)
is not TTSBackend.generate_batch
)
_native_batch_width = native_batch_width(backend) if _has_native_batch else 1
async def _prepare_local_dub():
"""Load local TTS only when gateway fallback actually needs it."""
nonlocal backend, _has_native_batch, _native_batch_width
if isinstance(backend, _RemoteDubBackend):
backend = await resolve_generation_backend(require_cloning=True)
_has_native_batch = (
getattr(type(backend), "generate_batch", TTSBackend.generate_batch)
is not TTSBackend.generate_batch
)
_native_batch_width = (
native_batch_width(backend) if _has_native_batch else 1
)
return gpu_gateway.LocalCall(fn=lambda: {})
def _segment_generation_args(index, segment) -> dict:
"""Resolve the per-row controls shared by serial and native batches."""
current_id = seg_ids[index] if index < len(seg_ids) else f"seg_{index}"
duration = segment.end - segment.start
profile_id = segment.profile_id or None
speed = segment.speed if segment.speed is not None else req.speed
language = segment.target_lang or req.language
instruct = segment.instruct or req.instruct
direction_text = getattr(segment, "direction", None)
if direction_text and direction_text.strip():
try:
from services.director import parse as _parse_direction
direction = _parse_direction(direction_text)
extra = direction.instruct_prompt()
if extra:
instruct = f"{instruct}, {extra}" if instruct else extra
bias = direction.rate_bias()
if (
bias
and abs(bias - 1.0) > 0.01
and strategy == "strict_slot"
):
speed = (speed or 1.0) * bias
except Exception as error:
logger.debug("direction parse skipped for %s: %s", current_id, error)
return {
"seg_id": current_id,
"text": segment.text,
"language": language,
"instruct": instruct,
"duration": duration if strategy == "strict_slot" else None,
"num_step": 8 if req.preview else _job_num_step,
"guidance_scale": req.guidance_scale,
"speed": speed,
"profile_id": profile_id,
"effect_preset": getattr(segment, "effect_preset", None) or "broadcast",
}
# One coarse remote lease for every segment that actually needs fresh
# synthesis. Assembly, fitting and the separately-pooled RVC pass stay
# here; the worker returns a single verified bundle of segment WAVs.
decision = gpu_gateway.decide("dub_segments")
if decision.remote:
remote_rows: list[dict] = []
remote_refs: list[str | None] = []
@@ -740,7 +849,8 @@ async def dub_generate(job_id: str, req: DubRequest):
"ref_text": ref_text, "ref_single_use": ref_single_use,
"instruct": seg_instruct,
"duration": (seg.end - seg.start) if strategy == "strict_slot" else None,
"num_step": 8 if req.preview else req.num_step,
"num_step": 8 if req.preview else _job_num_step,
"postprocess_output": _job_postprocess,
"guidance_scale": req.guidance_scale, "speed": seg_speed,
"effect_preset": seg.effect_preset or "broadcast",
"seed": seed,
@@ -752,13 +862,13 @@ async def dub_generate(job_id: str, req: DubRequest):
if remote_rows:
states: asyncio.Queue = asyncio.Queue()
call = gpu_gateway.RemoteCall(
engine=active_backend_id(), operation="dub_segments",
engine=engine_id, operation="dub_segments",
params={"segments": remote_rows, "ref_audio": remote_refs},
decode=_decode_remote_dub,
)
dub_run = gpu_gateway.JobRun("dub_segments")
run = asyncio.create_task(gpu_gateway.run(
"dub_segments", local=gpu_gateway.LocalCall(fn=lambda: {}),
"dub_segments", local=gpu_gateway.LocalCall(prepare=_prepare_local_dub),
remote=call, decision=decision, job=dub_run,
on_state=states.put_nowait,
))
@@ -777,11 +887,34 @@ async def dub_generate(job_id: str, req: DubRequest):
continue
fraction = float(state.get("progress") or 0.0)
yield f"data: {json.dumps({'type': 'progress', 'current': round(fraction * total, 2), 'total': total, 'text': state.get('stage') or state.get('phase')})}\n\n"
remote_audio = await run
try:
remote_audio = await run
except Exception as error:
from core.public_errors import stream_generation_failure
detail = stream_generation_failure(error)["detail"]
yield f"data: {json.dumps({'type': 'error', 'error': detail})}\n\n"
return
notice = dub_run.notice()
if notice is not None:
yield f"data: {json.dumps({'type': 'routing_notice', 'status': notice[0], 'reason': notice[1]})}\n\n"
if remote_audio and isinstance(backend, _RemoteDubBackend):
first_remote = next(iter(remote_audio.values()))
backend.sample_rate = int(torchaudio.info(first_remote).sample_rate)
elif isinstance(backend, _RemoteDubBackend):
# Fit-only / cache-only reruns synthesize nothing. Keep the cached
# track's native rate when one exists instead of resampling it to
# the carrier's conservative 24 kHz default.
for cached_id in seg_ids:
cached_path = _seg_lang_path(cached_id)
if os.path.exists(cached_path):
try:
backend.sample_rate = int(torchaudio.info(cached_path).sample_rate)
break
except Exception:
continue
for i, seg in enumerate(req.segments):
seg_id = seg_ids[i] if i < len(seg_ids) else f"seg_{i}"
@@ -848,7 +981,7 @@ async def dub_generate(job_id: str, req: DubRequest):
all_segment_wavs.append(
(seg.start, seg.end, seg_wav_path, backend.sample_rate)
)
sync_scores.append(getattr(seg, 'sync_ratio', None) or 1.0)
sync_scores.append(round(cached_info.num_frames / cached_info.sample_rate / max(seg_duration, 0.01), 3))
_t_cache += time.perf_counter() - _t_cache_0
continue
@@ -856,48 +989,30 @@ async def dub_generate(job_id: str, req: DubRequest):
if cached_sr != backend.sample_rate:
import torchaudio.functional as AF
cached_wav = AF.resample(cached_wav, cached_sr, backend.sample_rate)
# strict_slot persists slot-sized buffers. Every other
# strategy consumes natural-rate audio and lets the mix
# loop fit it to the current timeline.
if strategy == "strict_slot":
target_samples = int(seg_duration * backend.sample_rate)
current_samples = cached_wav.shape[-1]
if target_samples > current_samples:
cached_wav = torch.nn.functional.pad(cached_wav, (0, target_samples - current_samples))
elif current_samples > target_samples:
cached_wav = cached_wav[..., :target_samples]
cached_ratio = round(cached_wav.shape[-1] / backend.sample_rate / max(seg_duration, 0.01), 3)
all_segment_wavs.append(_store_mix_wav(seg.start, seg.end, cached_wav, backend.sample_rate, f"mix_{seg_id}"))
try:
del cached_wav
except Exception:
pass
_release_audio_tensors()
sync_scores.append(getattr(seg, 'sync_ratio', None) or 1.0)
sync_scores.append(cached_ratio)
_t_cache += time.perf_counter() - _t_cache_0
continue
except Exception as e:
# Fall through to a silent placeholder if the cached WAV
# is broken — cleaner than aborting the whole mix.
yield f"data: {json.dumps({'type': 'warning', 'segment': i, 'message': f'cached seg lost, padding silence: {str(e)[:120]}'})}\n\n"
sr = backend.sample_rate
silence = torch.zeros(1, max(0, int(seg_duration * sr)))
all_segment_wavs.append(_store_mix_wav(seg.start, seg.end, silence, sr, f"mix_{seg_id}"))
try:
del silence
except Exception:
pass
_release_audio_tensors()
sync_scores.append(1.0)
continue
except Exception:
logger.exception("Dub cached segment could not be read: %s", seg_id)
yield f"data: {json.dumps({'type': 'error', 'segment': i, 'segment_id': seg_id, 'error_code': 'dub_speech_missing', 'error': 'Cached speech could not be read. Regenerate this segment before exporting.'})}\n\n"
return
def _gen(text, lang, instruct_str, dur_s, nstep, cfg, spd, profile_id, effect_preset,
*, execution_target="local"):
*, execution_target="local", prepare_only=False, current_seg_id=None):
# Normalize once at the segment's text→engine choke point
# (covers the OOM-retry generate below too, which reuses this
# closure's `text`). Pref-gated, idempotent, never raises.
from services.text_normalization import normalize_for_tts
text = normalize_for_tts(text, lang)
effective_seg_id = seg_id if current_seg_id is None else current_seg_id
ref_audio = None
ref_text = None
used_seed = None
@@ -928,7 +1043,7 @@ async def dub_generate(job_id: str, req: DubRequest):
# CROSS binding (sid != this segment) can only come from an
# explicit request — honour its clip unchanged.
_consistent_alt = None
if voice_match == "consistent" and sid == str(seg_id):
if voice_match == "consistent" and sid == str(effective_seg_id):
_spk_key = _speaker_key_for_segment(job, sid)
if _spk_key:
_consistent_alt = resolve_consistent_ref(
@@ -969,7 +1084,9 @@ async def dub_generate(job_id: str, req: DubRequest):
# editor's Voice dropdown can actually render ("From
# Video → Speaker N"). `seg_id` is closed over from
# the per-segment loop below.
segment_speaker_key = _speaker_key_for_segment(job, seg_id)
segment_speaker_key = _speaker_key_for_segment(
job, effective_seg_id
)
# Legacy jobs may not persist diarized segment rows.
# Preserve their established per-line preference; only
# suppress it when current metadata proves the user
@@ -978,11 +1095,11 @@ async def dub_generate(job_id: str, req: DubRequest):
segment_speaker_key is None or segment_speaker_key == key
)
seg_ref = (
(job.get("segment_clones") or {}).get(str(seg_id))
(job.get("segment_clones") or {}).get(str(effective_seg_id))
if selected_is_segment_speaker
else None
)
if seg_ref:
if _ref_within_limit(seg_ref):
ref_audio = seg_ref.get("ref_audio")
ref_text = seg_ref.get("ref_text")
ref_single_use = True
@@ -990,7 +1107,7 @@ async def dub_generate(job_id: str, req: DubRequest):
auto = _find_speaker_clone(
job.get("speaker_clones") or {}, key
)
if auto is None:
if not _ref_within_limit(auto):
# Short lines may have no line-specific clip.
# Reuse this speaker's best source instead of
# silently reverting to the engine default.
@@ -1003,8 +1120,13 @@ async def dub_generate(job_id: str, req: DubRequest):
profile_id = None # prevent the voice_profiles lookup below
if profile_id:
with db_conn() as conn:
row = conn.execute("SELECT * FROM voice_profiles WHERE id=?", (profile_id,)).fetchone()
if profile_id not in _profile_row_cache:
with db_conn() as conn:
_profile_row_cache[profile_id] = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?",
(profile_id,),
).fetchone()
row = _profile_row_cache[profile_id]
if row:
if row["is_locked"] and row["locked_audio_path"]:
ref_audio = os.path.join(VOICES_DIR, row["locked_audio_path"])
@@ -1024,15 +1146,33 @@ async def dub_generate(job_id: str, req: DubRequest):
_vd = None
instruct_str = heal_design_instruct(row["instruct"], _vd)
if used_seed is not None:
if used_seed is not None and not prepare_only:
torch.manual_seed(used_seed)
# Last gate before the engine: every resolution branch above
# produces a PATH, and none of them can know it still exists.
ref_audio = warn_if_ref_missing(
ref_audio, job_id=job_id, seg_id=seg_id, where="dub render",
ref_audio, job_id=job_id, seg_id=effective_seg_id, where="dub render",
)
if prepare_only:
return {
"text": text,
"language": lang if lang != "Auto" else None,
"ref_audio": ref_audio,
"ref_text": ref_text,
"cache_ref": not ref_single_use,
"instruct": instruct_str if instruct_str else None,
"duration": dur_s,
"num_step": nstep,
"guidance_scale": cfg,
"speed": spd,
"denoise": True,
"postprocess_output": _job_postprocess,
"effect_preset": effect_preset or "broadcast",
"seed": used_seed,
}
try:
audio_out = backend.generate(
text=text, language=lang if lang != "Auto" else None,
@@ -1040,7 +1180,7 @@ async def dub_generate(job_id: str, req: DubRequest):
cache_ref=not ref_single_use,
instruct=instruct_str if instruct_str else None,
duration=dur_s, num_step=nstep, guidance_scale=cfg,
speed=spd, denoise=True, postprocess_output=True,
speed=spd, denoise=True, postprocess_output=_job_postprocess,
)
sr = backend.sample_rate
@@ -1081,7 +1221,7 @@ async def dub_generate(job_id: str, req: DubRequest):
cache_ref=not ref_single_use,
instruct=instruct_str if instruct_str else None,
duration=dur_s, num_step=retry_steps, guidance_scale=cfg,
speed=spd, denoise=True, postprocess_output=True,
speed=spd, denoise=True, postprocess_output=_job_postprocess,
)
sr = backend.sample_rate
@@ -1109,6 +1249,134 @@ async def dub_generate(job_id: str, req: DubRequest):
f"Underlying error: {retry_err}"
) from retry_err
async def _prefetch_native_batch(first_index: int) -> None:
"""Render one bounded batch and retain only its small output window."""
if _native_batch_width < 2 or remote_audio:
return
batch: list[tuple[int, dict]] = []
compatibility = None
for candidate_index in range(first_index, len(req.segments)):
candidate = req.segments[candidate_index]
candidate_id = (
seg_ids[candidate_index]
if candidate_index < len(seg_ids)
else f"seg_{candidate_index}"
)
if (
candidate_index in _batched_audio
or candidate.end - candidate.start <= 0.05
or not candidate.text.strip()
or (
regen_only is not None
and candidate_id not in regen_only
)
):
continue
args = _segment_generation_args(candidate_index, candidate)
try:
prepared = _gen(
args["text"],
args["language"],
args["instruct"],
args["duration"],
args["num_step"],
args["guidance_scale"],
args["speed"],
args["profile_id"],
args["effect_preset"],
prepare_only=True,
current_seg_id=args["seg_id"],
)
except Exception:
if candidate_index == first_index:
raise
break
# Fixed-seed profiles deliberately keep their established
# one-row deterministic RNG contract.
if prepared["seed"] is not None:
if candidate_index == first_index:
return
break
candidate_compatibility = (
prepared["cache_ref"],
bool(prepared["ref_audio"]),
prepared["num_step"],
prepared["guidance_scale"],
prepared["postprocess_output"],
)
if compatibility is None:
compatibility = candidate_compatibility
elif candidate_compatibility != compatibility:
break
batch.append((candidate_index, prepared))
if len(batch) >= _native_batch_width:
break
if len(batch) < 2:
return
def _render_batch() -> list[torch.Tensor]:
prepared_rows = [prepared for _, prepared in batch]
outputs = backend.generate_batch(
[prepared["text"] for prepared in prepared_rows],
language=[prepared["language"] for prepared in prepared_rows],
ref_audio=[prepared["ref_audio"] for prepared in prepared_rows],
ref_text=[prepared["ref_text"] for prepared in prepared_rows],
cache_ref=prepared_rows[0]["cache_ref"],
instruct=[prepared["instruct"] for prepared in prepared_rows],
duration=[prepared["duration"] for prepared in prepared_rows],
num_step=prepared_rows[0]["num_step"],
guidance_scale=prepared_rows[0]["guidance_scale"],
speed=[prepared["speed"] for prepared in prepared_rows],
denoise=True,
postprocess_output=prepared_rows[0]["postprocess_output"],
)
if len(outputs) != len(prepared_rows):
raise RuntimeError(
f"native batch returned {len(outputs)} outputs for "
f"{len(prepared_rows)} segments"
)
rendered = []
for output, prepared in zip(outputs, prepared_rows):
preset = prepared["effect_preset"]
if preset == "raw":
rendered.append(output)
continue
mastered = output
if not getattr(backend, "applies_own_mastering", False):
mastered = apply_mastering(mastered, sample_rate=backend.sample_rate)
effect_chain = get_effect_chain(preset)
if effect_chain:
mastered = apply_effects_chain(
mastered,
sample_rate=backend.sample_rate,
chain=effect_chain,
)
rendered.append(normalize_audio(mastered, target_dBFS=-2.0))
return rendered
try:
outputs = await run_on_gpu_pool_guarded(
_render_batch,
what="Dub generate batch",
timeout=batch_timeout_s(
[prepared["text"] for _, prepared in batch], backend
),
)
except TimeoutError:
raise
except Exception as error:
_prepare_oom_retry(error, execution_target="local")
logger.warning(
"Native dub batch failed for segments %s-%s; falling back: %s",
batch[0][0] + 1,
batch[-1][0] + 1,
error,
)
return
_batched_audio.update(
(index, output) for (index, _), output in zip(batch, outputs)
)
seg_profile = seg.profile_id or None
seg_speed = seg.speed if hasattr(seg, 'speed') and seg.speed is not None else req.speed
seg_lang = seg.target_lang if getattr(seg, 'target_lang', None) else req.language
@@ -1149,8 +1417,8 @@ async def dub_generate(job_id: str, req: DubRequest):
# quality for ~2× speed by dropping flow-matching steps.
# Client sends `preview=true` when the user is iterating;
# before final export the client should re-call without the
# flag to restore num_step=req.num_step quality.
_num_step = 8 if req.preview else req.num_step
# flag to restore the explicit override or shared profile.
_num_step = 8 if req.preview else _job_num_step
_t_tts_0 = time.perf_counter()
seg_effect_preset = getattr(seg, "effect_preset", None) or "broadcast"
@@ -1177,14 +1445,19 @@ async def dub_generate(job_id: str, req: DubRequest):
import torchaudio.functional as AF
audio_tensor = AF.resample(audio_tensor, remote_sr, backend.sample_rate)
else:
audio_tensor = await run_on_gpu_pool_guarded(
lambda: _gen(
seg.text, seg_lang, seg_instruct, _dur_for_tts,
_num_step, req.guidance_scale, seg_speed, seg_profile, seg_effect_preset,
),
what="Dub generate",
timeout=generate_timeout_s(seg.text, engine=backend),
)
if i not in _batched_audio:
await _prefetch_native_batch(i)
if i in _batched_audio:
audio_tensor = _batched_audio.pop(i)
else:
audio_tensor = await run_on_gpu_pool_guarded(
lambda: _gen(
seg.text, seg_lang, seg_instruct, _dur_for_tts,
_num_step, req.guidance_scale, seg_speed, seg_profile, seg_effect_preset,
),
what="Dub generate",
timeout=generate_timeout_s(seg.text, engine=backend),
)
_t_tts += time.perf_counter() - _t_tts_0
# Check abort immediately after GPU work completes
@@ -1192,25 +1465,19 @@ async def dub_generate(job_id: str, req: DubRequest):
yield f"data: {json.dumps({'type': 'cancelled', 'segments_processed': i + 1})}\n\n"
return
target_samples = int(seg_duration * backend.sample_rate)
current_samples = audio_tensor.shape[-1]
# Capture the real spoken duration before assembly fitting.
# This is the evidence used by Agent timing and
# keeps sync badges truthful for every timing strategy.
natural_generated_dur = current_samples / backend.sample_rate
if strategy == "strict_slot":
# Legacy: pad short audio + trim long audio so the mix
# loop receives slot-sized buffers. The atempo squeeze
# in the mix loop never fires here because we already
# forced size = target_samples.
if target_samples > current_samples:
pad_amount = target_samples - current_samples
audio_tensor = torch.nn.functional.pad(audio_tensor, (0, pad_amount))
elif current_samples > target_samples:
audio_tensor = audio_tensor[..., :target_samples]
# concise / stretch_video / smart_fit: keep audio at its
# natural length. The mix loop decides per-mode whether to
# trim, slip, stretch the video, or split audio/video
# retiming (smart_fit) to accommodate it.
# Keep the complete waveform in every cache. Fitting happens
# once during assembly; pre-trimming here destroyed words before
# the pitch-preserving stretcher could see them.
if current_samples == 0 or not torch.isfinite(audio_tensor).all() or not torch.any(audio_tensor.abs() > 1e-6):
raise ValueError("The speech engine returned empty or silent audio")
generated_dur = audio_tensor.shape[-1] / backend.sample_rate
generated_dur = natural_generated_dur
sync_ratio = round(generated_dur / max(seg_duration, 0.01), 3)
sync_scores.append(sync_ratio)
@@ -1218,7 +1485,7 @@ async def dub_generate(job_id: str, req: DubRequest):
# Duration-planner calibration sample: this text length spoke
# for this long at natural rate. Keyed by stable seg id and
# merged into the per-language job map after the loop.
if strategy != "strict_slot" and seg.text.strip() and generated_dur > 0:
if seg.text.strip() and generated_dur > 0:
_natural_dur_records[str(seg_id)] = {
"chars": len(seg.text.strip()),
"dur": round(generated_dur, 4),
@@ -1256,15 +1523,6 @@ async def dub_generate(job_id: str, req: DubRequest):
if rvc_sr == backend.sample_rate:
audio_tensor = rvc_wav
if strategy == "strict_slot":
target_samples = int(seg_duration * backend.sample_rate)
current_samples = audio_tensor.shape[-1]
if target_samples > current_samples:
audio_tensor = torch.nn.functional.pad(
audio_tensor, (0, target_samples - current_samples)
)
elif current_samples > target_samples:
audio_tensor = audio_tensor[..., :target_samples]
except Exception as e:
yield f"data: {json.dumps({'type': 'warning', 'segment': i, 'message': f'RVC skipped: {str(e)[:120]}'})}\n\n"
@@ -1311,10 +1569,9 @@ async def dub_generate(job_id: str, req: DubRequest):
from core.public_errors import stream_generation_failure
error_detail = stream_generation_failure(e)["detail"]
yield f"data: {json.dumps({'type': 'error', 'segment': i, 'error': error_detail})}\n\n"
sr = backend.sample_rate
all_segment_wavs.append(_store_mix_wav(seg.start, seg.end, torch.zeros(1, max(0, int(seg_duration * sr))), sr, f"mix_{seg_id}"))
sync_scores.append(1.0)
logger.exception("Dub generation failed for segment %s", seg_id)
yield f"data: {json.dumps({'type': 'error', 'segment': i, 'segment_id': seg_id, 'error': error_detail})}\n\n"
return
_t_loop_end = time.perf_counter()
@@ -1461,19 +1718,16 @@ async def dub_generate(job_id: str, req: DubRequest):
seg_gain = max(0.0, min(2.0, seg_gain))
try:
wav = _load_entry_wav((start, end, wav_path, sr), sr)
except Exception as e:
# A WAV header can be readable while its payload is
# truncated. Direct cache reuse deliberately defers the
# decode to assembly, so preserve the old recovery contract
# here: warn and fill this slot with silence instead of
# aborting the entire dub.
warning = {
"type": "warning",
"segment": i,
"message": f"cached seg lost, padding silence: {str(e)[:120]}",
}
yield f"data: {json.dumps(warning)}\n\n"
wav = torch.zeros(1, max(0, int((end - start) * sr)))
except Exception:
logger.exception("Dub assembly could not read segment %d", i)
yield f"data: {json.dumps({'type': 'error', 'segment': i, 'error_code': 'dub_speech_missing', 'error': 'A speech segment could not be read. Regenerate it before exporting.'})}\n\n"
return
from services.audio_dsp import trim_speech_padding
if seg_ref is not None and seg_ref.text.strip():
if wav.numel() == 0 or not torch.isfinite(wav).all() or not torch.any(wav.abs() > 1e-6):
yield f"data: {json.dumps({'type': 'error', 'segment': i, 'error_code': 'dub_speech_missing', 'error': 'A speech segment is empty or silent. Regenerate it before exporting.'})}\n\n"
return
wav = trim_speech_padding(wav, sr)
adjusted = wav * seg_gain
if adjusted.ndim == 2 and adjusted.shape[0] > 1:
adjusted = adjusted.mean(dim=0, keepdim=True)
@@ -1521,12 +1775,12 @@ async def dub_generate(job_id: str, req: DubRequest):
align_corners=False,
).squeeze(0)
wl = adjusted.shape[-1]
# Residual overflow → hard-trim to the segment's new video
# slot (fade below keeps the cut pop-free).
# Never publish a complete track with speech discarded by
# the fit caps. The user can shorten text or relax the caps.
new_slot_samples = int(max(0.0, sf.new_end - sf.new_start) * sr)
if new_slot_samples > 0 and wl > new_slot_samples:
adjusted = adjusted[..., :new_slot_samples]
wl = adjusted.shape[-1]
if new_slot_samples > 0 and wl > new_slot_samples + int(sr * 0.02):
yield f"data: {json.dumps({'type': 'error', 'segment': i, 'error_code': 'dub_timing_overflow', 'error': 'Speech exceeds the fitting limits. Shorten the translation or choose Strict Slot or Stretch Video before exporting.'})}\n\n"
return
# Truthful per-segment verdict for the UI badge.
entry = {"status": sf.status}
if abs(sf.audio_rate - 1.0) > 1e-6:
@@ -1546,8 +1800,8 @@ async def dub_generate(job_id: str, req: DubRequest):
elif strategy == "concise":
# Mode A: never compress. Allow the audio to extend into the
# silent gap before the next seg (existing heuristic) plus
# any extra `overflow_budget_s`. Beyond that, hard-trim with
# a short fade so we never overlap the next speaker.
# any extra `overflow_budget_s`. Beyond that, require a
# timing/text adjustment instead of discarding speech.
place_at = start
effective_end = end
if i + 1 < len(all_segment_wavs):
@@ -1561,47 +1815,25 @@ async def dub_generate(job_id: str, req: DubRequest):
slot_samples_eff = int(max(0.0, (effective_end - start)) * sr)
if slot_samples_eff > 0 and wl > slot_samples_eff:
overflow_s = (wl - slot_samples_eff) / sr
adjusted = adjusted[..., :slot_samples_eff]
wl = adjusted.shape[-1]
fit_status.append({
"status": "overflows",
"overflow_s": round(overflow_s, 3),
})
yield f"data: {json.dumps({'type': 'error', 'segment': i, 'segment_id': job['seg_order'][i], 'error_code': 'dub_timing_overflow', 'error': 'Speech exceeds its time slot. Shorten the translation or choose Strict Slot or Stretch Video before exporting.', 'overflow_s': round(overflow_s, 3)})}\n\n"
return
else:
fit_status.append({"status": "fits"})
else:
# strict_slot (legacy): preserve the previous atempo / trim /
# off semantics so existing callers and back-compat tests
# keep passing.
# Strict Slot fits the complete speech to the original
# slot. Explicit legacy trim/off choices remain available.
place_at = start
effective_end = end
slowed_rate = None
if i + 1 < len(all_segment_wavs):
next_start = all_segment_wavs[i + 1][0]
gap = next_start - end
if gap > GAP_OVERFLOW_BUFFER_S:
effective_end = end + min(
gap - GAP_OVERFLOW_BUFFER_S, GAP_OVERFLOW_MAX_S,
)
slot_samples = int(max(0.0, (effective_end - start)) * sr)
if slot_fit != "off" and slot_samples > 0 and wl > slot_samples:
if slot_fit == "time_stretch":
ratio = wl / slot_samples
capped_ratio = min(ratio, MAX_STRETCH_RATIO)
capped_target = int(wl / capped_ratio)
try:
adjusted = await _pitch_preserving_stretch(
adjusted, capped_target, sr,
adjusted, slot_samples, sr,
)
if adjusted.shape[-1] > slot_samples:
adjusted = adjusted[..., :slot_samples]
if ratio > MAX_STRETCH_RATIO:
logger.info(
"seg %d compression %.2f× exceeded cap; "
"stretched to %.2f×, tail trimmed",
i, ratio, capped_ratio,
)
except Exception as e:
logger.warning(
"atempo stretch failed for seg %d (%.2f×), "
@@ -1621,15 +1853,14 @@ async def dub_generate(job_id: str, req: DubRequest):
slot_fit == "time_stretch"
and slot_samples > 0
and wl > 0
and wl < slot_samples * UNDERRUN_TOLERANCE
and _underrun_min_rate() < 1.0 - 1e-6
and wl < slot_samples
):
# Underrun fill (mirror of the compression above): the
# dub finished early, leaving the on-screen mouth moving
# over the thin under-speech bed residue — perceived as
# dead air. Slow toward the slot, never below the floor.
rate = max(wl / slot_samples, _underrun_min_rate())
target = min(slot_samples, int(round(wl / rate)))
rate = wl / slot_samples
target = slot_samples
try:
adjusted = await _pitch_preserving_stretch(
adjusted, target, sr,
@@ -1730,6 +1961,7 @@ async def dub_generate(job_id: str, req: DubRequest):
"language_code": lang_code,
"duration": round(track_dur, 4),
"timing_strategy": strategy,
"source_segments": [{"start": seg.start, "end": seg.end} for seg in req.segments],
}
# Persist the timing strategy + (for Mode B) the per-segment stretch
@@ -1773,12 +2005,9 @@ async def dub_generate(job_id: str, req: DubRequest):
"fit_fp": fit_fp,
}
job["dubbed_tracks"][lang_code]["fit_fp"] = fit_fp
# Record what kind of per-segment WAVs are on disk so a later
# smart_fit run knows whether partial regen / fit-only re-mix can
# reuse them ("natural") or must regen once ("slotted"). Per-track
# (P1.3) — each language renders under its own strategy; the flat
# field stays in lock-step for older readers.
_kind = "slotted" if strategy == "strict_slot" else "natural"
# Every new cache preserves natural speech. Old slotted caches must
# be regenerated once because their missing tails cannot be recovered.
_kind = "natural"
job.setdefault("seg_wav_kind_by_lang", {})[lang_code] = _kind
job["seg_wav_kind"] = _kind
_save_job(job_id, job)
+339 -65
View File
@@ -1,12 +1,13 @@
import json
import os
import time
import asyncio
import logging
from typing import Optional
from fastapi import APIRouter
from fastapi import APIRouter, HTTPException
from fastapi.responses import JSONResponse
from schemas.requests import TranslateRequest
from schemas.requests import AgentFitRequest, TranslateRequest
from services.model_manager import _cpu_pool, _gpu_pool
from services.hf_revisions import revision_for
from services.translator import cinematic_available, cinematic_refine_many, _cinematic_budget
@@ -19,10 +20,11 @@ _NLLB_REPO_ID = "facebook/nllb-200-distilled-600M"
def _load_nllb_component(factory):
"""Load a curated NLLB component from its reviewed immutable revision."""
"""Load explicitly installed NLLB weights at their reviewed revision."""
return factory.from_pretrained(
_NLLB_REPO_ID,
revision=revision_for(_NLLB_REPO_ID),
local_files_only=True,
)
TRANSLATE_CODES = {
@@ -39,8 +41,33 @@ FLORES_CODES = {
"hi": "hin_Deva", "tr": "tur_Latn", "pl": "pol_Latn", "nl": "nld_Latn",
"sv": "swe_Latn", "th": "tha_Thai", "vi": "vie_Latn", "id": "ind_Latn",
"uk": "ukr_Cyrl",
"zh-TW": "zho_Hant", "zh-Hant": "zho_Hant", "cmn-Hant": "zho_Hant",
"zh-Hans": "zho_Hans", "yue": "yue_Hant",
"bn": "ben_Beng", "ta": "tam_Taml", "te": "tel_Telu", "ml": "mal_Mlym",
"kn": "kan_Knda", "gu": "guj_Gujr", "mr": "mar_Deva", "ur": "urd_Arab",
"fa": "pes_Arab", "he": "heb_Hebr", "el": "ell_Grek", "cs": "ces_Latn",
"da": "dan_Latn", "fi": "fin_Latn", "nb": "nob_Latn", "nn": "nno_Latn",
"ro": "ron_Latn", "hu": "hun_Latn", "bg": "bul_Cyrl", "sk": "slk_Latn",
"sl": "slv_Latn", "hr": "hrv_Latn", "sr": "srp_Cyrl", "lt": "lit_Latn",
"et": "est_Latn", "sw": "swh_Latn", "af": "afr_Latn", "ms": "zsm_Latn",
}
def _nllb_language(code: str) -> str | None:
"""Resolve aliases or tokenizer-supported FLORES codes without loading weights."""
from transformers.models.nllb.tokenization_nllb import FAIRSEQ_LANGUAGE_CODES
normalized = code.strip().replace("_", "-").lower()
aliases = {key.lower(): value for key, value in FLORES_CODES.items()}
if normalized in aliases:
return aliases[normalized]
exact = [value for value in FAIRSEQ_LANGUAGE_CODES if value.replace("_", "-").lower() == normalized]
if exact:
return exact[0]
# Bare ISO-639-3 codes are safe only when the tokenizer has one script.
matches = [value for value in FAIRSEQ_LANGUAGE_CODES if value.split("_")[0] == normalized]
return matches[0] if len(matches) == 1 else None
# Human-readable language names for LLM prompts. Empirically a tiny / 7B
# local LLM produces Devanagari Hindi reliably when told "translate into
# Hindi" but drifts to German / English / phonetic-Latin when told
@@ -163,9 +190,77 @@ def _looks_like_target(text: str, code: str, threshold: float = 0.5) -> bool:
codepoints alone."""
return _script_ratio(text, code) >= threshold
def _translation_output_error(text: object) -> str | None:
"""Reject provider error pages that arrive with HTTP 200.
Google's mobile endpoint occasionally returns its generic HTML error copy
inside the element deep-translator treats as a successful translation.
Passing that through would replace the user's transcript with the error
page, so treat it like any other transient provider failure and retry.
"""
if not isinstance(text, str) or not text.strip():
return "empty translation"
normalized = " ".join(text.split()).casefold()
error_markers = (
"error 500 (server error)",
"that's an error",
"thats an error",
"there was an error. please try again later",
"no translation was found using the current translator",
)
if "\ufffd" in text or any(marker in normalized for marker in error_markers):
return "translation provider returned invalid output"
return None
_nllb_model = None
_nllb_tokenizer = None
_nllb_device = None
_NLLB_BATCH_SIZE_ENV = "OMNIVOICE_NLLB_BATCH_SIZE"
_NLLB_MAX_BATCH_SIZE = 32
def _nllb_batch_size() -> int:
"""Bound NLLB forward-pass width; explicit overrides remain available."""
configured = os.environ.get(_NLLB_BATCH_SIZE_ENV, "").strip()
if configured:
try:
return max(1, min(_NLLB_MAX_BATCH_SIZE, int(configured)))
except (TypeError, ValueError):
logger.warning("%s=%r is not an integer; using the safe default", _NLLB_BATCH_SIZE_ENV, configured)
# The 600M checkpoint leaves ample room on modern discrete GPUs. Scale the
# forward-pass width there; CPU and unified-memory MPS keep the conservative
# width because their failure recovery moves the whole model.
if _nllb_device == "cuda":
try:
import torch
free_gib = int(torch.cuda.mem_get_info()[0]) / 1024**3
if free_gib >= 16:
return 24
if free_gib >= 8:
return 12
except Exception:
pass
return 8
return 4
def _nllb_hypothesis_budget() -> int:
"""Bound batch × beam hypotheses by currently available device memory."""
if _nllb_device != "cuda":
return 16
try:
import torch
free_gib = int(torch.cuda.mem_get_info()[0]) / 1024**3
if free_gib >= 16:
return 64
if free_gib >= 8:
return 32
except Exception:
pass
return 16
def _dialect_flags(req, applied: bool) -> dict:
@@ -255,10 +350,11 @@ def _resolve_translation_context(req, client, model_name: str, timeout: float,
def _unload_nllb():
"""Release NLLB VRAM so TTS model can reload."""
global _nllb_model, _nllb_tokenizer
global _nllb_device, _nllb_model, _nllb_tokenizer
import gc
_nllb_model = None
_nllb_tokenizer = None
_nllb_device = None
gc.collect()
try:
import torch
@@ -270,10 +366,33 @@ def _unload_nllb():
pass
def _should_unload_nllb() -> bool:
"""Retain a warm local translator only when the accelerator has safe headroom."""
override = os.environ.get("OMNIVOICE_UNLOAD_NLLB")
if override is not None:
return override.strip().lower() not in {"0", "false", "no", "off"}
if _nllb_device != "cuda":
return True
try:
import torch
free_bytes, total_bytes = torch.cuda.mem_get_info()
return total_bytes < 16 * 1024**3 or free_bytes < 8 * 1024**3
except Exception:
return True
@router.post("/dub/translate")
async def dub_translate(req: TranslateRequest):
try:
provider = (req.provider if req.provider else os.environ.get("TRANSLATE_PROVIDER", "google")).lower()
from services import translation_engines
if not translation_engines.get_engine(provider):
return JSONResponse(
status_code=400,
content={"error": "Choose a supported translation engine."},
)
lang_code = TRANSLATE_CODES.get(req.target_lang, req.target_lang)
api_key = os.environ.get("TRANSLATE_API_KEY", "")
loop = asyncio.get_running_loop()
@@ -281,8 +400,16 @@ async def dub_translate(req: TranslateRequest):
# Offline NLLB Transformer Translation
if provider == "nllb":
flores_tgt = FLORES_CODES.get(req.target_lang, "eng_Latn")
flores_src = FLORES_CODES.get(src_lang, "eng_Latn")
requested = [src_lang, req.target_lang, *(seg.target_lang for seg in req.segments if seg.target_lang)]
resolved = {code: _nllb_language(code) for code in requested}
unsupported = [code for code, language in resolved.items() if language is None]
if unsupported:
return JSONResponse(status_code=400, content={
"error": "NLLB does not support the requested language.",
"code": "unsupported_translation_language", "languages": unsupported,
})
flores_tgt = resolved[req.target_lang]
flores_src = resolved[src_lang]
def _translate_nllb():
global _nllb_model, _nllb_tokenizer, _nllb_device
@@ -314,44 +441,112 @@ async def dub_translate(req: TranslateRequest):
logger.exception("NLLB model load failed")
return [{"id": seg.id, "text": seg.text, "error": f"Model load error: {str(e)}"} for seg in req.segments]
results = []
for seg in req.segments:
from services.performance_profiles import translation_decode_defaults
# Snapshot once so every segment and device fallback in this
# job uses the same decoding effort even if preferences change.
decode_options = translation_decode_defaults()
def _generate_rows(rows, target_language):
global _nllb_device
_nllb_tokenizer.src_lang = flores_src
inputs = _nllb_tokenizer(
[seg.text for _, seg in rows],
return_tensors="pt",
padding=True,
)
if _nllb_device and _nllb_device != "cpu":
inputs = {key: value.to(_nllb_device) for key, value in inputs.items()}
forced_bos_token_id = _nllb_tokenizer.convert_tokens_to_ids(target_language)
try:
if not seg.text or not seg.text.strip():
results.append({"id": seg.id, "text": seg.text})
continue
tokens = _nllb_model.generate(
**inputs,
forced_bos_token_id=forced_bos_token_id,
max_length=400,
**decode_options,
)
except (RuntimeError, NotImplementedError) as error:
if _nllb_device != "mps":
raise
logger.warning("MPS generate failed, retrying on CPU: %s", error)
_nllb_model.to("cpu")
_nllb_device = "cpu"
inputs = {key: value.to("cpu") for key, value in inputs.items()}
tokens = _nllb_model.generate(
**inputs,
forced_bos_token_id=forced_bos_token_id,
max_length=400,
**decode_options,
)
decoded = _nllb_tokenizer.batch_decode(tokens, skip_special_tokens=True)
if len(decoded) != len(rows):
raise RuntimeError(
f"NLLB returned {len(decoded)} translations for {len(rows)} segments"
)
return decoded
tgt = FLORES_CODES.get(seg.target_lang, flores_tgt) if seg.target_lang else flores_tgt
# A target-language BOS token is shared by a forward pass, so
# group mixed-language rows first. Preserve request order in
# the final response even though groups render independently.
grouped: dict[str, list[tuple[int, object]]] = {}
results_by_index: dict[int, dict] = {}
for index, seg in enumerate(req.segments):
if not seg.text or not seg.text.strip():
results_by_index[index] = {"id": seg.id, "text": seg.text}
continue
target = resolved[seg.target_lang] if seg.target_lang else flores_tgt
grouped.setdefault(target, []).append((index, seg))
_nllb_tokenizer.src_lang = flores_src
inputs = _nllb_tokenizer(seg.text, return_tensors="pt")
if _nllb_device and _nllb_device != "cpu":
inputs = {k: v.to(_nllb_device) for k, v in inputs.items()}
forced_bos_token_id = _nllb_tokenizer.convert_tokens_to_ids(tgt)
# Beam search multiplies decoder memory per row. Keep the
# effective hypothesis count bounded while still widening the
# Fast path aggressively.
beam_count = max(1, int(decode_options.get("num_beams", 1)))
width = min(
_nllb_batch_size(),
max(1, _nllb_hypothesis_budget() // beam_count),
)
for target, rows in grouped.items():
for start in range(0, len(rows), width):
batch = rows[start : start + width]
try:
translated_tokens = _nllb_model.generate(
**inputs, forced_bos_token_id=forced_bos_token_id, max_length=400
translated_texts = _generate_rows(batch, target)
except Exception as batch_error:
if len(batch) == 1:
index, seg = batch[0]
results_by_index[index] = {
"id": seg.id,
"text": seg.text,
"error": str(batch_error),
}
continue
# A single unusually long row must not sink its
# neighbours. Clear a failed device allocation and
# retain the established per-segment degradation.
if torch.cuda.is_available():
torch.cuda.empty_cache()
logger.warning(
"NLLB batch of %d failed; retrying rows individually: %s",
len(batch),
batch_error,
)
except (RuntimeError, NotImplementedError) as e:
if _nllb_device == "mps":
logger.warning("MPS generate failed, retrying on CPU: %s", e)
_nllb_model.to("cpu")
_nllb_device = "cpu"
inputs = {k: v.to("cpu") for k, v in inputs.items()}
translated_tokens = _nllb_model.generate(
**inputs, forced_bos_token_id=forced_bos_token_id, max_length=400
)
else:
raise
translated_text = _nllb_tokenizer.batch_decode(translated_tokens, skip_special_tokens=True)[0]
results.append({"id": seg.id, "text": translated_text})
except Exception as e:
results.append({"id": seg.id, "text": seg.text, "error": str(e)})
return results
for index, seg in batch:
try:
translated_text = _generate_rows([(index, seg)], target)[0]
results_by_index[index] = {"id": seg.id, "text": translated_text}
except Exception as row_error:
results_by_index[index] = {
"id": seg.id,
"text": seg.text,
"error": str(row_error),
}
continue
for (index, seg), translated_text in zip(batch, translated_texts):
results_by_index[index] = {"id": seg.id, "text": translated_text}
return [results_by_index[index] for index in range(len(req.segments))]
translated = await loop.run_in_executor(_gpu_pool, _translate_nllb)
if os.environ.get("OMNIVOICE_UNLOAD_NLLB", "1") == "1":
if _should_unload_nllb():
_unload_nllb()
# Cinematic/Autofit refine + rate-ratio badges must run for NLLB too
# (previously this returned before _maybe_cinematic, so a Cinematic
@@ -472,6 +667,7 @@ async def dub_translate(req: TranslateRequest):
f"You are a professional dubbing translator. "
f"Translate the user's text from {src_name} into "
f"{tgt_name}.{script_clause}{dia_clause} "
f"{translation_style_brief(req)} "
f"Reply ONLY with the translated {tgt_name} text, do not "
f"add quotes, notes, headers, explanations, or commentary."
)
@@ -543,7 +739,7 @@ async def dub_translate(req: TranslateRequest):
source_lang=src_lang,
target_lang=tgt_code,
target_name=LANG_NAMES.get(tgt_code, tgt_code),
extra_clause=context_extra,
extra_clause="\n".join(filter(None, [context_extra, translation_style_brief(req)])),
)
except Exception as e: # noqa: BLE001
logger.warning("reflect pass skipped for %s: %s",
@@ -594,18 +790,35 @@ async def dub_translate(req: TranslateRequest):
f"switch the Engine dropdown to another provider."
)
return JSONResponse(status_code=400, content={"error": friendly})
target_codes = list(dict.fromkeys(
seg.target_lang if seg.target_lang else req.target_lang
for seg in req.segments
))
try:
pack_status = translation_engines.argos_pack_status(src_lang, target_codes)
except (ImportError, ValueError) as exc:
return JSONResponse(status_code=422, content={"error": str(exc)})
missing_packs = [
pair for pair in pack_status["pairs"] if not pair["installed"]
]
if missing_packs:
pairs = ", ".join(
f'{pair["source_lang"]}{pair["target_lang"]}'
for pair in missing_packs
)
return JSONResponse(
status_code=409,
content={
"error": f"Install the Argos language pack for {pairs} before translating.",
"code": "argos_pack_missing",
"pairs": missing_packs,
},
)
def _translate_argos():
cache_dir = os.environ.get("OMNIVOICE_CACHE_DIR")
if cache_dir:
argos_cache = os.path.join(cache_dir, "argos-translate")
os.makedirs(argos_cache, exist_ok=True)
os.environ.setdefault("ARGOS_PACKAGES_DIR", argos_cache)
os.environ.setdefault("ARGOS_DATA_DIR", argos_cache)
import argostranslate.package
import argostranslate.translate
from_code = src_lang
available_packages = argostranslate.package.get_installed_packages()
from_code = pack_status["source_lang"]
results = []
for seg in req.segments:
@@ -614,19 +827,12 @@ async def dub_translate(req: TranslateRequest):
results.append({"id": seg.id, "text": seg.text})
continue
to_code = seg.target_lang if seg.target_lang else req.target_lang
installed_pkg = next(filter(lambda x: x.from_code == from_code and x.to_code == to_code, available_packages), None)
if installed_pkg is None:
argostranslate.package.update_package_index()
all_packages = argostranslate.package.get_available_packages()
package_to_install = next(filter(lambda x: x.from_code == from_code and x.to_code == to_code, all_packages), None)
if package_to_install:
argostranslate.package.install_from_path(package_to_install.download())
available_packages = argostranslate.package.get_installed_packages()
else:
raise Exception(f"No Argos package available for {from_code} -> {to_code}")
translated_text = argostranslate.translate.translate(seg.text, from_code, to_code)
to_code = translation_engines.argos_lang_code(to_code)
translated_text = (
seg.text
if from_code == to_code
else argostranslate.translate.translate(seg.text, from_code, to_code)
)
results.append({"id": seg.id, "text": translated_text})
except Exception as e:
results.append({"id": seg.id, "text": seg.text, "error": str(e)})
@@ -700,9 +906,10 @@ async def dub_translate(req: TranslateRequest):
for attempt, src in enumerate([src_arg, src_arg, "auto"]):
try:
out = _build_translator(src, seg_lc).translate(seg.text)
if out and out.strip():
output_error = _translation_output_error(out)
if output_error is None:
return {"id": seg.id, "text": out}
last_err = "empty translation"
last_err = output_error
except Exception as e:
last_err = f"{type(e).__name__}: {e}"
logger.warning(
@@ -895,7 +1102,7 @@ async def _apply_fit_pass(rows, req, slots_by_id, source_by_id, quality, loop, d
their current text and get ``rate_error='fit-budget'``. Only rows with a
slot + text + no prior error participate.
"""
strict = (quality == "autofit")
strict = quality in ("autofit", "agent")
items = []
for row in rows:
seg_id = str(row["id"])
@@ -958,7 +1165,7 @@ async def _maybe_cinematic(translated, req, src_lang, loop, *, already_llm=False
# Fast (and anything unrecognised) returns the plain translation unchanged
# (plus the pre-synthesis duration-plan badges — no LLM needed for those).
if quality not in ("cinematic", "autofit"):
if quality not in ("cinematic", "autofit", "agent"):
await _finalize_duration_plan(translated, req, loop)
return base
@@ -1029,7 +1236,7 @@ async def _maybe_cinematic(translated, req, src_lang, loop, *, already_llm=False
target_lang=req.target_lang,
glossary=req.glossary,
directions=directions,
dialect_hint=dialect_hint,
dialect_hint="\n".join(filter(None, [dialect_hint, translation_style_brief(req)])),
executor=_cpu_pool,
)
refined_by_id = {r["id"]: r for r in refined}
@@ -1076,3 +1283,70 @@ async def _maybe_cinematic(translated, req, src_lang, loop, *, already_llm=False
"quality_used": quality,
**_dialect_flags(req, applied=bool(dialect_hint)),
}
def translation_style_brief(req) -> str:
instructions = (getattr(req, "translation_instructions", None) or "").strip()
return ("User translation style brief (tone and wording only; preserve meaning, timing and output format): "
+ json.dumps(instructions, ensure_ascii=False)) if instructions else ""
@router.post("/dub/agent-fit")
async def dub_agent_fit(req: AgentFitRequest):
"""Rewrite rendered lines from real duration evidence.
Synthesis stays in the normal Dubbing pipeline. The client renders each
candidate, measures it, and may request one more bounded correction.
"""
from services import llm_skills
from services.speech_rate import adjust_for_measured_slot_many
readiness = llm_skills.resolve_skill("slot_fitting")
if not readiness.ready:
raise HTTPException(
status_code=409,
detail={
"error": "llm_skill_unavailable",
"skill": "slot_fitting",
"reason": readiness.reason or "unavailable",
},
)
items = [
(
segment.id,
segment.text,
segment.slot_seconds,
segment.measured_seconds,
req.target_lang,
segment.source_text,
segment.context_before,
segment.context_after,
)
for segment in req.segments
]
budget = _cinematic_budget()
try:
call = adjust_for_measured_slot_many(items, executor=_cpu_pool, translation_instructions=req.translation_instructions)
rows = await asyncio.wait_for(call, timeout=budget) if budget and budget > 0 else await call
except asyncio.TimeoutError:
rows = {
segment.id: {
"text": segment.text,
"changed": False,
"measured_seconds": round(segment.measured_seconds, 3),
"target_seconds": round(segment.slot_seconds, 3),
"measured_ratio": round(
segment.measured_seconds / max(segment.slot_seconds, 0.001), 3
),
"error": "fit-budget",
}
for segment in req.segments
}
return {
"target_lang": req.target_lang,
"segments": [
{"id": segment.id, **rows[str(segment.id)]}
for segment in req.segments
],
}
+307 -7
View File
@@ -15,6 +15,7 @@ Environment variables (`OMNIVOICE_TTS_BACKEND`, `OMNIVOICE_ASR_BACKEND`,
`OMNIVOICE_LLM_BACKEND`) still win over the UI choice so power-users can pin
a backend without Settings silently undoing it.
"""
import asyncio
import logging
import os
import threading
@@ -23,10 +24,11 @@ from time import perf_counter
from fastapi import APIRouter, Depends, HTTPException
from huggingface_hub import utils as hf_utils
from huggingface_hub.errors import HFValidationError
from pydantic import BaseModel
from pydantic import BaseModel, Field
from api.dependencies import require_admin, require_admin_action, require_desktop
from core import prefs
from core.engine_licenses import LICENSE_GATED_ENGINES
from services import tts_backend, asr_backend, llm_backend, translation_engines
from services.audio_dsp import list_effect_presets
from api.schemas import EffectPresetsResponse
@@ -42,12 +44,65 @@ _FAMILIES = {
}
def _catalogue_active_id(family: str, module) -> str:
"""Return the active id represented by the public engine catalogue."""
active = module.active_backend_id()
if family != "tts" or active != "omnivoice-subprocess":
return active
from core.device_caps import detect_host_caps
try:
return "omnivoice" if detect_host_caps().family == "mps" else active
except Exception:
return active
def _family_payload(family: str, module):
"""Public inventory plus whether an environment pin owns this family."""
active = _catalogue_active_id(family, module)
model = None
if family == "asr":
model = asr_backend._offline_asr_repo(active)
elif family == "llm" and active != "off":
model = llm_backend.get_active_llm_backend().model_name
elif family == "tts":
if active in {"omnivoice", "omnivoice-subprocess"}:
from services.model_manager import resolve_omnivoice_checkpoint
model = resolve_omnivoice_checkpoint()
elif active == "mlx-audio":
from core import prefs
cls = tts_backend.MLXAudioBackend
key = prefs.resolve("mlx_audio_model_id", env="OMNIVOICE_MLX_AUDIO_MODEL", default=cls.DEFAULT_MODEL_KEY)
model = cls.CURATED_MODELS.get(key, key)
else:
instance = getattr(tts_backend, "_active_instance", None)
if instance is not None and getattr(tts_backend, "_active_instance_id", None) == active:
model = instance.model_identity()
backends = public_backends(module.list_backends())
if family == "tts":
from services import settings_store
for backend in backends:
engine_id = backend.get("id")
if engine_id in LICENSE_GATED_ENGINES:
backend["license_required"] = True
try:
backend["license_accepted"] = settings_store.get_license_accepted(engine_id)
except Exception:
logger.warning(
"Could not read license acceptance for %s",
engine_id,
exc_info=True,
)
backend["license_accepted"] = False
return {
"active": module.active_backend_id(),
# MPS hides the explicit compatibility row, so legacy configs report
# the visible canonical equivalent as active to picker consumers.
"active": active,
"active_model": model,
"env_override": bool(os.environ.get(f"OMNIVOICE_{family.upper()}_BACKEND")),
"backends": public_backends(module.list_backends()),
"backends": backends,
}
def _is_hf_repo_id(value: str) -> bool:
@@ -110,6 +165,90 @@ def list_effects_presets():
return {"presets": list_effect_presets()}
@router.get("/engines/diarisation")
def diarisation_status():
"""Describe the selected local diarisation runtime without loading weights."""
from services.diarization_runtime import (
PYANNOTE,
SORTFORMER,
selected_backend,
sortformer_status,
)
selected = selected_backend()
native = selected == SORTFORMER
options = []
from api.routers.setup.models import KNOWN_MODELS, cache_is_complete, is_cached
pyannote_repo = "pyannote/speaker-diarization-3.1"
spec = next(model for model in KNOWN_MODELS if model["repo_id"] == pyannote_repo)
pyannote_installed = is_cached(pyannote_repo) and cache_is_complete(spec)
pyannote_reason = None if pyannote_installed else "Install the pyannote model bundle"
options.append({
"id": PYANNOTE,
"label": "pyannote 3.1",
"model": pyannote_repo,
"installed": pyannote_installed,
"reason": pyannote_reason,
})
native_status = sortformer_status()
native_installed = native_status["installed"]
native_model = native_status["model"]
native_reason = native_status["reason"]
options.append({
"id": SORTFORMER,
"label": "Sortformer v1 (audio.cpp)",
"model": native_model,
"model_installed": native_status["model_installed"],
"runtime_installed": native_status["runtime_installed"],
"installed": native_installed,
"reason": native_reason,
})
if native:
from services.diarization_native import is_running
return {"active": SORTFORMER, "label": "Sortformer v1 (audio.cpp)",
"model": native_model, "installed": native_installed, "loaded": False,
"model_installed": native_status["model_installed"],
"runtime_installed": native_status["runtime_installed"],
"busy": is_running(), "reason": native_reason, "options": options}
from services import model_manager
return {"active": PYANNOTE, "label": "pyannote 3.1", "model": pyannote_repo,
"installed": pyannote_installed,
"loaded": model_manager._diar_pipeline is not None, "reason": pyannote_reason,
"options": options}
class DiarisationSelection(BaseModel):
engine_id: str
@router.post("/engines/diarisation/select", dependencies=[Depends(require_admin)])
def select_diarisation_engine(request: DiarisationSelection):
"""Persist an installed diarisation runtime; environment overrides still win."""
from services.diarization_runtime import SORTFORMER, select_backend, selected_backend
status = diarisation_status()
option = next(
(item for item in status["options"] if item["id"] == request.engine_id),
None,
)
if option is None:
raise HTTPException(404, "Unknown diarisation engine")
if not option["installed"]:
raise HTTPException(409, option.get("reason") or "Install this diarisation engine first")
select_backend(request.engine_id)
if request.engine_id == SORTFORMER:
# Native Sortformer is stateless. Release a previously loaded pyannote
# pipeline so Engine Ready cannot hide stale accelerator memory.
from services import model_manager
model_manager.unload_diarization_pipeline()
return {
"active": selected_backend(),
"env_override": bool(os.environ.get("OMNIVOICE_DIARIZATION_BACKEND")),
}
@router.get("/engines/translation")
def list_translation_engines():
"""Translation engines with per-engine pip-package availability.
@@ -120,6 +259,7 @@ def list_translation_engines():
an engine whose Python dependency isn't importable yet.
"""
return {
"active": prefs.get("translation_backend", "argos"),
"engines": [
{**entry, "availability_reason": public_unavailability(entry.get("availability_reason"))}
for entry in translation_engines.list_engines()
@@ -128,6 +268,69 @@ def list_translation_engines():
}
class TranslationSelection(BaseModel):
engine_id: str
class ArgosPackRequest(BaseModel):
source_lang: str | None = None
target_langs: list[str] = Field(min_length=1, max_length=32)
job_id: str | None = None
def _argos_pack_request(request: ArgosPackRequest) -> tuple[str, list[str]]:
source = request.source_lang
if not source and request.job_id:
from api.routers.dub_core import _get_job
job = _get_job(request.job_id)
source = job.get("source_lang") if job else None
if not source:
raise HTTPException(422, "Transcribe the source before installing its language pack")
return source, request.target_langs
@router.post(
"/engines/translation/argos/packs/status",
dependencies=[Depends(require_admin)],
)
def argos_pack_status(request: ArgosPackRequest):
source, targets = _argos_pack_request(request)
try:
return translation_engines.argos_pack_status(source, targets)
except (ImportError, ValueError) as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
@router.post(
"/engines/translation/argos/packs/install",
dependencies=[Depends(require_admin)],
)
async def install_argos_packs(request: ArgosPackRequest):
source, targets = _argos_pack_request(request)
try:
return await asyncio.to_thread(
translation_engines.install_argos_packs,
source,
targets,
)
except (ImportError, ValueError) as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
@router.post("/engines/translation/select", dependencies=[Depends(require_admin)])
def select_translation_engine(request: TranslationSelection):
entry = translation_engines.get_engine(request.engine_id)
if not entry:
raise HTTPException(404, "Unknown translation engine")
if not translation_engines.is_installed(request.engine_id):
raise HTTPException(409, "Install this translation engine before selecting it")
if not translation_engines.is_ready(request.engine_id):
raise HTTPException(409, "Configure this translation provider before selecting it")
prefs.set_("translation_backend", request.engine_id)
return {"active": request.engine_id}
@router.post(
"/engines/translation/{engine_id}/install",
dependencies=[Depends(require_admin)],
@@ -188,18 +391,49 @@ async def uninstall_translation_engine(engine_id: str):
pkg = entry.get("pip_package")
if not pkg:
return {"status": "no_op", "engine": engine_id}
# The builtin flag is a promise someone has to remember to make; this
# check does not depend on it (#2019).
blocked = translation_engines.uninstall_blocker(engine_id)
if blocked:
raise HTTPException(status_code=blocked[0], detail=blocked[1])
rc, out = await translation_engines.run_pip(["uninstall", "-y", pkg])
if rc != 0:
raise HTTPException(status_code=500, detail=f"pip uninstall {pkg} failed ({rc}): {out[-1000:]}")
return {"status": "uninstalled", "engine": engine_id, "package": pkg, "log_tail": out[-800:]}
# ── Checksummed native audio.cpp runtime install ───────────────────────────
@router.get(
"/engines/audiocpp/runtime/install/status",
dependencies=[Depends(require_admin)],
)
def audiocpp_runtime_install_status():
from services import audiocpp_runtime_install
return audiocpp_runtime_install.status()
@router.post(
"/engines/audiocpp/runtime/install",
dependencies=[Depends(require_admin), Depends(require_desktop)],
)
def install_audiocpp_runtime():
from services import audiocpp_runtime_install
try:
return audiocpp_runtime_install.start_install()
except RuntimeError as exc:
raise HTTPException(status_code=409, detail=str(exc)) from exc
# ── One-click sidecar-engine install (IndexTTS-2 & friends) ────────────────
#
# Sidecar engines (dedicated venv + source checkout + weights, isolated from
# the parent's transformers>=5.3) used to require four manual terminal steps.
# These routes drive services.sidecar_install: POST starts a resumable
# background job, GET polls its step-by-step status (the Model Catalogue → Engines
# background job, GET polls its step-by-step status (the Model Catalogue
# Install button polls this), DELETE removes an app-managed install.
#
# Path namespace: /engines/sidecar/{engine_id}/… — NOT /engines/{engine_id}/…
@@ -230,6 +464,11 @@ def install_sidecar_engine(engine_id: str):
from services import sidecar_install
try:
return sidecar_install.start_install(engine_id)
except sidecar_install.HostUnsupported as exc:
# The engine has an installer, but not one that can work on this
# machine. 409, not 404: the route is right, the host is the problem,
# and the message (a VoiceStudio-owned sentence) says what to do.
raise HTTPException(status_code=409, detail=str(exc))
except KeyError:
raise HTTPException(
status_code=404,
@@ -354,6 +593,9 @@ def engine_health(engine_id: str):
)
t0 = perf_counter()
# Stable exception class when the probe itself raised, None when it merely
# returned not-available. Never the exception text — see the log line below.
raised_class: str | None = None
if hasattr(cls, "health_check"):
# SubprocessBackend path — spawn sidecar (if not running) and ping.
# ``health_check`` already swallows its own exceptions per Plan
@@ -364,6 +606,7 @@ def engine_health(engine_id: str):
ok, msg = instance.health_check()
except Exception as exc:
ok, msg = False, f"{type(exc).__name__}: {exc}"
raised_class = type(exc).__name__
else:
# In-process backend — `is_available()` is the classmethod-level
# liveness check. Cheap and side-effect-free for every shipping
@@ -372,6 +615,7 @@ def engine_health(engine_id: str):
ok, msg = cls.is_available()
except Exception as exc:
ok, msg = False, f"{type(exc).__name__}: {exc}"
raised_class = type(exc).__name__
# Engine-owned output can contain much more than shaped HF tokens: local
# paths, arbitrary credentials, source lines, or a nested traceback.
@@ -379,7 +623,38 @@ def engine_health(engine_id: str):
latency_ms = (perf_counter() - t0) * 1000.0
if not ok:
logger.warning("Engine health check failed; details withheld")
# The response tells the user to "check the backend log for details",
# and docs/engines/*.md asks a user diagnosing an unavailable engine to
# copy that engine's log lines. The old line named neither the engine
# nor anything about the probe, so neither instruction could be
# followed (#1866).
#
# `probe=` reports what the PROBE DID, not what went wrong. It cannot
# classify the cause: SubprocessBackend.health_check() swallows its own
# exceptions per Plan 02-01's contract, so a dead sidecar and a package
# that was never installed both arrive here as `returned-unavailable`.
# Separating those needs structured failure metadata from the probes
# themselves, which is a wider change than this one.
#
# Still no diagnostic text and still not the caller-supplied id: the
# engine id comes off the resolved registry class and a raised probe
# contributes only its exception class, the same shape
# core.public_errors.public_failure() logs as `class=`.
# tests/test_response_safety.py pins that boundary and passes
# unchanged.
#
# The id is a class attribute off the registry rather than caller
# input, but this line is a log-injection surface either way, so it is
# flattened to a single token before it goes in.
engine_label = str(getattr(cls, "id", None) or cls.__name__)
engine_label = "".join(
c if (c.isalnum() or c in "-_.") else "-" for c in engine_label
)[:64]
logger.warning(
"Engine health check failed; engine=%s probe=%s, details withheld",
engine_label or "unknown",
f"raised:{raised_class}" if raised_class else "returned-unavailable",
)
return {
"id": engine_id,
"ok": bool(ok),
@@ -589,7 +864,15 @@ def select_engine(req: SelectEngineRequest):
if not family:
raise HTTPException(400, f"Unknown family: {req.family}. Expected one of tts/asr/llm.")
module, pref_key = family
available = {b["id"]: b for b in module.list_backends()}
# MPS intentionally hides the redundant explicit OmniVoice sidecar from
# the picker, but existing scripts and saved preferences may still submit
# that supported compatibility id directly.
rows = (
module.list_backends(include_hidden=True)
if req.family == "tts"
else module.list_backends()
)
available = {b["id"]: b for b in rows}
if req.backend_id not in available:
raise HTTPException(400, f"Unknown {req.family} backend: {req.backend_id!r}")
entry = available[req.backend_id]
@@ -608,7 +891,7 @@ def select_engine(req: SelectEngineRequest):
# #981: mlx-audio multiplexes 7+ curated models behind one backend id —
# persist the model pick alongside the backend id so the UI can actually
# select which curated model gets loaded (previously it always defaulted
# to Kokoro no matter what the user downloaded in Model Catalogue → Models).
# to Kokoro no matter what the user downloaded in the engine's Weights list in Model Catalogue).
if req.family == "tts" and req.backend_id == "mlx-audio" and req.model_id is not None:
known_keys = tts_backend.MLXAudioBackend.CURATED_MODELS
# Accept a curated key OR a raw HF repo id ("owner/name") — the same
@@ -623,6 +906,23 @@ def select_engine(req: SelectEngineRequest):
"Hugging Face repo ID like 'owner/name'.",
)
prefs.set_("mlx_audio_model_id", req.model_id)
if req.family == "asr" and req.model_id is not None:
if req.backend_id not in {"faster-whisper", "faster-whisper-isolated"}:
raise HTTPException(400, "This ASR engine does not accept a CTranslate2 model")
from api.routers.setup.models import KNOWN_MODELS, is_cached
model = next((item for item in KNOWN_MODELS if item["repo_id"] == req.model_id), None)
compatible = req.model_id.startswith("Systran/faster-") or req.model_id == (
"deepdml/faster-whisper-large-v3-turbo-ct2"
)
if model is None or str(model.get("role", "")).lower() != "asr" or not compatible:
raise HTTPException(400, "This model is not compatible with Faster-Whisper")
if not is_cached(req.model_id):
raise HTTPException(409, "Install this ASR model before selecting it")
try:
asr_backend.select_faster_whisper_model(req.model_id)
except ValueError as exc:
raise HTTPException(409, str(exc)) from exc
prefs.set_(pref_key, req.backend_id)
return {
"family": req.family,
+89 -18
View File
@@ -736,7 +736,7 @@ def _oom_friendly_reraise(e):
# the OOM catch-all, telling a user with 63 GB of RAM to press Flush. Point
# at the real fix — set the variable — and never mention memory or Flush.
# The underlying error already names the exact variable + what to point it
# at (and Model Catalogue → Engines shows a copy-paste setup line), so keep it
# at (and Model Catalogue shows a copy-paste setup line), so keep it
# front-and-center. Checked before the OOM branch so a config error can
# never be mislabeled as memory.
if _is_config_failure(e):
@@ -745,7 +745,7 @@ def _oom_friendly_reraise(e):
f"environment variable that isn't configured, so nothing was "
f"generated. Set it as the underlying error describes (it names the "
f"exact variable and what to point it at), then restart VoiceStudio — "
f"or pick a ready engine in Model Catalogue → Engines. This is a setup "
f"or pick a ready engine in Model Catalogue. This is a setup "
f"problem, not a memory one. Underlying error: {e}"
) from e
# #880 (the class bug): the OOM hint used to be the catch-all fallback,
@@ -804,16 +804,34 @@ def _oom_friendly_reraise(e):
) from e
def _generate_timeout_s(text: str, *, execution_device=None) -> float:
def _generate_timeout_s(
text: str,
*,
execution_device=None,
min_vram_gb=0.0,
hardware_family=None,
vram_gb=None,
) -> float:
"""Wall-clock budget for one generate, scaled to the request.
Thin alias for the canonical helper, which moved to
``services.model_manager.generate_timeout_s`` (#1190) so /v1/audio/speech,
batch, dub and archetype previews share it instead of each re-deriving (or,
as they did, silently keeping the flat 300s).
``min_vram_gb`` is the engine's declared VRAM floor. A GPU below it pages to
system RAM and renders slower than this machine's CPU, so it must not be
budgeted as fast hardware (#1804) — the same figure the dispatch already
hands the guard so a timeout message can name the card (#1226/#1222).
"""
from services.model_manager import generate_timeout_s
return generate_timeout_s(text, execution_device=execution_device)
return generate_timeout_s(
text,
execution_device=execution_device,
min_vram_gb=min_vram_gb,
hardware_family=hardware_family,
vram_gb=vram_gb,
)
def _run_inference(
@@ -1053,7 +1071,7 @@ def _language_rejection_or(e: BaseException, backend, language):
f"The {engine} engine can't speak{requested}. VoiceStudio offers every "
f"language its default engine supports, but each engine covers a "
f"different set — pick one this engine supports, or switch engine in "
f"Model Catalogue → Engines (the VoiceStudio engine has the widest coverage) "
f"Model Catalogue (the VoiceStudio engine has the widest coverage) "
f"and generate again. Engine's own message: {e}"
)
@@ -1348,12 +1366,12 @@ async def generate_speech(
ref_text: Optional[str] = Form(None),
instruct: Optional[str] = Form(None),
duration: Optional[float] = Form(None),
num_step: int = Form(16),
num_step: Optional[int] = Form(None),
guidance_scale: float = Form(2.0),
speed: float = Form(1.0),
t_shift: Optional[float] = Form(None),
denoise: bool = Form(True),
postprocess_output: bool = Form(True),
postprocess_output: Optional[bool] = Form(None),
layer_penalty_factor: Optional[float] = Form(None),
position_temperature: Optional[float] = Form(None),
class_temperature: Optional[float] = Form(None),
@@ -1399,6 +1417,13 @@ async def generate_speech(
)
engine_id = engine or active_backend_id()
from services.performance_profiles import tts_defaults
sampling_defaults = tts_defaults(engine_id)
if num_step is None:
num_step = sampling_defaults.get("num_step", 16)
if postprocess_output is None:
postprocess_output = sampling_defaults.get("postprocess_output", True)
try:
backend_cls = get_backend_class(engine_id)
except ValueError:
@@ -1434,6 +1459,8 @@ async def generate_speech(
# local fallback call's timeout device-neutral so the closure is valid
# without pretending the control plane describes the remote worker.
_routing = {"effective_device": None}
_routing_hardware_family = None
_routing_vram_gb = None
if not _remote:
# Single-active-engine memory discipline: hand back any OTHER resident
@@ -1490,11 +1517,16 @@ async def generate_speech(
# 4090 from a Mac control plane would be refused by a gate describing
# a machine that is about to do nothing.
from core.device_caps import detect_host_caps
from services.engine_routing import resolve_routing, routing_notice
_routing = resolve_routing(
getattr(backend_cls, "gpu_compat", ("cpu",)), detect_host_caps(),
_engine_min_vram_gb,
from services.engine_routing import (
routing_notice,
runtime_compute_profile_async,
)
_routing = await runtime_compute_profile_async(
backend_cls, detect_host_caps()
)
_engine_min_vram_gb = _routing["min_vram_gb"]
_routing_hardware_family = _routing.get("runtime_hardware_family")
_routing_vram_gb = _routing.get("runtime_vram_gb")
if _routing["routing_status"] == "unavailable":
# The engine needs an accelerator this host lacks and has no CPU path.
raise HTTPException(status_code=400, detail=_routing["routing_reason"])
@@ -1525,7 +1557,7 @@ async def generate_speech(
detail=(
f"TTS engine '{engine_id}' did not finish loading within its "
f"model-load budget — on a first run this usually means the "
f"weight download is slow or stalled (check Model Catalogue → Models "
f"weight download is slow or stalled (check the engine's Weights list in Model Catalogue "
f"for progress), not that generation failed. Retry once the "
f"model shows as installed."
),
@@ -1718,7 +1750,13 @@ async def generate_speech(
local=gpu_gateway.LocalCall(
_remote_only_local_call(_target_label),
what="TTS generate",
timeout=_generate_timeout_s(text, execution_device=_routing["effective_device"]),
timeout=_generate_timeout_s(
text,
execution_device=_routing["effective_device"],
min_vram_gb=_engine_min_vram_gb,
hardware_family=_routing_hardware_family,
vram_gb=_routing_vram_gb,
),
min_vram_gb=_engine_min_vram_gb,
),
remote=_remote_call,
@@ -2012,7 +2050,13 @@ async def generate_speech(
),
what="TTS generate",
min_vram_gb=_engine_min_vram_gb,
timeout=_generate_timeout_s(text, execution_device=_routing["effective_device"]),
timeout=_generate_timeout_s(
text,
execution_device=_routing["effective_device"],
min_vram_gb=_engine_min_vram_gb,
hardware_family=_routing_hardware_family,
vram_gb=_routing_vram_gb,
),
on_abandon=release,
)
)
@@ -2032,7 +2076,13 @@ async def generate_speech(
),
what="TTS generate",
min_vram_gb=_engine_min_vram_gb,
timeout=_generate_timeout_s(text, execution_device=_routing["effective_device"]),
timeout=_generate_timeout_s(
text,
execution_device=_routing["effective_device"],
min_vram_gb=_engine_min_vram_gb,
hardware_family=_routing_hardware_family,
vram_gb=_routing_vram_gb,
),
on_abandon=release,
)
)
@@ -2072,7 +2122,13 @@ async def generate_speech(
# Budget scaled to THIS chunk (#1190) — the flat
# 300s here is what made long streamed renders fail
# even after the v0.3.22 scaled budget shipped.
timeout=_generate_timeout_s(chunk_text, execution_device=_routing["effective_device"]),
timeout=_generate_timeout_s(
chunk_text,
execution_device=_routing["effective_device"],
min_vram_gb=_engine_min_vram_gb,
hardware_family=_routing_hardware_family,
vram_gb=_routing_vram_gb,
),
on_abandon=release,
)
)
@@ -2232,7 +2288,13 @@ async def generate_speech(
_REMOTE_OP,
local=gpu_gateway.LocalCall(
_local_render, what="TTS generate",
timeout=_generate_timeout_s(text, execution_device=_routing["effective_device"]),
timeout=_generate_timeout_s(
text,
execution_device=_routing["effective_device"],
min_vram_gb=_engine_min_vram_gb,
hardware_family=_routing_hardware_family,
vram_gb=_routing_vram_gb,
),
min_vram_gb=_engine_min_vram_gb,
on_abandon=release,
),
@@ -2357,7 +2419,16 @@ async def generate_speech(
raise HTTPException(status_code=503, detail=str(e)) from e
except ValueError as e:
logger.error("Validation failed: %s", e)
raise HTTPException(status_code=400, detail=str(e)) from e
# Most ValueErrors here are VoiceStudio's own validation messages and
# are exactly what the user should read. A few are raw library text
# naming parameters and files the user cannot act on — those get the
# owned remedy for their class instead (#1879). Unclassified ones keep
# passing through, so this cannot swallow a good message.
from core.failure import classify, public_hint_for_topic
_topic = classify(str(e))
_owned = public_hint_for_topic(_topic) if _topic else ""
raise HTTPException(status_code=400, detail=_owned or str(e)) from e
except Exception as e:
tb = traceback.format_exc()
logger.error("Inference failed: %s\n%s", e, tb)
+4 -5
View File
@@ -325,10 +325,9 @@ async def create_speech(req: SpeechRequest):
# Routing gate (#21 — no silent CPU fallback), identical to REST /generate.
from core.device_caps import detect_host_caps
from services.engine_routing import resolve_routing, routing_notice
_routing = resolve_routing(
getattr(backend, "gpu_compat", ("cpu",)), detect_host_caps(),
getattr(backend, "min_vram_gb", 0.0),
from services.engine_routing import routing_notice, runtime_compute_profile_async
_routing = await runtime_compute_profile_async(
backend, detect_host_caps()
)
if _routing["routing_status"] == "unavailable":
raise HTTPException(status_code=400, detail=_routing["routing_reason"])
@@ -436,7 +435,7 @@ async def create_speech(req: SpeechRequest):
detail=(
f"TTS engine '{backend.id}' did not finish loading within its "
f"model-load budget — on a first run this usually means the weight "
f"download is slow or stalled (check Model Catalogue → Models for "
f"download is slow or stalled (check the engine's Weights list in Model Catalogue for "
f"progress), not that generation failed. Retry once the model "
f"shows as installed."
),
+182
View File
@@ -0,0 +1,182 @@
"""Keyless portrait search with bounded, normalized, safe thumbnails."""
import asyncio
import base64
import html
import re
from html.parser import HTMLParser
from urllib.parse import urlsplit
import httpx
from fastapi import APIRouter, HTTPException, Query
from core.profile_images import MAX_IMAGE_BYTES, normalize_portrait
router = APIRouter()
def trusted_thumbnail(url: str) -> bool:
try:
parsed = urlsplit(url)
google = parsed.hostname in {
f"encrypted-tbn{i}.gstatic.com" for i in range(4)
}
openverse = (
parsed.hostname == "api.openverse.org"
and re.fullmatch(r"/v1/images/[0-9a-f-]+/thumb/?", parsed.path) is not None
)
return (
parsed.scheme == "https"
and not parsed.username
and not parsed.password
and parsed.port in (None, 443)
and (google or openverse)
)
except ValueError:
return False
def google_thumbnails(document: str) -> list[tuple[str, str]]:
"""Extract result thumbnails only; never download third-party originals."""
results = []
seen = set()
def add(title, source):
if not isinstance(source, str) or source in seen:
return
if not (trusted_thumbnail(source) or source.startswith("data:image/jpeg;base64,")):
return
seen.add(source)
if len(results) < 20:
results.append((title or "", source))
class Images(HTMLParser):
def handle_starttag(self, tag, attrs):
if tag == "img":
values = dict(attrs)
add(values.get("alt"), values.get("src") or values.get("data-src"))
Images().feed(document)
# Google also assigns thumbnails from script strings after rendering.
decoded = html.unescape(document)
for escaped, literal in ((r"\u003d", "="), (r"\u0026", "&"), (r"\/", "/")):
decoded = decoded.replace(escaped, literal)
for match in re.finditer(r'https://encrypted-tbn[0-3]\.gstatic\.com/[^\s"\'<>\\]+|data:image/jpeg;base64,[A-Za-z0-9+/=]+', decoded):
add("", match.group())
return results
async def openverse_thumbnails(client: httpx.AsyncClient, name: str) -> list[tuple[str, str]]:
"""Public-domain/CC portrait fallback when Google returns its JS-only shell.
Openverse requires no user credential, excludes sensitive results by
default, and can restrict results to licenses that allow modification and
commercial use. We still fetch only its own thumbnail proxy.
"""
response = await client.get(
"https://api.openverse.org/v1/images/",
headers={
"User-Agent": "VoiceStudio/0.5 (+https://github.com/debpalash/VoiceStudio)",
"Accept": "application/json",
},
params={
"q": name,
"page_size": 20,
"mature": "false",
"extension": "jpg,png",
"aspect_ratio": "square",
"license_type": "commercial,modification",
},
)
response.raise_for_status()
if len(response.content) > 4 * 1024 * 1024:
raise ValueError("Search response too large")
payload = response.json()
rows = payload.get("results") if isinstance(payload, dict) else None
if not isinstance(rows, list):
raise ValueError("Invalid search response")
results = []
seen = set()
for row in rows:
if not isinstance(row, dict):
continue
source = row.get("thumbnail")
if not isinstance(source, str) or source in seen or not trusted_thumbnail(source):
continue
seen.add(source)
title = str(row.get("title") or name)
creator = str(row.get("creator") or "").strip()
license_name = str(row.get("license") or "").upper()
credit = " · ".join(value for value in (creator, license_name) if value)
results.append((f"{title}{credit}" if credit else title, source))
return results
@router.get("/profile-images/search")
async def search_profile_images(name: str = Query(min_length=1, max_length=100)):
if not name.strip():
raise HTTPException(422, detail={"code": "image_search_failed"})
async with httpx.AsyncClient(
timeout=15,
follow_redirects=False,
headers={
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 Chrome/140.0.0.0 Safari/537.36",
"Accept-Language": "en-US,en;q=0.9",
},
) as client:
results: list[tuple[str, str]] = []
try:
async with client.stream("GET", "https://www.google.com/search", params={
"q": name.strip(), "udm": "2", "safe": "active", "tbs": "ift:jpg",
}) as response:
response.raise_for_status()
document = bytearray()
async for chunk in response.aiter_bytes():
document.extend(chunk)
if len(document) > 4 * 1024 * 1024:
raise ValueError("Search page too large")
page = document.decode("utf-8", errors="replace")
results = google_thumbnails(page)
except (httpx.HTTPError, ValueError, TypeError):
# Search providers can change their anonymous HTML or reject a
# non-browser request. The fallback below keeps this explicit,
# user-triggered feature useful without requiring credentials.
results = []
if not results:
try:
results = await openverse_thumbnails(client, name.strip())
except (httpx.HTTPError, ValueError, TypeError):
results = []
if not results:
raise HTTPException(502, detail={"code": "image_search_failed"})
async def thumbnail(title, url):
try:
if url.startswith("data:image/jpeg;base64,"):
encoded = url.partition(",")[2]
if len(encoded) > MAX_IMAGE_BYTES * 4 // 3 + 4:
return None
data = base64.b64decode(encoded, validate=True)
else:
if not trusted_thumbnail(url):
return None
async with client.stream("GET", url) as image:
image.raise_for_status()
data = bytearray()
async for chunk in image.aiter_bytes():
data.extend(chunk)
if len(data) > MAX_IMAGE_BYTES:
return None
normalized = await asyncio.to_thread(normalize_portrait, bytes(data))
return {"title": title[:200], "data": base64.b64encode(normalized).decode("ascii")}
except (httpx.HTTPError, HTTPException, ValueError, TypeError):
return None
images = []
for start in range(0, len(results), 5):
batch = await asyncio.gather(*(thumbnail(*result) for result in results[start:start + 5]))
images.extend(image for image in batch if image)
if len(images) >= 5:
break
return {"images": images[:5]}
+109 -7
View File
@@ -1,3 +1,5 @@
import asyncio
import logging
import os
import re
import uuid
@@ -14,8 +16,21 @@ from core import event_bus
from core.personalities import get_personalities
from omnivoice.utils.voice_design import heal_design_instruct, sanitize_instruct
from core.path_security import UnsafePath, resolve_within
from core.profile_images import MAX_IMAGE_BYTES, normalize_portrait
from starlette.datastructures import UploadFile as StarletteUploadFile
router = APIRouter()
logger = logging.getLogger("omnivoice.profiles")
def _profile_record(row):
result = dict(row)
image_path = _voices_path(f"{result['id']}.portrait.jpg")
result["image_url"] = (
f"/profiles/{result['id']}/image?v={os.stat(image_path).st_mtime_ns}"
if image_path and os.path.isfile(image_path) else None
)
return result
class ProfileUpdate(BaseModel):
@@ -35,7 +50,7 @@ def list_personalities():
def list_profiles():
with db_conn() as conn:
rows = conn.execute("SELECT * FROM voice_profiles ORDER BY created_at DESC").fetchall()
return [dict(r) for r in rows]
return [_profile_record(r) for r in rows]
_DESIGN_SEED = 42 # deterministic sample render, same as archetype previews
@@ -51,6 +66,7 @@ async def create_profile(
personality: str = Form(""),
kind: str = Form("clone"),
vd_states: Optional[str] = Form(None),
image: Optional[UploadFile] = File(None),
):
"""Create a voice profile (spec: docs/specs/voice-studio-unification.md §5).
@@ -60,6 +76,9 @@ async def create_profile(
archetype materialization) and stores it as the profile's
reference so the voice identity is stable across runs.
"""
name = name.strip()
if not name:
raise HTTPException(status_code=400, detail="A voice profile needs a name.")
if kind not in ("clone", "design"):
raise HTTPException(status_code=422, detail="kind must be 'clone' or 'design'")
if kind == "clone" and ref_audio is None:
@@ -106,13 +125,38 @@ async def create_profile(
instruct = sanitize_instruct(instruct)
profile_id = str(uuid.uuid4())[:8]
portrait = None
if isinstance(image, StarletteUploadFile):
portrait = normalize_portrait(await image.read(MAX_IMAGE_BYTES + 1))
portrait_path = os.path.join(VOICES_DIR, f"{profile_id}.portrait.jpg")
if kind == "clone":
ext = os.path.splitext(ref_audio.filename or ".wav")[1]
audio_filename = f"{profile_id}{ext}"
audio_path = os.path.join(VOICES_DIR, audio_filename)
# Storage can be removed after startup; recover before persisting uploads.
os.makedirs(VOICES_DIR, exist_ok=True)
with open(audio_path, "wb") as f:
f.write(await ref_audio.read())
# A matching transcript defines the boundary between the reference and
# the requested line. Saving a blank transcript and waiting until the
# first generation made that first take depend on the TTS model's
# internal ASR fallback; short lines could then start with stray words
# from the reference. Resolve it while the profile is being created so
# every synthesis, including the first, uses stable conditioning. This
# remains best-effort and local-only: transcribe_reference considers
# only already-installed ASR/dictation models.
if not ref_text.strip():
try:
from services.asr_backend import transcribe_reference
ref_text = (
await asyncio.to_thread(transcribe_reference, audio_path) or ""
).strip()
except Exception as exc: # noqa: BLE001 — profile save remains usable
logger.warning(
"reference transcription during profile save failed: %s", exc
)
used_seed = seed
else:
# Saving a design profile is a pure persistence operation — it must not
@@ -153,6 +197,10 @@ async def create_profile(
used_seed = seed if seed is not None else _DESIGN_SEED
try:
if portrait:
os.makedirs(VOICES_DIR, exist_ok=True)
with open(portrait_path, "wb") as out:
out.write(portrait)
with db_conn() as conn:
conn.execute(
"INSERT INTO voice_profiles (id, name, ref_audio_path, ref_text, instruct, "
@@ -162,12 +210,14 @@ async def create_profile(
used_seed, personality, kind, vd_states, time.time())
)
except Exception:
if os.path.exists(portrait_path):
os.remove(portrait_path)
# Clean up orphaned audio file if DB insert fails
if os.path.exists(audio_path):
os.remove(audio_path)
raise
event_bus.emit("profiles", {"action": "created", "id": profile_id})
return {"id": profile_id, "name": name, "kind": kind}
return get_profile(profile_id)
@router.get("/profiles/{profile_id}")
def get_profile(profile_id: str):
@@ -181,14 +231,47 @@ def get_profile(profile_id: str):
status_code=404,
detail="That voice profile doesn't exist. It may have been deleted from another tab.",
)
return dict(row)
return _profile_record(row)
@router.get("/profiles/{profile_id}/image")
def get_profile_image(profile_id: str):
get_profile(profile_id)
path = _voices_path(f"{profile_id}.portrait.jpg")
if not path or not os.path.isfile(path):
raise HTTPException(404, "Profile image not found")
return FileResponse(path, media_type="image/jpeg", headers={"Cache-Control": "no-cache"})
@router.put("/profiles/{profile_id}/image")
async def update_profile_image(profile_id: str, image: UploadFile = File(...)):
get_profile(profile_id)
path = _voices_path(f"{profile_id}.portrait.jpg")
if path is None:
raise HTTPException(404, "Profile not found")
portrait = normalize_portrait(await image.read(MAX_IMAGE_BYTES + 1))
os.makedirs(VOICES_DIR, exist_ok=True)
with open(path, "wb") as out:
out.write(portrait)
event_bus.emit("profiles", {"action": "updated", "id": profile_id})
return get_profile(profile_id)
@router.put("/profiles/{profile_id}")
def update_profile(profile_id: str, patch: ProfileUpdate):
"""Partial update — only fields set on the payload are changed."""
with db_conn() as conn:
existing = conn.execute(
"SELECT kind FROM voice_profiles WHERE id = ?", (profile_id,),
).fetchone()
if not existing:
raise HTTPException(
status_code=404,
detail="That voice profile doesn't exist. It may have been deleted from another tab.",
)
fields = []
params = []
edited_instruct = None
for col in ("name", "ref_text", "instruct", "language", "personality"):
val = getattr(patch, col)
if val is None:
@@ -199,12 +282,22 @@ def update_profile(profile_id: str, patch: ProfileUpdate):
# Never let an edit persist a validator-rejecting instruct (prose /
# "[object Object]"); keep only whitelist tags (#550 #571 #594 #596).
val = sanitize_instruct(val)
edited_instruct = val
fields.append(f"{col} = ?")
params.append(val.strip() if col in ("name", "language") else val)
if edited_instruct is not None and existing["kind"] == "design":
# Keep the complete recipe synchronized with the editable instruct.
# Otherwise clients restore a stale vd_states snapshot and a successful
# style edit has no effect on the next generation.
import json
from core.describe_voice import instruct_to_vd_states
fields.append("vd_states = ?")
params.append(json.dumps(instruct_to_vd_states(edited_instruct)))
if not fields:
raise HTTPException(
status_code=400,
detail="PUT /profiles/{id} body contained no editable fields. Include at least one of: name, language, instruct, description.",
detail="PUT /profiles/{id} body contained no editable fields. Include at least one of: name, language, ref_text, instruct, personality.",
)
params.append(profile_id)
with db_conn() as conn:
@@ -221,7 +314,7 @@ def update_profile(profile_id: str, patch: ProfileUpdate):
"SELECT * FROM voice_profiles WHERE id = ?", (profile_id,),
).fetchone()
event_bus.emit("profiles", {"action": "updated", "id": profile_id})
return dict(row)
return _profile_record(row)
@router.get("/profiles/{profile_id}/usage")
@@ -252,8 +345,14 @@ def get_profile_usage(profile_id: str):
state = json.loads(r["state_json"] or "{}")
except Exception:
continue
segs = state.get("segments") or []
n = sum(1 for s in segs if s.get("profile_id") == profile_id)
if not isinstance(state, dict):
continue
# Current desktop snapshots use dubSegments. An explicit empty list
# supersedes legacy segments retained in an older snapshot.
segs = state.get("dubSegments", state.get("segments", []))
if not isinstance(segs, list):
continue
n = sum(1 for s in segs if isinstance(s, dict) and s.get("profile_id") == profile_id)
if n:
project_hits.append({
"project_id": r["id"],
@@ -539,6 +638,9 @@ def delete_profile(profile_id: str):
path = _voices_path(row[col])
if path and os.path.exists(path):
os.remove(path)
portrait_path = _voices_path(f"{profile_id}.portrait.jpg")
if portrait_path and os.path.isfile(portrait_path):
os.remove(portrait_path)
# Prevent FOREIGN KEY constraint failure
conn.execute("UPDATE generation_history SET profile_id = NULL WHERE profile_id=?", (profile_id,))
conn.execute("DELETE FROM voice_profiles WHERE id=?", (profile_id,))
+12 -1
View File
@@ -32,7 +32,11 @@ from pydantic import BaseModel
from api.dependencies import require_admin
from core.db import db_conn
from services.pronunciation import apply_pronunciation, entries_for_language
from services.pronunciation import (
apply_pronunciation,
entries_for_language,
inert_entries_for_language,
)
logger = logging.getLogger("omnivoice.pronunciation")
router = APIRouter(dependencies=[Depends(require_admin)])
@@ -250,11 +254,18 @@ def test_substitution(req: PronTestRequest):
).fetchall()
substituted = apply_pronunciation(req.text, rows, req.language)
applied = entries_for_language(rows, req.language)
# IPA/CMU rows are validated and stored but not applied yet, so a term that
# DOES match can still change nothing. Reporting them separately keeps the
# dry run honest — otherwise it says "no entries match", which is wrong and
# sends the user to re-type an entry that was already correct (#1949).
inert = inert_entries_for_language(rows, req.language)
return {
"input": req.text,
"substituted": substituted,
"changed": substituted != req.text,
"applied_terms": sorted(applied.keys(), key=len, reverse=True),
# Present but not honoured: [{term, type}, …]. Empty on the happy path.
"inert_entries": inert,
}
+75 -13
View File
@@ -20,6 +20,7 @@ from fastapi import APIRouter, Depends, HTTPException, Query
from pydantic import BaseModel, Field
from core.logging_utils import log_safe
from core.engine_licenses import LICENSE_GATED_ENGINES
from api.dependencies import require_admin, require_admin_action
logger = logging.getLogger("omnivoice.api.settings")
@@ -35,11 +36,11 @@ class _HFTokenBody(BaseModel):
token: str = Field(..., min_length=1, description="HuggingFace access token")
def _state_response() -> dict:
def _state_response(*, validate: bool = False) -> dict:
"""Return the same shape the React panel renders. Never includes raw token."""
from services import token_resolver
s = token_resolver.state()
s = token_resolver.state(validate=validate)
return {
"active": s["active"],
"sources": [asdict(row) for row in s["sources"]],
@@ -65,8 +66,7 @@ def save_hf_token(body: _HFTokenBody):
@router.delete("/hf-token")
def clear_hf_token(also_clear_hf_cli: bool = Query(False)):
"""Clear the App-source token. Optionally also call huggingface_hub.logout
to clear the canonical HF file. Returns the updated cascade state."""
"""Clear the App token and optionally recognized local Hub token files."""
from services import token_resolver
try:
token_resolver.clear_app_token(also_clear_hf_cli=also_clear_hf_cli)
@@ -82,13 +82,12 @@ def get_hf_token_state(fresh: bool = Query(False)):
``fresh=1`` drops the resolver's whoami validation cache first so the
response re-runs whoami for every source this is what the panel's
"Test now" button sends. Plain GETs (panel mounts) keep the 300s cache
so repeat Settings visits don't hammer the HF API.
"Test now" button sends. Plain GETs only inspect local token presence.
"""
from services import token_resolver
if fresh:
token_resolver.invalidate_cache()
return _state_response()
return _state_response(validate=fresh)
# ── Performance settings (INST-12) ────────────────────────────────────────
@@ -98,9 +97,66 @@ def get_hf_token_state(fresh: bool = Query(False)):
_TORCH_COMPILE_KEY = "perf.torch_compile_disabled"
from services.performance_profiles import (
_PERFORMANCE_PROFILE_KEY, _PERFORMANCE_TIERS, _PERFORMANCE_FAMILIES,
activate_performance_tier,
profile_state as _performance_profile_state,
)
class _PerformanceProfileBody(BaseModel):
tier: str = Field(..., description="fast | balanced | quality | max")
family: str | None = Field(None, description="Engine family, or null to set the global tier")
@router.get("/performance-profile")
def get_performance_profile():
"""Return the global speed/quality preference and per-engine overrides."""
return _performance_profile_state()
@router.put("/performance-profile")
def set_performance_profile(body: _PerformanceProfileBody):
"""Persist a performance preference and apply installed Max-capacity picks."""
from core import prefs
tier = body.tier.strip().lower()
if tier not in _PERFORMANCE_TIERS:
raise HTTPException(status_code=400, detail="Unknown performance tier")
family = body.family.strip().lower() if body.family else None
if family is not None and family not in _PERFORMANCE_FAMILIES:
raise HTTPException(status_code=400, detail="Unknown engine family")
state = _performance_profile_state()
applicable = state["applicable_families"]
if (family is not None and family not in applicable) or (family is None and not applicable):
raise HTTPException(status_code=409, detail="The selected engines do not support this performance preset")
from core import job_store
from api.routers.batch import list_batch_jobs
if job_store.list_jobs(status="active", limit=1) or list_batch_jobs(status="active", limit=1):
raise HTTPException(status_code=409, detail="Wait for queued or running jobs to finish before changing performance presets")
try:
if family is None:
# One atomic write clears family overrides together with the global
# choice, so a crash cannot leave half of a global change persisted.
prefs.update_mapping(_PERFORMANCE_PROFILE_KEY, {"global": tier}, replace=True)
else:
prefs.update_mapping(_PERFORMANCE_PROFILE_KEY, {family: tier})
except Exception:
logger.exception("set_performance_profile failed")
raise HTTPException(status_code=500, detail="Failed to persist performance profile")
activations = activate_performance_tier(tier, family)
result = _performance_profile_state()
if activations:
result["runtime_activations"] = activations
if tier == "max":
result["capacity_activations"] = activations
return result
class _TorchCompileBody(BaseModel):
enabled: bool = Field(..., description="True to set TORCH_COMPILE_DISABLE=1 on engine subprocesses")
enabled: bool = Field(..., description="True to disable torch.compile (eager mode) for the engine")
def _torch_compile_state() -> dict:
@@ -114,15 +170,21 @@ def _torch_compile_state() -> dict:
@router.get("/perf/torch-compile-disabled")
def get_torch_compile_disabled():
"""Return the current torch.compile-disabled toggle + the runtime platform.
UI uses the platform to render the toggle disabled (with an explainer)
on non-Windows hosts, since the OOM is Windows-specific (issue #65)."""
`platform` is still reported (clients may show it), but since #2135 the
toggle is live on every host: it used to be rendered disabled off Windows
on the assumption that only #65's Windows OOM needed it, which left the
Linux/CUDA reporter of #2135 with no way to switch off the compile that
was killing their backend.
"""
return _torch_compile_state()
@router.put("/perf/torch-compile-disabled")
def set_torch_compile_disabled(body: _TorchCompileBody):
"""Persist the toggle. Honoured by `services.engine_env.build_engine_env()`
which injects TORCH_COMPILE_DISABLE=1 on Windows when enabled."""
(subprocess engines) and `services.engine_env.should_torch_compile()`
(in-process), on every platform since #2135."""
from services import settings_store
try:
@@ -150,7 +212,7 @@ def _compute_device_state() -> dict:
caps = device_caps.detect_host_caps()
env_pin = (os.environ.get("OMNIVOICE_DEVICE") or "").strip().lower()
auto_family = next(
(f for f in ("cuda", "rocm", "xpu", "mps") if f in caps.available_families),
(f for f in device_caps.ACCELERATOR_PRIORITY if f in caps.available_families),
"cpu",
)
value = device_caps.requested_device_override()
@@ -655,7 +717,7 @@ def set_llm_skill(skill_id: str, body: _LLMSkillBody):
#: Engines that have an in-tree acceptance dialog. Adding a new engine
#: here means adding a corresponding frontend dialog + a license URLs
#: dict in its constants module. Until that, the API refuses the write.
_LICENSE_ALLOWED_ENGINES: frozenset[str] = frozenset({"supertonic3", "pockettts"})
_LICENSE_ALLOWED_ENGINES = LICENSE_GATED_ENGINES
class _LicenseAcceptBody(BaseModel):
+293 -21
View File
@@ -14,6 +14,7 @@ import logging
import os
import sys
import threading
import time
from fastapi import APIRouter, HTTPException
from fastapi.responses import StreamingResponse
@@ -31,7 +32,7 @@ from utils import download_aggregator
from .models import ( # noqa: F401
KNOWN_MODELS,
invalidate_cache,
snapshot_has_weights,
snapshot_is_complete,
disk_space_error,
_MIN_WEIGHT_BYTES,
_WEIGHT_FLOORS,
@@ -42,6 +43,9 @@ router = APIRouter()
# Cooldown: prevent rapid re-install after a failure. Maps repo_id → last_fail_time.
_install_cooldowns: dict[str, float] = {}
# Last classified failure per repo. The SSE stream carries the same detail live;
# retaining it here keeps recovery useful after navigation or renderer reconnect.
_install_failures: dict[str, dict] = {}
_COOLDOWN_SECS = 60.0
# Evict cooldown entries older than this so the dict can't grow unbounded across
# a long-lived process (MM2-06). Anything past the cooldown window is dead state.
@@ -54,6 +58,14 @@ def _sweep_cooldowns(now: float) -> None:
stale = [k for k, t in _install_cooldowns.items() if (now - t) > _COOLDOWN_TTL_SECS]
for k in stale:
_install_cooldowns.pop(k, None)
_install_failures.pop(k, None)
stale_failures = [
repo_id
for repo_id, failure in _install_failures.items()
if (now - float(failure.get("failed_at") or 0)) > _COOLDOWN_TTL_SECS
]
for repo_id in stale_failures:
_install_failures.pop(repo_id, None)
def clear_install_cooldowns() -> None:
@@ -63,6 +75,7 @@ def clear_install_cooldowns() -> None:
very next action is "retry the failed download on the new mirror", and a
429 there would dead-end the wizard's switch-and-retry flow."""
_install_cooldowns.clear()
_install_failures.clear()
# Repo_ids the user asked to cancel (FDL-11). Checked between retry attempts.
# Note: a single in-flight snapshot_download/Xet fetch is not interruptible
@@ -180,6 +193,28 @@ def _repo_cancelled(repo_id: str) -> bool:
return repo_id in _cancelled
def _create_cache_pointer(blob_path: str, pointer: str) -> None:
"""Keep the canonical blob while exposing it from the snapshot tree.
huggingface_hub's ``new_blob=True`` fallback moves the blob into the
snapshot when Windows symlinks are unavailable. The next model load then
sees a missing blob and downloads the same multi-gigabyte weight again.
NTFS hardlinks preserve both cache paths without doubling disk usage; other
filesystems fall back to Hugging Face's copy/symlink path.
"""
from huggingface_hub.file_download import _create_symlink
if os.name == "nt":
try:
os.link(blob_path, pointer)
return
except FileExistsError:
return
except OSError:
pass
_create_symlink(blob_path, pointer, new_blob=False)
def _segmented_snapshot(repo_id: str, *, endpoint: "str | None", revision: str) -> str:
"""Fetch every file of a repo via the segmented downloader into the HF
cache, mirroring hf_hub_download's blob+snapshot+refs layout so the result
@@ -190,7 +225,7 @@ def _segmented_snapshot(repo_id: str, *, endpoint: "str | None", revision: str)
import asyncio as _asyncio
from huggingface_hub import HfApi, constants as _C
from huggingface_hub.file_download import (
hf_hub_url, get_hf_file_metadata, repo_folder_name, _create_symlink,
hf_hub_url, get_hf_file_metadata, repo_folder_name,
)
from services.segmented_download import segmented_download
from services.token_resolver import resolve as _resolve_token
@@ -229,7 +264,7 @@ def _segmented_snapshot(repo_id: str, *, endpoint: "str | None", revision: str)
cancel_check=lambda: _repo_cancelled(repo_id),
))
if not os.path.lexists(pointer):
_create_symlink(blob_path, pointer, new_blob=True)
_create_cache_pointer(blob_path, pointer)
# refs/main → commit so scan_cache_dir maps the revision correctly.
ref_path = os.path.join(refs_dir, "main")
@@ -279,10 +314,17 @@ def _validate_snapshot_has_weights(repo_id: str, snapshot_path: str) -> None:
retry loop and the UI's re-download path can deal with it, instead of at
first synthesis with an opaque transformers error.
Delegates the weight check to ``models.snapshot_has_weights`` (single source of
the floors); only the install-time error message lives here."""
if snapshot_has_weights(snapshot_path):
Delegates to ``models.snapshot_is_complete`` so configuration-only pipeline
repositories use their declared required files instead of a weight floor."""
model = next((m for m in KNOWN_MODELS if m["repo_id"] == repo_id), {"repo_id": repo_id})
if snapshot_is_complete(model, snapshot_path):
return
if model.get("config_only"):
required = ", ".join(model.get("config_required_files") or ())
raise OSError(f"{repo_id}: download is incomplete; required configuration files: {required}")
if model.get("required_files"):
required = ", ".join(model["required_files"])
raise OSError(f"{repo_id}: required model files are missing or incomplete: {required}")
biggest = 0
try:
for root, _dirs, files in os.walk(snapshot_path, followlinks=True):
@@ -297,7 +339,7 @@ def _validate_snapshot_has_weights(repo_id: str, snapshot_path: str) -> None:
f"{repo_id}: download finished but no model weights were found in the "
"snapshot (largest file "
f"{biggest} bytes). The download was likely interrupted — delete the "
"model in Model Catalogue → Models and install it again."
"model in Model Catalogue and install it again."
)
@@ -346,6 +388,65 @@ class InstallModelRequest(BaseModel):
target: str | None = None
@router.get("/models/install/status")
def model_install_status():
"""Read local and remote jobs after navigation without starting downloads."""
from services import gpu_gateway # noqa: PLC0415
now = time.time()
_sweep_cooldowns(now)
with _active_installs_lock:
active = tuple(_active_installs)
jobs = []
for repo_id in active:
aggregate = download_aggregator._get(repo_id)
jobs.append(
{
"repo_id": repo_id,
"target": "local",
"state": "downloading",
**(aggregate.snapshot() if aggregate else {}),
}
)
detailed = set()
for repo_id, failure in tuple(_install_failures.items()):
failed_at = float(failure.get("failed_at") or 0)
if repo_id in active or now - failed_at >= _COOLDOWN_SECS:
continue
detailed.add(repo_id)
cooldown_at = _install_cooldowns.get(repo_id)
retry_after = (
max(0, int(_COOLDOWN_SECS - (now - cooldown_at) + 0.999))
if cooldown_at is not None
else 0
)
jobs.append(
{
"repo_id": repo_id,
"target": "local",
"state": "failed",
"retry_after_seconds": retry_after,
**failure,
}
)
# Preserve status for callers/tests that seed the legacy cooldown map alone.
jobs.extend(
{
"repo_id": repo_id,
"target": "local",
"state": "failed",
"retry_after_seconds": max(
0, int(_COOLDOWN_SECS - (now - failed_at) + 0.999)
),
}
for repo_id, failed_at in tuple(_install_cooldowns.items())
if repo_id not in active
and repo_id not in detailed
and now - failed_at < _COOLDOWN_SECS
)
jobs.extend(gpu_gateway.remote_download_jobs())
return {"jobs": jobs}
def _is_retryable_download_error(exc: BaseException) -> bool:
"""Whether a failed download attempt is worth retrying.
@@ -381,11 +482,60 @@ def _is_retryable_download_error(exc: BaseException) -> bool:
return is_hf_connectivity_error(str(exc))
def _segmented_retry_plan(
exc: BaseException, attempt: int, max_attempts: int
) -> tuple[bool, bool]:
"""What to do after the segmented accelerator failed on ``attempt``.
Returns ``(disable_accelerator, reraise)``.
A dropped connection is not the accelerator's fault, so the error is
re-raised for the outer retry: the next attempt re-enters
:func:`_segmented_snapshot`, which resumes from the ``.part`` manifest.
Falling straight through to ``snapshot_download`` instead would finish the
install from a separate ``.incomplete`` file and strand that manifest the
restart-from-zero this exists to prevent.
The final attempt is always reserved for the plain path, so the accelerator
can never be the reason an install fails outright. The two flags are
decoupled for that handover: the attempt that exhausts the accelerator still
re-raises, so the plain path starts on the LAST attempt rather than the
second-to-last. Disabling and falling through in the same attempt would
abandon the resumable manifest one attempt early and restart through a
separate file which is the failure this whole helper exists to avoid.
"""
if not _is_retryable_download_error(exc):
return True, False # the accelerator cannot work here at all
if attempt >= max_attempts:
# Nothing left to hand over to: take the plain path now rather than
# re-raising out of the loop with no fallback ever tried.
return True, False
return attempt >= max_attempts - 1, True
def _segmented_retry_note(disable: bool, reraise: bool) -> str:
"""How to describe the outcome of :func:`_segmented_retry_plan` in the log.
Three distinct states, and reading only ``disable`` conflates two of them:
the attempt that exhausts the accelerator is disabled AND re-raises, so the
fallback starts on the NEXT attempt, not this one.
"""
if not disable:
return "kept for the next attempt (resumes from its manifest)"
if reraise:
return "exhausted — retrying once more, then snapshot_download takes over"
return "disabled for this install — falling back to snapshot_download now"
@router.post("/models/install")
async def install_model(req: InstallModelRequest):
"""Download one HF repo snapshot; progress goes through the shared
``/setup/download-stream`` SSE feed."""
if req.repo_id not in [m["repo_id"] for m in KNOWN_MODELS]:
model_spec = next(
(model for model in KNOWN_MODELS if model["repo_id"] == req.repo_id),
None,
)
if model_spec is None:
raise HTTPException(
status_code=400,
detail=(
@@ -393,6 +543,7 @@ async def install_model(req: InstallModelRequest):
+ ", ".join(m["repo_id"] for m in KNOWN_MODELS)
),
)
allow_patterns = list(model_spec.get("allow_patterns") or []) or None
target = (req.target or "").strip()
if target != "local":
from services import gpu_gateway # noqa: PLC0415
@@ -424,6 +575,9 @@ async def install_model(req: InstallModelRequest):
loop = asyncio.get_running_loop()
def _do():
# Failure handling must work even when imports, token resolution or
# revision lookup fail before the heartbeat thread is started.
_resolving = threading.Event()
token = hf_progress.current_repo_id.set(req.repo_id)
target_token = hf_progress.current_target.set("local")
hf_progress.emit({
@@ -450,6 +604,12 @@ async def install_model(req: InstallModelRequest):
"revision": revision_for(req.repo_id),
"max_workers": _download_max_workers(),
}
from services.token_resolver import resolve as resolve_token
resolved_token = resolve_token()
if resolved_token:
dl_kwargs["token"] = resolved_token.token
if allow_patterns:
dl_kwargs["allow_patterns"] = allow_patterns
_tqdm_cls = hf_progress.tracked_tqdm_class()
if _tqdm_cls is not None:
dl_kwargs["tqdm_class"] = _tqdm_cls
@@ -461,9 +621,7 @@ async def install_model(req: InstallModelRequest):
# Emit a 'resolving' heartbeat every 2s while snapshot_download
# resolves repo metadata (before any tqdm bars appear).
import threading
import time as _t
_resolving = threading.Event()
def _heartbeat():
_step = 0
@@ -493,10 +651,26 @@ async def install_model(req: InstallModelRequest):
"revision": dl_kwargs["revision"],
"dry_run": True,
}
if allow_patterns:
_preflight_kwargs["allow_patterns"] = allow_patterns
if _endpoint:
_preflight_kwargs["endpoint"] = _endpoint
if resolved_token:
_preflight_kwargs["token"] = resolved_token.token
try:
_plan = snapshot_download(**_preflight_kwargs) # nosec B615 -- immutable revision_for pin
_plan = list(snapshot_download(**_preflight_kwargs)) # nosec B615 -- immutable revision_for pin
for dependency in model_spec.get("dependencies") or ():
if req.repo_id in _cancelled:
raise _InstallCancelled()
dependency_plan_kwargs = {
**_preflight_kwargs,
"repo_id": dependency["repo_id"],
"revision": revision_for(dependency["repo_id"]),
}
dependency_plan_kwargs.pop("allow_patterns", None)
if dependency.get("allow_patterns"):
dependency_plan_kwargs["allow_patterns"] = dependency["allow_patterns"]
_plan.extend(snapshot_download(**dependency_plan_kwargs)) # nosec B615 -- immutable revision_for pin
_summary = compute_plan(_plan)
# Disk-space guard (before a single byte flows): the preflight
# gives an exact "to download" size, so reject an install that
@@ -514,6 +688,11 @@ async def install_model(req: InstallModelRequest):
"phase": "install_error",
"error": _disk_err,
})
_install_failures[req.repo_id] = {
"failed_at": time.time(),
"error": _disk_err,
"docs_topic": "DISK_SPACE_LOW",
}
# A disk-full is not a transient network failure — don't set
# a cooldown (freeing space, not waiting, is the fix). The
# outer finally still cleans up the aggregator + context.
@@ -530,6 +709,8 @@ async def install_model(req: InstallModelRequest):
"phase": "install_plan",
**_summary,
})
except _InstallCancelled:
raise
except Exception as _pf_err:
# No preflight (older/gated repo, mirror without dry-run, etc.):
# fall back to today's fill-in-as-files-appear behaviour.
@@ -548,6 +729,11 @@ async def install_model(req: InstallModelRequest):
_max_attempts = 5
_attempt = 0
# The accelerator is retried across attempts so its manifest-based
# resume actually gets used; it is disabled for the rest of the
# install only when it fails for a reason that is NOT transient
# network trouble (i.e. the accelerator itself is unusable here).
_segmented_off = False
while True:
if req.repo_id in _cancelled:
raise _InstallCancelled()
@@ -555,11 +741,17 @@ async def install_model(req: InstallModelRequest):
try:
# Segmented accelerator (FDL-09, default ON): parallel
# byte-range fetch with real live progress, for the
# legacy-LFS path. Any failure falls through to
# snapshot_download — the accelerator can never compromise a
# correct install.
# legacy-LFS path. A failure that is not transient network
# trouble falls through to snapshot_download, and so does the
# install's last attempt — the accelerator can never
# compromise a correct install (see _segmented_retry_plan).
_snapshot_path = None
if _attempt == 1 and _segmented_enabled() and not _xet_active():
if (
not _segmented_off
and not allow_patterns
and _segmented_enabled()
and not _xet_active()
):
try:
_snapshot_path = _segmented_snapshot(
req.repo_id,
@@ -569,10 +761,16 @@ async def install_model(req: InstallModelRequest):
except _InstallCancelled:
raise
except Exception as _seg_err:
logger.info(
"segmented download for %s failed (%s); falling back to snapshot_download",
req.repo_id, _seg_err,
_segmented_off, _seg_reraise = _segmented_retry_plan(
_seg_err, _attempt, _max_attempts
)
logger.info(
"segmented download for %s failed (%s); accelerator %s",
req.repo_id, _seg_err,
_segmented_retry_note(_segmented_off, _seg_reraise),
)
if _seg_reraise:
raise
_snapshot_path = None
if _snapshot_path is None:
_snapshot_path = snapshot_download(**dl_kwargs) # nosec B615 -- immutable revision_for pin
@@ -580,6 +778,27 @@ async def install_model(req: InstallModelRequest):
from huggingface_hub.constants import HF_HUB_CACHE
from services.hf_revisions import remember_revision
remember_revision(req.repo_id, dl_kwargs["revision"], HF_HUB_CACHE)
# A pipeline config is not a runnable installation by itself.
# Download its reviewed dependencies only inside this explicit
# install action, retaining the parent cancellation/retry flow.
for dependency in model_spec.get("dependencies") or ():
if req.repo_id in _cancelled:
raise _InstallCancelled()
dependency_id = dependency["repo_id"]
dependency_kwargs = {
**dl_kwargs,
"repo_id": dependency_id,
"revision": revision_for(dependency_id),
}
dependency_kwargs.pop("allow_patterns", None)
if dependency.get("allow_patterns"):
dependency_kwargs["allow_patterns"] = dependency["allow_patterns"]
dependency_path = snapshot_download(**dependency_kwargs) # nosec B615 -- immutable revision_for pin
if not snapshot_is_complete(dependency, dependency_path):
raise OSError(f"{dependency_id}: required model files are missing or incomplete")
remember_revision(dependency_id, dependency_kwargs["revision"], HF_HUB_CACHE)
if req.repo_id in _cancelled:
raise _InstallCancelled()
break
except Exception as net_err:
# #1224: a truncated body ("peer closed connection without
@@ -643,12 +862,28 @@ async def install_model(req: InstallModelRequest):
"phase": "install_done",
})
_install_cooldowns.pop(req.repo_id, None) # success clears any cooldown (MM2-06)
_install_failures.pop(req.repo_id, None)
invalidate_cache()
# A saved performance pack owns the desired engine/model policy.
# Reconcile after every successful local install so the new model
# becomes usable without a restart or a second manual selection.
try:
from services.performance_profiles import reconcile_active_profile
activated = reconcile_active_profile()
if activated:
logger.info("model install activated performance profile: %s", activated)
except Exception:
# The model is fully installed even if optional preference
# reconciliation fails; readiness refresh and manual selection
# remain available instead of misreporting the download.
logger.exception("performance profile reconciliation failed after model install")
except _InstallCancelled:
_resolving.set()
logger.info("model install cancelled: %s", req.repo_id)
# A cancel is user intent, not a failure — don't set a cooldown.
_install_cooldowns.pop(req.repo_id, None)
_install_failures.pop(req.repo_id, None)
hf_progress.emit({
"repo_id": req.repo_id,
"filename": req.repo_id,
@@ -659,7 +894,8 @@ async def install_model(req: InstallModelRequest):
_resolving.set()
logger.info("model install failed for %s: %s", req.repo_id, e)
import time as _time_fail
_install_cooldowns[req.repo_id] = _time_fail.time()
_failed_at = _time_fail.time()
_install_cooldowns[req.repo_id] = _failed_at
# #874: when the install failed because the configured HF mirror is
# unreachable, name the mirror + the setting instead of leaking the
# raw connectivity error. #959: likewise for the SOCKS-proxy class
@@ -668,15 +904,41 @@ async def install_model(req: InstallModelRequest):
# class so the wizard can react structurally (HF_MIRROR_UNREACHABLE
# raises the inline mirror picker) without string-matching.
from core.failure import append_hint, classify
_error = append_hint(str(e))
_docs_topic = classify(str(e))
# Gated catalogue entries own their recovery topic. Hugging Face
# uses several exception wordings for the same access verdict, so
# the UI must not depend on parsing an English 401/403 message.
_catalogue_topic = str(model_spec.get("failure_topic") or "")
if _catalogue_topic and _docs_topic in {
"",
"HF_AUTH_FAILED",
"PYANNOTE_LICENSE_REQUIRED",
}:
_docs_topic = _catalogue_topic
# Waiting cannot fix an access/token verdict. Let the user accept
# the terms or update the token and retry immediately.
if _docs_topic in {
"HF_AUTH_FAILED",
"PYANNOTE_LICENSE_REQUIRED",
"POCKETTTS_GATED_WEIGHTS",
}:
_install_cooldowns.pop(req.repo_id, None)
_install_failures[req.repo_id] = {
"failed_at": _failed_at,
"error": _error,
"docs_topic": _docs_topic,
}
hf_progress.emit({
"repo_id": req.repo_id,
"filename": req.repo_id,
"downloaded": 0, "total": 0, "pct": 0.0,
"phase": "install_error",
"error": append_hint(str(e)),
"docs_topic": classify(str(e)),
"error": _error,
"docs_topic": _docs_topic,
})
finally:
_resolving.set()
_cancelled.discard(req.repo_id)
download_aggregator.finish(req.repo_id, target=target or "local")
hf_progress.current_repo_id.reset(token)
@@ -691,6 +953,7 @@ async def install_model(req: InstallModelRequest):
# Admission and task publication are one atomic generation boundary:
# cancellation can never observe an admitted install without its task.
_cancelled.discard(req.repo_id)
_install_failures.pop(req.repo_id, None)
try:
task = loop.create_task(asyncio.to_thread(_do))
_install_tasks.add(task)
@@ -740,8 +1003,17 @@ async def cancel_install(req: InstallModelRequest):
in hf_hub 1.7.2, so an already-streaming file finishes; the cancel takes
effect at the next retry boundary. Clears the cooldown so the user can
immediately restart."""
target = (req.target or "local").strip() or "local"
if target != "local":
from services import gpu_gateway # noqa: PLC0415
try:
return await gpu_gateway.cancel_download(req.repo_id, target=target)
except gpu_gateway.GatewayError as exc:
raise HTTPException(status_code=409, detail=str(exc)) from exc
_cancelled.add(req.repo_id)
_install_cooldowns.pop(req.repo_id, None)
_install_failures.pop(req.repo_id, None)
return {"cancelling": req.repo_id}
+122 -24
View File
@@ -15,7 +15,7 @@ import sys
import time
from pathlib import Path
from fastapi import APIRouter
from fastapi import APIRouter, HTTPException, Query
logger = logging.getLogger("omnivoice.setup.models")
router = APIRouter()
@@ -117,7 +117,7 @@ def _target_repo_inventory() -> tuple[str, set[str]] | None:
for capability in live.record.capabilities or []:
if capability.get("downloaded"):
downloaded.update(str(repo) for repo in capability.get("repo_ids") or [])
return live.id, downloaded
return live.worker_id, downloaded
def _current_platform_tags() -> list[str]:
@@ -368,23 +368,44 @@ def _snapshot_dirs(repo_id: str) -> list[str]:
return dirs
def snapshot_is_complete(model: dict, snapshot_path: str) -> bool:
"""Apply the same catalogue requirements during installation and listing."""
config_only = bool(model.get("config_only"))
required = tuple(str(name) for name in (
model.get("config_required_files") if config_only else model.get("required_files")
) or ())
if config_only and not required:
return False
try:
present = all(
os.path.isfile(os.path.join(snapshot_path, name))
and os.path.getsize(os.path.join(snapshot_path, name)) >= (
1 if config_only else _WEIGHT_FLOORS.get(os.path.splitext(name)[1].lower(), 1)
)
for name in required
)
return present and (config_only or snapshot_has_weights(snapshot_path))
except OSError:
return False
def cache_is_complete(model: dict) -> bool:
"""True when this model's on-disk cache is usable (not a truncated download).
Config-only repos (``config_only: true`` in models.yaml e.g. pyannote's
diarisation pipeline, whose real weights live in referenced sub-repos) carry no
weight file of their own, so the weight check would false-positive them as
incomplete (#622 caveat). They're exempt: cache presence alone means complete.
A weight-bearing repo is complete only if at least one of its snapshots has
weights; if no snapshot dir is found on disk we can't prove truncation, so we
don't downgrade (the size-based caller already decided it's cached).
Config-only repos carry no weight file of their own. Their catalogue entry
declares the small files that make the pipeline usable, so a README left by a
gated 403 is not mistaken for a completed install. A weight-bearing repo is
complete only if at least one snapshot has weights; if no snapshot directory
is found, the size-based caller's cached result is preserved.
"""
if model.get("config_only"):
return True
for dependency in model.get("dependencies") or ():
snapshots = _snapshot_dirs(dependency["repo_id"])
if not any(snapshot_is_complete(dependency, path) for path in snapshots):
return False
dirs = _snapshot_dirs(model["repo_id"])
if not dirs:
return True
return any(snapshot_has_weights(d) for d in dirs)
return any(snapshot_is_complete(model, snapshot) for snapshot in dirs)
def _is_cached_on_disk(repo_id: str) -> bool:
@@ -449,6 +470,17 @@ def _scan_cache_on_disk() -> dict[str, dict]:
return out
def _cache_dir_missing(exc: Exception) -> bool:
"""Whether Hugging Face is reporting the normal empty-cache state.
``CacheNotFound`` is expected on a clean installation before the first
download. Treating it like a damaged Windows cache makes every model probe
perform a redundant filesystem fallback and fills the first-run log with
warnings. Unexpected scan failures remain visible and recoverable below.
"""
return type(exc).__name__ == "CacheNotFound"
def is_cached(repo_id: str) -> bool:
"""Best-effort check: does HF have this repo in its cache on disk?"""
try:
@@ -459,6 +491,8 @@ def is_cached(repo_id: str) -> bool:
return True
return False
except Exception as e:
if _cache_dir_missing(e):
return False
# scan_cache_dir can raise on Windows (WinError 448 'untrusted mount
# point'); fall back to a direct disk check so a cached model isn't
# mistaken for missing and re-downloaded in a loop (#117/#118). Logged
@@ -497,6 +531,62 @@ def invalidate_cache() -> None:
# ── Endpoints ──────────────────────────────────────────────────────────────
@router.get("/models/access/status")
def model_access_status(repo_id: str = Query(...)):
"""Check gated Hub access without downloading model files.
This route runs only after an explicit UI action. It never returns the
token or a raw Hub exception; callers need only the per-repository verdict.
"""
model = _catalog.get(repo_id)
if model is None:
raise HTTPException(status_code=404, detail="Unknown model")
if not model.get("gated"):
return {
"repo_id": repo_id,
"token_present": False,
"ready": True,
"repositories": [],
}
from services import token_resolver
resolved = token_resolver.resolve()
repositories = [repo_id]
prerequisite = str(model.get("prerequisite_repo_id") or "").strip()
if prerequisite:
repositories.append(prerequisite)
if not resolved:
return {
"repo_id": repo_id,
"token_present": False,
"ready": False,
"repositories": [
{"repo_id": current, "access": "token_missing"}
for current in repositories
],
}
from huggingface_hub import get_hf_file_metadata, hf_hub_url
results = []
for current in repositories:
try:
url = hf_hub_url(current, filename=".gitattributes")
get_hf_file_metadata(url, token=resolved.token)
access = "granted"
except Exception as exc: # Hub exception types vary across releases.
status = getattr(getattr(exc, "response", None), "status_code", None)
access = "required" if status in {401, 403, 404} else "unavailable"
results.append({"repo_id": current, "access": access})
return {
"repo_id": repo_id,
"token_present": True,
"ready": all(item["access"] == "granted" for item in results),
"repositories": results,
}
@router.get("/models")
def list_models():
"""Catalogue every known model + its on-disk install state.
@@ -532,10 +622,13 @@ def list_models():
"nb_files": entry.nb_files,
}
except Exception as e:
# WinError-448 fallback (#117/#118): use a direct disk scan so installed
# models still show as installed instead of offering a re-download.
logger.warning("scan_cache_dir failed (%s); using disk fallback", e)
cached_by_repo = _scan_cache_on_disk()
if _cache_dir_missing(e):
cached_by_repo = {}
else:
# WinError-448 fallback (#117/#118): use a direct disk scan so installed
# models still show as installed instead of offering a re-download.
logger.warning("scan_cache_dir failed (%s); using disk fallback", e)
cached_by_repo = _scan_cache_on_disk()
out = []
host_tags = set(platform_tags)
@@ -562,6 +655,7 @@ def list_models():
"curated": _model_curated(m, host_tags),
})
response = {
"target": target_key,
"models": out,
"total_installed_bytes": sum(m["size_on_disk_bytes"] for m in out),
"hf_cache_dir": "" if remote_inventory is not None else hf_cache_dir(),
@@ -623,21 +717,21 @@ def recommendations():
rationale = (
"NVIDIA preset: VoiceStudio (required) runs standalone. Optional ASR picks "
"are CUDA-accelerated via CTranslate2 — Whisper large-v3 for dubbing "
"(best word timestamps), Turbo for 5× faster transcription, Parakeet TDT "
"v3 for live dictation. KittenTTS adds CPU-realtime English."
"(best word timestamps) and Turbo for 5× faster transcription. Whisper "
"Tiny provides broad-language local dictation. KittenTTS adds CPU-realtime English."
)
elif has_rocm:
rationale = (
"AMD/ROCm preset: VoiceStudio (required) runs standalone. CTranslate2 has "
"no ROCm backend, so the PyTorch Whisper large-v3 build is the "
"GPU-accelerated ASR route; faster-whisper works on CPU, and Parakeet "
"TDT v3 handles live dictation."
"GPU-accelerated ASR route; faster-whisper works on CPU, and Whisper Tiny "
"provides broad-language local dictation."
)
else:
rationale = (
"CPU preset: VoiceStudio (required) runs standalone. Optional picks favour "
"speed on CPU — Whisper large-v3 (int8) for accuracy, Turbo when speed "
"matters, Parakeet TDT v3 (int8 ONNX) for live dictation, KittenTTS for "
"matters, Whisper Tiny (ONNX) for live dictation, KittenTTS for "
"instant English TTS."
)
@@ -653,9 +747,12 @@ def recommendations():
entry.repo_id for entry in info.repos if entry.size_on_disk > 0
}
except Exception as e:
# WinError-448 fallback (#117/#118): recommend based on the disk scan.
logger.debug("scan_cache_dir failed (%s); using disk fallback", e)
cached_ids = set(_scan_cache_on_disk().keys())
if _cache_dir_missing(e):
cached_ids = set()
else:
# WinError-448 fallback (#117/#118): recommend based on the disk scan.
logger.debug("scan_cache_dir failed (%s); using disk fallback", e)
cached_ids = set(_scan_cache_on_disk().keys())
entries = []
for meta in curated:
@@ -679,6 +776,7 @@ def recommendations():
all_installed = all(e["installed"] for e in entries)
return {
"target": remote_inventory[0] if remote_inventory is not None else "local",
"device": {
"os": target_os,
"arch": target_arch,
+12 -7
View File
@@ -19,6 +19,7 @@ import sys
from fastapi import APIRouter
from api.schemas import SetupStatusResponse, PreflightResponse
from core.device_caps import KERNEL_RISK_MARKER
# MIN_FREE_GB + disk_free_bytes are single-sourced in ``.models`` (the lowest
# module in the setup import graph) so the wizard gate, the /models header, and
# the per-install disk guard can't drift apart.
@@ -174,8 +175,8 @@ def _detect_gpu() -> dict:
return info
def _probe_network(host: str = "huggingface.co", port: int = 443, timeout: float = 2.0) -> bool:
"""Tiny TCP connect test."""
def _probe_network(host: str = "huggingface.co", port: int = 443, timeout: float = 8.0) -> bool:
"""Tiny TCP connect test. 8s default — high-latency / China paths often exceed 23s."""
import socket
try:
with socket.create_connection((host, port), timeout=timeout):
@@ -188,7 +189,7 @@ def _hf_endpoint_host() -> tuple[str, int]:
"""Host/port of the Hugging Face endpoint actually in effect.
Mirror-aware: restricted-network users (e.g. behind the Great Firewall)
point HF_ENDPOINT at a mirror via Model Catalogue Models Hugging Face
point HF_ENDPOINT at a mirror via Settings Network Hugging Face
mirror. Probing hardcoded huggingface.co would fail them even when their
configured mirror works fine.
"""
@@ -286,7 +287,7 @@ def _network_check() -> dict:
"id": "network", "label": "Network (configured endpoint)",
"status": "warn",
"detail": "The configured Hugging Face endpoint could not be validated.",
"fix": "Review the endpoint in Model Catalogue → Models, then re-check.",
"fix": "Review the endpoint in Settings → Network, then re-check.",
"mirror_reachable": False,
}
net_ok = _probe_network(net_host, net_port)
@@ -498,10 +499,14 @@ def preflight():
_why = gpu_routing.get("routing_reason")
if _rs == "accelerated" and not _why:
r_status, r_detail, r_fix = "pass", f"{_eng}{_dev} (accelerated)", None
elif _rs == "accelerated": # driver/arch caveat
elif _rs == "accelerated" and KERNEL_RISK_MARKER in (_why or ""):
r_status, r_detail, r_fix = "warn", f"{_eng}{_dev}: {_why}", (
"GPU selected but may fail at kernel launch — update drivers / "
"reinstall torch for this GPU architecture.")
elif _rs == "accelerated": # low-VRAM caveat — not a driver/arch issue
r_status, r_detail, r_fix = "warn", f"{_eng}{_dev}: {_why}", (
"Unload other models before generating, keep the text short, "
"or pick a lighter engine.")
elif _rs == "cpu_fallback":
r_status, r_detail, r_fix = "warn", (
f"{_eng} runs on CPU here: {_why or 'no GPU path for this host'}"), (
@@ -512,10 +517,10 @@ def preflight():
elif _rs == "unavailable":
r_status, r_detail, r_fix = "fail", (
f"{_eng} can't run on this host: {_why or 'needs a GPU this machine lacks'}"), (
"Select an engine with a CPU path in Model Catalogue → Engines.")
"Select an engine with a CPU path in Model Catalogue.")
else: # "none" / unknown
r_status, r_detail, r_fix = "warn", "No active TTS engine resolved for routing.", (
"Pick an engine in Model Catalogue → Engines.")
"Pick an engine in Model Catalogue.")
checks.append({
"id": "gpu_routing", "label": "Active engine routing",
"status": r_status, "detail": r_detail, "fix": r_fix,
+362 -57
View File
@@ -15,6 +15,8 @@ from api.dependencies import is_loopback, require_admin, require_admin_action
from fastapi.responses import FileResponse, StreamingResponse
import torch
import shutil
import subprocess
import shlex
from core.config import OUTPUTS_DIR, DATA_DIR, CRASH_LOG_PATH, LOG_PATH, IDLE_TIMEOUT_SECONDS
from core.version import APP_VERSION
@@ -38,6 +40,10 @@ logger = logging.getLogger("omnivoice.api")
# Cache device checks at module load — they don't change at runtime
_is_mac = hasattr(torch.backends, "mps") and torch.backends.mps.is_available()
_is_cuda = torch.cuda.is_available()
try:
_is_xpu = hasattr(torch, "xpu") and torch.xpu.is_available()
except Exception:
_is_xpu = False
# Prime psutil's internal CPU counter so the first non-blocking call returns useful data
psutil.cpu_percent(interval=None)
@@ -51,6 +57,14 @@ def _detect_cpu_model() -> str:
for line in f:
if line.lower().startswith("model name"):
return line.split(":", 1)[1].strip()
if sys.platform == "win32":
import winreg
with winreg.OpenKey(
winreg.HKEY_LOCAL_MACHINE,
r"HARDWARE\DESCRIPTION\System\CentralProcessor\0",
) as key:
return str(winreg.QueryValueEx(key, "ProcessorNameString")[0]).strip()
if sys.platform == "darwin":
import subprocess
return subprocess.check_output(
@@ -61,6 +75,77 @@ def _detect_cpu_model() -> str:
return platform.processor() or ""
def _gpu_name_priority(name: str) -> tuple[int, int]:
lowered = name.lower()
if any(token in lowered for token in ("remote", "virtual", "basic display")):
return (-1, len(name))
if any(token in lowered for token in ("nvidia", "radeon", "amd", "intel arc")):
return (2, len(name))
return (1, len(name))
def _detect_os_gpu_name() -> str:
"""Best-effort display-adapter identity when the active torch build is CPU-only."""
try:
if sys.platform == "win32":
executable = shutil.which("powershell.exe") or shutil.which("powershell")
if not executable:
return ""
result = subprocess.run(
[
executable,
"-NoProfile",
"-NonInteractive",
"-Command",
"Get-CimInstance Win32_VideoController | Select-Object -ExpandProperty Name",
],
capture_output=True,
text=True,
timeout=3,
check=False,
creationflags=getattr(subprocess, "CREATE_NO_WINDOW", 0),
)
names = [line.strip() for line in result.stdout.splitlines() if line.strip()]
return max(names, key=_gpu_name_priority, default="")
if sys.platform.startswith("linux"):
executable = shutil.which("lspci")
if not executable:
return ""
result = subprocess.run(
[executable, "-mm"],
capture_output=True,
text=True,
timeout=2,
check=False,
)
names = []
for line in result.stdout.splitlines():
parts = shlex.split(line)
if len(parts) >= 4 and parts[1] in {"VGA compatible controller", "3D controller"}:
names.append(" ".join(parts[2:4]))
return max(names, key=_gpu_name_priority, default="")
if sys.platform == "darwin":
executable = shutil.which("system_profiler")
if not executable:
return ""
result = subprocess.run(
[executable, "SPDisplaysDataType"],
capture_output=True,
text=True,
timeout=3,
check=False,
)
names = [
line.split(":", 1)[1].strip()
for line in result.stdout.splitlines()
if "Chipset Model:" in line
]
return max(names, key=_gpu_name_priority, default="")
except (OSError, ValueError, subprocess.SubprocessError):
return ""
return ""
def _detect_gpu() -> tuple[str, float]:
"""(gpu_name, vram_total_gb) — static for the process lifetime.
@@ -71,11 +156,15 @@ def _detect_gpu() -> tuple[str, float]:
if _is_cuda:
props = torch.cuda.get_device_properties(0)
return torch.cuda.get_device_name(0), round(props.total_memory / (1024 ** 3), 1)
if _is_xpu:
props = torch.xpu.get_device_properties(0)
total_memory = float(getattr(props, "total_memory", 0.0))
return torch.xpu.get_device_name(0), round(total_memory / (1024 ** 3), 1)
if _is_mac:
return "Apple Silicon (MPS)", 0.0
except Exception:
pass
return "", 0.0
return _detect_os_gpu_name(), 0.0
# Static hardware facts, captured once — /system/info is hit on every
@@ -93,6 +182,37 @@ def _disk_free_gb() -> float:
return 0.0
def _nvidia_live_stats() -> tuple[float, float, float] | None:
"""Return GPU%, used VRAM GiB, total VRAM GiB without an optional Python dependency."""
executable = shutil.which("nvidia-smi")
if not executable:
return None
try:
creationflags = subprocess.CREATE_NO_WINDOW if sys.platform == "win32" else 0
result = subprocess.run(
[
executable,
"--query-gpu=utilization.gpu,memory.used,memory.total",
"--format=csv,noheader,nounits",
"--id=0",
],
capture_output=True,
text=True,
timeout=1.5,
check=False,
creationflags=creationflags,
)
if result.returncode != 0:
return None
values = [float(value.strip()) for value in result.stdout.splitlines()[0].split(",")]
if len(values) != 3:
return None
utilization, used_mib, total_mib = values
return utilization, used_mib / 1024, total_mib / 1024
except (OSError, ValueError, IndexError, subprocess.SubprocessError):
return None
def _ui_port() -> int:
"""The Vite UI dev-server port, single-sourced from OMNIVOICE_UI_PORT.
@@ -185,8 +305,8 @@ def loaded_models():
@router.post("/model/unload/{model_id}")
async def unload_model(model_id: str):
"""Unload a specific model by id (MM2-04). Delegates to model_lifecycle;
an unknown id maps to HTTP 400. ``tts`` | ``diarization`` |
``sidecar:<id>`` | ``sidecars``."""
an unknown id maps to HTTP 400. Supports every id returned by
``GET /model/loaded`` plus the aggregate ``sidecars`` id."""
from services import model_lifecycle
try:
return await model_lifecycle.unload(model_id)
@@ -204,7 +324,18 @@ def system_info():
try:
_ffmpeg = find_ffmpeg()
from services import model_manager as _mm
from services import asr_backend as _asr_backend
from core import prefs as _prefs_mod
_asr_engine = _asr_backend.active_backend_id()
_asr_model = (
_asr_backend._offline_asr_repo(_asr_engine)
or os.environ.get("ASR_MODEL")
or _asr_engine
)
_translation_provider = (
os.environ.get("TRANSLATE_PROVIDER")
or _prefs_mod.get("translation_backend", "argos")
)
return {
"app_version": APP_VERSION,
"generate_timeout_s": _mm.GPU_JOB_TIMEOUT_S,
@@ -224,8 +355,8 @@ def system_info():
"crash_log_path": CRASH_LOG_PATH,
"idle_timeout_seconds": IDLE_TIMEOUT_SECONDS,
"model_checkpoint": resolve_omnivoice_checkpoint(), # #693: show the effective checkpoint, not a leaked raw value
"asr_model": os.environ.get("ASR_MODEL", "Systran/faster-whisper-large-v3"),
"translate_provider": os.environ.get("TRANSLATE_PROVIDER", "google"),
"asr_model": _asr_model,
"translate_provider": _translation_provider,
"has_hf_token": _has_hf_token(),
"fast_download": _fast_download_status(),
"device": get_best_device(),
@@ -297,6 +428,142 @@ def _tail_file(path: str, tail: int):
return all_lines[-tail:], len(all_lines)
# Must track main.py's _WindowsSafeRotatingFileHandler(backupCount=3). The
# handler rolls omnivoice.log at 2 MB into .1/.2/.3, so up to 6 MB of history
# lives in files this module used to ignore entirely.
_LOG_BACKUP_COUNT = 3
def _rotated_log_paths(base: str) -> list[str]:
"""Existing `<base>.1 … .N`, newest first."""
return [p for p in (f"{base}.{i}" for i in range(1, _LOG_BACKUP_COUNT + 1)) if os.path.exists(p)]
def _tail_rolling(base: str, tail: int):
"""Tail `base`, reaching into its rotated siblings when it runs short.
A rollover leaves omnivoice.log nearly empty, and the Backend tab then
showed a handful of lines or none while the failure the user was asked
to copy sat in omnivoice.log.1. Reading the current file first keeps the
common case at one file read; the backups are only touched when they are
the only place the requested lines can come from.
Returns (lines oldest-first, total lines across the files read, paths read
oldest-first). The total counts only the files it had to open it stops as
soon as `tail` is satisfied, so it is "how much is behind these lines",
not the size of the whole rotation set.
"""
chunks: list[list[str]] = []
paths: list[str] = []
total = 0
remaining = tail
candidates = [p for p in [base, *_rotated_log_paths(base)] if os.path.exists(p)]
for path in candidates:
if remaining <= 0:
break
try:
lines, count = _tail_file(path, remaining)
except FileNotFoundError:
# A rollover can rename a candidate between the existence check
# above and this open, and the handler holds no lock we can take
# from a route. Skip the vanished file rather than 500 the whole
# panel over one member of the set — the previous single-file
# version failed the request outright in the same situation.
#
# A roll landing mid-walk can also shift which chunk a file holds,
# so a tail taken at that instant may repeat or miss a block. The
# panel re-polls every 5s and the next read is clean; buying strict
# consistency here would mean reaching into logging's internals.
continue
except PermissionError as exc:
# Windows only, and only the sharing violation: the handler still
# holds the file it is rolling. Any other permission failure is a
# real misconfiguration and must not be hidden.
if os.name == "nt" and getattr(exc, "winerror", None) == 32:
continue
raise
if count == 0:
continue
chunks.append(lines)
paths.append(path)
total += count
remaining -= len(lines)
# Files were visited newest-first; the reader wants oldest-first.
out: list[str] = []
for chunk in reversed(chunks):
out.extend(chunk)
return out, total, list(reversed(paths))
def _tauri_plugin_log_candidates():
"""The `tauri-plugin-log` files — the shell's own log, and the only thing
the Tauri tab actually displays.
Split out from :func:`_tauri_log_candidates` so Clear can touch these and
leave the backend stdout/stderr redirect alone. See
:func:`clear_tauri_logs`.
"""
home = os.path.expanduser("~")
bid = "com.debpalash.omnivoice-studio"
if sys.platform == "darwin":
return [
os.path.join(home, "Library/Logs", bid, "tauri.log"),
os.path.join(home, "Library/Logs", bid, "VoiceStudio.log"),
]
if sys.platform.startswith("linux"):
data_dir = os.environ.get("XDG_DATA_HOME") or os.path.join(home, ".local/share")
return [
os.path.join(data_dir, bid, "logs", "tauri.log"),
os.path.join(home, ".config", bid, "logs", "tauri.log"),
]
if sys.platform.startswith("win"):
appdata = os.environ.get("APPDATA", home)
localappdata = os.environ.get("LOCALAPPDATA") or os.path.join(home, "AppData", "Local")
return [
os.path.join(localappdata, bid, "logs", "tauri.log"),
os.path.join(appdata, bid, "logs", "tauri.log"),
]
return []
def _backend_redirect_log_candidates():
"""`backend.log` / `backend_err.log` — the spawned backend's stdout and
stderr, written by `src-tauri/src/backend.rs::backend_log_path()`.
Deliberately NOT cleared by the Tauri tab's Clear button.
`open_err_log_for_run()` opens `backend_err.log` **append-only** so "a
respawn must not destroy the previous run's evidence" (#1510), rotates it
to `.1` rather than truncating, and its spawn diagnostics are described
there as "retained in backend_err.log across runs and lands verbatim in bug
reports". A native death (a Windows access violation, a SIGSEGV) writes
nothing to the Python log by construction, so this file is the only record
of it.
`OMNIVOICE_LOG_DIR` is honoured first, in the same precedence
`backend_log_path()` uses. The backend is a child of the shell, so an
ambient override reaches both and a resolver that ignored it would look
in the per-OS default while the writer wrote somewhere else, which is the
divergence class this file already has one of (see #1782).
"""
override = (os.environ.get("OMNIVOICE_LOG_DIR") or "").strip()
if override:
return [
os.path.join(override, "backend.log"),
os.path.join(override, "backend_err.log"),
]
home = os.path.expanduser("~")
if sys.platform == "darwin":
base = os.path.join(home, "Library/Logs/OmniVoice")
elif sys.platform.startswith("linux"):
state_dir = os.environ.get("XDG_STATE_HOME") or os.path.join(home, ".local/state")
base = os.path.join(state_dir, "OmniVoice")
elif sys.platform.startswith("win"):
localappdata = os.environ.get("LOCALAPPDATA") or os.path.join(home, "AppData", "Local")
base = os.path.join(localappdata, "OmniVoice", "Logs")
else:
return []
return [os.path.join(base, "backend.log"), os.path.join(base, "backend_err.log")]
def _tauri_log_candidates():
"""Likely paths for Tauri-side logs, most useful first.
@@ -308,40 +575,15 @@ def _tauri_log_candidates():
`com.debpalash.omnivoice-studio` (frontend/src-tauri/tauri.conf.json).
- backend.rs::backend_log_path() redirects the spawned backend's
stdout/stderr to `backend.log` / `backend_err.log` under
`~/Library/Logs/OmniVoice` (macOS), `$XDG_STATE_HOME/VoiceStudio` falling
`~/Library/Logs/OmniVoice` (macOS), `$XDG_STATE_HOME/OmniVoice` falling
back to `~/.local/state/OmniVoice` (Linux), and
`%LOCALAPPDATA%\\OmniVoice\\Logs` (Windows). This is where uvicorn
startup banners and hard-crash tracebacks land keep all three OS
shapes listed or sidecar crashes become invisible off-macOS.
"""
home = os.path.expanduser("~")
bid = "com.debpalash.omnivoice-studio"
if sys.platform == "darwin":
return [
os.path.join(home, "Library/Logs", bid, "tauri.log"),
os.path.join(home, "Library/Logs", bid, "VoiceStudio.log"),
os.path.join(home, "Library/Logs/OmniVoice/backend.log"),
os.path.join(home, "Library/Logs/OmniVoice/backend_err.log"),
]
if sys.platform.startswith("linux"):
data_dir = os.environ.get("XDG_DATA_HOME") or os.path.join(home, ".local/share")
state_dir = os.environ.get("XDG_STATE_HOME") or os.path.join(home, ".local/state")
return [
os.path.join(data_dir, bid, "logs", "tauri.log"),
os.path.join(home, ".config", bid, "logs", "tauri.log"),
os.path.join(state_dir, "OmniVoice", "backend.log"),
os.path.join(state_dir, "OmniVoice", "backend_err.log"),
]
if sys.platform.startswith("win"):
appdata = os.environ.get("APPDATA", home)
localappdata = os.environ.get("LOCALAPPDATA") or os.path.join(home, "AppData", "Local")
return [
os.path.join(localappdata, bid, "logs", "tauri.log"),
os.path.join(appdata, bid, "logs", "tauri.log"),
os.path.join(localappdata, "OmniVoice", "Logs", "backend.log"),
os.path.join(localappdata, "OmniVoice", "Logs", "backend_err.log"),
]
return []
# Composed from the two halves so the read path keeps seeing every file
# while Clear can be narrowed to the shell's own log.
return _tauri_plugin_log_candidates() + _backend_redirect_log_candidates()
@router.get("/system/logs")
@@ -356,12 +598,24 @@ async def system_logs(tail: int = 200):
except Exception:
tail = 200
path = LOG_PATH if os.path.exists(LOG_PATH) else CRASH_LOG_PATH
if not os.path.exists(path):
if os.path.exists(LOG_PATH) or _rotated_log_paths(LOG_PATH):
base = LOG_PATH
else:
base = CRASH_LOG_PATH
if not os.path.exists(base) and not _rotated_log_paths(base):
return {"lines": [], "path": LOG_PATH, "exists": False}
path = base
try:
lines, total = await asyncio.to_thread(_tail_file, path, tail)
return {"lines": lines, "path": path, "exists": True, "total_lines": total}
lines, total, paths = await asyncio.to_thread(_tail_rolling, base, tail)
return {
"lines": lines,
"path": path,
"exists": True,
"total_lines": total,
# Which files the tail actually came from, oldest first. A bug
# report can then say whether it crossed a rollover.
"paths": paths,
}
except Exception as e:
raise HTTPException(
status_code=500,
@@ -467,9 +721,23 @@ def _read_from_pos(path: str, pos: int) -> list[str]:
@router.post("/system/logs/clear")
async def clear_system_logs():
"""Truncate the rolling runtime log and the crash log (what the Backend tab reads)."""
"""Truncate the rolling runtime log and the crash log (what the Backend tab reads).
Includes the rotated siblings. Truncating only omnivoice.log left up to
6 MB in .1/.2/.3, so Clear freed almost nothing and now that the tail
reaches into those files would have looked like it did nothing at all.
"""
cleared_any = False
for p in (LOG_PATH, CRASH_LOG_PATH):
# The full fixed name set rather than a snapshot of what exists: enumerating
# first leaves a window where a rollover creates a backup after the scan and
# its history survives a Clear that reported success. Names the handler can
# ever write are known up front, so there is nothing to enumerate.
targets = [
LOG_PATH,
*(f"{LOG_PATH}.{i}" for i in range(1, _LOG_BACKUP_COUNT + 1)),
CRASH_LOG_PATH,
]
for p in targets:
if os.path.exists(p):
try:
await asyncio.to_thread(_truncate_file, p)
@@ -502,10 +770,20 @@ def _truncate_file(path: str):
@router.post("/system/logs/tauri/clear")
async def clear_tauri_logs():
"""Truncate whichever Tauri-side log files we know about. OS-level rotation may recreate them."""
"""Truncate the shell's own log files. OS-level rotation may recreate them.
The backend stdout/stderr redirect is deliberately excluded. This button
lives on a tab that shows `tauri.log`, and truncating `backend_err.log`
from it destroyed evidence the user was never shown the one record of a
native death, which writes nothing to the Python log. `backend.rs`'s
`open_err_log_for_run()` opens that file append-only precisely so "a
respawn must not destroy the previous run's evidence" (#1510) and rotates
it to `.1` instead of truncating, so it manages its own size and does not
need clearing from here.
"""
cleared = []
failed = 0
for p in _tauri_log_candidates():
for p in _tauri_plugin_log_candidates():
if os.path.exists(p):
try:
await asyncio.to_thread(_truncate_file, p)
@@ -522,6 +800,7 @@ async def clear_tauri_logs():
@router.get("/sysinfo", response_model=SysinfoResponse)
def get_sys_info():
vram = 0.0
total_vram = 0.0
gpu_active = False
try:
@@ -534,18 +813,38 @@ def get_sys_info():
vram = alloc() / (1024**3)
elif _is_cuda:
vram = torch.cuda.memory_allocated() / (1024**3)
total_vram = torch.cuda.get_device_properties(torch.cuda.current_device()).total_memory / (1024**3)
elif _is_xpu:
vram = torch.xpu.memory_allocated() / (1024**3)
total_vram = float(
getattr(torch.xpu.get_device_properties(0), "total_memory", 0.0)
) / (1024**3)
except Exception:
pass
if vram > 0.01:
gpu_active = True
gpu_utilization = None
nvidia_stats = _nvidia_live_stats() if _is_cuda else None
if nvidia_stats:
gpu_utilization, vram, total_vram = nvidia_stats
gpu_active = gpu_active or gpu_utilization > 0 or vram > 0.01
vm = psutil.virtual_memory()
cpu_frequency = psutil.cpu_freq()
return {
"cpu": psutil.cpu_percent(interval=None),
"cpu_model": _CPU_MODEL,
"cpu_physical_cores": psutil.cpu_count(logical=False) or 0,
"cpu_logical_cores": psutil.cpu_count(logical=True) or 0,
"cpu_frequency_ghz": round((cpu_frequency.current if cpu_frequency else 0.0) / 1000, 2),
"ram": vm.used / (1024**3),
"total_ram": vm.total / (1024**3),
"gpu_name": _GPU_NAME,
"gpu_utilization": gpu_utilization,
"vram": round(vram, 2),
"total_vram": round(total_vram, 2),
"gpu_active": gpu_active
}
@@ -561,12 +860,18 @@ async def flush_memory(unload_model: bool = False):
freed_model = False
if unload_model:
import services.model_manager as mm
async with mm._model_lock:
# Also drops the clone-prompt side cache, which this path used to
# leave resident — an "unload" that kept the encoded reference
# tensors belonging to the model it just released (#1495).
freed_model = mm.unload_shared_model()
from services import model_lifecycle
# The user-facing action has always promised "Unload all". Route it
# through the lifecycle facade so alternate TTS engines, dictation,
# diarisation, translation and sidecars are released as well as the
# shared OmniVoice model. Individual runtimes still decline while
# leased by active work.
released = await model_lifecycle.unload_all()
freed_model = any(
bool(result.get("success"))
for result in released.get("results", {}).values()
)
# Multi-pass GC to break reference cycles
gc.collect(generation=2)
@@ -736,7 +1041,7 @@ def system_notifications():
from core import run_sentinel
rec = run_sentinel.newest_record()
if rec is not None and not rec[1]:
if rec is not None and not rec[1] and run_sentinel.warrants_user_notice(rec[0]):
record = rec[0]
last = record.get("last_activity") or {}
doing = f" Last activity: {last.get('kind')}." if last.get("kind") else ""
@@ -910,7 +1215,7 @@ async def set_env_var(body: dict):
Persistent keys (proxy, FFMPEG_PATH, translation provider keys, ) are
saved to ``prefs.json`` so they survive backend restarts (restored at
startup in ``main.py``). HF_TOKEN is persisted via
``huggingface_hub.login()`` (and cleared via ``logout()``). Other keys
``huggingface_hub.login()`` (and cleared with the shared token-file helper). Other keys
are set on ``os.environ`` for the running process.
The loopback-origin gate that previously lived inline here is now applied
@@ -985,14 +1290,14 @@ async def set_env_var(body: dict):
# Mirror the persistence on clear — wipe the saved token file too.
if key == "HF_TOKEN":
try:
from huggingface_hub import logout as _hf_logout
_hf_logout()
logger.info("HF token cleared from $HF_HOME/token via logout()")
except Exception as e:
logger.warning("Could not clear HF token file: %s", e)
from services.token_resolver import clear_hf_cli_tokens
clear_hf_cli_tokens()
logger.info("Local Hugging Face token files cleared")
except Exception:
raise HTTPException(status_code=500, detail="Could not clear local Hugging Face token files") from None
# HF_TOKEN persistence is handled above via huggingface_hub.login()/
# logout() — it never touches prefs.json. Everything else in
# clear_hf_cli_tokens() — it never touches prefs.json. Everything else in
# PERSISTENT_KEYS (proxy, FFMPEG_PATH, translation provider keys, …) is
# saved to prefs.json so it survives backend restarts (restored at
# startup in main.py). Non-persistent keys stay process-local.
@@ -1100,7 +1405,7 @@ def asr_backends():
def hf_token_state():
"""Return the 3-source HF token cascade state for the Settings UI
(Wave 2 React panel consumes this). Never returns the raw token
only a masked preview, whoami username, and per-source validity.
only a masked preview and local presence; no outbound validation.
"""
from dataclasses import asdict
from services import token_resolver
+7 -4
View File
@@ -174,11 +174,14 @@ async def ws_tts(websocket: WebSocket):
# close on `unavailable`, a one-time `routing` frame on
# cpu_fallback / accelerated-with-caveat (before any audio).
from core.device_caps import detect_host_caps
from services.engine_routing import resolve_routing, routing_notice
from services.engine_routing import (
routing_notice,
runtime_compute_profile_async,
)
from core.scrub import scrub_text
_routing = resolve_routing(
getattr(backend, "gpu_compat", ("cpu",)), detect_host_caps(),
getattr(backend, "min_vram_gb", 0.0))
_routing = await runtime_compute_profile_async(
backend, detect_host_caps()
)
if _routing["routing_status"] == "unavailable":
await websocket.send_json({
"type": "error",
+23 -5
View File
@@ -217,7 +217,7 @@ async def convert_speech(
# clone-less engine with the actionable switch-engine message (→ 400),
# and a backend mid-shutdown raises ModelLoadInterruptedByShutdown out
# of the model load → the global 503 [shutting_down] handler.
from services.tts_backend import resolve_generation_backend
from services.tts_backend import active_backend_id, resolve_generation_backend
try:
backend = await resolve_generation_backend(
require_cloning=True, cloning_purpose="voice conversion",
@@ -307,24 +307,42 @@ async def convert_speech(
GpuPoolBusyError,
run_on_gpu_pool_guarded,
)
from core.device_caps import detect_host_caps
from services.engine_routing import runtime_compute_profile_async
compute_profile = await runtime_compute_profile_async(
backend, detect_host_caps()
)
if compute_profile["routing_status"] == "unavailable":
raise HTTPException(
status_code=400,
detail=compute_profile["routing_reason"],
)
start_time = time.time()
from services.performance_profiles import tts_defaults
_profile_defaults = tts_defaults(active_backend_id())
_render = functools.partial(
_run_backend_inference,
backend, text, language, cond["ref_audio_path"], cond["ref_text"],
cond["instruct"],
None, # duration — the model picks; match_duration owns pacing
16, 2.0, # num_step / guidance_scale (the /generate defaults)
_profile_defaults.get("num_step", 16), 2.0,
1.0, # speed
True, True, # denoise / postprocess_output
True, _profile_defaults.get("postprocess_output", True),
used_seed,
)
try:
audio_tensor = await run_on_gpu_pool_guarded(
_render,
what="Voice convert",
timeout=_generate_timeout_s(text),
min_vram_gb=getattr(type(backend), "min_vram_gb", 0.0),
timeout=_generate_timeout_s(
text,
execution_device=compute_profile["effective_device"],
min_vram_gb=compute_profile["min_vram_gb"],
hardware_family=compute_profile.get("runtime_hardware_family"),
vram_gb=compute_profile.get("runtime_vram_gb"),
),
min_vram_gb=compute_profile["min_vram_gb"],
)
except GpuPoolBusyError as e:
raise HTTPException(
+12
View File
@@ -116,6 +116,18 @@ def get_target(op: str = "") -> dict:
return routing.status(op=op.strip() or None)
@router.get("/runtime")
async def get_runtime(engine: str = "", op: str = "tts") -> dict:
"""Runtime/model facts for the machine that will execute this operation."""
from services import gpu_gateway # noqa: PLC0415
return await gpu_gateway.status(
engine=engine.strip() or None,
op=op.strip() or "tts",
control_plane=service.control_plane,
)
@router.post("/target")
def set_target(request: TargetRequest) -> dict:
"""Choose where work runs. Exactly one target is active at a time."""
+8 -1
View File
@@ -15,9 +15,16 @@ class SysinfoResponse(BaseModel):
model_config = ConfigDict(extra="allow")
cpu: float = Field(description="CPU usage percentage (0100)")
cpu_model: str = ""
cpu_physical_cores: int = 0
cpu_logical_cores: int = 0
cpu_frequency_ghz: float = 0.0
ram: float = Field(description="Used RAM in GiB")
total_ram: float = Field(description="Total RAM in GiB")
gpu_name: str = ""
gpu_utilization: float | None = None
vram: float = Field(0.0, description="Used VRAM in GiB")
total_vram: float = Field(0.0, description="Total VRAM in GiB when reported by the runtime")
gpu_active: bool = Field(False, description="Whether a GPU is actively used")
@@ -82,7 +89,7 @@ class ModelStatusResponse(BaseModel):
status: str = Field(description="idle | loading | ready")
checkpoint: str | None = None
loaded_at: str | None = None
sub_stage: str | None = Field(None, description="Current loading sub-stage: importing | loading_weights | loading_asr | compiling | ready | error")
sub_stage: str | None = Field(None, description="Current TTS loading sub-stage: importing | loading_weights | compiling | ready | error")
detail: str | None = Field(None, description="Human-readable detail of current loading phase")
error: str | None = Field(None, description="Error message if loading failed")
+79 -2
View File
@@ -8,8 +8,9 @@
#
# Fields:
# repo_id (required) — HuggingFace repository ID
# engines (required) — backend ids that load this repo; [] for pipeline weights no single engine owns (they list under "Other weights")
# label (required) — Human-readable display name
# role (required) — TTS | ASR | Diarisation
# role (required) — TTS | ASR | Translation | Diarisation
# size_gb (required) — Approximate download size in GiB
# required (optional) — true if the app needs this model to function.
# Only the TTS model is required: the app boots and
@@ -28,6 +29,9 @@
# their own (weights live in referenced sub-repos). Such
# a cache is legitimately tiny, so the truncated-download
# (weights-missing) detector must NOT flag it incomplete.
# allow_patterns (optional) — restrict installation to these repository paths.
# Use for multi-package repos so an explicit install
# never downloads unrelated model variants.
# ─────────────────────────────────────────────────────────────────────────
models:
@@ -36,10 +40,25 @@ models:
- repo_id: "k2-fsa/OmniVoice"
label: "VoiceStudio TTS (k2-fsa/OmniVoice, 600+ languages, zero-shot)"
role: TTS
engines: [omnivoice, omnivoice-subprocess]
size_gb: 2.4
required: true
curated_on: [all]
- repo_id: "audio-cpp/audio.cpp-gguf"
label: "audio.cpp native bundle (Breeze-TTS-2 + Sortformer diarisation)"
role: TTS
engines: [audiocpp]
families: [tts, diarisation]
size_gb: 4.98
required_files:
- "Breeze-TTS-2-GGUF/breeze-tts-2-q8_0.gguf"
- "Sortformer-Diar-4spk-v1-GGUF/sortformer-diar-4spk-v1-q8_0.gguf"
allow_patterns:
- "Breeze-TTS-2-GGUF/breeze-tts-2-q8_0.gguf"
- "Sortformer-Diar-4spk-v1-GGUF/sortformer-diar-4spk-v1-q8_0.gguf"
note: "Optional audio.cpp bundle for native voice cloning and up-to-four-speaker diarisation. Research/non-commercial weights and self-hosted outputs; install only after reviewing the licenses."
# ── ASR (optional — curated per platform) ─────────────────────────────
# No ASR model is required to boot: TTS-only installs work. Dubbing,
# dictation, and clone-reference transcription prompt for the curated
@@ -48,6 +67,7 @@ models:
- repo_id: "Systran/faster-whisper-large-v3"
label: "Whisper large-v3 (faster-whisper — cross-platform, 99 langs)"
role: ASR
engines: [faster-whisper, faster-whisper-isolated, whisperx]
size_gb: 2.9
curated_on: [cuda, rocm, cpu, darwin-x86_64]
note: "The universal pick: best word-timestamp robustness for dubbing, runs on CUDA and CPU everywhere. On Apple Silicon prefer the MLX build."
@@ -55,6 +75,7 @@ models:
- repo_id: "mlx-community/whisper-large-v3-mlx"
label: "Whisper large-v3 (MLX — best for Apple Silicon)"
role: ASR
engines: [mlx-whisper]
size_gb: 3.0
platforms: [darwin-arm64]
curated_on: [darwin-arm64]
@@ -63,6 +84,7 @@ models:
- repo_id: "mlx-community/whisper-large-v3-turbo"
label: "Whisper large-v3 Turbo (MLX — fastest dictation)"
role: ASR
engines: [mlx-whisper]
size_gb: 1.6
platforms: [darwin-arm64]
curated_on: [darwin-arm64]
@@ -71,6 +93,7 @@ models:
- repo_id: "openai/whisper-large-v3"
label: "Whisper large-v3 (PyTorch — GPU path for AMD/ROCm)"
role: ASR
engines: [pytorch-whisper]
size_gb: 3.1
platforms: [cuda, rocm]
curated_on: [rocm]
@@ -79,12 +102,14 @@ models:
- repo_id: "mlx-community/whisper-tiny-mlx"
label: "Whisper tiny (MLX ASR — fast fallback)"
role: ASR
engines: [mlx-whisper]
size_gb: 0.08
platforms: [darwin-arm64]
- repo_id: "deepdml/faster-whisper-large-v3-turbo-ct2"
label: "Whisper large-v3 Turbo (5× faster, 0.8B)"
role: ASR
engines: [faster-whisper, faster-whisper-isolated, whisperx]
size_gb: 1.6
curated_on: [cuda, cpu]
note: "Best speed/quality tradeoff. 5× faster than large-v3 with minimal WER loss. Community CTranslate2 conversion (no official Systran/OpenAI turbo repo) — re-verify availability on catalog audits."
@@ -92,24 +117,28 @@ models:
- repo_id: "Systran/faster-distil-whisper-large-v3"
label: "Distil-Whisper large-v3 (distilled, fast)"
role: ASR
engines: [faster-whisper, faster-whisper-isolated, whisperx]
size_gb: 1.5
note: "Knowledge-distilled from large-v3. Good accuracy at higher speed."
- repo_id: "Systran/faster-whisper-medium"
label: "Whisper medium (balanced, lower VRAM)"
role: ASR
engines: [faster-whisper, faster-whisper-isolated, whisperx]
size_gb: 1.5
note: "Good balance of speed and accuracy. Half the VRAM of large-v3."
- repo_id: "Systran/faster-whisper-small"
label: "Whisper small (fast preview, low VRAM)"
role: ASR
engines: [faster-whisper, faster-whisper-isolated, whisperx]
size_gb: 0.5
note: "Quick previews and testing. ~2× faster than medium."
- repo_id: "Systran/faster-whisper-base"
label: "Whisper base (minimal, fastest Whisper)"
role: ASR
engines: [faster-whisper, faster-whisper-isolated, whisperx]
size_gb: 0.15
note: "Lowest accuracy but near-instant. Good for rapid iteration."
@@ -118,6 +147,7 @@ models:
- repo_id: "nvidia/parakeet-tdt-0.6b-v3"
label: "Parakeet TDT 0.6B v3 (NVIDIA — SOTA, 25+ langs)"
role: ASR
engines: [nemo-parakeet]
size_gb: 1.2
platforms: [cuda]
note: "Beats Whisper large-v3 on English benchmarks. Requires nemo_toolkit[asr]."
@@ -125,6 +155,7 @@ models:
- repo_id: "nvidia/parakeet-tdt-0.6b-v2"
label: "Parakeet TDT 0.6B v2 (NVIDIA — English + punctuation)"
role: ASR
engines: [nemo-parakeet]
size_gb: 1.2
platforms: [cuda]
note: "English-optimized with punctuation/capitalization. Requires nemo_toolkit[asr]."
@@ -132,6 +163,7 @@ models:
- repo_id: "mlx-community/parakeet-tdt-0.6b-v3"
label: "Parakeet TDT 0.6B v3 (MLX — Apple Silicon, 25 EU langs)"
role: ASR
engines: [parakeet-mlx]
size_gb: 1.2
platforms: [darwin-arm64]
curated_on: [darwin-arm64]
@@ -140,12 +172,14 @@ models:
- repo_id: "UsefulSensors/moonshine-base"
label: "Moonshine base (edge-optimized, 61M, ONNX)"
role: ASR
engines: [moonshine]
size_gb: 0.12
note: "Variable-length processing, sub-200ms latency. Great for CPU/edge. Requires moonshine-onnx."
- repo_id: "UsefulSensors/moonshine-tiny"
label: "Moonshine tiny (edge-optimized, 27M, ONNX)"
role: ASR
engines: [moonshine]
size_gb: 0.05
note: "Smallest/fastest Moonshine, sub-200ms latency. Lower accuracy than base. Requires moonshine-onnx."
@@ -159,6 +193,7 @@ models:
- repo_id: "csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8"
label: "Parakeet TDT v3 (sherpa-onnx — dictation, 25 EU langs)"
role: ASR
engines: [sherpa-onnx-asr]
size_gb: 0.67
engine: sherpa-onnx
dictation_id: sherpa-parakeet-tdt-v3
@@ -168,6 +203,7 @@ models:
- repo_id: "csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8"
label: "Parakeet TDT v2 (sherpa-onnx — dictation, English)"
role: ASR
engines: [sherpa-onnx-asr]
size_gb: 0.66
engine: sherpa-onnx
dictation_id: sherpa-parakeet-tdt-v2
@@ -177,6 +213,7 @@ models:
- repo_id: "csukuangfj/sherpa-onnx-streaming-zipformer-bilingual-zh-en-2023-02-20"
label: "Zipformer Bilingual (sherpa-onnx — streaming, zh+en)"
role: ASR
engines: [sherpa-onnx-asr]
size_gb: 0.2
engine: sherpa-onnx
dictation_id: sherpa-zipformer-bilingual-zh-en
@@ -186,6 +223,7 @@ models:
- repo_id: "csukuangfj/sherpa-onnx-streaming-paraformer-bilingual-zh-en"
label: "Paraformer Bilingual (sherpa-onnx — streaming, zh+en)"
role: ASR
engines: [sherpa-onnx-asr]
size_gb: 0.24
engine: sherpa-onnx
dictation_id: sherpa-paraformer-bilingual-zh-en
@@ -195,6 +233,7 @@ models:
- repo_id: "csukuangfj/sherpa-onnx-streaming-zipformer-en-20M-2023-02-17"
label: "Zipformer Streaming EN 20M (sherpa-onnx — streaming, English)"
role: ASR
engines: [sherpa-onnx-asr]
size_gb: 0.044
engine: sherpa-onnx
dictation_id: sherpa-zipformer-en-20m
@@ -204,6 +243,7 @@ models:
- repo_id: "csukuangfj/sherpa-onnx-streaming-zipformer-zh-14M-2023-02-23"
label: "Zipformer Streaming ZH 14M (sherpa-onnx — streaming, Chinese)"
role: ASR
engines: [sherpa-onnx-asr]
size_gb: 0.025
engine: sherpa-onnx
dictation_id: sherpa-zipformer-zh-14m
@@ -213,6 +253,7 @@ models:
- repo_id: "csukuangfj/sherpa-onnx-whisper-tiny"
label: "Whisper Tiny (sherpa-onnx — dictation, 90+ langs)"
role: ASR
engines: [sherpa-onnx-asr]
size_gb: 0.104
engine: sherpa-onnx
dictation_id: sherpa-whisper-tiny
@@ -220,43 +261,72 @@ models:
curated_on: [all]
note: "Recommended cross-platform dictation default (auto-detect). CPU, int8 ONNX. Requires sherpa-onnx."
# ── Translation ──────────────────────────────────────────────────────
- repo_id: "facebook/nllb-200-distilled-600M"
label: "NLLB-200 distilled 600M (local, 200 languages)"
role: Translation
engines: []
size_gb: 2.4
note: "Best fully-local translation quality. Install explicitly before selecting NLLB; translation never downloads these weights in the background."
# ── Diarisation ───────────────────────────────────────────────────────
- repo_id: "pyannote/speaker-diarization-3.1"
label: "pyannote speaker diarisation (multi-speaker videos)"
role: Diarisation
engines: []
size_gb: 0.8
config_only: true # pipeline repo; real weights live in referenced sub-repos
note: "Needs an HF_TOKEN with license accepted."
config_required_files: ["config.yaml"]
dependencies:
- repo_id: "pyannote/segmentation-3.0"
required_files: ["pytorch_model.bin"]
allow_patterns: ["config.yaml", "pytorch_model.bin"]
- repo_id: "pyannote/wespeaker-voxceleb-resnet34-LM"
required_files: ["pytorch_model.bin"]
allow_patterns: ["config.yaml", "pytorch_model.bin"]
gated: true
requires_hf_token: true
access_url: "https://huggingface.co/pyannote/speaker-diarization-3.1"
prerequisite_repo_id: "pyannote/segmentation-3.0"
prerequisite_access_url: "https://huggingface.co/pyannote/segmentation-3.0"
failure_topic: "PYANNOTE_LICENSE_REQUIRED"
note: "Requires access to both pyannote repositories and an HF token."
# ── Optional TTS ──────────────────────────────────────────────────────
- repo_id: "OpenMOSS-Team/MOSS-TTS-Nano-100M"
label: "MOSS-TTS-Nano 100M (20 langs, CPU-realtime)"
role: TTS
engines: [moss-tts-nano]
size_gb: 0.4
- repo_id: "KittenML/kitten-tts-mini-0.8"
label: "KittenTTS (English, 8 preset voices, CPU realtime)"
role: TTS
engines: [kittentts]
size_gb: 0.08
curated_on: [all]
- repo_id: "openbmb/VoxCPM2"
label: "VoxCPM2 (30 languages, voice cloning and design)"
role: TTS
engines: [voxcpm2]
size_gb: 5.0
curated_on: [cuda]
- repo_id: "FunAudioLLM/Fun-CosyVoice3-0.5B-2512"
label: "CosyVoice 3 0.5B (multilingual zero-shot)"
role: TTS
engines: [cosyvoice]
size_gb: 9.8
curated_on: [cuda]
- repo_id: "lj1995/GPT-SoVITS"
label: "GPT-SoVITS pretrained weights"
role: TTS
engines: [gpt-sovits]
size_gb: 2.0
curated_on: [cuda]
@@ -265,6 +335,7 @@ models:
- repo_id: "mlx-community/Kokoro-82M-bf16"
label: "Kokoro 82M (8 langs, small, mlx-audio default)"
role: TTS
engines: [mlx-audio]
size_gb: 0.15
curated_on: [darwin-arm64]
note: "Apple Silicon only — via mlx-audio backend."
@@ -273,6 +344,7 @@ models:
- repo_id: "mlx-community/csm-1b-8bit"
label: "CSM 1B (voice cloning, mlx-audio)"
role: TTS
engines: [mlx-audio]
size_gb: 1.1
note: "Apple Silicon only — via mlx-audio backend."
platforms: [darwin-arm64]
@@ -280,6 +352,7 @@ models:
- repo_id: "mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit"
label: "Qwen3-TTS 1.7B 4bit (voice design, mlx-audio)"
role: TTS
engines: [mlx-audio]
size_gb: 1.4
note: "Apple Silicon only — via mlx-audio backend."
platforms: [darwin-arm64]
@@ -287,6 +360,7 @@ models:
- repo_id: "mlx-community/Dia-1.6B"
label: "Dia 1.6B (expressive, mlx-audio)"
role: TTS
engines: [mlx-audio]
size_gb: 3.2
note: "Apple Silicon only — via mlx-audio backend."
platforms: [darwin-arm64]
@@ -294,6 +368,7 @@ models:
- repo_id: "mlx-community/Llama-OuteTTS-1.0-1B-4bit"
label: "Llama-OuteTTS 1.0 1B 4bit (voice clone, mlx-audio)"
role: TTS
engines: [mlx-audio]
size_gb: 0.8
note: "Apple Silicon only — via mlx-audio backend."
platforms: [darwin-arm64]
@@ -301,6 +376,7 @@ models:
- repo_id: "mlx-community/Chatterbox-TTS-4bit"
label: "Chatterbox TTS 4bit (mlx-audio)"
role: TTS
engines: [mlx-audio]
size_gb: 0.5
note: "Apple Silicon only — via mlx-audio backend."
platforms: [darwin-arm64]
@@ -308,6 +384,7 @@ models:
- repo_id: "mlx-community/MeloTTS-English-v3-MLX"
label: "MeloTTS English v3 (mlx-audio)"
role: TTS
engines: [mlx-audio]
size_gb: 0.2
note: "Apple Silicon only — via mlx-audio backend."
platforms: [darwin-arm64]
+20
View File
@@ -15,6 +15,18 @@ def get_app_data_dir():
return os.path.expanduser("~/.omnivoice")
def _configured_hf_token_path():
"""Match Hub's token location without importing or refreshing credentials."""
default_cache = os.path.join(os.path.expanduser("~"), ".cache")
hf_home = os.environ.get("HF_HOME", os.path.join(os.environ.get("XDG_CACHE_HOME", default_cache), "huggingface"))
return os.path.expandvars(os.path.expanduser(os.environ.get("HF_TOKEN_PATH", os.path.join(hf_home, "token"))))
# Snapshot recognized locations before automatic model-cache redirection.
# Explicit cache/token overrides restrict clearing to their selected location.
HF_CLI_TOKEN_PATHS = (_configured_hf_token_path(),)
def _ensure_short_hf_cache_on_windows():
"""Redirect HuggingFace cache to a short path on Windows.
@@ -38,6 +50,14 @@ def _ensure_short_hf_cache_on_windows():
return
short_cache = os.path.join(local_app, "OmniVoice", "hf_cache")
os.makedirs(short_cache, exist_ok=True)
if "HF_TOKEN_PATH" not in os.environ:
global HF_CLI_TOKEN_PATHS
canonical = HF_CLI_TOKEN_PATHS[0]
legacy = os.path.join(short_cache, "token")
HF_CLI_TOKEN_PATHS = tuple(dict.fromkeys((canonical, legacy)))
# Keep existing app-written logins usable without copying credentials.
selected = canonical if os.path.exists(canonical) or not os.path.exists(legacy) else legacy
os.environ.setdefault("HF_TOKEN_PATH", selected)
os.environ["HF_HOME"] = short_cache
os.environ["HF_HUB_CACHE"] = short_cache
+59
View File
@@ -0,0 +1,59 @@
"""Native-crash diagnostics for the backend process (#2135).
A crash inside torch/CUDA graph capture, a driver fault, an allocator abort
kills the interpreter below the level any ``except`` can reach. #2135's reporter
saw exactly that: the backend "simply exited" mid-``/generate`` with no Python
traceback, no HTTP response, and ``ConnectionRefused`` on the next ``/health``.
There was nothing in the logs to diagnose because nothing in Python ever ran
again.
``faulthandler`` installs handlers for the fatal signals (SIGSEGV, SIGABRT,
SIGBUS, SIGFPE, SIGILL) that print every thread's Python stack to stderr on the
way down. That is the difference between "the process vanished" and a named
frame pointing at the engine call that killed it.
This is strictly a diagnostic: it does not prevent the crash, and it must never
be the reason startup fails.
"""
from __future__ import annotations
import os
_DISABLE_ENV = "OMNIVOICE_DISABLE_FAULTHANDLER"
_TRUTHY = frozenset({"1", "true", "yes", "on"})
def _disabled() -> bool:
return os.environ.get(_DISABLE_ENV, "").strip().lower() in _TRUTHY
def enable_fault_handler(stderr=None) -> bool:
"""Arm fatal-signal tracebacks. Returns True when armed.
Call as early as possible before torch is imported so a crash during
model load is covered too. Honours ``OMNIVOICE_DISABLE_FAULTHANDLER=1`` for
hosts whose outer supervisor installs its own handlers.
Args:
stderr: optional file object to write dumps to. Defaults to the real
``sys.stderr`` ( ``backend_err.log``). faulthandler keeps the
underlying fd, so the object must stay open for the process
lifetime.
Never raises: a frozen build with a detached stderr, or a platform without
the signals, degrades to "no crash dump" rather than a failed boot.
"""
if _disabled():
return False
try:
import faulthandler
# all_threads=True: the fatal frame is routinely on a GPU-pool or
# compile worker, not whichever thread happens to take the signal.
if stderr is not None:
faulthandler.enable(file=stderr, all_threads=True)
else:
faulthandler.enable(all_threads=True)
return True
except Exception:
return False
+4 -2
View File
@@ -73,7 +73,7 @@ def _origin_tuple(value: str | None) -> tuple[str, str, int | None] | None:
):
return None
scheme = parsed.scheme.lower()
if scheme not in {"http", "https", "tauri"}:
if scheme not in {"http", "https", "tauri", "app"}:
return None
if port is None:
if scheme == "http":
@@ -83,6 +83,8 @@ def _origin_tuple(value: str | None) -> tuple[str, str, int | None] | None:
return scheme, parsed.hostname.lower(), port
DEFAULT_DESKTOP_ORIGINS = ("tauri://localhost", "http://tauri.localhost", "app://voicestudio")
def configured_allowed_origins() -> frozenset[tuple[str, str, int | None]]:
raw_port = os.environ.get("OMNIVOICE_UI_PORT", "3901")
try:
@@ -92,7 +94,7 @@ def configured_allowed_origins() -> frozenset[tuple[str, str, int | None]]:
values = os.environ.get(
"OMNIVOICE_ALLOWED_ORIGINS",
f"http://localhost:{ui_port},http://127.0.0.1:{ui_port},"
"tauri://localhost,http://tauri.localhost",
+ ",".join(DEFAULT_DESKTOP_ORIGINS),
).split(",")
return frozenset(
origin
+2 -1
View File
@@ -274,7 +274,8 @@ _BASE_SCHEMA = """
started_at REAL,
finished_at REAL,
lease_expires_at REAL,
grace_expires_at REAL
grace_expires_at REAL,
deadlines_json TEXT
);
CREATE INDEX IF NOT EXISTS idx_remote_attempts_task ON remote_task_attempts(task_id);
CREATE INDEX IF NOT EXISTS idx_remote_attempts_worker ON remote_task_attempts(worker_id, state);
+17
View File
@@ -51,6 +51,23 @@ _DIALECTS = set(_VD._INSTRUCT_CATEGORIES[5]) # the 12 Chinese dialect tokens
# the archetype ``attrs`` shape, so the response drops straight into vdStates.
CATEGORY_ORDER = ("Gender", "Age", "Pitch", "Style", "EnglishAccent", "ChineseDialect")
def instruct_to_vd_states(instruct: str | None) -> dict[str, str]:
"""Project a saved validator-token instruct onto the complete UI recipe."""
attrs = {category: "Auto" for category in CATEGORY_ORDER}
sanitized = _VD.sanitize_instruct(instruct)
if not sanitized:
return attrs
for token in sanitized.split(", "):
category_index = _VD._instruct_category_index(token)
if category_index < 0 or category_index >= len(CATEGORY_ORDER):
continue
# The first four frontend categories use the English canonical token;
# dialects and accents already use their engine-native form.
canonical = _VD._INSTRUCT_ZH_TO_EN.get(token, token)
attrs[CATEGORY_ORDER[category_index]] = canonical
return attrs
# ── Pinyin / romanized names → Chinese-dialect tokens (functional vocabulary) ─
DIALECT_PINYIN = {
"henan": "河南话",
+25 -4
View File
@@ -35,7 +35,8 @@ import sys
from dataclasses import dataclass
from typing import Literal
DeviceFamily = Literal["cuda", "rocm", "mps", "xpu", "cpu"]
DeviceFamily = Literal["cuda", "rocm", "mps", "xpu", "npu", "cpu"]
ACCELERATOR_PRIORITY = ("cuda", "rocm", "xpu", "npu", "mps")
# Stable substring stamped onto notes that represent a real kernel-launch risk
# (arch/driver mismatch) — as opposed to advisory notes (multi-GPU, VRAM query
@@ -533,9 +534,13 @@ def _probe() -> HostCaps:
# is the whole truth in that case (CodeRabbit, #1425).
notes.extend(why_no_gpu(torch))
# ── Intel XPU via IPEX ───────────────────────────────────────────────
# Older builds register XPU through IPEX; modern torch exposes it directly.
try:
import intel_extension_for_pytorch # noqa: F401
except Exception:
# Optional IPEX may be absent or incompatible; still probe native torch XPU.
pass
try:
if hasattr(torch, "xpu") and torch.xpu.is_available():
detected.append("xpu")
if not device_name:
@@ -546,7 +551,23 @@ def _probe() -> HostCaps:
pass
notes.append("XPU VRAM not queried (unreliable across IPEX versions)")
except Exception:
# IPEX absent or XPU probe failed — no XPU on this host.
# XPU probe failed — no usable XPU on this host.
pass
# Vendor extensions may register an NPU with torch. Probe only an already
# registered backend; never install or import an optional vendor package.
try:
if hasattr(torch, "npu") and torch.npu.is_available():
detected.append("npu")
if not device_name:
try:
device_name = torch.npu.get_device_name(0)
except Exception:
# An unavailable display name does not invalidate a usable NPU.
pass
notes.append("NPU VRAM not queried")
except Exception:
# Missing or broken vendor backends mean no usable NPU; continue probing.
pass
# ── Apple Silicon MPS ────────────────────────────────────────────────
@@ -579,7 +600,7 @@ def _probe() -> HostCaps:
# Preferred family by priority; cpu when nothing accelerated was detected.
family: DeviceFamily = "cpu"
for pref in ("cuda", "rocm", "xpu", "mps"):
for pref in ACCELERATOR_PRIORITY:
if pref in detected:
family = pref # type: ignore[assignment]
break
+14 -3
View File
@@ -28,6 +28,7 @@ import shutil
import sys
from core.config import DATA_DIR
from core.device_caps import KERNEL_RISK_MARKER
from core.scrub import scrub_text
from core.version import APP_VERSION
@@ -189,7 +190,7 @@ def _check_ram() -> dict:
def _check_engines() -> dict:
try:
from services.tts_backend import list_backends, active_backend_id
backends = list_backends()
backends = list_backends(include_hidden=True)
active = active_backend_id()
except Exception as e:
return _check("engines", "TTS engines", WARN, f"could not enumerate: {e}")
@@ -230,11 +231,16 @@ def _check_gpu_routing() -> dict:
host = v.get("host_family", "cpu")
if status == "accelerated":
if reason: # driver/arch caveat — accelerated but at risk
if reason and KERNEL_RISK_MARKER in reason: # driver/arch caveat — at risk
return _check("gpu_routing", "GPU routing", WARN,
f"{engine} -> {dev}: {reason}",
"The GPU is selected but may fail at kernel launch — "
"update drivers / reinstall torch for this GPU arch.")
if reason: # low-VRAM caveat — not a driver/arch issue
return _check("gpu_routing", "GPU routing", WARN,
f"{engine} -> {dev}: {reason}",
"Unload other models before generating, keep the text "
"short, or pick a lighter engine.")
return _check("gpu_routing", "GPU routing", OK, f"{engine} -> {dev} (accelerated)")
if status == "cpu_fallback":
return _check("gpu_routing", "GPU routing", WARN,
@@ -374,7 +380,12 @@ def run_diagnostics(include_network: bool = True, deep: bool = False) -> dict:
try:
module = importlib.import_module(f"services.{family}_backend")
active = module.active_backend_id()
row = next((item for item in module.list_backends() if item.get("id") == active), None)
rows = (
module.list_backends(include_hidden=True)
if family == "tts"
else module.list_backends()
)
row = next((item for item in rows if item.get("id") == active), None)
if row is not None:
engine_execution.append({
"family": family,
+3
View File
@@ -0,0 +1,3 @@
"""Stable engine IDs whose first use requires local license acceptance."""
LICENSE_GATED_ENGINES: frozenset[str] = frozenset({"supertonic3", "pockettts"})
+3 -1
View File
@@ -4,7 +4,7 @@ Used by the React ErrorBoundary's "Open docs for this error" button (via the
TypeScript mirror at `frontend/src/utils/errorDocsMap.ts`) and by the Phase 5
bug-reporter for "this error has a docs page" links.
The 5-class taxonomy below is the contract Phase 5 reporter consumes it,
The error taxonomy below is the contract Phase 5 reporter consumes it,
the TS map mirrors it, and `test_error_docs_map.test_keys_match_taxonomy`
locks the key set. To add a new class:
@@ -20,6 +20,8 @@ from core import links
_BASE = links.PROJECT_REPO_BLOB_MAIN
ERROR_DOCS: dict[str, str] = {
"DIARIZATION_LOAD_FAILED": f"{_BASE}/docs/features/diarization.md#troubleshooting",
"DIARIZATION_MODEL_MISSING": f"{_BASE}/docs/features/diarization.md#local-installation-and-repair",
"GATEKEEPER_QUARANTINE": f"{_BASE}/docs/install/macos.md#gatekeeper-quarantine",
"APPIMAGE_WEBKIT_WHITESCREEN": f"{_BASE}/docs/install/linux.md#appimage-white-screen-on-fedora-44--ubuntu-2404",
"PKG_RESOURCES_MISSING": f"{_BASE}/docs/install/troubleshooting.md#pkg_resources-missing",
+76 -3
View File
@@ -85,6 +85,8 @@ _HINTS: dict[str, str] = {
"GATEKEEPER_QUARANTINE": "Clear the macOS quarantine flag (xattr -cr the app), then reopen.",
"APPIMAGE_WEBKIT_WHITESCREEN": "Launch with WEBKIT_DISABLE_DMABUF_RENDERER=1 set.",
"HF_AUTH_FAILED": "Set a valid HF_TOKEN in Settings → Hugging Face and retry.",
"DIARIZATION_MODEL_MISSING": "Install or repair the selected diarisation model in Settings > Models > Diarisation, then retry transcription.",
"DIARIZATION_LOAD_FAILED": "Open Settings > Logs > Backend for the model load error, then retry transcription after correcting it.",
"PYANNOTE_LICENSE_REQUIRED": "Accept the pyannote model licenses on Hugging Face, then retry.",
"POCKETTTS_GATED_WEIGHTS": "PocketTTS weights are gated on HuggingFace. Accept the access agreement at huggingface.co/kyutai/pocket-tts, then set HF_TOKEN in Settings → Hugging Face and retry.",
"COMPUTE_TYPE_UNSUPPORTED": "Your GPU doesn't support float16 — VoiceStudio retried on int8. If transcription still fails, set OMNIVOICE/ASR_COMPUTE_TYPE=int8 or use CPU.",
@@ -94,9 +96,12 @@ _HINTS: dict[str, str] = {
# fail with "file not found" for exactly the users most likely to need it
# (greptile on #1377). tests/test_failure_classify.py pins these literals
# to the constraint file so they cannot drift when the pins bump.
"TRANSFORMERS_IMPORT": "Your transformers install is incomplete, or a package it loads models through (torchaudio, torchvision) is missing or mismatched with your torch — a torch/torchvision version mismatch fails with exactly this wording. Reinstall them together at the pinned versions (`uv pip install --python .venv --reinstall torch==2.8.0 torchaudio==2.8.0 torchvision==0.23.0 transformers` in the project folder), then restart the backend. If only transcription is affected, switching ASR to faster-whisper (Model Catalogue → Models) also works around it.",
"TRANSFORMERS_IMPORT": "Your transformers install is incomplete, or a package it loads models through (torchaudio, torchvision) is missing or mismatched with your torch — a torch/torchvision version mismatch fails with exactly this wording. Reinstall them together at the pinned versions (`uv pip install --python .venv --reinstall torch==2.8.0 torchaudio==2.8.0 torchvision==0.23.0 transformers` in the project folder), then restart the backend. If only transcription is affected, switching ASR to faster-whisper (Use on its row in Model Catalogue) also works around it.",
"WINDOWS_APP_CONTROL_BLOCKED": "Windows refused to load a file VoiceStudio needs — an Application Control policy (Smart App Control, WDAC, or AppLocker) blocked it. On a personal PC: Windows Security → App & browser control → Smart App Control → Off (Windows only lets you turn it off once — re-enabling requires a Windows reset), then restart VoiceStudio. On a managed/work PC, ask IT to allow the VoiceStudio install folder.",
"WINDOWS_PAGING_FILE_TOO_SMALL": "Windows ran out of virtual memory while mapping the model into memory — its paging file is smaller than the model needs. This is not the same as your RAM being full, and closing other apps usually won't fix it: Windows has to be allowed to back the mapping. Set a bigger paging file — Settings → System → About → Advanced system settings → Performance → Settings → Advanced → Virtual memory → Change: untick \"Automatically manage\", pick your system drive, choose \"Custom size\" and set both Initial and Maximum to at least 32768 MB (more than the model's size), then OK and restart Windows. A smaller/quantized engine (OmniVoice GGUF, Supertonic-3) also avoids the large mapping entirely.",
"WINDOWS_UNTRUSTED_MOUNT": "Windows refused to walk a folder on the way to this file because the path crosses a mount point it does not trust (WinError 448). That is a Windows rule about the VOLUME, not about VoiceStudio or the file itself — it turns up on Dev Drives, on mounted VHD/ReFS volumes, and on junctions pointing into another user profile, so retrying the same link cannot help. Point VoiceStudio at a folder on an ordinary local drive instead: Settings → Storage → data directory, or the download/output folder named in the message. If that folder has to stay where it is, trust the volume with `fsutil devdrv trust <drive>:` from an elevated prompt and restart.",
"INPUT_TOO_SHORT": "The input was too short for this engine to process — its first convolution needs more frames than the text (or the reference clip) produced. This is a hard limit of the model, not a transient failure, so retrying the same input will fail the same way. Give it a few more words, or a longer reference clip: a short phrase rather than one or two characters, and about a second of speech rather than a fragment.",
"CLONE_REFERENCE_MISSING": "This engine was asked to clone a voice but got no reference audio to clone FROM, and the model folder carries no built-in voice either. Pick a voice profile that has a saved reference clip, or record/upload a few seconds of clean speech as the reference, then generate again. A designed voice with no saved reference cannot be cloned from — synthesize with it directly instead.",
"MEDIA_TOOL_MISSING": "VoiceStudio's media engine (ffmpeg/ffprobe) wasn't on the system path when a component went looking for it. Open Settings → Audio tools and use Download/Repair to fetch the bundled copy, then retry — a restart picks it up for everything. If you'd rather use a system install, install ffmpeg (macOS: `brew install ffmpeg`; Windows: `winget install Gyan.FFmpeg`; Linux: your package manager) and restart VoiceStudio, or point FFMPEG_PATH / OMNIVOICE_FFPROBE_PATH at the binaries in Settings.",
"AUDIO_IO_FAILED": "An audio file couldn't be read or written at the OS level. Check the drive isn't full, that the output and temp folders exist and are writable, and that antivirus or OneDrive isn't locking them (add a VoiceStudio exclusion if you use one).",
"VIDEO_DOWNLOAD_OS_ERROR": "The OS refused a file operation while saving the downloaded video — this is a disk/folder problem, not a network one, so retrying the same link won't help. The download is written to a job folder under your VoiceStudio data directory (Settings → Storage shows the path): check that drive isn't full, that the folder exists and is writable, and that antivirus or a cloud-sync client (OneDrive, Dropbox) isn't locking it — add a VoiceStudio exclusion if you use one. If your data directory sits on a synced or network drive, move it to a local one.",
@@ -104,6 +109,9 @@ _HINTS: dict[str, str] = {
"SOCKS_PROXY_SUPPORT_MISSING": "A SOCKS proxy is configured in your environment (ALL_PROXY/HTTPS_PROXY=socks5://…) and the backend's HTTP client is missing SOCKS support. Newer VoiceStudio builds ship SOCKS support (the socksio package) — update the app. If you still see this, unset ALL_PROXY/HTTPS_PROXY for VoiceStudio, or run `uv pip install 'httpx[socks]'` in the backend venv, then restart.",
"SSL_HANDSHAKE_FAILURE": "A corporate or antivirus proxy is intercepting HTTPS traffic and re-signing certificates with its own CA — your OS trusts that CA, but Python's bundled certifi CA list doesn't, so the TLS handshake fails even though the connection reached the server. Newer VoiceStudio builds trust the OS certificate store at startup (the truststore package), which should already fix this — update the app and retry. If you still see this, add an HTTPS-scanning exclusion for VoiceStudio/Python in your antivirus, or ask IT for the proxy's CA bundle and set SSL_CERT_FILE to it, then restart.",
"UNSUPPORTED_VIDEO_URL": "This link isn't a directly downloadable video. Paste a direct video page (e.g. a youtube.com/watch?v=… or douyin.com/video/<id> link), not a share/profile/feed link — or download the file and drop it in directly.",
# #2034: yt-dlp's own advice names CLI flags (--cookies-from-browser)
# that a VoiceStudio user has no way to pass.
"VIDEO_DOWNLOAD_BOT_CHECK": "YouTube refused this download until it can confirm a signed-in person is asking (its “Sign in to confirm youre not a bot” check), so retrying the same link wont help. Sign in to YouTube in your browser, export its cookies as a Netscape cookies.txt file (a cookies.txt browser extension does this), and attach it in Dub with “Choose a cookies.txt export” before importing the link again. The desktop app accepts cookie exports; a browser connection needs HTTPS. Or download the video yourself and upload the file.",
"VIDEO_DRM_PROTECTED": "The video host only offered VoiceStudio a DRM-protected copy, which can't be downloaded. This is often not a property of the video itself — the host serves a different format set to different clients, and VoiceStudio already retried through every client it has. Try the link again in a minute, or download the video with a browser extension / the host's own download button and drop the file into Dubbing directly.",
# #1301: distinct from SSL_HANDSHAKE_FAILURE. The handshake did not fail on
# trust — the connection was CUT while TLS was in progress, so the certifi /
@@ -118,7 +126,7 @@ _HINTS: dict[str, str] = {
# told the reporter to reinstall transformers — advice that cannot work,
# because nothing is wrong with their install. Checked first so the cause
# wins over the symptom.
"MODEL_DOWNLOAD_INTERRUPTED": "A model download was cut off mid-request, and the component it was fetching then failed to load. Nothing is wrong with your install — reinstalling won't help, and the partial download is resumed rather than restarted. Just retry. If it keeps happening, check your connection (and any VPN, proxy or HF mirror setting); if only transcription is affected, switching ASR to faster-whisper in Model Catalogue → Models avoids the pipeline that downloads this component.",
"MODEL_DOWNLOAD_INTERRUPTED": "A model download was cut off mid-request, and the component it was fetching then failed to load. Nothing is wrong with your install — reinstalling won't help, and the partial download is resumed rather than restarted. Just retry. If it keeps happening, check your connection (and any VPN, proxy or HF mirror setting); if only transcription is affected, switching ASR to faster-whisper in Model Catalogue avoids the pipeline that downloads this component.",
"BROKEN_VENV": "The Python backend environment was moved or damaged. VoiceStudio rebuilds it automatically on the next launch; if it keeps failing, use Clean & Retry on the setup screen.",
"MODEL_CACHE_CORRUPT": "A model file is missing or damaged — a download that stopped part-way, a broken link to downloaded data, or a file changed on disk after it arrived (interrupted renames and antivirus interference both cause this). VoiceStudio repairs it automatically and retries the load once, re-downloading the damaged file where a resume would not have replaced it. If the error persists, quit VoiceStudio, delete the model's models--<org>--<name> folder inside the Hugging Face cache, and restart — the model re-downloads automatically.",
# HF_MIRROR_UNREACHABLE has a DYNAMIC hint (it names the configured mirror)
@@ -308,6 +316,23 @@ _CONTEXT_FREE_HINT_CLASSES = frozenset({
# a Windows virtual-memory setting rather than a connectivity problem, and
# the detailed hint we already had for it never reached them.
"WINDOWS_PAGING_FILE_TOO_SMALL",
# #1957: triggered by WinError 448 or the literal "untrusted mount
# point" — both unmistakable, and it reaches the user as a bare
# download failure with only the OS sentence attached.
"WINDOWS_UNTRUSTED_MOUNT",
# #1826: torch's own conv wording, which nothing else produces, and it
# reaches the user through the generic 500.
"INPUT_TOO_SHORT",
# #1879: matched on wording no other failure produces, and it reaches the
# user as a bare 400 carrying only the library sentence.
"CLONE_REFERENCE_MISSING",
# Its trigger is a VoiceStudio-authored sentence — "the TTS model cache
# for … is incomplete" plus "could not be auto-repaired" / "weights
# missing" — so it cannot be produced by an unrelated library. The 500
# handler is the surface a corrupt cache actually reaches, and dropping
# its hint there would leave the user with no way to know a redownload
# is the fix.
"MODEL_CACHE_CORRUPT",
})
@@ -381,7 +406,22 @@ def classify(reason: str) -> str:
or "access conditions" in low
) and ("pocket" in low or "kyutai" in low):
return "POCKETTTS_GATED_WEIGHTS"
if "pyannote" in low or ("gated" in low and "model" in low) or "accept the" in low:
diarisation = any(marker in low for marker in (
"pyannote", "diarization", "diarisation", "sortformer",
))
access_failure = any(marker in low for marker in (
"gated", "unauthorized", "forbidden", "401", "403",
"accept the", "license", "user conditions",
))
if diarisation and not access_failure:
if any(marker in low for marker in (
"files are missing", "files are missing or incomplete",
"filenotfounderror", "localentrynotfounderror", "model is missing",
)):
return "DIARIZATION_MODEL_MISSING"
if any(marker in low for marker in ("failed to load", "load failed", "runtime failed")):
return "DIARIZATION_LOAD_FAILED"
if (diarisation and access_failure) or ("gated" in low and "model" in low) or "accept the" in low:
return "PYANNOTE_LICENSE_REQUIRED"
# ASR robustness (#551 / #549): name the class so the no-segments toast is
# actionable. Place before the generic returns so a compute-type/transformers
@@ -549,6 +589,11 @@ def classify(reason: str) -> str:
# download path now escalates the client the way it does for a 403. If
# every client still says DRM, the video genuinely can't be fetched and the
# user needs to hear that rather than retry a fourth time.
# #2034: YouTube's anti-automation wall. Not transient and not a
# player-client format set, so it must reach neither the network retry
# nor the 403 client escalation; the remedy is signed-in cookies.
if "not a bot" in low and ("sign in" in low or "cookies" in low):
return "VIDEO_DOWNLOAD_BOT_CHECK"
if "drm protected" in low or "drm-protected" in low:
return "VIDEO_DRM_PROTECTED"
if (
@@ -569,6 +614,34 @@ def classify(reason: str) -> str:
or "application control policy" in low
):
return "WINDOWS_APP_CONTROL_BLOCKED"
# #1957: the path to a download or output file crosses a mount point
# Windows will not traverse (Dev Drive, mounted VHD/ReFS, a junction into
# another profile). Matched on the numeric code first because the OS
# translates the sentence, with the English phrase as a fallback.
if "[winerror 448]" in low or "untrusted mount point" in low:
return "WINDOWS_UNTRUSTED_MOUNT"
# #1826: a degenerate-length input reaches a conv layer whose kernel is
# wider than the tensor, and torch says so in its own terms — "Calculated
# padded input size per channel: (1). Kernel size: (2). Kernel size can't
# be greater than actual input size". That arrived doubly wrapped in
# "Underlying error:" and told the user nothing they could act on, when
# the fix is simply "type more than one character".
if "kernel size can't be greater than actual input size" in low or (
"calculated padded input size per channel" in low
):
return "INPUT_TOO_SHORT"
# #1879: mlx-audio (and the Chatterbox-family models under it) raise a
# bare ValueError naming their own parameters — "No conditionals
# available. Either provide audio_prompt/audio_prompt_sr ... or ensure
# conds.safetensors is in the model directory." The generate route passed
# that straight through as the 400 detail, so the user was told to supply
# an argument they have no way to name and to check for a file they have
# never heard of. What actually happened is "you asked to clone without a
# reference clip".
if "no conditionals available" in low or (
"audio_prompt" in low and "conds.safetensors" in low
):
return "CLONE_REFERENCE_MISSING"
# #1221: libsndfile failed an OS-level audio read/write. Its own wording is
# a bare "System error.", so match the library name — audio_io already
# prefixes the target path and free space onto the write-path failures.
+56
View File
@@ -57,6 +57,48 @@ class HFTokenRedactor(logging.Filter):
return True
class RoutineHealthAccessFilter(logging.Filter):
"""Drop only successful routine liveness access lines.
The desktop supervisor probes every two seconds. Startup/not-ready responses
and every other request remain visible, while the steady-state 200 line no
longer consumes the small rotating diagnostic log.
"""
def filter(self, record: logging.LogRecord) -> bool:
try:
args = record.args
if not isinstance(args, tuple) or len(args) < 5:
return True
_client, method, path, _http_version, status = args[:5]
return not (
method == "GET"
and str(path).partition("?")[0] == "/health"
and int(status) == 200
)
except (TypeError, ValueError):
return True
class RoutineAsyncioTransportFilter(logging.Filter):
"""Drop only expected socket-close noise from asyncio's transport layer."""
def filter(self, record: logging.LogRecord) -> bool:
try:
message = record.getMessage()
if record.levelno == logging.WARNING and "socket.send() raised exception" in message:
return False
exception = record.exc_info[1] if record.exc_info else None
return not (
message.startswith(
"Exception in callback _ProactorBasePipeTransport._call_connection_lost"
)
and isinstance(exception, (BrokenPipeError, ConnectionResetError))
)
except Exception:
return True
def install_redaction_filter(root_logger: logging.Logger | None = None) -> None:
"""Attach a single HFTokenRedactor to the root logger and to every
existing handler. Idempotent repeated calls do not stack up duplicate
@@ -69,3 +111,17 @@ def install_redaction_filter(root_logger: logging.Logger | None = None) -> None:
for handler in list(target.handlers):
if not any(isinstance(f, HFTokenRedactor) for f in handler.filters):
handler.addFilter(HFTokenRedactor())
def install_access_log_filter(logger: logging.Logger | None = None) -> None:
"""Install the routine-health filter on Uvicorn's access logger once."""
target = logger or logging.getLogger("uvicorn.access")
if not any(isinstance(item, RoutineHealthAccessFilter) for item in target.filters):
target.addFilter(RoutineHealthAccessFilter())
def install_asyncio_transport_filter(logger: logging.Logger | None = None) -> None:
"""Install the expected transport-close filter on asyncio once."""
target = logger or logging.getLogger("asyncio")
if not any(isinstance(item, RoutineAsyncioTransportFilter) for item in target.filters):
target.addFilter(RoutineAsyncioTransportFilter())
+94 -3
View File
@@ -4,8 +4,32 @@ from __future__ import annotations
import os
import sys
import threading
from typing import BinaryIO, Callable
import time
from typing import Any, BinaryIO, Callable, Optional
# Poll cadence for the Windows pipe watcher. Exit latency after the desktop
# closes its end is bounded by this; the desktop's own kill-on-close Job is the
# hard backstop, so a quarter second is plenty and costs nothing measurable.
WINDOWS_PIPE_POLL_INTERVAL_S = 0.25
_FILE_TYPE_PIPE = 3 # winbase.h FILE_TYPE_PIPE
def _exit_after_parent_loss(code: int) -> None:
"""Retire backend-only crash forensics before the desktop-owned exit.
Losing the containment pipe means the desktop process ended, including an
Electron development reload. That is not a backend crash: the shell owns
this child and the watchdog is deliberately terminating it. ``os._exit``
skips FastAPI lifespan cleanup, so clear the run sentinel here first. A
real backend abort/OOM never reaches this callback and remains detectable
on the next start.
"""
try:
from core import run_sentinel
run_sentinel.clear_sentinel()
except Exception:
pass
os._exit(code)
def _watch_parent_pipe(reader: BinaryIO, exit_process: Callable[[int], None]) -> None:
"""Block until the desktop-owned stdin pipe closes, then exit immediately."""
@@ -18,6 +42,64 @@ def _watch_parent_pipe(reader: BinaryIO, exit_process: Callable[[int], None]) ->
exit_process(0)
def _watch_parent_pipe_handle(
handle: int,
exit_process: Callable[[int], None],
*,
peek: Optional[Callable[[int], Any]] = None,
read_file: Optional[Callable[[int, int], Any]] = None,
sleep: Callable[[float], None] = time.sleep,
interval: float = WINDOWS_PIPE_POLL_INTERVAL_S,
) -> None:
"""Windows twin of :func:`_watch_parent_pipe` that never leaves a read
pending on the pipe.
A synchronous ``ReadFile`` parked on the stdin pipe whether issued through
the C runtime's ``read()`` or straight to the kernel — deadlocks the
OpenBLAS DLL initializer that ``import torch`` reaches (numpy's
``_multiarray_umath``) in the startup worker: every desktop-spawned backend
on Windows froze at "Loading ML runtime (PyTorch)" while the identical
command from a terminal, with no stdin pipe and no watchdog, started in
seconds. A thread that merely sleeps does not trigger it; only the pending
read on that pipe does. So instead of blocking in a read, poll with
``PeekNamedPipe``: it returns immediately, holds no I/O on the file object,
drains any keepalive bytes the desktop might write, and fails with
``ERROR_BROKEN_PIPE`` the moment the desktop closes its end which is the
same EOF signal the POSIX reader gets.
"""
if peek is None or read_file is None:
import _winapi # Windows-only stdlib module; the caller gates on the platform
peek = peek or _winapi.PeekNamedPipe
read_file = read_file or _winapi.ReadFile
try:
while True:
available, _ = peek(handle)
if available:
# Bytes are already buffered, so this read cannot block.
read_file(handle, available)
else:
sleep(interval)
except OSError:
# ERROR_BROKEN_PIPE (109) is how the closed parent end surfaces here.
pass
exit_process(0)
def _windows_pipe_handle(reader: Any) -> Optional[int]:
"""The OS handle behind ``reader`` when it is a pipe, else None."""
try:
import msvcrt
import _winapi
handle = msvcrt.get_osfhandle(reader.fileno())
if _winapi.GetFileType(handle) != _FILE_TYPE_PIPE:
return None
return handle
except (OSError, ValueError, AttributeError, ImportError):
return None
def arm_desktop_parent_watchdog() -> bool:
"""Use stdin EOF as an unforgeable parent-liveness signal for desktop runs."""
if os.environ.get("OMNIVOICE_DESKTOP_CONTAINED") != "1":
@@ -25,9 +107,18 @@ def arm_desktop_parent_watchdog() -> bool:
reader = getattr(sys.stdin, "buffer", None)
if reader is None:
return False
target: Callable[..., None] = _watch_parent_pipe
args: tuple = (reader, _exit_after_parent_loss)
if os.name == "nt":
handle = _windows_pipe_handle(reader)
if handle is not None:
target = _watch_parent_pipe_handle
args = (handle, _exit_after_parent_loss)
# A non-pipe stdin (file, NUL) cannot have a read pending against a
# pipe file object, so the blocking reader stays correct there.
threading.Thread(
target=_watch_parent_pipe,
args=(reader, os._exit),
target=target,
args=args,
name="desktop-parent-watchdog",
daemon=True,
).start()
+11
View File
@@ -75,6 +75,17 @@ def set_(key: str, value: Any) -> None:
_save(data)
def update_mapping(key: str, changes: dict, *, replace: bool = False) -> None:
"""Atomically update one preference object without losing concurrent edits."""
with _MUTATE_LOCK:
data = _load()
current = data.get(key)
value = dict(current) if isinstance(current, dict) and not replace else {}
value.update(changes)
data[key] = value
_save(data)
def delete(key: str) -> None:
"""Remove *key* from prefs.json if present."""
with _MUTATE_LOCK:
+27
View File
@@ -0,0 +1,27 @@
"""Small, metadata-free local profile portraits."""
import io
import warnings
from fastapi import HTTPException
from PIL import Image, ImageOps, UnidentifiedImageError
MAX_IMAGE_BYTES = 5 * 1024 * 1024
def normalize_portrait(data: bytes) -> bytes:
if len(data) > MAX_IMAGE_BYTES:
raise HTTPException(413, "Profile image exceeds 5 MB")
try:
with warnings.catch_warnings():
warnings.simplefilter("error", Image.DecompressionBombWarning)
with Image.open(io.BytesIO(data)) as source:
if source.format not in {"JPEG", "PNG", "WEBP"}:
raise ValueError("unsupported image format")
if source.width * source.height > 16_000_000:
raise ValueError("image dimensions too large")
portrait = ImageOps.fit(ImageOps.exif_transpose(source).convert("RGB"), (256, 256))
output = io.BytesIO()
portrait.save(output, format="JPEG", quality=88)
return output.getvalue()
except (UnidentifiedImageError, OSError, ValueError, Image.DecompressionBombWarning, Image.DecompressionBombError) as exc:
raise HTTPException(422, "Use a valid JPEG, PNG or WebP image up to 16 megapixels") from exc
+31 -2
View File
@@ -94,6 +94,16 @@ def stream_generation_failure(error: BaseException | object) -> dict[str, object
replace the failure being diagnosed.
"""
payload = stream_failure("generation_failed")
if isinstance(error, BaseException):
# The exception's TYPE NAME, never its message. Two failures that both
# render the floor message "Generation failed. Check the selected
# engine and try again." are indistinguishable in an auto-filed report,
# so every unclassified streaming failure arrives as the same issue and
# none of them can be triaged (#1800). A class name is VoiceStudio-safe
# by the same reasoning that already puts it on the wire as
# `error_class` in the dub routes and on the analytics allowlist: it is
# a Python type, not user text, and no substring of `error` is copied.
payload["error_class"] = type(error).__name__
try:
enriched = public_exception_response(error, fallback=str(payload["detail"]))
except Exception:
@@ -149,12 +159,31 @@ def public_exception_response(error: BaseException, *, fallback: str) -> dict[st
Classification may inspect the private diagnostic locally, but response
values come exclusively from VoiceStudio-owned constants. No substring of
``error`` is copied into the payload.
Every caller is a CONTEXT-FREE surface the global 500 handler, the
streaming generate error frame, the dub GPU-OOM 503 so the topic is
filtered through ``failure._CONTEXT_FREE_HINT_CLASSES`` before its hint is
attached. Without that filter a topic whose trigger is a generic phrase
stamps a confidently wrong remediation on an unrelated failure: #1943 is a
macOS mlx-audio TTS 500 that came back advising the user that "the
connection to the video server dropped mid-download", because
VIDEO_DOWNLOAD_NETWORK triggers on a bare "timed out" / "connection
reset". The allowlist already existed and already named that class as the
example of what must not appear here; only :func:`failure.append_hint`
honoured it, and this helper replaced ``append_hint`` on the 500 path
without carrying the rule across.
HF_MIRROR_UNREACHABLE is allowed alongside it: its hint is dynamic (it
names the configured mirror) and its trigger requires that a mirror is
configured at all, so it cannot fire on an unrelated failure (#874).
"""
from core.failure import classify, public_hint_for_topic
from core.failure import _CONTEXT_FREE_HINT_CLASSES, classify, public_hint_for_topic
try:
topic = classify(str(error))
hint = public_hint_for_topic(topic)
if topic and topic not in _CONTEXT_FREE_HINT_CLASSES and topic != "HF_MIRROR_UNREACHABLE":
topic = ""
hint = public_hint_for_topic(topic) if topic else ""
except Exception:
topic = ""
hint = ""
+32
View File
@@ -70,6 +70,20 @@ LOG_TAIL_LINES = 40
#: burst instead of one per request.
ACTIVITY_THROTTLE_S = 2.0
# An idle desktop process can disappear with its owning shell during an OS
# shutdown, package replacement, or a forced development relaunch. Keep that
# forensic record, but do not nag the user unless there is evidence that work
# was interrupted or the backend itself logged a fatal failure.
_ACTIONABLE_LOG_MARKERS = (
"traceback (most recent call last)",
"critical",
"fatal error",
"out of memory",
"memoryerror",
"segmentation fault",
"access violation",
)
# In-memory run state. `owns` guards clear_sentinel()/touch_activity() so an
# instance that skipped writing (another live instance holds the sentinel)
# can never clobber or delete the other instance's sentinel.
@@ -298,6 +312,24 @@ def _build_crash_record(sentinel: dict, now: float) -> dict:
}
def warrants_user_notice(record: dict) -> bool:
"""Whether an unclean record is actionable enough to interrupt the user.
The record remains available to diagnostics either way. A meaningful
activity marker means a generation/transcription/task may have been lost;
a strict fatal-log marker catches startup/native crashes that happened
before an activity could be recorded. Idle shell-owned exits stay quiet.
"""
activity = record.get("last_activity")
if isinstance(activity, dict) and str(activity.get("kind") or "").strip():
return True
tail = record.get("log_tail")
if not isinstance(tail, list):
return False
joined = "\n".join(str(line).lower() for line in tail[-LOG_TAIL_LINES:])
return any(marker in joined for marker in _ACTIONABLE_LOG_MARKERS)
def _load_store() -> dict:
store = _read_json(CRASH_RECORD_PATH) or {}
records = store.get("records")
+32 -1
View File
@@ -10,6 +10,27 @@ from core import run_sentinel
logger = logging.getLogger("omnivoice.tasks")
def _stream_failure(update):
"""Recognize terminal SSE failures, including generators that do not raise."""
if isinstance(update, bytes):
update = update.decode("utf-8", errors="replace")
if not isinstance(update, str):
return None
lines = update.splitlines()
try:
payload = json.loads("\n".join(line[5:].strip() for line in lines if line.startswith("data:")))
except (ValueError, TypeError):
return None
if not isinstance(payload, dict):
return None
if payload.get("type") != "error" and not any(line.strip() == "event: error" for line in lines):
return None
detail = payload.get("reason") or payload.get("error") or payload.get("detail")
if isinstance(detail, dict):
detail = detail.get("message") or detail.get("reason")
return detail if isinstance(detail, str) and detail else "Task failed"
class TaskManager:
"""In-memory task dispatcher with SQLite-backed metadata.
@@ -125,9 +146,19 @@ class TaskManager:
except Exception: logger.exception("job_store.mark_cancelled failed")
break
await self._push_event(task_id, update)
stream_error = _stream_failure(update)
if stream_error is not None:
t["status"] = "failed"
t["error"] = stream_error
try:
job_store.mark_failed(task_id, stream_error)
except Exception:
logger.exception("job_store.mark_failed failed")
await res.aclose()
break
elif inspect.iscoroutine(res):
await res
if t["status"] != "cancelled":
if t["status"] not in {"cancelled", "failed"}:
t["status"] = "done"
try: job_store.mark_done(task_id)
except Exception: logger.exception("job_store.mark_done failed")
+43
View File
@@ -0,0 +1,43 @@
"""The PyTorch wheel index VoiceStudio installs CUDA builds from.
A local-version pin such as ``torch==2.9.1+cu128`` exists only on PyTorch's
own index, never on PyPI. The app's own ``pyproject.toml`` routes torch there
through ``[tool.uv.sources]``, but a sidecar engine is installed with
``uv pip install`` into its own venv, which knows nothing about that config
so every CUDA-pinned sidecar install has to name the index itself.
MOSS-TTS-v1.5's install did not, and its ``[torch-runtime]`` extra
(``torch==2.9.1+cu128``) could never resolve: ``uv pip compile`` reports it
unsatisfiable without this index and resolves it with it. One definition here,
imported by the one-click installer and by the engine's own bootstrap, so the
two cannot drift apart again. ``tests/test_sidecar_install.py`` pins the URL
to the ``pytorch-cuda`` index declared in the app's ``pyproject.toml``.
"""
PYTORCH_CU128_INDEX_URL = "https://download.pytorch.org/whl/cu128"
# `unsafe-best-match`: the PyTorch index also mirrors common dependencies
# (numpy, pillow, sympy, …) at a narrower range of versions than PyPI. uv's
# default first-index strategy would stop at whichever index lists a name first
# and could pin an old mirror copy or fail outright. The index is PyTorch's
# official one, so the dependency-confusion risk the name warns about does not
# apply to it.
UV_PIP_CU128_ARGS: tuple[str, ...] = (
"--extra-index-url",
PYTORCH_CU128_INDEX_URL,
"--index-strategy",
"unsafe-best-match",
)
PYTORCH_CPU_INDEX_URL = "https://download.pytorch.org/whl/cpu"
# For an engine that runs torch only on the CPU (PocketTTS). On Linux, PyPI's
# torch is the CUDA build and pulls ~15 NVIDIA packages the engine never uses;
# this index serves `+cpu` builds for Linux and Windows and the regular build
# for macOS.
UV_PIP_CPU_ARGS: tuple[str, ...] = (
"--extra-index-url",
PYTORCH_CPU_INDEX_URL,
"--index-strategy",
"unsafe-best-match",
)
+3 -1
View File
@@ -144,6 +144,8 @@ def _drop_invalid_path_keys() -> None:
logger.warning(
"%s from the saved env file points at an unusable path (%s) — "
"ignoring it for this run and falling back to the default "
"location. Fix or clear it in Model Catalogue → Models.", key, val,
"location. For the models folder, choose it again in Settings → "
"Storage; otherwise fix or remove the entry in the saved env file.",
key, val,
)
os.environ.pop(key, None)
+1 -1
View File
@@ -24,7 +24,7 @@ from pathlib import Path
# tests/test_app_version.py::test_all_version_files_in_lockstep and bumped by
# release.yml's version-bump job, so it stays equal to
# pyproject/tauri.conf/Cargo/package.json.
_FALLBACK_VERSION = "0.5.2"
_FALLBACK_VERSION = "0.5.3"
def _fallback_version() -> str:
+10 -3
View File
@@ -124,9 +124,16 @@ def _get_model():
return _model
def _transcribe(audio_path, word_timestamps):
def _transcribe(audio_path, word_timestamps, decode_options=None):
options = decode_options or {}
if not isinstance(options, dict) or any(
key not in {"beam_size", "best_of"}
or type(value) is not int or not 1 <= value <= 8
for key, value in options.items()
):
raise ValueError("Invalid ASR decoding options")
model = _get_model()
segments, info = model.transcribe(audio_path, word_timestamps=word_timestamps)
segments, info = model.transcribe(audio_path, word_timestamps=word_timestamps, **options)
out = []
for s in segments:
seg = {"start": float(s.start), "end": float(s.end), "text": s.text}
@@ -178,7 +185,7 @@ def main() -> int:
if op == "ping":
_send(stdout, {"op": "pong"})
elif op == "transcribe":
result = _transcribe(msg.get("audio_path"), bool(msg.get("word_timestamps", True)))
result = _transcribe(msg.get("audio_path"), bool(msg.get("word_timestamps", True)), msg.get("decode_options"))
_send(stdout, {"op": "segments", "result": result})
elif op == "shutdown":
return 0
+26
View File
@@ -27,6 +27,12 @@ Test-only crash hook (only when OMNIVOICE_ECHO_CRASH=1): the sidecar will
self-`os._exit(1)` after dispatching exactly one frame, to exercise the
parent's "sidecar died mid-generate" recovery path.
Test-only handshake hooks (only with OMNIVOICE_ECHO_TEST_MODE=1), checked
before the ready frame: OMNIVOICE_ECHO_EXIT_BEFORE_READY=<code> exits with
that code, OMNIVOICE_ECHO_STALL_BEFORE_READY=1 sleeps past any test deadline,
OMNIVOICE_ECHO_ERROR_BEFORE_READY=<message> sends an error frame, and
OMNIVOICE_ECHO_WRONG_READY=1 sends a pong instead of ready (#2026).
This script is stdlib-only on purpose no torch, no numpy. The whole point
of the echo sidecar is that it can spawn under the bare system Python
interpreter without any engine venv.
@@ -100,6 +106,26 @@ def main() -> int:
test_mode = os.environ.get("OMNIVOICE_ECHO_TEST_MODE") == "1"
crash_after_one = os.environ.get("OMNIVOICE_ECHO_CRASH") == "1"
# Test-only ready-handshake failures (#2026): exit, stall, or answer
# with the wrong op before the ready frame, so the parent's report of
# each can be checked.
if test_mode:
early_exit = os.environ.get("OMNIVOICE_ECHO_EXIT_BEFORE_READY")
if early_exit:
print("echo: exiting before ready on purpose", file=sys.stderr, flush=True)
return int(early_exit)
if os.environ.get("OMNIVOICE_ECHO_STALL_BEFORE_READY") == "1":
import time
print("echo: stalling before ready on purpose", file=sys.stderr, flush=True)
time.sleep(60)
error_message = os.environ.get("OMNIVOICE_ECHO_ERROR_BEFORE_READY")
if error_message:
_send(stdout, {"op": "error", "stage": "startup", "message": error_message})
return 1
if os.environ.get("OMNIVOICE_ECHO_WRONG_READY") == "1":
_send(stdout, {"op": "pong"})
_send(stdout, {"op": "ready", "engine": "_echo"})
frames_handled = 0
+590
View File
@@ -0,0 +1,590 @@
"""audio.cpp TTS backend — Breeze-TTS-2 via a managed native server.
audio.cpp (0xShug0/audio.cpp) is a pure-C++ ggml runtime: prebuilt
``audiocpp_server`` binaries for Windows/macOS/Linux, no Python venv, no
``transformers`` pin so this engine needs neither the venv-isolation
(``engines.dots_tts``) nor the per-generate CLI-spawn (``engines
.omnivoice_gguf``) patterns. The parent instead:
1. resolves the binary + GGUF model (``bootstrap.py``),
2. spawns ONE long-lived ``audiocpp_server`` on 127.0.0.1 (lazy model load,
so model memory is only held after the first generate), and
3. speaks its OpenAI-style ``POST /v1/audio/speech`` per generate.
v1 serves the ``breeze_tts`` family only (Breeze-TTS-2, en+zh, voice clone
+ voice design + voice direction). The server is task-agnostic on the
speech route reference-audio presence selects clone/direction vs design
so a single ``task: tts`` model entry covers all three modes.
License honesty: Breeze-TTS-2 weights (``BreezeBlue/Breeze-TTS-2`` and the
audio.cpp GGUF repack) are RESEARCH AND NON-COMMERCIAL ONLY
(``BreezeBlue Research and Non-Commercial License``); only the audio.cpp
code is Apache-2.0. There is no in-tree acceptance dialog for this engine
yet (settings ``/license`` allow-list), so the restriction is surfaced in
the display name, the install hint, and ``docs/engines/audio-cpp.md``
not silently.
"""
from __future__ import annotations
import atexit
import base64
import io
import json
import logging
import os
import secrets
# Used only for stream constants; spawn_owned performs the process launch.
import subprocess # nosec B404
import threading
import time
import urllib.error
import urllib.request
from pathlib import Path
from typing import TYPE_CHECKING, Any
from core.contained_subprocess import spawn_owned
from services.tts_backend import TTSBackend, TTSInputError
if TYPE_CHECKING:
import torch
logger = logging.getLogger("omnivoice.audiocpp")
#: Engine id in the TTS registry.
ENGINE_ID = "audiocpp"
#: How long to wait for ``/health`` after spawning the server (first spawn
#: extracts nothing heavy — the model loads lazily on first generate).
_HEALTH_TIMEOUT_S = 120.0
#: Finish the inner HTTP request before the canonical generation guard can
#: abandon its worker thread. This leaves enough time to terminate the owned
#: native process and release its model memory synchronously.
_TERMINATE_GRACE_S = 5.0
_TERMINATE_KILL_S = 5.0
_GENERATE_TIMEOUT_MARGIN_S = (
_TERMINATE_GRACE_S + _TERMINATE_KILL_S + 5.0
)
# ── pure request/config builders (unit-tested, no I/O) ──────────────────────
def _cpu_thread_count() -> int:
"""Use up to 16 physical cores, with a stdlib fallback."""
try:
import psutil
cores = psutil.cpu_count(logical=False)
except (ImportError, OSError):
cores = None
return min(16, max(1, cores or os.cpu_count() or 1))
def _device_min_vram_gb(device) -> float:
"""Dedicated-memory comfort floor for one discovered native device."""
return 6.0 if (
device
and device.kind == "GPU"
and (
device.backend == "vulkan"
or device.hardware_family in {"cuda", "rocm"}
)
) else 0.0
def build_server_config(
*, model_id: str, family: str, model_path: str, port: int,
backend: str = "cpu", device: int = 0,
execution_target: str | None = None,
) -> dict:
"""``server.json`` dict for the managed ``audiocpp_server``.
``lazy_load`` defers the ~4.73 GiB GGUF load to the first generate;
``max_loaded_models: 1`` bounds residency to the one model we serve.
"""
return {
"host": "127.0.0.1",
"port": port,
"backend": backend,
"device": device,
# The pinned CPU runtime scales strongly through 16 workers while
# producing byte-identical audio.
"threads": _cpu_thread_count()
if (execution_target or backend) == "cpu" else 1,
"lazy_load": True,
"max_loaded_models": 1,
"models": [
{
"id": model_id,
"family": family,
"path": model_path,
"task": "tts",
"mode": "offline",
}
],
}
def build_speech_payload(
*, model_id: str, text: str, ref_audio: str | None = None,
ref_text: str | None = None, instructions: str | None = None,
guidance_scale: float | None = None, seed: int | None = None,
) -> dict:
"""``POST /v1/audio/speech`` JSON body.
Field spellings verified against ``app/server/runtime.cpp``
(``build_speech_request``): ``instructions`` (plural, OpenAI spelling)
feeds the ``instruction`` request option; ``reference_text`` and
``guidance_scale``/``seed`` pass through top-level; ``voice_ref`` takes
a ``{"type": "path", ...}`` object so the reference stays on disk
(the 5 MiB base64 cap never bites). ``response_format: json`` returns
the WAV base64-in-JSON one round trip, no binary framing.
"""
payload: dict[str, Any] = {
"model": model_id,
"input": text,
"response_format": "json",
}
if instructions:
payload["instructions"] = instructions
if ref_audio:
payload["voice_ref"] = {"type": "path", "path": str(ref_audio)}
if ref_text:
payload["reference_text"] = ref_text
if guidance_scale is not None:
payload["guidance_scale"] = float(guidance_scale)
if seed is not None:
payload["seed"] = int(seed)
return payload
def decode_speech_json(obj: dict) -> tuple[int, object]:
"""``(sample_rate, mono float32 numpy)`` from a ``response_format=json``
speech body. Raises ``ValueError`` on a server error payload."""
if not isinstance(obj, dict):
raise TypeError(f"audio.cpp speech reply is not JSON: {obj!r:.120}")
if "audio" not in obj:
raise ValueError(f"audio.cpp speech failed: {obj.get('error', obj)!r:.300}")
import numpy as np
import soundfile as sf
wav_bytes = base64.b64decode(obj["audio"])
wav, sr = sf.read(io.BytesIO(wav_bytes), dtype="float32", always_2d=False)
wav = np.asarray(wav, dtype=np.float32)
if wav.ndim > 1:
wav = wav.mean(axis=-1)
return int(sr), wav
# ── backend ─────────────────────────────────────────────────────────────────
class AudioCPPBackend(TTSBackend):
"""Breeze-TTS-2 through a parent-managed ``audiocpp_server``."""
id = ENGINE_ID
display_name = (
"audio.cpp · Breeze-TTS-2 (native GGUF, en+zh, clone+design; "
"weights research/non-commercial)"
)
supports_voice_design = True
applies_own_mastering = True # model-decoded 24 kHz studio output
gpu_compat = ("cpu",)
runs_out_of_process = True
# Same marker SubprocessBackend sets: this engine lives in another OS
# process. Consumers only branch the matrix label and the self-test
# route (spawn-and-ping instead of in-process synth) — both correct
# here; nothing assumes the stdio protocol from it.
_is_subprocess_isolated = True
_DEFAULT_SAMPLE_RATE = 24000 # Breeze-TTS-2 native rate
def __init__(self) -> None:
self._proc: Any | None = None
self._port: int | None = None
self._server_model_id: str | None = None
self._sr = self._DEFAULT_SAMPLE_RATE
self._lock = threading.RLock()
self._server_json: Path | None = None
self._selection = None
self._device = None
self._provider = None
# ── availability ────────────────────────────────────────────────────
@classmethod
def is_available(cls) -> tuple[bool, str]:
from engines.audiocpp import bootstrap
try:
bootstrap.resolve_server_binary()
bootstrap.resolve_model_file()
except RuntimeError as exc:
return False, str(exc)
return True, "ready"
@classmethod
def runtime_compute_profile(cls, caps) -> dict:
from dataclasses import replace
from engines.audiocpp import bootstrap
from services.engine_routing import low_vram_caveat
try:
selection = bootstrap.resolve_compute_selection(caps)
targets = bootstrap.runtime_targets()
except RuntimeError as exc:
return {
"gpu_compat": cls.gpu_compat,
"min_vram_gb": 0.0,
"effective_device": "cpu",
"routing_status": "unavailable",
"routing_reason": str(exc),
"runtime_backend": None,
"runtime_device_index": None,
"runtime_device_name": None,
"runtime_hardware_family": None,
"runtime_vram_gb": None,
"runtime_device_verified": False,
}
selected = selection.device
accelerated = selected.target != "cpu"
min_vram_gb = _device_min_vram_gb(selected)
dedicated = min_vram_gb > 0
reason = selection.fallback_reason
if accelerated and dedicated and reason is None:
selected_caps = replace(
caps,
device_name=selected.name,
vram_gb=selection.verified_vram_gb,
)
reason = low_vram_caveat(
selected_caps,
min_vram_gb,
family=selected.hardware_family,
vram_gb=selection.verified_vram_gb,
)
status = "accelerated" if accelerated else (
"cpu_fallback" if selection.fallback_reason else "cpu_only"
)
return {
"gpu_compat": targets,
"min_vram_gb": min_vram_gb,
"effective_device": selected.target,
"routing_status": status,
"routing_reason": reason,
"runtime_backend": selected.backend,
"runtime_device_index": selected.index,
"runtime_device_name": selected.name,
"runtime_hardware_family": selected.hardware_family,
"runtime_vram_gb": selection.verified_vram_gb,
"runtime_device_verified": selection.verified_vram_gb > 0,
}
# ── TTSBackend protocol ─────────────────────────────────────────────
@property
def sample_rate(self) -> int:
return self._sr
@property
def supported_languages(self) -> list[str]:
return ["en", "zh"]
def model_identity(self) -> str | None:
from engines.audiocpp import bootstrap
return f"{bootstrap.FAMILY}/{bootstrap.package_filename()}"
# ── server lifecycle ────────────────────────────────────────────────
def _base_url(self) -> str:
return f"http://127.0.0.1:{self._port}"
def _ensure_loaded(self) -> None:
"""Spawn the server (once) and wait for ``/health``. Idempotent."""
with self._lock:
if self._proc is not None and self._proc.poll() is None:
return
self._proc = None # stale handle — respawn below
from engines.audiocpp import bootstrap
binary = bootstrap.resolve_server_binary()
selection = bootstrap.resolve_compute_selection()
model_file = bootstrap.resolve_model_file()
self._port = bootstrap.server_port()
# The random model id is a per-launch challenge. Before sending
# speech text or a reference path, _verify_server_identity asks
# /v1/models to prove this is the child configured by this process,
# not an unrelated listener that pre-bound the loopback port.
self._server_model_id = f"{bootstrap.MODEL_ID}-{secrets.token_hex(16)}"
config = build_server_config(
model_id=self._server_model_id,
family=bootstrap.FAMILY,
model_path=str(model_file),
port=self._port,
backend=selection.device.backend,
device=selection.device.index,
execution_target=selection.device.target,
)
self._selection = selection
self._device = selection.device.target
self._provider = selection.device.backend
from core.config import DATA_DIR
workdir = Path(str(DATA_DIR)) / "audiocpp"
workdir.mkdir(parents=True, exist_ok=True)
self._server_json = workdir / "server.json"
flags = os.O_WRONLY | os.O_CREAT | os.O_TRUNC
config_fd = os.open(self._server_json, flags, 0o600)
try:
if os.name != "nt":
os.fchmod(config_fd, 0o600)
with os.fdopen(config_fd, "w", encoding="utf-8") as config_fh:
config_fd = -1
json.dump(config, config_fh, indent=2)
finally:
if config_fd >= 0:
os.close(config_fd)
log_path = workdir / "server.log"
logger.info(
"audio.cpp: starting %s (backend=%s, device=%d, port=%d, model=%s)",
binary.name, selection.device.backend, selection.device.index,
self._port, model_file.name,
)
with open(log_path, "ab") as log_fh:
self._proc = spawn_owned(
[str(binary), "--config", str(self._server_json)],
stdout=log_fh,
stderr=subprocess.STDOUT,
stdin=subprocess.DEVNULL,
)
atexit.register(self._terminate_server)
self._wait_for_health()
def _wait_for_health(self) -> None:
if self._proc is None or self._port is None:
raise RuntimeError("managed audio.cpp server was not started")
deadline = time.monotonic() + _HEALTH_TIMEOUT_S
last_err = "unknown"
url = self._base_url() + "/health"
while time.monotonic() < deadline:
if self._proc.poll() is not None:
raise RuntimeError(
"audiocpp_server exited during startup "
f"(code {self._proc.returncode}). See the server log next "
"to server.json under the app data audiocpp/ directory — "
"the managed port may already be in use."
)
try:
# ``url`` is always the hard-coded loopback host plus a
# validated integer port; arbitrary schemes are impossible.
with urllib.request.urlopen(url, timeout=5) as resp: # nosec B310
if resp.status == 200:
self._verify_server_identity()
if self._proc.poll() is None:
logger.info(
"audio.cpp: managed server is healthy on loopback"
)
return
last_err = f"HTTP {resp.status}"
except Exception as exc: # noqa: BLE001 — still starting; retry
last_err = f"{type(exc).__name__}: {exc}"
time.sleep(1.0)
self._terminate_server()
raise RuntimeError(
f"audiocpp_server did not become healthy within "
f"{_HEALTH_TIMEOUT_S:.0f}s (last: {last_err})."
)
def _get_json(self, path: str, timeout: float = 5.0) -> dict:
"""GET one loopback JSON endpoint without sending request content."""
if self._port is None:
raise RuntimeError("managed audio.cpp server port is missing")
req = urllib.request.Request(self._base_url() + path, method="GET")
with urllib.request.urlopen(req, timeout=timeout) as resp: # nosec B310
obj = json.loads(resp.read().decode("utf-8"))
if not isinstance(obj, dict):
raise TypeError("audio.cpp returned an invalid JSON response")
return obj
def _verify_server_identity(self) -> None:
"""Prove the loopback listener owns this launch's random model id."""
if self._proc is None or self._proc.poll() is not None:
raise RuntimeError("managed audio.cpp server is not running")
expected = self._server_model_id
if not expected:
raise RuntimeError("managed audio.cpp server identity is missing")
obj = self._get_json("/v1/models")
data = obj.get("data", [])
if not isinstance(data, list):
raise TypeError("managed audio.cpp server identity is invalid")
model_ids = {
item.get("id") for item in data
if isinstance(item, dict)
}
if expected not in model_ids or self._proc.poll() is not None:
raise RuntimeError(
"loopback listener did not prove managed audio.cpp ownership"
)
def _post_json(self, path: str, payload: dict, timeout: float) -> dict:
"""Verify child ownership, then POST JSON to the managed server."""
if self._port is None:
raise RuntimeError("managed audio.cpp server port is missing")
self._verify_server_identity()
body = json.dumps(payload).encode("utf-8")
req = urllib.request.Request(
self._base_url() + path,
data=body,
headers={"Content-Type": "application/json"},
method="POST",
)
try:
# ``req`` targets only ``_base_url()`` (127.0.0.1 + validated
# integer port), never a caller-provided URL.
with urllib.request.urlopen(req, timeout=timeout) as resp: # nosec B310
return json.loads(resp.read().decode("utf-8"))
except urllib.error.HTTPError as exc:
detail = exc.read().decode("utf-8", errors="replace")[:500]
raise RuntimeError(
f"audio.cpp {path} failed (HTTP {exc.code}): {detail}"
) from exc
except urllib.error.URLError as exc:
if isinstance(exc.reason, TimeoutError):
raise TimeoutError("audio.cpp request timed out") from exc
raise
def _terminate_server(self) -> None:
proc, self._proc = self._proc, None
self._server_model_id = None
if proc is None:
return
try:
proc.terminate()
proc.wait(timeout=_TERMINATE_GRACE_S)
except Exception: # noqa: BLE001 — kill as last resort, never raise
try:
proc.kill()
proc.wait(timeout=_TERMINATE_KILL_S)
except Exception as exc: # noqa: BLE001 — process is already failing
logger.debug("audio.cpp: final server kill failed: %s", exc)
# ── generate ────────────────────────────────────────────────────────
def generate(self, text: str, **kw) -> torch.Tensor:
import torch
from services.model_manager import (
GENERATE_PROGRESS_GRACE_S,
generate_timeout_s,
report_generate_progress,
)
if not text or not text.strip():
raise TTSInputError(
"audio.cpp: the input contains no speakable text — "
"send at least one word."
)
ref_audio = kw.get("ref_audio")
ref_text = kw.get("ref_text")
if ref_text and not ref_audio:
logger.info(
"audio.cpp: ref_text supplied without ref_audio; ignoring."
)
ref_text = None
# Voice design: our `description=` (no ref) and voice direction
# (`instruct=` + ref) both ride the server's `instructions` field —
# verified spelling against app/server/runtime.cpp.
instruct = kw.get("instruct") or kw.get("description") or None
language = kw.get("language")
if language and str(language).strip().lower() not in {
"auto", "en", "english", "zh", "chinese",
}:
logger.info(
"audio.cpp (Breeze-TTS-2) is en+zh only; ignoring "
"language=%r.", language,
)
if kw.get("speed", 1.0) != 1.0:
logger.info("audio.cpp: speed is not supported; ignoring.")
request_started = time.monotonic()
with self._lock:
self._ensure_loaded()
selected = self._selection.device if self._selection else None
min_vram_gb = _device_min_vram_gb(selected)
request_budget = generate_timeout_s(
text,
execution_device=selected.target if selected else "cpu",
min_vram_gb=min_vram_gb,
hardware_family=selected.hardware_family if selected else None,
vram_gb=self._selection.verified_vram_gb
if self._selection else 0.0,
)
if not self._server_model_id:
raise RuntimeError("managed audio.cpp server identity is missing")
payload = build_speech_payload(
model_id=self._server_model_id,
text=text,
ref_audio=str(ref_audio) if ref_audio else None,
ref_text=ref_text,
instructions=instruct,
guidance_scale=kw.get("guidance_scale", 1.0),
seed=kw.get("seed"),
)
# Device discovery and server startup can consume part of the soft
# budget. This fresh synthesis lease gives the lazy model load and
# request a bounded window. The inner request always expires early
# enough to reap the owned server before the outer guard abandons us.
report_generate_progress()
soft_remaining = request_budget - (time.monotonic() - request_started)
timeout = (
max(soft_remaining, GENERATE_PROGRESS_GRACE_S)
- _GENERATE_TIMEOUT_MARGIN_S
)
if timeout <= 0:
self._terminate_server()
raise TimeoutError(
"audio.cpp startup exhausted the generation time budget"
)
try:
obj = self._post_json(
"/v1/audio/speech", payload, timeout=timeout,
)
except TimeoutError:
self._terminate_server()
raise RuntimeError(
"audio.cpp generation timed out; its managed server was reset"
) from None
sr, wav_np = decode_speech_json(obj)
self._sr = sr
wav = torch.from_numpy(wav_np).float()
if wav.ndim == 0:
raise RuntimeError("audio.cpp produced empty audio")
return wav.unsqueeze(0)
# ── lifecycle ───────────────────────────────────────────────────────
def unload(self) -> None:
"""Free the model server-side, then stop it. Idempotent."""
with self._lock:
if self._port is not None and self._proc is not None \
and self._proc.poll() is None:
try:
self._post_json("/v1/tasks/unload_all_models", {}, timeout=30)
except Exception as exc: # noqa: BLE001 — best effort
logger.warning("audio.cpp: server unload failed: %s", exc)
self._port = None
self._terminate_server()
super().unload()
__all__ = [
"ENGINE_ID",
"AudioCPPBackend",
"build_server_config",
"build_speech_payload",
"decode_speech_json",
]
+765
View File
@@ -0,0 +1,765 @@
"""audio.cpp binary probe + model resolution.
audio.cpp (0xShug0/audio.cpp) is a pure-C++ ggml inference engine with
prebuilt release binaries no Python venv, no ``transformers`` pin, so
none of the dependency-isolation machinery in ``engines._venv_probe`` or
``services.subprocess_backend`` applies. The parent instead:
1. locates a user-installed ``audiocpp_server`` (env var, user dir, or this
package's ``bin/``), and
2. resolves an explicitly installed GGUF model file from a direct path or
the shared Hugging Face cache.
Probe order for the server binary (existing installs win, zero migration):
1. ``${OMNIVOICE_AUDIOCPP_BIN}`` absolute path to the binary itself.
2. ``${OMNIVOICE_AUDIOCPP_DIR}/audiocpp_server[.exe]`` a user-managed
install dir (e.g. an extracted release zip, or a self-built tree).
3. ``backend/engines/audiocpp/bin/audiocpp_server[.exe]`` an explicitly
installed local copy.
``is_installed()`` is a cheap file-existence check no spawn, no network.
VoiceStudio never downloads executable code for this engine.
"""
from __future__ import annotations
import errno
import functools
import logging
import os
import platform
import subprocess # nosec B404 -- fixed argv probes a user-selected executable
import sys
from dataclasses import dataclass
from pathlib import Path
logger = logging.getLogger("omnivoice.audiocpp.bootstrap")
#: Pinned audio.cpp release. BreezeTTS-2 support landed in 0.7.2; 0.7.4 adds
#: the current native fixes and Sortformer v2.1 streaming runtime.
VERSION = "v0.7.4"
#: GitHub repo serving the prebuilt binaries.
GH_REPO = "0xShug0/audio.cpp"
#: HuggingFace repo serving the GGUF model packages (not gated).
HF_MODEL_REPO = "audio-cpp/audio.cpp-gguf"
# Immutable repository revision used for the Breeze-TTS-2 package.
# Pinning prevents a later upstream file replacement from silently changing
# the model exercised by this backend.
HF_MODEL_REVISION = "dc6fecccc2b0c6bdda0a8b2f38fa61394fee0b9c"
#: Model id used in the generated ``server.json`` and in speech requests.
MODEL_ID = "breeze-tts-2"
#: audio.cpp family name for BreezeTTS 2 (``--family`` / server ``family``).
FAMILY = "breeze_tts"
#: GGUF package directory inside :data:`HF_MODEL_REPO`.
PACKAGE_DIR = "Breeze-TTS-2-GGUF"
#: Default package (Q8_0, the upstream-recommended GGUF). ``bf16`` is
#: available via ``OMNIVOICE_AUDIOCPP_PACKAGE``.
DEFAULT_PACKAGE = "breeze-tts-2-q8_0.gguf"
#: Env var pointing directly at the ``audiocpp_server`` binary.
BIN_ENV = "OMNIVOICE_AUDIOCPP_BIN"
#: Env var pointing at a directory containing ``audiocpp_server``.
DIR_ENV = "OMNIVOICE_AUDIOCPP_DIR"
#: Env var overriding the GGUF package filename (e.g. the bf16 package).
PACKAGE_ENV = "OMNIVOICE_AUDIOCPP_PACKAGE"
#: Optional advanced overrides for a binary that exposes several runtimes or
#: devices. Device indices are local to the selected backend registry.
BACKEND_ENV = "OMNIVOICE_AUDIOCPP_BACKEND"
DEVICE_ENV = "OMNIVOICE_AUDIOCPP_DEVICE"
#: Env var overriding the loopback port the managed server binds.
PORT_ENV = "OMNIVOICE_AUDIOCPP_PORT"
#: Default loopback port. High and engine-specific to avoid clashing with
#: the app itself or a user-run ``audiocpp_server`` (default 8080).
DEFAULT_PORT = 17860
#: This package's owned binary dir (probe 3).
_PKG_BIN_DIR: Path = Path(__file__).parent / "bin"
# Recommended (asset filename, sha256) per platform slug, from the v0.7.4
# release. Windows and Linux use the vendor-neutral Vulkan build, which also
# exposes the native CPU backend. Upstream publishes the macOS builds under
# the Metal package name. No linux-aarch64 prebuilt exists in v0.7.4.
_ASSETS: dict[str, tuple[str, str]] = {
"windows-x64": (
"audio-v0.7.4-bin-windows-x64-vulkan.zip",
"057332f9e3fb37706a8ecb5075ac1797efcd85fdccd739f7b65761a5920f2828",
),
"linux-x64": (
"audio-v0.7.4-bin-ubuntu-x64-vulkan.tar.gz",
"e0ef3123a9f94e130ad463db0db5a69b65485ef8db1b46edead00c03a86fa787",
),
"darwin-arm64": (
"audio-v0.7.4-bin-macos-arm64-metal.tar.gz",
"639926715b1cb537f82aa31656aabbae5d9a85ac36568c402026968f3072e2b3",
),
"darwin-x64": (
"audio-v0.7.4-bin-macos-x64-metal.tar.gz",
"bdb797d54dcf8416bd5ac0fac282ce5500dd08843f8f22e20e9fc378ebc24c1f",
),
}
_ASSET_SIZES = {
"windows-x64": 56_818_905,
"linux-x64": 71_551_673,
"darwin-arm64": 25_270_657,
"darwin-x64": 26_718_959,
}
#: Binary filename per platform.
_BINARY_NAMES = {"windows-x64": "audiocpp_server.exe"}
_REGISTRY_BACKENDS = {
"CPU": "cpu",
"CUDA": "cuda",
"MUSA": "cuda",
"HIP": "hip",
"ROCm": "hip",
"Vulkan": "vulkan",
"Metal": "metal",
"MTL": "metal",
}
_BACKEND_ALIASES = {
"cpu": "cpu",
"cuda": "cuda",
"hip": "hip",
"rocm": "hip",
"vulkan": "vulkan",
"metal": "metal",
}
@dataclass(frozen=True)
class AudioCPPDevice:
"""One immutable device from audio.cpp's backend-local registry."""
registry: str
backend: str
index: int
name: str
kind: str
target: str
hardware_family: str
@dataclass(frozen=True)
class AudioCPPSelection:
"""The runtime/device chosen for the next managed server."""
device: AudioCPPDevice
fallback_reason: str | None = None
verified_vram_gb: float = 0.0
@dataclass(frozen=True)
class _ProbeOutcome:
devices: tuple[AudioCPPDevice, ...] = ()
error: str | None = None
def _cpu_probe_fallback(error: RuntimeError) -> AudioCPPSelection:
"""A usable automatic fallback when native device discovery fails."""
return AudioCPPSelection(
AudioCPPDevice(
registry="CPU",
backend="cpu",
index=0,
name="Host CPU",
kind="CPU",
target="cpu",
hardware_family="cpu",
),
f"{error}; running on CPU",
)
def _vulkan_hardware_family(name: str) -> str:
low = name.casefold()
if any(token in low for token in ("nvidia", "geforce", "quadro", "tesla")):
return "cuda"
if any(token in low for token in ("amd", "radeon")):
return "rocm"
if any(token in low for token in ("intel", "arc ")):
return "xpu"
return "vulkan"
def _device_families(registry: str, name: str, kind: str) -> tuple[str, str]:
# Software adapters such as Vulkan llvmpipe may be listed by a GPU
# registry but still execute on the CPU. Keep their runtime backend for
# explicit overrides while reporting and routing them as CPU work.
if kind == "CPU":
return "cpu", "cpu"
if registry in {"CUDA", "MUSA"}:
return "cuda", "cuda"
if registry in {"HIP", "ROCm"}:
return "rocm", "rocm"
if registry in {"Metal", "MTL"}:
return "mps", "mps"
if registry == "Vulkan":
return "vulkan", _vulkan_hardware_family(name)
return "cpu", "cpu"
def parse_device_list(output: str) -> tuple[AudioCPPDevice, ...]:
"""Parse the stable stdout contract of ``--list-devices``.
Backend diagnostics are emitted on stderr and deliberately never enter
this parser. Unknown future registries are ignored; malformed entries for
a registry we understand fail closed instead of selecting the wrong GPU.
"""
devices: list[AudioCPPDevice] = []
seen: set[tuple[str, int]] = set()
for raw in str(output or "").splitlines():
line = raw.strip()
registry, colon, detail = line.partition(":")
if not colon or registry not in _REGISTRY_BACKENDS:
continue
index_text, space, remainder = detail.strip().partition(" ")
if not space or not index_text.isascii() or not index_text.isdecimal():
raise RuntimeError(
f"malformed audio.cpp {registry} device entry"
)
index = int(index_text)
remainder = remainder.strip()
kind_start = remainder.rfind("[")
if kind_start < 0 or not remainder.endswith("]"):
raise RuntimeError(
f"malformed audio.cpp {registry} device entry"
)
name_field = remainder[:kind_start].strip()
if name_field:
if len(name_field) < 2 or name_field[0] != '"' or name_field[-1] != '"':
raise RuntimeError(
f"malformed audio.cpp {registry} device entry"
)
name = name_field[1:-1]
else:
name = ""
kind = remainder[kind_start + 1:-1].strip().upper()
if kind not in {"CPU", "GPU", "IGPU", "ACCEL", "META"}:
raise RuntimeError("unknown audio.cpp device kind")
# Registry aliases such as HIP/ROCm share one backend-local index
# namespace and therefore cannot safely describe different devices.
key = (_REGISTRY_BACKENDS[registry], index)
if key in seen:
raise RuntimeError(
f"duplicate audio.cpp device entry: {registry}:{index}"
)
seen.add(key)
target, hardware_family = _device_families(registry, name, kind)
devices.append(AudioCPPDevice(
registry=registry,
backend=_REGISTRY_BACKENDS[registry],
index=index,
name=name,
kind=kind,
target=target,
hardware_family=hardware_family,
))
if not devices:
raise RuntimeError("audio.cpp reported no recognized compute devices")
return tuple(devices)
def _platform_slug() -> str:
system = sys.platform
machine = platform.machine().lower()
if system == "win32":
return "windows-x64"
if system == "darwin":
return "darwin-arm64" if machine in ("arm64", "aarch64") else "darwin-x64"
if machine in ("x86_64", "amd64"):
return "linux-x64"
return f"linux-{machine}"
def binary_name(slug: str | None = None) -> str:
"""``audiocpp_server`` filename for ``slug`` (``.exe`` on Windows)."""
return _BINARY_NAMES.get(slug or _platform_slug(), "audiocpp_server")
def _probe_paths() -> list[Path]:
out: list[Path] = []
direct = os.environ.get(BIN_ENV, "").strip()
if direct:
out.append(Path(direct))
user_dir = os.environ.get(DIR_ENV, "").strip()
if user_dir:
out.append(Path(user_dir) / binary_name())
out.append(managed_runtime_dir() / binary_name())
out.append(_PKG_BIN_DIR / binary_name())
return out
def platform_slug() -> str:
"""Stable release-platform key used by the managed runtime installer."""
return _platform_slug()
def managed_runtime_dir() -> Path:
"""Update-surviving location for the checksummed app-managed runtime."""
from core.config import DATA_DIR
return Path(DATA_DIR) / "engines" / "audio-cpp" / VERSION.lstrip("v") / _platform_slug()
def is_installed() -> bool:
"""Cheap precedence-aware check for a usable server binary."""
try:
resolve_server_binary()
except RuntimeError:
return False
return True
def resolve_server_binary() -> Path:
"""Resolve the ``audiocpp_server`` binary. Raises ``RuntimeError`` with
install instructions when none is found."""
for cand in _probe_paths():
if cand.is_file():
if os.name == "nt" or os.access(cand, os.X_OK):
return cand
raise RuntimeError(
"audiocpp_server is not executable. Run `chmod +x "
"audiocpp_server` on the configured binary, then restart "
"VoiceStudio. See docs/engines/audio-cpp.md."
)
slug = _platform_slug()
asset = _ASSETS.get(slug)
if asset is None:
raise RuntimeError(
f"audio.cpp ships no prebuilt binary for this platform ({slug}). "
"Build from https://github.com/0xShug0/audio.cpp and set "
f"{BIN_ENV} to your audiocpp_server binary. See "
"docs/engines/audio-cpp.md."
)
raise RuntimeError(
"audiocpp_server not found. Download "
f"https://github.com/{GH_REPO}/releases/download/{VERSION}/{asset[0]} "
f"(SHA-256 {asset[1]}), verify and extract it, and set {BIN_ENV} to the "
"audiocpp_server binary (or "
f"{DIR_ENV} to its directory). See docs/engines/audio-cpp.md."
)
@functools.lru_cache(maxsize=4)
def _probe_device_outcome(binary: str) -> _ProbeOutcome:
try:
proc = subprocess.run( # nosec B603 -- executable is the resolved engine binary
[binary, "--list-devices"],
capture_output=True,
text=True,
timeout=10,
check=False,
)
except subprocess.TimeoutExpired:
return _ProbeOutcome(
error="audiocpp_server device discovery timed out after 10 seconds"
)
except OSError as exc:
return _ProbeOutcome(
error=(
"audiocpp_server device discovery could not start: "
f"{type(exc).__name__}"
)
)
if proc.returncode != 0:
return _ProbeOutcome(
error=(
"audiocpp_server device discovery failed "
f"(code {proc.returncode}). Check the audio.cpp server log "
"for details."
)
)
try:
return _ProbeOutcome(devices=parse_device_list(proc.stdout))
except RuntimeError as exc:
return _ProbeOutcome(error=str(exc))
def _probe_devices(binary: str) -> tuple[AudioCPPDevice, ...]:
outcome = _probe_device_outcome(binary)
if outcome.error:
raise RuntimeError(outcome.error)
return outcome.devices
def probe_devices() -> tuple[AudioCPPDevice, ...]:
"""Return the installed binary's devices without loading a model."""
return _probe_devices(str(resolve_server_binary()))
def _priority(device: AudioCPPDevice) -> tuple[int, int]:
if device.kind == "META":
# Tensor-parallel meta devices are valid explicit targets, but their
# resource footprint is not safe to choose implicitly over CPU.
rank = 8
elif device.backend != "cpu" and device.kind == "CPU":
# Native CPU is the predictable fallback. Software adapters remain
# available to an explicit backend override but never win auto mode.
rank = 7
elif device.backend == "cuda":
rank = 0
elif device.backend == "hip":
rank = 1
elif device.backend == "metal":
rank = 2
elif device.backend == "vulkan" and device.kind == "GPU":
rank = 3
elif device.backend == "vulkan" and device.kind in {"IGPU", "ACCEL"}:
rank = 4
elif device.backend == "cpu":
rank = 6
else:
rank = 5
return rank, device.index
def select_device(
devices: tuple[AudioCPPDevice, ...],
*,
requested_family: str = "auto",
backend_override: str | None = None,
device_override: int | None = None,
preferred_name: str = "",
) -> AudioCPPSelection:
"""Resolve one device with explicit overrides and discrete-GPU priority."""
if backend_override:
normalized = _BACKEND_ALIASES.get(backend_override.strip().lower())
if normalized is None:
valid = ", ".join(_BACKEND_ALIASES)
raise RuntimeError(
f"unknown audio.cpp backend '{backend_override}' (valid: {valid})"
)
candidates = [device for device in devices if device.backend == normalized]
if device_override is not None:
candidates = [
device for device in candidates if device.index == device_override
]
if not candidates:
suffix = "" if device_override is None else f" device {device_override}"
available = ", ".join(
f"{device.backend}:{device.index}" for device in devices
)
raise RuntimeError(
f"audio.cpp backend '{backend_override}'{suffix} is unavailable "
f"(available: {available})"
)
# An explicit runtime request should still prefer a compute device to
# a software adapter when no backend-local index was supplied. META is
# valid here because the user explicitly chose this registry.
return AudioCPPSelection(min(
candidates,
key=lambda device: (device.kind == "CPU", _priority(device)),
))
if device_override is not None:
raise RuntimeError(
f"{DEVICE_ENV} requires {BACKEND_ENV} because device indices are "
"backend-local"
)
family = (requested_family or "auto").strip().lower()
if family != "auto":
candidates = [
device for device in devices if device.hardware_family == family
]
if candidates:
preferred = preferred_name.casefold().strip()
if preferred:
named = [
device for device in candidates
if device.name
and (
preferred in device.name.casefold()
or device.name.casefold() in preferred
)
]
if named:
candidates = named
return AudioCPPSelection(min(candidates, key=_priority))
cpu = [device for device in devices if device.backend == "cpu"]
if cpu:
return AudioCPPSelection(
min(cpu, key=_priority),
f"requested {family.upper()} device is not exposed by the "
"installed audio.cpp binary; running on CPU",
)
raise RuntimeError(
f"requested {family.upper()} device is not exposed by the "
"installed audio.cpp binary"
)
return AudioCPPSelection(min(devices, key=_priority))
def resolve_compute_selection(caps=None) -> AudioCPPSelection:
"""Select the runtime from engine env overrides, Settings, then auto."""
backend_override = os.environ.get(BACKEND_ENV, "").strip() or None
raw_device = os.environ.get(DEVICE_ENV, "").strip()
device_override: int | None = None
if raw_device:
try:
device_override = int(raw_device)
except ValueError as exc:
raise RuntimeError(
f"{DEVICE_ENV} must be a non-negative integer"
) from exc
if device_override < 0:
raise RuntimeError(f"{DEVICE_ENV} must be a non-negative integer")
if caps is None:
from core.device_caps import detect_host_caps
caps = detect_host_caps()
requested = getattr(caps, "requested_family", "auto") or "auto"
try:
devices = probe_devices()
except RuntimeError as exc:
if backend_override or raw_device or requested != "auto":
raise
return _cpu_probe_fallback(exc)
selection = select_device(
devices,
requested_family=requested,
backend_override=backend_override,
device_override=device_override,
preferred_name=getattr(caps, "device_name", "") or "",
)
# HostCaps measures the preferred accelerator's device 0. Reuse that VRAM
# only when the selected native registry has exactly one device with the
# same normalized name. Multi-GPU peers with identical names stay unknown.
selected_name = " ".join(selection.device.name.casefold().split())
host_name = " ".join(
str(getattr(caps, "device_name", "") or "").casefold().split()
)
peers = [
device for device in devices
if device.backend == selection.device.backend
and " ".join(device.name.casefold().split()) == host_name
]
if (
selected_name
and selected_name == host_name
and len(peers) == 1
and float(getattr(caps, "vram_gb", 0.0) or 0.0) > 0
):
return AudioCPPSelection(
selection.device,
selection.fallback_reason,
float(caps.vram_gb),
)
return selection
def runtime_targets(devices: tuple[AudioCPPDevice, ...] | None = None) -> tuple[str, ...]:
"""Actual compute backends compiled into the selected binary."""
if devices is not None:
found = devices
else:
try:
found = probe_devices()
except RuntimeError:
if (
os.environ.get(BACKEND_ENV, "").strip()
or os.environ.get(DEVICE_ENV, "").strip()
):
raise
return ("cpu",)
ordered: list[str] = []
for device in sorted(found, key=_priority):
if device.target not in ordered:
ordered.append(device.target)
return tuple(ordered)
def invalidate() -> None:
"""Forget cached binary capability discovery after an install change."""
_probe_device_outcome.cache_clear()
def default_asset() -> tuple[str, str] | None:
"""``(filename, sha256)`` of the release asset for this host, or None
when upstream ships no prebuilt for it."""
return _ASSETS.get(_platform_slug())
def default_asset_size() -> int | None:
"""Published byte size of this host's pinned release archive."""
return _ASSET_SIZES.get(_platform_slug())
def server_port() -> int:
"""Loopback port for the managed server (env override or default)."""
raw = os.environ.get(PORT_ENV, "").strip()
if raw:
try:
port = int(raw)
if 1 <= port <= 65535:
return port
logger.warning("Ignoring %s=%r: out of range.", PORT_ENV, raw)
except ValueError:
logger.warning("Ignoring %s=%r: not a number.", PORT_ENV, raw)
return DEFAULT_PORT
def package_filename() -> str:
"""GGUF package filename (env override or the Q8_0 default)."""
return os.environ.get(PACKAGE_ENV, "").strip() or DEFAULT_PACKAGE
def _materialize_gguf_cache_path(model_file: Path) -> Path:
"""Return a real ``.gguf`` path when the HF snapshot is a symlink.
audio.cpp canonicalizes model paths before inspecting the suffix. The
Hugging Face cache points the friendly ``.gguf`` snapshot name at an
extensionless content-addressed blob, so passing that symlink makes the
server reject a valid model. A hard link beside the snapshot keeps the
required suffix without copying a multi-gigabyte model or escaping the
snapshot's cleanup lifecycle.
"""
resolved = model_file.resolve()
if resolved.suffix.lower() == ".gguf":
return model_file
if model_file.suffix.lower() != ".gguf":
raise RuntimeError(f"audio.cpp model must be a .gguf file: {model_file}")
def _link(alias: Path) -> Path:
for attempt in range(2):
try:
os.link(resolved, alias)
except FileExistsError:
if (
not alias.is_symlink()
and alias.is_file()
and os.path.samefile(resolved, alias)
):
return alias
if attempt == 0 and alias.is_symlink():
alias.unlink()
continue
raise RuntimeError(
f"audio.cpp model alias points at a different file: {alias}"
) from None
return alias
raise RuntimeError(f"audio.cpp model alias could not be created: {alias}")
alias = model_file.with_name(
f".{model_file.stem}-{HF_MODEL_REVISION[:12]}.audiocpp.gguf"
)
try:
return _link(alias)
except OSError as exc:
if exc.errno == errno.EXDEV:
# An explicit symlink may live on a different filesystem from its
# target. Put the suffix-preserving hard link beside the resolved
# file so no multi-gigabyte copy is needed.
target_alias = resolved.with_name(
f".{resolved.name}-{HF_MODEL_REVISION[:12]}.audiocpp.gguf"
)
try:
return _link(target_alias)
except OSError as target_exc:
exc = target_exc
raise RuntimeError(
"audio.cpp cannot materialize the Hugging Face cache symlink as "
f"a .gguf hard link: {exc}"
) from exc
def resolve_model_file() -> Path:
"""Resolve an explicitly installed Breeze-TTS-2 GGUF file.
An explicit ``OMNIVOICE_AUDIOCPP_MODEL`` path wins (file or directory
containing the package file). Otherwise only the local Hugging Face cache
is inspected. Downloads must be started explicitly from Model Catalogue
Models, so generation can never silently transfer the 4.73 GiB package.
"""
override = os.environ.get("OMNIVOICE_AUDIOCPP_MODEL", "").strip()
if override:
cand = Path(override)
if cand.is_file():
return _materialize_gguf_cache_path(cand)
if cand.is_dir():
inner = cand / package_filename()
if inner.is_file():
return _materialize_gguf_cache_path(inner)
raise RuntimeError(
f"OMNIVOICE_AUDIOCPP_MODEL={override} is not a GGUF file or a "
"directory containing one."
)
from huggingface_hub import snapshot_download
from huggingface_hub.utils import LocalEntryNotFoundError
try:
cached = Path(
snapshot_download(
repo_id=HF_MODEL_REPO,
# Full immutable commit SHA declared above; Bandit cannot follow
# the module constant through this call.
revision=HF_MODEL_REVISION, # nosec B615
allow_patterns=[f"{PACKAGE_DIR}/{package_filename()}"],
local_files_only=True,
)
)
except (LocalEntryNotFoundError, OSError) as exc:
raise RuntimeError(
"Breeze-TTS-2 is not installed. Install the audio.cpp Breeze-TTS-2 "
"model from the engine's Weights list in Model Catalogue, or set "
"OMNIVOICE_AUDIOCPP_MODEL to an existing GGUF file."
) from exc
model_file = cached / PACKAGE_DIR / package_filename()
if not model_file.is_file():
raise RuntimeError(
f"Breeze-TTS-2 package {package_filename()} is not completely "
"installed. Reinstall it from the engine's Weights list in Model Catalogue."
)
return _materialize_gguf_cache_path(model_file)
__all__ = [
"AudioCPPDevice",
"AudioCPPSelection",
"BACKEND_ENV",
"BIN_ENV",
"DEFAULT_PACKAGE",
"DEFAULT_PORT",
"DEVICE_ENV",
"DIR_ENV",
"FAMILY",
"HF_MODEL_REPO",
"HF_MODEL_REVISION",
"MODEL_ID",
"PACKAGE_DIR",
"PACKAGE_ENV",
"PORT_ENV",
"VERSION",
"_materialize_gguf_cache_path",
"binary_name",
"default_asset",
"default_asset_size",
"invalidate",
"is_installed",
"managed_runtime_dir",
"package_filename",
"platform_slug",
"parse_device_list",
"probe_devices",
"resolve_compute_selection",
"resolve_model_file",
"resolve_server_binary",
"runtime_targets",
"server_port",
]
+4 -4
View File
@@ -56,15 +56,15 @@ class Confucius4Backend(SubprocessBackend):
id = "confucius4-tts"
display_name = (
"Confucius4-TTS (LLM, 14 langs, cross-lingual zero-shot clone, CUDA/CPU, Apache-2.0)"
"Confucius4-TTS (LLM, 14 langs, cross-lingual zero-shot clone, Apache-2.0)"
)
supports_voice_design = False # timbre comes from a reference clip
# Upstream vocoder rate (config target_sample_rate) — confirmed 22 050 Hz by
# a live run (2026-07-02); still re-read from the sidecar's ready/audio frames.
_DEFAULT_SAMPLE_RATE = 22050
# CUDA fast path + CPU fallback, both exercised (CPU end-to-end validated).
# No MPS claim — upstream has no Metal path.
gpu_compat = ("cuda", "cpu")
# Match device propagation into upstream .to(device). XPU/NPU routing is
# contract-tested, not a claim of physical-hardware synthesis validation.
gpu_compat = ("cuda", "rocm", "xpu", "npu", "cpu")
@classmethod
def is_available(cls) -> tuple[bool, str]:
+13 -2
View File
@@ -104,7 +104,7 @@ def _ensure_clone_on_sys_path() -> None:
def _load_model(stdout):
"""Cold-construct the Confucius4 model (CUDA, else CPU — both validated)."""
"""Cold-construct using an available torch accelerator, with CPU fallback."""
global _model
if _model is not None:
return _model
@@ -115,7 +115,18 @@ def _load_model(stdout):
import torch
from confuciustts.cli.inference import ConfuciusTTS # type: ignore[import-not-found]
device = "cuda" if torch.cuda.is_available() else "cpu"
try:
# Existing manually provisioned venvs may predate torch.accelerator.
current_accelerator = getattr(getattr(torch, "accelerator", None), "current_accelerator", None)
if current_accelerator is None:
device = torch.device("cuda") if torch.cuda.is_available() else None
else:
device = current_accelerator(check_available=True)
device = device.type if device is not None else "cpu" # 'cuda', 'npu', 'mps', 'xpu', 'cpu'
except Exception:
device = "cpu" # Broken accelerator drivers must not block CPU loading.
if device == "mps":
device = "cpu" # MPS was slower than CPU in the existing validation run
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 50})
_model = ConfuciusTTS(config_path=_config_path(), device=device)
@@ -0,0 +1,102 @@
"""cosyvoice-subprocess: CosyVoice 3 from its own venv (one-click install).
The in-process engine needs CosyVoice importable from VoiceStudio's own
interpreter, which upstream's setup (its own Python 3.10 environment, pins
that conflict with the app's) never provides. The one-click installer clones a
reviewed CosyVoice commit and the Matcha-TTS code it vendors as a submodule
into ``DATA_DIR/engines/cosyvoice/``, builds a Python 3.10 venv from a trimmed
requirements list (``requirements.txt`` beside this file), downloads the
CosyVoice 3 weights, and this class runs the model there in a sidecar.
The engine id stays ``cosyvoice``. ``tts_backend._effective_backend_class``
resolves to this class once that venv exists, and to the in-process
``CosyVoiceBackend`` otherwise, so an existing source installation keeps
working as it did.
"""
from __future__ import annotations
import math
import os
from pathlib import Path
from services.subprocess_backend import SubprocessBackend
VENV_ENV_VAR = "OMNIVOICE_COSYVOICE_DIR"
#: Where the installer puts the CosyVoice 3 weights, inside the checkout.
MANAGED_MODEL_SUBDIR = "pretrained_models/Fun-CosyVoice3-0.5B"
def own_venv_python() -> "Path | None":
"""The venv the one-click installer made for CosyVoice, if any."""
from services.sidecar_install import engine_venv_python
return engine_venv_python(VENV_ENV_VAR)
class CosyVoiceSubprocessBackend(SubprocessBackend):
"""CosyVoice in a killable sidecar running the engine's own venv."""
id = "cosyvoice"
display_name = "CosyVoice 3 (9 langs, zero-shot, instruct, Apache-2.0)"
gpu_compat = ("cuda", "cpu")
_DEFAULT_SAMPLE_RATE = 24_000
@classmethod
def is_available(cls) -> tuple[bool, str]:
if own_venv_python() is None:
return False, (
"cosyvoice package not installed. Install it from "
"Model Catalogue → Engines."
)
return True, "ready"
@classmethod
def venv_python(cls) -> Path:
py = own_venv_python()
if py is None:
raise RuntimeError(
"CosyVoice's environment is missing. Reinstall it from "
"Model Catalogue → Engines."
)
return py
@classmethod
def sidecar_script(cls) -> Path:
return Path(__file__).resolve().parent / "main.py"
@property
def recv_timeout_s(self) -> float:
# Loading the model and its text normalizers takes a while on a cold
# start; the sidecar heartbeats progress frames meanwhile.
try:
v = float(os.environ.get("OMNIVOICE_COSYVOICE_RECV_TIMEOUT_S", "900"))
except (TypeError, ValueError):
return 900.0
if not math.isfinite(v): # reject inf/nan so the deadline can't be disabled
return 900.0
return max(30.0, v)
@property
def sample_rate(self) -> int:
# The sidecar resamples to this rate if a model ever reports another.
return self._DEFAULT_SAMPLE_RATE
@property
def supported_languages(self) -> list[str]:
from services.tts_backend import CosyVoiceBackend
return CosyVoiceBackend.supported_languages.fget(self)
def model_identity(self) -> str:
# v1/v2/v3 share the one "cosyvoice" id; the folder name tells them
# apart, as it does for the in-process engine.
model_dir = os.environ.get("OMNIVOICE_COSYVOICE_MODEL") or MANAGED_MODEL_SUBDIR
return os.path.basename(os.path.normpath(model_dir))
__all__ = [
"CosyVoiceSubprocessBackend",
"MANAGED_MODEL_SUBDIR",
"VENV_ENV_VAR",
"own_venv_python",
]
@@ -0,0 +1,299 @@
"""cosyvoice sidecar: CosyVoice in the engine's own venv (one-click install).
Launched as ``<engine venv python> main.py`` by CosyVoiceSubprocessBackend.
The venv holds only CosyVoice's own dependencies, so this file imports nothing
from the app. CosyVoice is not a package: like upstream's own ``example.py``,
this puts the checkout and its ``third_party/Matcha-TTS`` on ``sys.path``.
The mode mapping mirrors the in-process CosyVoiceBackend. For a CosyVoice 3
model, prompts also take the system-prompt prefix upstream's v3 examples use
(``You are a helpful assistant.<|endofprompt|>``), and with no reference clip
the model speaks in the voice of upstream's own sample prompt, because v3
ships no built-in speakers.
Wire protocol: identical to the other sidecars (engines/pockettts/main.py).
"""
from __future__ import annotations
import base64
import contextlib
import json
import os
import re
import struct
import sys
import threading
import traceback
# Mirrors services/subprocess_backend.py::MAX_FRAME_BYTES.
MAX_FRAME_BYTES = 64 * 1024 * 1024
#: The rate the engine reports; a model's output is resampled to it if needed.
COSYVOICE_SAMPLE_RATE = 24_000
_HEARTBEAT_S = 5.0
#: ref_audio must be a local file path, not a URL (local-first; no SSRF).
_URL_RE = re.compile(r"^[a-z][a-z0-9+.\-]*://", re.IGNORECASE)
#: Where the installer puts the CosyVoice 3 weights, inside the checkout.
_MANAGED_MODEL_SUBDIR = ("pretrained_models", "Fun-CosyVoice3-0.5B")
#: Upstream's sample prompt, the default voice for a v3 model.
_DEFAULT_PROMPT_CLIP = ("asset", "zero_shot_prompt.wav")
_ENDOFPROMPT = "<|endofprompt|>"
_SYSTEM_PROMPT = "You are a helpful assistant."
#: Cross-lingual language tags for v1/v2 models (the in-process mapping).
_LANG_TAGS = {
"zh": "<|zh|>", "en": "<|en|>", "ja": "<|ja|>",
"ko": "<|ko|>", "yue": "<|yue|>", "de": "<|de|>",
"es": "<|es|>", "fr": "<|fr|>", "it": "<|it|>",
"ru": "<|ru|>",
}
_MODEL = None
# -- wire protocol -----------------------------------------------------------
_send_lock = threading.Lock()
def _send(stream, obj: dict) -> None:
body = json.dumps(obj, separators=(",", ":")).encode("utf-8")
with _send_lock:
stream.write(struct.pack("!I", len(body)))
stream.write(body)
stream.flush()
def _recv(stream):
header = stream.read(4)
if len(header) < 4:
return None # EOF
(n,) = struct.unpack("!I", header)
if n > MAX_FRAME_BYTES:
raise IOError(f"frame too large: {n}")
body = bytearray()
while len(body) < n:
chunk = stream.read(n - len(body))
if not chunk:
raise IOError("short read")
body.extend(chunk)
return json.loads(bytes(body).decode("utf-8"))
def _measure_vram_mb() -> float:
try:
import torch # noqa: PLC0415
if torch.cuda.is_available():
return float(torch.cuda.memory_allocated()) / (1024 * 1024)
except Exception: # noqa: BLE001 — a probe, never fatal
pass
return 0.0
@contextlib.contextmanager
def _heartbeat(stdout, stage: str):
stop = threading.Event()
def beat() -> None:
pct = 1
while not stop.wait(_HEARTBEAT_S):
pct = min(pct + 1, 99)
_send(stdout, {"op": "progress", "stage": stage, "percent": pct})
thread = threading.Thread(target=beat, daemon=True)
thread.start()
try:
yield
finally:
stop.set()
thread.join(timeout=_HEARTBEAT_S + 1)
# -- loading -----------------------------------------------------------------
def _checkout() -> str:
path = os.environ.get("OMNIVOICE_COSYVOICE_DIR")
if not path:
raise RuntimeError(
"OMNIVOICE_COSYVOICE_DIR is not set. Reinstall CosyVoice from "
"Model Catalogue → Engines."
)
return path
def _model_dir(checkout: str) -> str:
override = os.environ.get("OMNIVOICE_COSYVOICE_MODEL")
if not override:
return os.path.join(checkout, *_MANAGED_MODEL_SUBDIR)
if not os.path.isdir(override):
# Falling back to the installed model would synthesize with a model
# and voice the user did not choose, while model_identity() still
# named theirs. Say what is wrong instead.
raise RuntimeError(
f"OMNIVOICE_COSYVOICE_MODEL points at {override}, which is not a "
"folder. Point it at a CosyVoice model folder, or clear it to use "
"the model the one-click install downloaded."
)
return override
def _load_model(stdout):
global _MODEL
if _MODEL is not None:
return _MODEL
checkout = _checkout()
model_dir = _model_dir(checkout)
# Never hand AutoModel a folder that is missing: it would treat the name as
# a ModelScope id and start a download the user never asked for.
if not os.path.isdir(model_dir):
raise RuntimeError(
f"CosyVoice model folder is missing ({model_dir}). Reinstall "
"CosyVoice from Model Catalogue → Engines."
)
for path in (os.path.join(checkout, "third_party", "Matcha-TTS"), checkout):
if path not in sys.path:
sys.path.insert(0, path)
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 0})
with _heartbeat(stdout, "loading_model"):
from cosyvoice.cli.cosyvoice import AutoModel # type: ignore[import-not-found] # noqa: PLC0415
_MODEL = AutoModel(model_dir=model_dir)
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 100})
return _MODEL
# -- synthesis ---------------------------------------------------------------
def _is_v3(model) -> bool:
return type(model).__name__ == "CosyVoice3"
def _v3_prompt(text: str) -> str:
"""v3 wants the system prompt ahead of the prompt transcript or text."""
return text if _ENDOFPROMPT in text else f"{_SYSTEM_PROMPT}{_ENDOFPROMPT}{text}"
def _instruct(text: str, v3: bool) -> str:
if not text.endswith(_ENDOFPROMPT):
text = f"{text}{_ENDOFPROMPT}"
if v3 and not text.startswith(_SYSTEM_PROMPT):
text = f"{_SYSTEM_PROMPT} {text}"
return text
def _lang_tag(language) -> str:
if not language:
return ""
full = str(language).lower()
return _LANG_TAGS.get(full) or _LANG_TAGS.get(full[:2], "")
def _run_model(model, msg: dict):
text = msg.get("text")
ref_audio = msg.get("ref_audio") or None
ref_text = msg.get("ref_text") or None
instruct = msg.get("instruct") or None
v3 = _is_v3(model)
if not ref_audio and v3:
ref_audio = os.path.join(_checkout(), *_DEFAULT_PROMPT_CLIP)
ref_text = None # the sample's transcript is not ours to supply
if instruct and ref_audio:
return model.inference_instruct2(text, _instruct(instruct, v3), ref_audio, stream=False)
if ref_audio and ref_text:
prompt_text = _v3_prompt(ref_text) if v3 else ref_text
return model.inference_zero_shot(text, prompt_text, ref_audio, stream=False)
if ref_audio:
tts_text = _v3_prompt(text) if v3 else f"{_lang_tag(msg.get('language'))}{text}"
return model.inference_cross_lingual(tts_text, ref_audio, stream=False)
speakers = model.list_available_spks()
if not speakers:
raise ValueError("This CosyVoice model has no built-in voices; pass a reference clip.")
return model.inference_sft(text, speakers[0], stream=False)
def _mono_pcm_b64(chunks, sample_rate: int) -> tuple[str, int]:
import numpy as np # noqa: PLC0415
import torch # noqa: PLC0415
pieces = [c["tts_speech"] for c in chunks if c.get("tts_speech") is not None]
if not pieces:
raise RuntimeError("CosyVoice produced no audio")
wav = torch.cat([torch.as_tensor(p, dtype=torch.float32).reshape(-1) for p in pieces])
if sample_rate != COSYVOICE_SAMPLE_RATE:
import torchaudio # noqa: PLC0415
wav = torchaudio.functional.resample(wav, sample_rate, COSYVOICE_SAMPLE_RATE)
arr = np.clip(wav.detach().cpu().numpy(), -1.0, 1.0)
pcm = (arr * 32767.0).astype(np.int16).tobytes()
return base64.b64encode(pcm).decode("ascii"), int(arr.shape[-1])
def _handle_synthesize(msg: dict, stdout) -> None:
text = msg.get("text")
if not text or not isinstance(text, str):
raise ValueError("synthesize: missing or non-string 'text'")
ref_audio = msg.get("ref_audio") or None
if ref_audio and _URL_RE.match(str(ref_audio)):
raise ValueError(
"ref_audio must be a local file path; URLs are not accepted (local-first)."
)
model = _load_model(stdout)
chunks = list(_run_model(model, msg))
sample_rate = int(getattr(model, "sample_rate", COSYVOICE_SAMPLE_RATE) or COSYVOICE_SAMPLE_RATE)
pcm_b64, n_samples = _mono_pcm_b64(chunks, sample_rate)
_send(stdout, {
"op": "audio",
"audio_pcm_b64": pcm_b64,
"sample_rate": COSYVOICE_SAMPLE_RATE,
"n_samples": n_samples,
})
# -- main loop ---------------------------------------------------------------
def main() -> int:
stdin = sys.stdin.buffer
# Frames go down a PRIVATE fd, and fd 1 is pointed at stderr (#1428): the
# libraries this loads print to fd 1, and those bytes would otherwise
# interleave with the length-prefixed frames.
_frame_fd = os.dup(1)
os.dup2(2, 1)
stdout = os.fdopen(_frame_fd, "wb")
_send(stdout, {"op": "ready", "engine": "cosyvoice", "sample_rate": COSYVOICE_SAMPLE_RATE})
while True:
try:
msg = _recv(stdin)
except Exception as exc: # noqa: BLE001
_send(stdout, {
"op": "error",
"stage": "recv",
"message": f"{type(exc).__name__}: {exc}",
"traceback": traceback.format_exc(),
})
return 1
if msg is None:
return 0
op = msg.get("op") if isinstance(msg, dict) else None
try:
if op == "ping":
_send(stdout, {"op": "pong", "vram_mb": _measure_vram_mb()})
elif op == "synthesize":
_handle_synthesize(msg, stdout)
elif op == "shutdown":
return 0
else:
_send(stdout, {"op": "error", "stage": "dispatch", "message": f"unknown op: {op!r}"})
except Exception as exc: # noqa: BLE001
_send(stdout, {
"op": "error",
"stage": op or "unknown",
"message": f"{type(exc).__name__}: {exc}",
"traceback": traceback.format_exc(),
})
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,55 @@
# CosyVoice 3 inference requirements for VoiceStudio's one-click install.
#
# Derived from upstream's requirements.txt at FunAudioLLM/CosyVoice@074ca6dc
# (2026-05-25) and trimmed to what synthesis needs. What differs, and why:
#
# - No --extra-index-url lines. Upstream adds PyTorch's cu121 index and a
# third-party Azure DevOps feed for onnxruntime-gpu. The installer chooses
# the PyTorch build itself and uses no third-party index.
# - torch and torchaudio are not listed here. The installer pins the 2.7.0
# pair per host: +cu128 on NVIDIA GPUs, +cpu on other Windows and Linux
# machines, the regular build on macOS. Upstream's 2.3.1 exists only for
# CUDA 12.1 and cannot run on RTX 50-series GPUs.
# - Raised past published security advisories, where upstream pins an
# affected release: diffusers, hydra-core, lightning, modelscope, onnx,
# protobuf and transformers. This exact set was installed on Windows and
# passes the installer's import probe (tests pin the advisory floors).
# transformers 5 was also checked against CosyVoice's own tokenizer (token
# ids identical to 4.57.6) and its cached step-by-step decoding.
# - Dropped: deepspeed and tensorrt-cu12* (Linux-only acceleration);
# onnxruntime-gpu (the onnxruntime below serves every host); pyworld and
# pyarrow (not needed to import or run CosyVoice; pyworld has no Python 3.10
# wheels for Linux or macOS, pyarrow carried an advisory); wetext (its data
# is published only on ModelScope, which rate-limits downloads, and a
# half-downloaded normaliser failed silently; CosyVoice reads text as
# written without it); and the web UI, server, training and download tools
# (fastapi, fastapi-cli, gradio, grpcio, grpcio-tools, uvicorn, tensorboard,
# gdown, wget). Upstream's full list cannot be installed together anyway:
# its fastapi pin conflicts.
# - openai-whisper 20231117 -> 20250625: the older release needs
# pkg_resources at build time and fails to build. CosyVoice only uses its
# mel-spectrogram frontend.
#
# Everything else keeps upstream's exact pin. tests/test_cosyvoice_subprocess.py
# checks this file.
conformer==0.3.2
diffusers==0.39.0
hydra-core==1.3.6
HyperPyYAML==1.2.3
inflect==7.3.1
librosa==0.10.2
lightning==2.6.6
matplotlib==3.7.5
modelscope==1.40.0
networkx==3.1
numpy==1.26.4
omegaconf==2.3.0
onnx==1.22.0
onnxruntime==1.18.0
openai-whisper==20250625
protobuf==5.29.6
pydantic==2.7.0
rich==13.7.1
soundfile==0.12.1
transformers==5.10.1
x-transformers==2.11.24
+6 -1
View File
@@ -115,7 +115,12 @@ def _load_runtime(stdout):
from dots_tts.runtime import DotsTtsRuntime # type: ignore[import-not-found]
repo = os.environ.get("OMNIVOICE_DOTS_TTS_MODEL", _DEFAULT_REPO)
default_precision = "bfloat16" if torch.cuda.is_available() else "float32"
# Match DotsTtsRuntime's own CUDA/CPU selection. Its _check_torch_env
# rejects half precision without CUDA, even when an XPU/NPU is available.
try:
default_precision = "bfloat16" if torch.cuda.is_available() else "float32"
except Exception:
default_precision = "float32" # Probe failure must not force half precision.
precision = os.environ.get("OMNIVOICE_DOTS_TTS_PRECISION", default_precision)
optimize = os.environ.get("OMNIVOICE_DOTS_TTS_OPTIMIZE", "0") == "1"
@@ -0,0 +1,89 @@
"""moss-tts-nano-subprocess: MOSS-TTS-Nano from its own venv (one-click install).
The in-process engine needs ``moss_tts_nano`` installed into VoiceStudio's own
environment, with upstream's exact pins (torch 2.7.0, transformers 4.57.1)
landing there too. It also looks for a model class the package no longer
exports: at the commit pinned here, ``moss_tts_nano`` exports only
``__version__``, and the entry point is the top-level
``moss_tts_nano_runtime.NanoTTSService``. The one-click installer clones that
reviewed commit into ``DATA_DIR/engines/moss-tts-nano/`` with its own venv,
and this class runs upstream's runtime there in a sidecar.
The engine id stays ``moss-tts-nano``. ``tts_backend._effective_backend_class``
resolves to this class once that venv exists, and to the in-process
``MossTTSNanoBackend`` otherwise.
"""
from __future__ import annotations
import math
import os
from pathlib import Path
from services.subprocess_backend import SubprocessBackend
VENV_ENV_VAR = "OMNIVOICE_MOSS_TTS_NANO_DIR"
def own_venv_python() -> "Path | None":
"""The venv the one-click installer made for MOSS-TTS-Nano, if any."""
from services.sidecar_install import engine_venv_python
return engine_venv_python(VENV_ENV_VAR)
class MossTTSNanoSubprocessBackend(SubprocessBackend):
"""MOSS-TTS-Nano in a killable sidecar running the engine's own venv."""
id = "moss-tts-nano"
display_name = "MOSS-TTS-Nano (20 langs, CPU realtime, 48 kHz)"
gpu_compat = ("cuda", "cpu")
_DEFAULT_SAMPLE_RATE = 48_000
@classmethod
def is_available(cls) -> tuple[bool, str]:
if own_venv_python() is None:
return False, (
"moss_tts_nano package not installed. Install it from "
"Model Catalogue."
)
return True, "ready"
@classmethod
def venv_python(cls) -> Path:
py = own_venv_python()
if py is None:
raise RuntimeError(
"MOSS-TTS-Nano's environment is missing. Reinstall it from "
"Model Catalogue."
)
return py
@classmethod
def sidecar_script(cls) -> Path:
return Path(__file__).resolve().parent / "main.py"
@property
def recv_timeout_s(self) -> float:
# A cold load downloads the model and its audio tokenizer; the sidecar
# heartbeats progress frames meanwhile, and each re-arms this deadline.
try:
v = float(os.environ.get("OMNIVOICE_MOSS_TTS_NANO_RECV_TIMEOUT_S", "900"))
except (TypeError, ValueError):
return 900.0
if not math.isfinite(v): # reject inf/nan so the deadline can't be disabled
return 900.0
return max(30.0, v)
@property
def sample_rate(self) -> int:
# The sidecar resamples to this rate if upstream ever returns another.
return self._DEFAULT_SAMPLE_RATE
@property
def supported_languages(self) -> list[str]:
from services.tts_backend import MossTTSNanoBackend
return MossTTSNanoBackend.supported_languages.fget(self)
__all__ = ["MossTTSNanoSubprocessBackend", "VENV_ENV_VAR", "own_venv_python"]
@@ -0,0 +1,254 @@
"""moss-tts-nano sidecar: MOSS-TTS-Nano in the engine's own venv (one-click install).
Launched as ``<engine venv python> main.py`` by MossTTSNanoSubprocessBackend.
The venv holds the reviewed upstream checkout, installed editable, and its
pinned dependencies, so this file imports nothing from the app. It drives the
runtime that checkout ships, ``moss_tts_nano_runtime.NanoTTSService``.
Wire protocol: identical to the other sidecars (engines/pockettts/main.py).
Progress frames are sent while the model loads and through the first
synthesis, which is when upstream fetches its audio tokenizer. Later calls
send none, so the parent's watchdog still catches a generation that wedges.
"""
from __future__ import annotations
import base64
import contextlib
import json
import os
import re
import struct
import sys
import tempfile
import threading
import time
import traceback
# Mirrors services/subprocess_backend.py::MAX_FRAME_BYTES.
MAX_FRAME_BYTES = 64 * 1024 * 1024
#: The rate the engine reports; upstream's output is resampled to it if needed.
NANO_SAMPLE_RATE = 48_000
_HEARTBEAT_S = 5.0
#: ref_audio must be a local file path, not a URL (local-first; no SSRF).
_URL_RE = re.compile(r"^[a-z][a-z0-9+.\-]*://", re.IGNORECASE)
#: A download failure worth retrying; anything else propagates at once.
_TRANSIENT_MARKERS = (
"connection", "timed out", "timeout", "peer closed", "incomplete",
"remoteprotocolerror", "temporarily unavailable",
)
_SERVICE = None
_WARM = False
# -- wire protocol -----------------------------------------------------------
_send_lock = threading.Lock()
def _send(stream, obj: dict) -> None:
body = json.dumps(obj, separators=(",", ":")).encode("utf-8")
with _send_lock:
stream.write(struct.pack("!I", len(body)))
stream.write(body)
stream.flush()
def _recv(stream):
header = stream.read(4)
if len(header) < 4:
return None # EOF
(n,) = struct.unpack("!I", header)
if n > MAX_FRAME_BYTES:
raise IOError(f"frame too large: {n}")
body = bytearray()
while len(body) < n:
chunk = stream.read(n - len(body))
if not chunk:
raise IOError("short read")
body.extend(chunk)
return json.loads(bytes(body).decode("utf-8"))
def _measure_vram_mb() -> float:
try:
import torch # noqa: PLC0415
if torch.cuda.is_available():
return float(torch.cuda.memory_allocated()) / (1024 * 1024)
except Exception: # noqa: BLE001 — a probe, never fatal
pass
return 0.0
# -- loading -----------------------------------------------------------------
def _with_retries(action):
"""Run ``action``, retrying a transient download failure with a short
backoff, the way the app's own loader does for in-process engines."""
try:
attempts = max(1, int(os.environ.get("OMNIVOICE_MODEL_LOAD_RETRIES", "3")))
except ValueError:
attempts = 3
for attempt in range(1, attempts + 1):
try:
return action()
except Exception as exc: # noqa: BLE001 — classified below
text = f"{type(exc).__name__}: {exc}".lower()
if attempt == attempts or not any(m in text for m in _TRANSIENT_MARKERS):
raise
time.sleep(2.0 * attempt)
raise AssertionError("unreachable")
@contextlib.contextmanager
def _heartbeat(stdout, stage: str):
"""Progress frames every few seconds while a download may be running."""
stop = threading.Event()
def beat() -> None:
pct = 1
while not stop.wait(_HEARTBEAT_S):
pct = min(pct + 1, 99)
_send(stdout, {"op": "progress", "stage": stage, "percent": pct})
thread = threading.Thread(target=beat, daemon=True)
thread.start()
try:
yield
finally:
stop.set()
thread.join(timeout=_HEARTBEAT_S + 1)
def _load_service(stdout):
global _SERVICE
if _SERVICE is not None:
return _SERVICE
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 0})
with _heartbeat(stdout, "loading_model"):
from moss_tts_nano.defaults import ( # type: ignore[import-not-found] # noqa: PLC0415
DEFAULT_AUDIO_TOKENIZER_PATH,
DEFAULT_CHECKPOINT_PATH,
)
from moss_tts_nano_runtime import NanoTTSService # type: ignore[import-not-found] # noqa: PLC0415
service = NanoTTSService(
checkpoint_path=os.environ.get("OMNIVOICE_MOSS_TTS_MODEL", DEFAULT_CHECKPOINT_PATH),
audio_tokenizer_path=os.environ.get(
"OMNIVOICE_MOSS_TTS_TOKENIZER", DEFAULT_AUDIO_TOKENIZER_PATH
),
# Upstream writes every synthesis to a file; keep them out of the
# install folder (and see _handle_synthesize: one file, reused).
output_dir=tempfile.mkdtemp(prefix="moss-tts-nano-"),
)
_with_retries(lambda: service.preload(load_model=True))
_SERVICE = service
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 100})
return _SERVICE
# -- synthesis ---------------------------------------------------------------
def _mono_pcm_b64(result: dict) -> tuple[str, int]:
"""Upstream's waveform, downmixed to mono at NANO_SAMPLE_RATE, as base64
int16 PCM. The in-process engine downmixed the same way."""
import numpy as np # noqa: PLC0415
wav = np.asarray(result["waveform_numpy"], dtype=np.float32)
if wav.ndim == 2: # (samples, channels): the layout upstream returns
wav = wav.mean(axis=1)
sr = int(result.get("sample_rate") or NANO_SAMPLE_RATE)
if sr != NANO_SAMPLE_RATE:
import torch # noqa: PLC0415
import torchaudio # noqa: PLC0415
wav = torchaudio.functional.resample(torch.from_numpy(wav), sr, NANO_SAMPLE_RATE).numpy()
wav = np.clip(wav, -1.0, 1.0)
pcm = (wav * 32767.0).astype(np.int16).tobytes()
return base64.b64encode(pcm).decode("ascii"), int(wav.shape[-1])
def _handle_synthesize(msg: dict, stdout) -> None:
global _WARM
text = msg.get("text")
if not text or not isinstance(text, str):
raise ValueError("synthesize: missing or non-string 'text'")
ref_audio = msg.get("ref_audio") or None
if ref_audio and _URL_RE.match(str(ref_audio)):
raise ValueError(
"ref_audio must be a local file path; URLs are not accepted (local-first)."
)
service = _load_service(stdout)
kwargs = {
"text": text,
# Reference cloning, as the in-process engine did; with no clip,
# upstream uses its default voice preset.
"mode": "voice_clone",
"prompt_audio_path": ref_audio,
"output_audio_path": os.path.join(str(service.output_dir), "last.wav"),
}
if _WARM:
result = service.synthesize(**kwargs)
else:
with _heartbeat(stdout, "loading_model"):
result = _with_retries(lambda: service.synthesize(**kwargs))
_WARM = True
pcm_b64, n_samples = _mono_pcm_b64(result)
_send(stdout, {
"op": "audio",
"audio_pcm_b64": pcm_b64,
"sample_rate": NANO_SAMPLE_RATE,
"n_samples": n_samples,
})
# -- main loop ---------------------------------------------------------------
def main() -> int:
stdin = sys.stdin.buffer
# Frames go down a PRIVATE fd, and fd 1 is pointed at stderr (#1428): the
# libraries this loads print to fd 1, and those bytes would otherwise
# interleave with the length-prefixed frames.
_frame_fd = os.dup(1)
os.dup2(2, 1)
stdout = os.fdopen(_frame_fd, "wb")
_send(stdout, {"op": "ready", "engine": "moss-tts-nano", "sample_rate": NANO_SAMPLE_RATE})
while True:
try:
msg = _recv(stdin)
except Exception as exc: # noqa: BLE001
_send(stdout, {
"op": "error",
"stage": "recv",
"message": f"{type(exc).__name__}: {exc}",
"traceback": traceback.format_exc(),
})
return 1
if msg is None:
return 0
op = msg.get("op") if isinstance(msg, dict) else None
try:
if op == "ping":
_send(stdout, {"op": "pong", "vram_mb": _measure_vram_mb()})
elif op == "synthesize":
_handle_synthesize(msg, stdout)
elif op == "shutdown":
return 0
else:
_send(stdout, {"op": "error", "stage": "dispatch", "message": f"unknown op: {op!r}"})
except Exception as exc: # noqa: BLE001
_send(stdout, {
"op": "error",
"stage": op or "unknown",
"message": f"{type(exc).__name__}: {exc}",
"traceback": traceback.format_exc(),
})
if __name__ == "__main__":
sys.exit(main())
+11 -15
View File
@@ -29,15 +29,11 @@ Do NOT import ``main.py`` from the parent process — it runs under a
different venv (``transformers==5.0.0``) and importing it in-process would
re-introduce the exact conflict this isolation exists to avoid.
Hardware honesty (cross-platform rule): MOSS-TTS-v1.5's upstream documents
only CUDA and CPU. There is **no documented or tested MPS path** the
custom ``trust_remote_code`` modelling code and the separate audio
tokenizer are unverified on Apple Silicon. We therefore advertise
``gpu_compat = ("cuda", "cpu")`` and the sidecar selects ``cuda`` when
present else ``cpu`` it never silently routes to MPS where it might
crash. On Apple Silicon the engine honestly resolves to CPU (slow but
correct), and the engine is opt-in regardless, so it never becomes a
broken default on any platform.
Hardware routing follows the sidecar's runtime-available PyTorch accelerator:
CUDA/ROCm, XPU, or a registered NPU. MPS remains excluded; CPU is the fallback.
XPU/NPU routing is covered with mocked device contracts, not physical-hardware
synthesis certification; users need a compatible torch/vendor runtime in the
isolated engine venv.
"""
from __future__ import annotations
@@ -85,13 +81,13 @@ class MossTTSV15Backend(SubprocessBackend):
id = "moss-tts-v15"
display_name = (
"MOSS-TTS-v1.5 (8B, 31 langs, zero-shot clone, CUDA/CPU, Apache-2.0)"
"MOSS-TTS-v1.5 (8B, 31 langs, zero-shot clone, Apache-2.0)"
)
supports_voice_design = False # requires ref audio for timbre cloning
_DEFAULT_SAMPLE_RATE = 24000
# Honest hardware surface: upstream documents CUDA + CPU only. MPS is
# undocumented / untested, so we do NOT claim it (cross-platform rule).
gpu_compat = ("cuda", "cpu")
# Accelerator routing requires its matching runtime in the isolated venv.
# MPS remains untested and is deliberately excluded.
gpu_compat = ("cuda", "rocm", "xpu", "npu", "cpu")
# ── availability ───────────────────────────────────────────────────────
@@ -111,7 +107,7 @@ class MossTTSV15Backend(SubprocessBackend):
return False, (
"MOSS-TTS-v1.5 venv not found. Set OMNIVOICE_MOSS_TTS_V15_DIR "
"to your MOSS-TTS clone (the directory containing pyproject.toml) "
"and restart VoiceStudio. CUDA or CPU only (no MPS). See "
"and restart VoiceStudio. Install the matching PyTorch runtime. See "
"docs/engines/moss-tts-v15.md for the full install walk-through."
)
if not MOSS_TTS_V15_SIDECAR_SCRIPT.exists():
@@ -119,7 +115,7 @@ class MossTTSV15Backend(SubprocessBackend):
"MOSS-TTS-v1.5 sidecar script missing at "
f"{MOSS_TTS_V15_SIDECAR_SCRIPT} — reinstall VoiceStudio."
)
return True, "ok (CUDA when present, else CPU)"
return True, "ok (runtime-available accelerator or CPU; no MPS)"
@classmethod
def venv_python(cls):
+11 -5
View File
@@ -222,8 +222,7 @@ def _bootstrap_engines_venv(clone_dir: Path) -> Path:
Runs ``uv venv <engines_venv>`` then ``uv pip install --python
<engines_venv>/bin/python -e "<clone>[torch-runtime]"``. Verifies the
result by re-probing the import a successful uv invocation that still
can't import the stack indicates a deeper environment problem (e.g. the
``+cu128`` torch-runtime extra can't resolve on a non-CUDA host) and we
can't import the stack indicates a deeper environment problem, and we
raise with whatever stderr we captured plus a docs pointer.
"""
uv = _locate_uv()
@@ -254,6 +253,8 @@ def _bootstrap_engines_venv(clone_dir: Path) -> Path:
f"{exc.stderr.decode('utf-8', errors='replace') if exc.stderr else exc}"
) from exc
from core.torch_indexes import UV_PIP_CU128_ARGS
python_path = _venv_python_path(_ENGINES_VENV_DIR)
try:
subprocess.run(
@@ -261,6 +262,10 @@ def _bootstrap_engines_venv(clone_dir: Path) -> Path:
uv, "pip", "install",
"--python", str(python_path),
"-e", f"{clone_dir}[torch-runtime]",
# The extra pins torch==2.9.1+cu128, which exists only on
# PyTorch's index — without it this could never resolve, on
# any host (core.torch_indexes).
*UV_PIP_CU128_ARGS,
],
check=True,
timeout=_UV_PIP_INSTALL_TIMEOUT_S,
@@ -270,9 +275,10 @@ def _bootstrap_engines_venv(clone_dir: Path) -> Path:
except subprocess.CalledProcessError as exc:
raise RuntimeError(
"uv pip install -e failed during MOSS-TTS-v1.5 bootstrap "
f"({clone_dir}). On a non-CUDA host the upstream '[torch-runtime]' "
"extra (cu128) cannot resolve — set up the venv manually per "
"docs/engines/moss-tts-v15.md. Error: "
# uv's own error names what failed; the PyTorch index is always
# supplied now, so a guess about the host would only mislead.
f"({clone_dir}). See docs/engines/moss-tts-v15.md for the manual "
"install. Error: "
f"{exc.stderr.decode('utf-8', errors='replace') if exc.stderr else exc}"
) from exc
+22 -7
View File
@@ -138,11 +138,12 @@ _state = None
def _load_model(stdout):
"""Cold-construct the MOSS-TTS-v1.5 processor + model.
Device selection is CUDA-or-CPU only MOSS's upstream documents no MPS
path and the custom ``trust_remote_code`` modelling code is untested on
Apple Silicon, so we never route to MPS where it might crash. dtype is
bf16 on CUDA, fp32 on CPU (bf16 CPU ops are spotty). Emits progress
frames so the parent can surface the multi-GB cold-load latency.
Device selection uses the torch.accelerator API to support any backend
(CUDA, NPU, XPU, etc.) automatically. MPS is excluded MOSS's upstream
``trust_remote_code`` modelling code is untested on Apple Silicon. dtype is
bf16 on GPU-class accelerators, fp32 on CPU (bf16 CPU ops are spotty).
Emits progress frames so the parent can surface the multi-GB cold-load
latency.
"""
global _state
if _state is not None:
@@ -154,8 +155,22 @@ def _load_model(stdout):
from transformers import AutoModel, AutoProcessor
repo, revision = _model_source()
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
# current_accelerator() returns None on CPU-only builds (no accelerator
# compiled in) or when no accelerator is available; fall back to "cpu".
# Existing manually provisioned venvs may predate torch.accelerator.
current_accelerator = getattr(getattr(torch, "accelerator", None), "current_accelerator", None)
try:
if current_accelerator is None:
accel = torch.device("cuda") if torch.cuda.is_available() else None
else:
accel = current_accelerator(check_available=True)
except Exception:
# Optional drivers can fail during probing; CPU loading remains usable.
accel = None
device = accel.type if accel is not None else "cpu" # 'cuda', 'npu', 'mps', 'xpu', 'cpu'
if device == "mps":
device = "cpu" # MOSS is untested on MPS; fall back to CPU for safety
dtype = torch.bfloat16 if device != "cpu" else torch.float32
# "sdpa" works on CUDA + CPU and needs no extra dep. flash_attention_2
# (Ampere+ CUDA, optional flash-attn) is opt-in via env.
attn = os.environ.get("OMNIVOICE_MOSS_TTS_V15_ATTN", "sdpa")
+3 -3
View File
@@ -176,7 +176,7 @@ def _binary_repair_hint() -> str:
f"the bundled GGUF runtime is not usable on this machine — build it "
f"with `scripts/build-omnivoice-tts.sh --platform {_platform_slug()}`, "
f"reinstall VoiceStudio, or switch to the default in-process "
f"OmniVoice engine (Model Catalogue → Engines)"
f"OmniVoice engine (Model Catalogue)"
)
@@ -412,7 +412,7 @@ def _make_backend_class():
f"built — run `scripts/build-omnivoice-tts.sh "
f"--platform {_platform_slug()}`, reinstall "
f"VoiceStudio, or use the default in-process "
f"OmniVoice engine (Model Catalogue → Engines)."
f"OmniVoice engine (Model Catalogue)."
)
# Manifest-based SHA-256 verification (T-04-01).
manifest = _load_checksum_manifest()
@@ -861,7 +861,7 @@ def select_default_engine() -> str:
Returns ``"omnivoice"`` (the existing in-process default) on any
failure. The fallback is deliberately silent a user who hits this
code path still gets a working cloning engine; the failure surfaces
in the Model Catalogue Engines Compatibility Matrix (Plan 02-04) so the
in the Model Catalogue Compatibility Matrix (Plan 02-04) so the
user can investigate if they care to.
"""
cls = _make_backend_class()
+34 -4
View File
@@ -34,6 +34,7 @@ import json
import os
import struct
import sys
import threading
import traceback
# Mirrors backend/services/subprocess_backend.py::MAX_FRAME_BYTES (T-02-01).
@@ -55,6 +56,8 @@ _GEN_KW_ALLOWLIST = (
)
_model = None
_SEND_LOCK = threading.Lock()
_LOAD_HEARTBEAT_S = 5.0
# ── wire protocol ─────────────────────────────────────────────────────────
@@ -62,9 +65,12 @@ _model = None
def _send(stream, obj: dict) -> None:
body = json.dumps(obj, separators=(",", ":")).encode("utf-8")
stream.write(struct.pack("!I", len(body)))
stream.write(body)
stream.flush()
# Progress callbacks and the cold-load heartbeat can write from different
# threads. Keep each frame atomic or their header/body pairs can interleave.
with _SEND_LOCK:
stream.write(struct.pack("!I", len(body)))
stream.write(body)
stream.flush()
def _recv(stream):
@@ -133,11 +139,27 @@ def _load_model(stdout):
# Forward real HF download/weight progress so the parent's recv loop keeps
# its watchdog alive across a slow cold load (the parent consumes these
# {"op": "progress"} frames and re-arms its deadline on each one).
progress = {"percent": 0}
def _on_progress(ev):
pct = ev.get("pct", 0.0)
if pct:
progress["percent"] = min(round(pct * 100), 99)
_send(stdout, {"op": "progress", "stage": "loading_model",
"percent": min(round(pct * 100), 99)})
"percent": progress["percent"]})
# Cached checkpoints produce no download callbacks. Loading and moving a
# model onto MPS can still exceed the normal generation budget, so keep
# both bounded parent watchdogs informed that the child remains alive.
stop_heartbeat = threading.Event()
def _heartbeat():
while not stop_heartbeat.wait(_LOAD_HEARTBEAT_S):
_send(stdout, {
"op": "progress",
"stage": "loading_model",
"percent": progress["percent"],
})
torch = _lazy_torch()
OmniVoice = _lazy_omnivoice()
@@ -146,11 +168,19 @@ def _load_model(stdout):
preload_asr = should_preload_tts_asr()
lid = register_listener(_on_progress)
heartbeat = threading.Thread(
target=_heartbeat,
name="omnivoice-load-heartbeat",
daemon=True,
)
heartbeat.start()
try:
_model = OmniVoice.from_pretrained(
checkpoint, device_map=device, dtype=torch.float16, load_asr=preload_asr,
)
finally:
stop_heartbeat.set()
heartbeat.join()
unregister_listener(lid)
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 100})
return _model
+28 -18
View File
@@ -48,6 +48,15 @@ from services.subprocess_backend import SubprocessBackend
logger = logging.getLogger("omnivoice.engines.pockettts")
_VENV_ENV_VAR = "OMNIVOICE_POCKETTTS_DIR"
def _own_venv_python() -> "Path | None":
"""The venv the one-click installer made for this engine, if any."""
from services.sidecar_install import engine_venv_python
return engine_venv_python(_VENV_ENV_VAR)
if TYPE_CHECKING:
import torch # noqa: F401
@@ -94,7 +103,7 @@ class PocketTTSBackend(SubprocessBackend):
raise RuntimeError(platform_error)
if not self._license_accepted():
raise RuntimeError(
"PocketTTS license not accepted. Review it in Model Catalogue → Engines."
"PocketTTS license not accepted. Review it in Model Catalogue."
)
super().__init__()
@@ -104,7 +113,7 @@ class PocketTTSBackend(SubprocessBackend):
# relying on every caller to evict its cached instance.
if not self._license_accepted():
raise RuntimeError(
"PocketTTS license not accepted. Review it in Model Catalogue → Engines."
"PocketTTS license not accepted. Review it in Model Catalogue."
)
return super().generate(*args, **kwargs)
@@ -114,30 +123,31 @@ class PocketTTSBackend(SubprocessBackend):
# so revocation while waiting cannot reach the sidecar or return audio.
if not self._license_accepted():
raise RuntimeError(
"PocketTTS license not accepted. Review it in Model Catalogue → Engines."
"PocketTTS license not accepted. Review it in Model Catalogue."
)
@classmethod
def is_available(cls) -> tuple[bool, str]:
if platform_error := cls._platform_error():
return False, platform_error
# Optional-dep gate: the pocket-tts wheel is installed only when the user
# opted in. The interpreter is the parent's own (sys.executable), so
# there is no separate venv to validate.
try:
import pocket_tts # type: ignore[import-not-found] # noqa: F401
except Exception as e:
return False, (
f"pocket_tts package not installed or failed to import ({e}). "
f"Enable in Settings -> Engines (uv sync --extra pockettts)."
)
# Installed either into its own venv by the one-click installer, which
# verified `import pocket_tts` there before saving the path, or into the
# app's environment by `uv sync --extra pockettts`.
if _own_venv_python() is None:
try:
import pocket_tts # type: ignore[import-not-found] # noqa: F401
except Exception as e:
return False, (
f"pocket_tts package not installed or failed to import ({e}). "
"Install it from Model Catalogue."
)
# The model repository has an additional gated-access agreement and
# prohibited-use conditions beyond its CC-BY-4.0 license. Keep first
# use behind an explicit local acknowledgement, matching the dialog.
if not cls._license_accepted():
return False, (
"PocketTTS license not accepted. Open Model Catalogue → Engines → "
"PocketTTS license not accepted. Open Model Catalogue → "
"PocketTTS and review the MIT code license, CC-BY-4.0 model "
"license, and gated-access conditions before enabling it."
)
@@ -145,10 +155,10 @@ class PocketTTSBackend(SubprocessBackend):
@classmethod
def venv_python(cls) -> Path:
# Parent interpreter: pocket-tts deps (torch>=2.5, scipy, beartype) sit
# happily at the parent's pins, so this isolates for crash recovery, not
# dependency pins (same rationale as omnivoice-subprocess).
return Path(sys.executable)
# Its own venv when the one-click installer made one. Otherwise the
# parent interpreter, where `uv sync --extra pockettts` installs it
# (its deps sit happily at the parent's pins).
return _own_venv_python() or Path(sys.executable)
@classmethod
def sidecar_script(cls) -> Path:
+24 -13
View File
@@ -48,6 +48,15 @@ if TYPE_CHECKING:
logger = logging.getLogger("omnivoice.supertonic3")
_VENV_ENV_VAR = "OMNIVOICE_SUPERTONIC3_DIR"
def _own_venv_python() -> "Path | None":
"""The venv the one-click installer made for this engine, if any."""
from services.sidecar_install import engine_venv_python
return engine_venv_python(_VENV_ENV_VAR)
# Absolute path to the sidecar script ‑‑ same pattern as IndexTTS's
# ``INDEXTTS_SIDECAR_SCRIPT``. SubprocessBackend spawns it with the
@@ -80,11 +89,11 @@ class Supertonic3Backend(SubprocessBackend):
@classmethod
def venv_python(cls) -> Path:
"""Supertonic-3 lives in the main OmniVoice venv ‑‑ no dedicated
venv. ``sys.executable`` is the parent interpreter, which is the
same Python that ``uv sync --extra supertonic`` populated.
"""Its own venv when the one-click installer made one. Otherwise the
parent interpreter, the same Python ``uv sync --extra supertonic``
populated.
"""
return Path(sys.executable)
return _own_venv_python() or Path(sys.executable)
@classmethod
def sidecar_script(cls) -> Path:
@@ -96,14 +105,16 @@ class Supertonic3Backend(SubprocessBackend):
def is_available(cls) -> tuple[bool, str]:
# 1. Optional-dep gate (TTS-02). The ``supertonic`` wheel is only
# installed when the user opted in via ``--extra supertonic``.
try:
import supertonic # type: ignore[import-not-found] # noqa: F401
except ImportError:
return False, (
"supertonic package not installed. Enable in "
"Model Catalogue → Engines (installs `supertonic` via `uv add --optional "
"supertonic supertonic==1.3.1`)."
)
# Its own venv (made by the one-click installer, which verified the
# import there) or the app's environment (`uv sync --extra`).
if _own_venv_python() is None:
try:
import supertonic # type: ignore[import-not-found] # noqa: F401
except ImportError:
return False, (
"supertonic package not installed. Install it from "
"Model Catalogue."
)
# 2. License acceptance gate (TTS-05). Defence in depth: the
# settings_store helper handles the read; we just refuse
@@ -120,7 +131,7 @@ class Supertonic3Backend(SubprocessBackend):
accepted = False
if not accepted:
return False, (
"Supertonic-3 license not accepted. Open Model Catalogue → Engines → "
"Supertonic-3 license not accepted. Open Model Catalogue → "
"Supertonic-3 and click Accept to enable. "
"(MIT code license + OpenRAIL-M model license.)"
)
+11 -3
View File
@@ -137,9 +137,17 @@ def _resolve_pinned_sha() -> str:
# Final fallback ‑‑ relative import for when the file is invoked
# via ``python backend/engines/supertonic3/sidecar.py`` rather
# than via ``python -m backend.engines.supertonic3.sidecar``.
sys.path.insert(0, str(Path(__file__).resolve().parents[2]))
from engines.supertonic3.constants import PINNED_REVISION_SHA # type: ignore[import-not-found]
return PINNED_REVISION_SHA
# Load constants.py by path. Importing it as `engines.supertonic3…`
# runs the package __init__, which imports the app's backend, and that
# is absent from the engine's own venv (one-click install).
import importlib.util
spec = importlib.util.spec_from_file_location(
"_supertonic3_constants", Path(__file__).resolve().with_name("constants.py"),
)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module) # type: ignore[union-attr]
return module.PINNED_REVISION_SHA
# ── model loading (lazy, on first synthesize) ─────────────────────────────
@@ -0,0 +1,106 @@
"""voxcpm2-subprocess: VoxCPM2 from its own venv (one-click install).
VoxCPM2 used to run only in-process, which meant installing ``voxcpm``, and a
torch of its choosing, into VoiceStudio's own environment. The one-click
installer now gives it a venv under ``DATA_DIR/engines/voxcpm2/``, and this
class runs the model there in a sidecar, so nothing it installs can touch the
app or another engine.
The engine id stays ``voxcpm2``. ``tts_backend._effective_backend_class``
resolves to this class once that venv exists and to the in-process
``VoxCPM2Backend`` otherwise, so an install made with ``pip install voxcpm``
keeps working as it always has. What the app sees is the same: voice design,
48 kHz output, its own mastering, the same languages. The parent still
prepares the reference clip and trims the silent tail, as the in-process
engine does.
"""
from __future__ import annotations
import math
import os
from pathlib import Path
from typing import TYPE_CHECKING
from services.subprocess_backend import SubprocessBackend
if TYPE_CHECKING:
import torch # noqa: F401
VENV_ENV_VAR = "OMNIVOICE_VOXCPM2_DIR"
def own_venv_python() -> "Path | None":
"""The venv the one-click installer made for VoxCPM2, if any."""
from services.sidecar_install import engine_venv_python
return engine_venv_python(VENV_ENV_VAR)
class VoxCPM2SubprocessBackend(SubprocessBackend):
"""VoxCPM2 in a killable sidecar running the engine's own venv."""
id = "voxcpm2"
display_name = "VoxCPM2 (30 langs, studio 48 kHz, voice design)"
supports_voice_design = True
applies_own_mastering = True # native 48 kHz studio output — skip apply_mastering()
gpu_compat = ("cuda", "mps", "cpu")
_DEFAULT_SAMPLE_RATE = 48_000
@classmethod
def is_available(cls) -> tuple[bool, str]:
if own_venv_python() is None:
return False, (
"voxcpm package not installed. Install it from Model Catalogue."
)
return True, "ready"
@classmethod
def venv_python(cls) -> Path:
py = own_venv_python()
if py is None:
raise RuntimeError(
"VoxCPM2's environment is missing. Reinstall it from "
"Model Catalogue."
)
return py
@classmethod
def sidecar_script(cls) -> Path:
return Path(__file__).resolve().parent / "main.py"
@property
def recv_timeout_s(self) -> float:
# A cold load downloads several GB of weights; the sidecar heartbeats
# progress frames meanwhile, and each one re-arms this deadline.
try:
v = float(os.environ.get("OMNIVOICE_VOXCPM2_RECV_TIMEOUT_S", "900"))
except (TypeError, ValueError):
return 900.0
if not math.isfinite(v): # reject inf/nan so the deadline can't be disabled
return 900.0
return max(30.0, v)
@property
def sample_rate(self) -> int:
return self._DEFAULT_SAMPLE_RATE
@property
def supported_languages(self) -> list[str]:
from services.tts_backend import VoxCPM2Backend
return VoxCPM2Backend.supported_languages.fget(self)
def generate(self, text: str, **kw) -> "torch.Tensor":
# The same preparation and finishing as VoxCPM2Backend.generate: the
# reference clip is trimmed and capped here (the model no longer does
# it), and the output's long silent tail is cut.
from services.audio_dsp import trim_trailing_silence
from services.tts_backend import _prepare_voxcpm_ref
if kw.get("ref_audio"):
kw["ref_audio"] = _prepare_voxcpm_ref(kw["ref_audio"])
wav = super().generate(text, **kw)
return trim_trailing_silence(wav, self.sample_rate)
__all__ = ["VENV_ENV_VAR", "VoxCPM2SubprocessBackend", "own_venv_python"]
+263
View File
@@ -0,0 +1,263 @@
"""voxcpm2 sidecar: VoxCPM2 in the engine's own venv (one-click install).
Launched as ``<engine venv python> main.py`` by VoxCPM2SubprocessBackend. It
imports nothing from the app: the venv holds only ``voxcpm`` and what it
depends on (torch, torchaudio, numpy), so this file must stay importable with
the standard library plus those. The parent prepares the reference clip and
trims the output's silent tail, exactly as the in-process engine does; this
process only loads the model and synthesizes.
Wire protocol: length-prefixed JSON over stdio, identical to the other
sidecars (engines/pockettts/main.py). A ``ready`` frame comes first, then one
``audio`` (or ``error``) frame per ``synthesize``, with ``progress`` frames
while a cold load runs so the parent's watchdog stays armed.
"""
from __future__ import annotations
import base64
import json
import os
import re
import struct
import sys
import threading
import time
import traceback
# Mirrors services/subprocess_backend.py::MAX_FRAME_BYTES.
MAX_FRAME_BYTES = 64 * 1024 * 1024
#: VoxCPM2's studio output rate; the in-process engine assumes the same.
VOXCPM2_SAMPLE_RATE = 48_000
#: Emit a progress frame at least this often during a cold load (a multi-GB
#: first download) so the parent's recv watchdog doesn't kill a healthy sidecar.
_HEARTBEAT_S = 5.0
#: ref_audio must be a local file path, not a URL (local-first; no SSRF).
_URL_RE = re.compile(r"^[a-z][a-z0-9+.\-]*://", re.IGNORECASE)
#: A download failure worth retrying (the HF cache resumes, so a retry
#: continues rather than restarts). Anything else propagates at once.
_TRANSIENT_MARKERS = (
"connection", "timed out", "timeout", "peer closed", "incomplete",
"remoteprotocolerror", "temporarily unavailable",
)
_MODEL = None
# -- wire protocol -----------------------------------------------------------
#: Serializes _send across threads (the cold-load heartbeat + the main loop) so
#: concurrent length+body writes can't interleave and corrupt the framing.
_send_lock = threading.Lock()
def _send(stream, obj: dict) -> None:
body = json.dumps(obj, separators=(",", ":")).encode("utf-8")
with _send_lock:
stream.write(struct.pack("!I", len(body)))
stream.write(body)
stream.flush()
def _recv(stream):
header = stream.read(4)
if len(header) < 4:
return None # EOF
(n,) = struct.unpack("!I", header)
if n > MAX_FRAME_BYTES:
raise IOError(f"frame too large: {n}")
body = bytearray()
while len(body) < n:
chunk = stream.read(n - len(body))
if not chunk:
raise IOError("short read")
body.extend(chunk)
return json.loads(bytes(body).decode("utf-8"))
def _measure_vram_mb() -> float:
try:
import torch # noqa: PLC0415
if torch.cuda.is_available():
return float(torch.cuda.memory_allocated()) / (1024 * 1024)
except Exception: # noqa: BLE001 — a probe, never fatal
pass
return 0.0
# -- model loading (lazy, on the first synthesize) ---------------------------
def _with_retries(load):
"""Run ``load``, retrying a transient download failure with a short
backoff, the way the app's own loader does for in-process engines."""
try:
attempts = max(1, int(os.environ.get("OMNIVOICE_MODEL_LOAD_RETRIES", "3")))
except ValueError:
attempts = 3
for attempt in range(1, attempts + 1):
try:
return load()
except Exception as exc: # noqa: BLE001 — classified below
text = f"{type(exc).__name__}: {exc}".lower()
if attempt == attempts or not any(m in text for m in _TRANSIENT_MARKERS):
raise
time.sleep(2.0 * attempt)
raise AssertionError("unreachable")
def _load_model(stdout):
global _MODEL
if _MODEL is not None:
return _MODEL
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 0})
stop = threading.Event()
def _heartbeat() -> None:
pct = 1
while not stop.wait(_HEARTBEAT_S):
pct = min(pct + 1, 99)
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": pct})
hb = threading.Thread(target=_heartbeat, daemon=True)
hb.start()
try:
from voxcpm import VoxCPM # type: ignore[import-not-found] # noqa: PLC0415
checkpoint = os.environ.get("OMNIVOICE_VOXCPM_MODEL", "openbmb/VoxCPM2")
_MODEL = _with_retries(
lambda: VoxCPM.from_pretrained(checkpoint, load_denoiser=False)
)
finally:
stop.set()
hb.join(timeout=_HEARTBEAT_S + 1)
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 100})
return _MODEL
def _sample_rate(model) -> int:
for owner in (model, getattr(model, "tts_model", None)):
sr = getattr(owner, "sample_rate", None)
if isinstance(sr, int) and sr > 0:
return sr
return VOXCPM2_SAMPLE_RATE
def _at_engine_rate(wav, sample_rate: int):
"""The waveform at VOXCPM2_SAMPLE_RATE. The parent reads the PCM at that
fixed rate (it trims the tail and labels the audio with it), so a model
reporting another rate is resampled here rather than mislabelled."""
if sample_rate == VOXCPM2_SAMPLE_RATE:
return wav
import torch # noqa: PLC0415
import torchaudio # noqa: PLC0415
tensor = torch.as_tensor(wav.detach().cpu() if hasattr(wav, "detach") else wav,
dtype=torch.float32).reshape(-1)
return torchaudio.functional.resample(tensor, sample_rate, VOXCPM2_SAMPLE_RATE)
def _to_pcm_b64(wav) -> tuple[str, int]:
"""A float waveform in [-1, 1] (numpy or torch) as base64 int16 PCM."""
import numpy as np # noqa: PLC0415
if hasattr(wav, "detach"):
wav = wav.detach().float().cpu().numpy()
arr = np.asarray(wav, dtype=np.float32).squeeze()
if arr.ndim > 1:
raise ValueError(f"expected mono audio (1-D after squeeze), got shape {arr.shape}")
arr = np.clip(arr, -1.0, 1.0)
pcm = (arr * 32767.0).astype(np.int16).tobytes()
return base64.b64encode(pcm).decode("ascii"), int(arr.shape[-1])
def _handle_synthesize(msg: dict, stdout) -> None:
"""One synthesize request. The mapping mirrors VoxCPM2Backend.generate."""
text = msg.get("text")
if not text or not isinstance(text, str):
raise ValueError("synthesize: missing or non-string 'text'")
ref_audio = msg.get("ref_audio") or None
if ref_audio and _URL_RE.match(str(ref_audio)):
raise ValueError(
"ref_audio must be a local file path; URLs are not accepted (local-first)."
)
model = _load_model(stdout)
description = msg.get("description")
cfg_value = msg.get("guidance_scale", 2.0)
timesteps = msg.get("num_step", 10)
if description and not ref_audio:
# Voice design: a voice from a text description, no reference clip.
wav = model.generate(
text=text,
voice_description=description,
cfg_value=cfg_value,
inference_timesteps=timesteps,
)
else:
instruct = msg.get("instruct")
ref_text = msg.get("ref_text")
wav = model.generate(
text=f"({instruct}){text}" if instruct else text,
cfg_value=cfg_value,
inference_timesteps=timesteps,
reference_wav_path=ref_audio,
prompt_wav_path=ref_audio if ref_text else None,
prompt_text=ref_text,
)
pcm_b64, n_samples = _to_pcm_b64(_at_engine_rate(wav, _sample_rate(model)))
_send(stdout, {
"op": "audio",
"audio_pcm_b64": pcm_b64,
"sample_rate": VOXCPM2_SAMPLE_RATE,
"n_samples": n_samples,
})
# -- main loop ---------------------------------------------------------------
def main() -> int:
stdin = sys.stdin.buffer
# Frames go down a PRIVATE fd, and fd 1 is pointed at stderr (#1428): the
# libraries this loads print to fd 1 (tqdm, native torch output), and those
# bytes would otherwise interleave with the length-prefixed frames.
_frame_fd = os.dup(1)
os.dup2(2, 1)
stdout = os.fdopen(_frame_fd, "wb")
# Ready handshake fires BEFORE any heavy import.
_send(stdout, {"op": "ready", "engine": "voxcpm2", "sample_rate": VOXCPM2_SAMPLE_RATE})
while True:
try:
msg = _recv(stdin)
except Exception as exc: # noqa: BLE001
_send(stdout, {
"op": "error",
"stage": "recv",
"message": f"{type(exc).__name__}: {exc}",
"traceback": traceback.format_exc(),
})
return 1
if msg is None:
return 0
op = msg.get("op") if isinstance(msg, dict) else None
try:
if op == "ping":
_send(stdout, {"op": "pong", "vram_mb": _measure_vram_mb()})
elif op == "synthesize":
_handle_synthesize(msg, stdout)
elif op == "shutdown":
return 0
else:
_send(stdout, {"op": "error", "stage": "dispatch", "message": f"unknown op: {op!r}"})
except Exception as exc: # noqa: BLE001
_send(stdout, {
"op": "error",
"stage": op or "unknown",
"message": f"{type(exc).__name__}: {exc}",
"traceback": traceback.format_exc(),
})
if __name__ == "__main__":
sys.exit(main())
+109 -41
View File
@@ -9,6 +9,13 @@ _backend_dir = os.path.dirname(os.path.abspath(__file__))
if _backend_dir not in sys.path:
sys.path.insert(0, _backend_dir)
# #2135: arm fatal-signal tracebacks before anything heavy is imported, so a
# native crash inside torch/CUDA leaves a named frame in backend_err.log
# instead of a silently vanished process. See core/crash_diagnostics.py.
from core.crash_diagnostics import enable_fault_handler # noqa: E402
enable_fault_handler()
# PyInstaller re-executes this entry module when the frozen backend binary is
# launched. Nested operation supervisors therefore dispatch here, before math,
# logging, FastAPI, torch, or any application initialization. Source launches
@@ -274,16 +281,15 @@ logging.basicConfig(
# inherits the filter, so even handler-formatted output (file, stream,
# JSON) strips real HF tokens. Cheap (regex on each record) and
# idempotent — extra calls are no-ops.
from core.logging_filter import install_redaction_filter # noqa: E402
from core.logging_filter import ( # noqa: E402
install_access_log_filter,
install_asyncio_transport_filter,
install_redaction_filter,
)
install_redaction_filter()
class AsyncioExceptionFilter(logging.Filter):
def filter(self, record: logging.LogRecord) -> bool:
if record.levelno == logging.WARNING and "socket.send() raised exception" in record.getMessage():
return False
return True
logging.getLogger("asyncio").addFilter(AsyncioExceptionFilter())
install_access_log_filter()
install_asyncio_transport_filter()
# Silence HF Hub unauthenticated warnings unless specifically requested.
logging.getLogger("huggingface_hub.utils._http").setLevel(logging.ERROR)
@@ -537,8 +543,8 @@ async def _cancel_and_await_tasks(*tasks, timeout: float = 3.0) -> None:
# mutation, runs in an executor thread (deferred) or inline (eager).
# Phase A finalize: router/mount registration — mutates the app, so it runs
# ON the event loop (deferred) with no awaits inside, making it atomic with
# respect to in-flight requests; the StartupGate keeps everything but
# /health + /startup/progress out until ready regardless.
# respect to in-flight requests; the StartupGate keeps work routes out until
# ready while retaining readiness probes and deliberate desktop shutdown.
# Phase B: the old lifespan startup body (DB, background services).
_phase_a_built = False
@@ -624,7 +630,19 @@ def _phase_a_build_inner() -> None:
_startup_progress.begin_step("ml_imports")
import torchaudio
warnings.filterwarnings("ignore", category=UserWarning)
torchaudio.set_audio_backend("soundfile")
# torchaudio 2.9 REMOVED set_audio_backend(); soundfile has been the only
# backend since 2.0, so the call was already a no-op there and is simply
# absent now. Unguarded it raises AttributeError inside `ml_imports`, and a
# failure in that phase takes the whole backend down — the desktop app sits
# on "starting backend" forever and /health stays 503.
#
# That is not a hypothetical version: RTX 50-series (Blackwell, sm_120)
# users have no choice but to move off the pinned torch 2.8.0, which has no
# sm_120 kernels, and the torch 2.9.x they land on brings torchaudio 2.9
# with it. So the one group forced to upgrade hit a hard startup crash for
# a line that does nothing (#1931).
if hasattr(torchaudio, "set_audio_backend"):
torchaudio.set_audio_backend("soundfile")
from utils import hf_progress
# HF tqdm patch before any library import that can trigger
# hf_hub_download (transformers, mlx_whisper, …).
@@ -652,6 +670,7 @@ def _phase_a_build_inner() -> None:
from api.routers import (
system,
profiles,
profile_images,
exports,
generation,
dub_core,
@@ -691,7 +710,7 @@ def _phase_a_build_inner() -> None:
from api.routers import mcp_bindings as _mcp_bindings_router # noqa: E402
from api.routers import workers as workers_router # noqa: E402
_router_modules.extend([
system, profiles, exports, generation, voice_convert, dub_core, dub_generate,
system, profiles, profile_images, exports, generation, voice_convert, dub_core, dub_generate,
dub_export, dub_translate, projects, glossary, engines, tools,
stories, setup, gallery, archetypes, describe_voice, community,
batch, watermark, events, capture, capture_ws, speech_platform, dictation,
@@ -865,6 +884,18 @@ async def _phase_b(app: FastAPI) -> None:
logger.exception("Startup job-sweep failed (non-fatal).")
_startup_progress.begin_step("services_start")
# Reapply an explicitly saved speed/quality profile after the local model
# inventory is available. Older builds saved the slider but could leave
# ASR/Dictation pointing at missing models even when compatible weights
# were already installed. This path is local-cache-only and download-free.
try:
from services.performance_profiles import reconcile_active_profile
recovered = reconcile_active_profile()
if recovered:
logger.info("Startup performance selections reconciled: %s", recovered)
except Exception:
logger.exception("Performance-profile reconciliation failed (non-fatal).")
# Phase 1 Wave 3 — macOS Gatekeeper quarantine probe (#54). Informational
# only; we never auto-run `xattr -cr`.
try:
@@ -906,12 +937,8 @@ async def _phase_b(app: FastAPI) -> None:
"Capture ASR preload skipped: <4GB free RAM; "
"dictation ASR will load on first use.")
return
loading_detail = None
prev_loading_detail = None
try:
from services.model_manager import _gpu_pool, _loading_detail
loading_detail = _loading_detail
prev_loading_detail = dict(loading_detail)
from services.model_manager import _gpu_pool
loop = asyncio.get_running_loop()
def _warm():
from services.asr_backend import (
@@ -925,20 +952,12 @@ async def _phase_b(app: FastAPI) -> None:
"Capture ASR preload skipped: no ASR model installed; "
"dictation will offer a download on first use.")
return
loading_detail["sub_stage"] = "loading_asr"
loading_detail["detail"] = "Warming up ASR engine…"
backend = get_capture_asr_backend()
logger.info("Capture ASR backend selected: %s", backend.id)
if hasattr(backend, 'warmup'):
loading_detail["detail"] = f"Loading {backend.display_name}"
backend.warmup()
loading_detail["sub_stage"] = "ready"
loading_detail["detail"] = "ASR engine ready"
await loop.run_in_executor(_gpu_pool, _warm)
except Exception as e:
if loading_detail is not None and loading_detail.get("sub_stage") == "loading_asr":
loading_detail.clear()
loading_detail.update(prev_loading_detail or {})
logger.warning("Capture ASR preload skipped: %s", e)
app.state.capture_preload_task = asyncio.create_task(_preload_capture_asr())
else:
@@ -1080,6 +1099,30 @@ async def lifespan(app: FastAPI):
app.state.startup_task = asyncio.create_task(_deferred_startup(app))
yield
# ── Graceful shutdown (SIGTERM from Tauri, Ctrl+C, etc.) ────────────
# Retire the run sentinel FIRST, before any bounded wait below (#1895):
# once uvicorn has begun graceful shutdown the exit is deliberate by
# definition, so the sentinel has already done its job. This is one
# os.remove, against a ~50s worst-case tail of bounded waits plus model
# unload / free_vram() / gc.collect() below. Measured on macOS: a normal
# shutdown takes 5.25s end to end, while the desktop shell allows 2s
# (bootstrap.rs terminate_process_tree) before SIGKILL — so the old
# placement at the very end was killed every time on any run that had
# reached a working state. Doing the deadline-sensitive step first makes
# correctness independent of how much of that tail runs, instead of
# depending on the shell-side deadline being long enough to cover it.
#
# Desktop shells that must hard-kill a Windows process tree retire the
# sentinel through /system/shutdown-intent before termination. This
# remains the graceful-path fallback for every platform and direct server
# runs.
#
# sentinel_cleared feeds the truthful "Shutdown: done."/degraded log at
# the end of this function; nothing below re-clears the sentinel, so a
# later failure can't mask this result.
try:
sentinel_cleared = run_sentinel.clear_sentinel()
except Exception:
sentinel_cleared = False
# May run after a startup that never finished (SIGTERM mid-Phase-A/B), so
# every handle is read from app.state with a None default and every
# deferred-phase name is guarded.
@@ -1180,11 +1223,17 @@ async def lifespan(app: FastAPI):
# Best-effort drain: a failure here must not abort the remaining
# shutdown steps (model unload, MCP teardown) below.
logger.warning("Watermark pool drain failed at shutdown", exc_info=True)
# Unload the model and free GPU memory
# Release every runtime that can retain model memory, then free allocator
# caches. This includes alternate TTS engines, dictation and translation;
# limiting shutdown to the shared OmniVoice model left those runtimes to
# process-exit cleanup and made graceful restarts look like crashes.
try:
import services.model_manager as mm
if mm.unload_shared_model():
logger.info("Shutdown: model unloaded.")
from services import model_lifecycle
released = await model_lifecycle.unload_all()
if any(result.get("success") for result in released["results"].values()):
logger.info("Shutdown: model runtimes unloaded.")
# Still unconditional: there are allocator caches to hand back even when
# no model was resident.
mm.free_vram()
@@ -1206,13 +1255,9 @@ async def lifespan(app: FastAPI):
await close_http_client()
except Exception:
pass
# Last thing on a clean shutdown: retire the run sentinel so the next
# startup doesn't misread this exit as a crash (#1164). If clearing fails,
# retain the sentinel and report a degraded shutdown truthfully.
try:
sentinel_cleared = run_sentinel.clear_sentinel()
except Exception:
sentinel_cleared = False
# Sentinel was already retired at the TOP of this block (#1895) — report
# truthfully using that result rather than clearing (or re-checking) it
# again here, so a failure in the steps above can't mask it as "done."
if sentinel_cleared:
logger.info("Shutdown: done.")
else:
@@ -1230,6 +1275,23 @@ app = FastAPI(
)
@app.post("/system/shutdown-intent", include_in_schema=False)
def prepare_deliberate_shutdown_during_startup(request: Request):
"""Retire crash forensics even while deferred startup is still gated.
Electron must hard-kill a Windows process tree after a bounded wait. The
ordinary system router is registered only after native/ML imports finish,
so a quit during those imports previously received the startup 503 and left
a false crash sentinel behind. Keep this one tiny control route available
from socket bind; its authorization remains identical to the system router.
"""
from api.dependencies import require_admin
from core import run_sentinel
require_admin(request)
return {"prepared": run_sentinel.clear_sentinel()}
@app.get("/docs", include_in_schema=False)
async def scalar_docs():
"""Interactive API documentation powered by Scalar."""
@@ -1419,13 +1481,17 @@ async def global_exception_handler(request: Request, exc: Exception):
_SHELL_PATHS = {"/", "/index.html", "/favicon.ico", "/health"}
# Paths that answer while the deferred startup is still running.
_STARTUP_EXEMPT = {"/health", "/startup/progress"}
# Paths that answer while deferred startup is still running. The shutdown
# signal must exist before the ordinary system router so a bounded Windows
# process-tree stop cannot leave a false crash sentinel.
_STARTUP_EXEMPT = {"/health", "/startup/progress", "/system/shutdown-intent"}
class StartupGateMiddleware:
"""503 everything except /health + /startup/progress until the deferred
startup completes. Two jobs: honest not-ready signaling (the [starting]
"""503 work routes until deferred startup completes.
Readiness probes and deliberate desktop shutdown remain live. Two jobs:
honest not-ready signaling (the [starting]
marker keeps the UI from offering "Report" for it, same convention as
[shutting_down]), and route-mutation safety no request can reach the
router while _phase_a_finalize is still adding routes, because the ready
@@ -1658,10 +1724,12 @@ def _ui_port() -> int:
return 3901
from core.csrf import DEFAULT_DESKTOP_ORIGINS
_ui = _ui_port()
_allowed = os.environ.get(
"OMNIVOICE_ALLOWED_ORIGINS",
f"http://localhost:{_ui},http://127.0.0.1:{_ui},tauri://localhost,http://tauri.localhost",
f"http://localhost:{_ui},http://127.0.0.1:{_ui}," + ",".join(DEFAULT_DESKTOP_ORIGINS),
).split(",")
# Registered FIRST → innermost: the startup gate holds every request except
+62 -14
View File
@@ -271,19 +271,61 @@ def _write_output(audio_id: str, raw: bytes) -> str:
return path
def _post_timeout_s() -> float:
"""Seconds the tools wait on a backend POST (OMNIVOICE_MCP_TIMEOUT_S,
default 120). A CPU host renders a paragraph in minutes and serializes
generations, so an agent behind another render used to hit the fixed
budget with an empty-message timeout; the knob follows the backend's own
OMNIVOICE_GENERATE_TIMEOUT_S when a deployment raises that."""
raw = os.environ.get("OMNIVOICE_MCP_TIMEOUT_S", "").strip()
# Extra seconds a tool waits past the backend's own budget, so the backend's
# error (which says what ran out) reaches the agent instead of an empty
# client-side timeout (#2040).
_BACKEND_GRACE_S = 30.0
def _env_seconds(name: str, default: float) -> float:
raw = os.environ.get(name, "").strip()
if not raw:
return default
try:
value = float(raw) if raw else 120.0
value = float(raw)
except ValueError:
logger.warning("OMNIVOICE_MCP_TIMEOUT_S=%r is not a number; using 120", raw)
return 120.0
return value if value > 0 else 120.0
logger.warning("%s=%r is not a number; using %g", name, raw, default)
return default
return value if value > 0 else default
def _backend_budget_s(kind: str, text: str = "") -> float | None:
"""The backend's own execution budget for this kind of request, read from
the environment variables and defaults the backend itself uses.
The MCP server cannot see which device the backend runs on, so generation
assumes the larger CPU base. The backend still stops a job at its own
budget; this only keeps the tool from giving up first.
"""
if kind == "transcribe":
# run_transcribe_guarded starts this clock when the job is submitted,
# so time spent queued in the pool already counts against it.
return _env_seconds("OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S", 300.0)
if kind == "generate":
base = max(
_env_seconds("OMNIVOICE_GENERATE_TIMEOUT_S", 300.0),
_env_seconds("OMNIVOICE_CPU_GENERATE_TIMEOUT_S", 600.0),
)
# As model_manager.generate_timeout_s: +1 s per 40 characters past 1200.
execution = base + max(0, len(text or "") - 1200) / 40.0
# A generation first waits in the GPU pool's queue, on its own clock
# (model_manager.GPU_QUEUE_TIMEOUT_S), before that budget starts.
return _env_seconds("OMNIVOICE_GPU_QUEUE_TIMEOUT_S", 1800.0) + execution
return None
def _post_timeout_s(kind: str = "", text: str = "") -> float:
"""Seconds a tool waits on a backend POST.
An explicit OMNIVOICE_MCP_TIMEOUT_S wins. Otherwise the tool waits for the
backend's own budget for that request plus a grace period, and never less
than 120 s. A fixed 120 s used to cut off transcriptions the backend would
have finished (its ASR budget is 300 s) with an empty error (#2040).
"""
if os.environ.get("OMNIVOICE_MCP_TIMEOUT_S", "").strip():
return _env_seconds("OMNIVOICE_MCP_TIMEOUT_S", 120.0)
budget = _backend_budget_s(kind, text)
return 120.0 if budget is None else max(120.0, budget + _BACKEND_GRACE_S)
def _maybe_number(value):
@@ -397,9 +439,12 @@ def create_mcp_server():
r.raise_for_status()
return r.json()
async def _api_post_form(path: str, data: dict, files: dict | None = None):
async def _api_post_form(
path: str, data: dict, files: dict | None = None, *, timeout: float | None = None
):
import httpx
async with httpx.AsyncClient(base_url=_api_base(), timeout=_post_timeout_s()) as c:
wait = _post_timeout_s() if timeout is None else timeout
async with httpx.AsyncClient(base_url=_api_base(), timeout=wait) as c:
r = await c.post(path, data=data, files=files or {})
r.raise_for_status()
return r
@@ -471,7 +516,9 @@ def create_mcp_server():
if instruct:
form["instruct"] = instruct
r = await _api_post_form("/generate", data=form)
r = await _api_post_form(
"/generate", data=form, timeout=_post_timeout_s("generate", text)
)
audio_id = r.headers.get("X-Audio-Id", "unknown")
gen_time = _maybe_number(r.headers.get("X-Gen-Time", "?"))
@@ -547,6 +594,7 @@ def create_mcp_server():
"/transcribe", data=data,
files={"audio": (f"audio{_sniff_audio_ext(raw)}", raw,
"application/octet-stream")},
timeout=_post_timeout_s("transcribe"),
)
return str(r.json())
@@ -0,0 +1,18 @@
"""Retain the dispatch-time deadline policy across worker/control-plane loss."""
from alembic import op
import sqlalchemy as sa
revision = "0011_remote_attempt_deadlines"
down_revision = "0010_remote_worker_schema"
branch_labels = None
depends_on = None
def upgrade() -> None:
columns = op.get_bind().execute(sa.text("PRAGMA table_info(remote_task_attempts)"))
if not any(row[1] == "deadlines_json" for row in columns):
op.add_column("remote_task_attempts", sa.Column("deadlines_json", sa.Text(), nullable=True))
def downgrade() -> None:
op.drop_column("remote_task_attempts", "deadlines_json")
+36 -12
View File
@@ -1,4 +1,4 @@
from pydantic import BaseModel, field_validator
from pydantic import BaseModel, Field, field_validator
from typing import List, Literal, Optional
from services.audio_dsp import EFFECT_PRESETS
@@ -60,7 +60,9 @@ class DubRequest(BaseModel):
language: str = "Auto"
language_code: str = "und" # ISO 639-1 for ffmpeg metadata (e.g. "es", "fr", "de")
instruct: str = ""
num_step: int = 16
# None means "use the shared performance profile". An explicit value is
# still authoritative for Production overrides and existing API clients.
num_step: Optional[int] = None
guidance_scale: float = 2.0
speed: float = 1.0
# Phase 4.1 — partial regen. Parallel lists by index with `segments`.
@@ -71,7 +73,8 @@ class DubRequest(BaseModel):
regen_only: Optional[List[str]] = None
# Fast-preview mode for interactive edits. When true, TTS runs at
# num_step=8 (~2× faster, ~10-20% quality drop). Client is responsible
# for re-rendering preview segs at full quality before final export.
# for re-rendering preview segs with the explicit override or shared
# performance profile before final export.
preview: Optional[bool] = False
# How to handle segs whose TTS audio is longer than its slot (the
# "ghost lang" overlap bug otherwise). Options:
@@ -86,9 +89,9 @@ class DubRequest(BaseModel):
# (Bengali, Hindi, Arabic…). Three modes:
# "concise" — never compress TTS audio. Trim text up-front via
# speech_rate so it fits naturally; if it still
# overflows, hard-trim at slot with a short fade and
# surface fit_status="overflows" so the UI can prompt
# the user to shorten the segment. DEFAULT.
# overflows, fail without replacing the current track;
# the user must shorten it or choose another fit mode.
# DEFAULT.
# "stretch_video" — never compress TTS audio. Re-lay the timeline so
# each segment's video portion is stretched (via
# ffmpeg setpts) to fit the natural-rate dub audio.
@@ -97,13 +100,13 @@ class DubRequest(BaseModel):
# mild pitch-preserving audio speed-up (≤1.2× alone,
# ≤1.5× in hybrid) and a mild per-segment video
# slow-down (≤2.0×), per services/fit_planner.py.
# Residual overflow is trimmed and surfaced.
# "strict_slot" — legacy: keep `slot_fit` semantics (atempo squeeze
# when audio > slot). Kept for back-compat.
# Residual overflow fails without discarding words.
# "strict_slot" — pitch-preserving fit of the complete speech to
# the original start/end; may sound faster or slower.
timing_strategy: Optional[Literal["concise", "stretch_video", "strict_slot", "smart_fit"]] = "concise"
# Per-job slip budget for "concise" mode. Hard-trim only kicks in once
# gap absorption + this much extra time has been consumed.
# Per-job slip budget for "concise" mode. Overflow fails once gap
# absorption + this much extra time has been consumed.
overflow_budget_s: Optional[float] = 0.0
# Knob overrides for `smart_fit` (ignored by other strategies). Omitted
@@ -150,7 +153,7 @@ class TranslateRequest(BaseModel):
provider: Optional[str] = None
source_lang: Optional[str] = None # ISO 639-1; overrides job detection
job_id: Optional[str] = None # Dub job id, used to resolve detected source_lang
quality: Optional[str] = "fast" # "fast" (one-shot) | "cinematic" (reflect→adapt) | "autofit" (cinematic + strict fit-to-slot)
quality: Optional[str] = "fast" # fast | cinematic | autofit | agent (measured render/rewrite loop)
glossary: Optional[List[dict]] = None # [{"source": "...", "target": "...", "note": "..."}]
# Optional regional dialect (BCP-47, e.g. "es-AR", "pt-BR") — #280 item 2.
# Applied by LLM-backed paths (provider="openai" or quality="cinematic"):
@@ -158,6 +161,7 @@ class TranslateRequest(BaseModel):
# voseo: "vos sos" instead of "tú eres"). Non-LLM providers (Argos, NLLB,
# Google) can't honor it; the response then carries dialect_applied=false.
dialect: Optional[str] = None
translation_instructions: Optional[str] = Field(default=None, max_length=5000)
# Two-stage LLM translation quality (provider="openai" only; MT engines
# ignore both). None = default ON for the LLM engine.
# auto_glossary — one up-front LLM pass over the full transcript extracts
@@ -175,6 +179,26 @@ class TranslateRequest(BaseModel):
# No LLM configured / LLM failure → silently no suggestion.
condense: Optional[bool] = False
class AgentFitSegment(BaseModel):
"""One rendered translation and its measured timing evidence."""
id: str
text: str
source_text: Optional[str] = None
context_before: Optional[str] = None
context_after: Optional[str] = None
slot_seconds: float
measured_seconds: float
class AgentFitRequest(BaseModel):
"""Revise only rendered lines that missed their exact timeline slot."""
translation_instructions: Optional[str] = Field(default=None, max_length=5000)
segments: List[AgentFitSegment]
target_lang: str
class ParseSubtitleTextRequest(BaseModel):
"""Raw pasted subtitle text (SRT/VTT-ish) to be parsed into timed cues.
+322 -81
View File
@@ -133,13 +133,13 @@ def _isolated_engine_hint(streak: int) -> str:
"%d consecutive ASR transcribe timeouts this session — pool resets are "
"not recovering the hang. Recommend switching the ASR engine to "
"'Faster-Whisper (crash-isolated subprocess)' [faster-whisper-isolated] "
"in Model Catalogue → Engines. Not switching automatically (#730).", streak,
"in Model Catalogue. Not switching automatically (#730).", streak,
)
return (
f"This is {streak} transcribe timeouts in a row this session, so pool "
"resets aren't recovering the underlying hang. Recommended: switch the "
"ASR engine to 'Faster-Whisper (crash-isolated subprocess)' "
"(faster-whisper-isolated) in Model Catalogue → Engines — it runs "
"(faster-whisper-isolated) in Model Catalogue — it runs "
"transcription in a separate process that can be force-killed to "
"reclaim a hung transcribe and its VRAM. VoiceStudio never switches "
"engines automatically."
@@ -232,7 +232,7 @@ async def run_transcribe_guarded(executor, fn, *, what: str = "ASR",
"The native call cannot be killed safely, so its capacity remains "
"reserved until it exits. For a durable fix Flush the "
"TTS model to free VRAM, pick a smaller ASR model in "
f"Model Catalogue → Models, or set ASR to CPU. (Raise {timeout_env} "
f"the engine's Weights list in Model Catalogue, or set ASR to CPU. (Raise {timeout_env} "
"for very long transcribes.)"
)
hint = _isolated_engine_hint(streak)
@@ -648,7 +648,8 @@ class WhisperXBackend(ASRBackend):
logger.warning(
"whisperx VRAM preflight: %.1f GB free is too little for %s on CUDA "
"(needs ≥%.1f GB even at int8) — using CPU int8 instead. Free VRAM "
"(flush the TTS model, or close other GPU apps) for GPU-speed ASR. (#723)",
"(flush the TTS model, or close other GPU apps) for GPU-speed ASR, or "
"set OMNIVOICE_ASR_VRAM_PREFLIGHT=0 to skip this check. (#723)",
free, self._model_name,
self._CUDA_VRAM_BUDGET_GB["int8"] * scale,
)
@@ -1004,13 +1005,11 @@ class FasterWhisperBackend(ASRBackend):
# CTranslate2: CUDA or CPU (no upstream ROCm/HIP build — see WhisperX note).
gpu_compat = ("cuda", "cpu")
def __init__(self):
def __init__(self, model_name: str | None = None):
# Defaulting to the CTranslate2-converted large-v3 repo. Matches
# KNOWN_MODELS in api/routers/setup.py so the first-run wizard
# downloads what the backend will actually load.
self._model_name = os.environ.get(
"ASR_MODEL_FASTER", "Systran/faster-whisper-large-v3"
)
self._model_name = model_name or faster_whisper_model_id()
self._model = None # lazy — first transcribe() loads weights
# Set by _ensure_model() to the device/compute_type that actually loaded
# (after the #551 compute_type / #255 OOM→CPU fallback chain).
@@ -1113,10 +1112,13 @@ class FasterWhisperBackend(ASRBackend):
# faster-whisper returns a generator of Segment objects + an Info
# struct. Materialise the generator so downstream consumers can
# index / re-iterate.
from services.performance_profiles import asr_decode_defaults
segments_iter, info = self._model.transcribe(
audio_path,
word_timestamps=word_timestamps,
vad_filter=True, # built-in Silero VAD — cleaner segment starts
**asr_decode_defaults(),
)
segments = list(segments_iter)
# Normalise to the shape segment_transcript(...) expects: a dict with
@@ -1292,10 +1294,54 @@ class PyTorchWhisperBackend(ASRBackend):
# Reuses the `_asr_pipe` attached to the TTS model when available.
self._pipe = asr_pipe
# whisper-large-v3-turbo occupies roughly 3.2 GiB before generation adds
# its encoder/decoder workspace. Loading it onto a nearly full card works,
# then the first transcribe fails with a CUDA OOM and yields zero segments.
_CUDA_VRAM_BUDGET_GB = 5.0
# Free VRAM needed on CUDA, where _ensure_pipe loads fp16 weights: the
# weights (parameters x 2 bytes), the batch-16 decode workspace, and
# headroom. Loading onto a nearly full card works, then the first
# transcribe fails with a CUDA OOM and yields zero segments, so the device
# pick checks this against actually-free VRAM first.
#
# #2041: this was a flat 5.0 GB for every model, sized for full large-v3.
# The default is large-v3-turbo (0.81B parameters, about 1.6 GB in fp16),
# so a 6 GB card with nothing else resident reported 5.0 GB free and was
# sent to CPU every time, although CUDA transcribed the same audio in 37 s.
_CUDA_VRAM_BUDGET_GB = 5.0 # full large-v3, and any model not listed below
# fp16 weights (GB) of the OpenAI Whisper checkpoints, by exact repo id.
# Anything else, including a fine-tune or a custom repo whose name happens
# to contain "small" or "turbo", keeps the conservative 5.0 GB budget.
_FP16_WEIGHTS_GB = {
"openai/whisper-large-v3-turbo": 1.6,
"openai/whisper-large-v3": 3.1,
"openai/whisper-large-v2": 3.1,
"openai/whisper-large": 3.1,
"openai/whisper-medium": 1.5,
"openai/whisper-medium.en": 1.5,
"openai/whisper-small": 0.5,
"openai/whisper-small.en": 0.5,
"openai/whisper-base": 0.15,
"openai/whisper-base.en": 0.15,
"openai/whisper-tiny": 0.08,
"openai/whisper-tiny.en": 0.08,
}
_CUDA_WORKSPACE_GB = 1.5 # batch 16 x 15 s chunks
_CUDA_HEADROOM_GB = 0.5
@classmethod
def _cuda_budget_gb(cls, model_name: str) -> float:
"""Free VRAM (GB) this model needs on CUDA; never above the 5.0 GB
that full large-v3 was measured to need."""
weights_gb = cls._FP16_WEIGHTS_GB.get((model_name or "").strip().lower())
if weights_gb is None:
return cls._CUDA_VRAM_BUDGET_GB
return min(
cls._CUDA_VRAM_BUDGET_GB,
weights_gb + cls._CUDA_WORKSPACE_GB + cls._CUDA_HEADROOM_GB,
)
@staticmethod
def _model_name() -> str:
return os.environ.get(
"OMNIVOICE_PYTORCH_ASR_MODEL", "openai/whisper-large-v3-turbo"
)
@classmethod
def is_available(cls) -> tuple[bool, str]:
@@ -1306,7 +1352,7 @@ class PyTorchWhisperBackend(ASRBackend):
return False, f"transformers not installed: {e}"
@classmethod
def _pick_device(cls) -> str:
def _pick_device(cls, model_name: str | None = None) -> str:
from services.model_manager import get_best_device
device = str(get_best_device())
@@ -1321,14 +1367,18 @@ class PyTorchWhisperBackend(ASRBackend):
free_gb = free / 1024**3
except Exception: # noqa: BLE001 — an unavailable probe must not block ASR
return device
if free_gb >= cls._CUDA_VRAM_BUDGET_GB:
model_name = model_name or cls._model_name()
budget_gb = cls._cuda_budget_gb(model_name)
if free_gb >= budget_gb:
return device
logger.warning(
"PyTorch Whisper VRAM preflight: %.1f GB free < %.1f GB needed "
"for reliable CUDA transcription — using CPU instead. Close other "
"GPU apps or Flush models to restore GPU-speed ASR.",
"for %s on CUDA — using CPU instead. Close other GPU apps or Flush "
"models to restore GPU-speed ASR, or set "
"OMNIVOICE_ASR_VRAM_PREFLIGHT=0 to skip this check.",
free_gb,
cls._CUDA_VRAM_BUDGET_GB,
budget_gb,
model_name,
)
return "cpu"
@@ -1351,10 +1401,8 @@ class PyTorchWhisperBackend(ASRBackend):
# constructor and this path is skipped.
import torch
from transformers import pipeline as hf_pipeline
model_name = os.environ.get(
"OMNIVOICE_PYTORCH_ASR_MODEL", "openai/whisper-large-v3-turbo"
)
device = self._pick_device()
model_name = self._model_name()
device = self._pick_device(model_name)
asr_dtype = torch.float16 if str(device).startswith("cuda") else torch.float32
logger.info(
"PyTorchWhisperBackend: loading standalone ASR pipeline %s on %s",
@@ -1401,7 +1449,15 @@ class PyTorchWhisperBackend(ASRBackend):
import soundfile as sf
import torch
self._ensure_pipe()
audio_np, sr = sf.read(audio_path, dtype="float32")
# #2039: libsndfile cannot open MP4/M4A (AAC), which /transcribe and
# the MCP tool both accept. Those decode through the validated ffmpeg
# path, which resamples to 16 kHz properly. Anything soundfile can
# read keeps its native rate, so the pipeline's band-limited
# resampler does the conversion rather than a linear interpolation.
try:
audio_np, sr = sf.read(audio_path, dtype="float32")
except Exception:
audio_np, sr = _decode_audio_16k_mono(audio_path), 16000
if audio_np.ndim > 1:
audio_np = audio_np.mean(axis=1)
bs = 16 if torch.cuda.is_available() else 2
@@ -1801,9 +1857,7 @@ class SherpaDictationBackend(ASRBackend):
def __init__(self, model_id: str | None = None):
from services import sherpa_dictation as _sd
mid = model_id or os.environ.get(
"OMNIVOICE_SHERPA_ASR_MODEL", _sd.DEFAULT_MODEL_ID
)
mid = model_id or sherpa_engine_model_id()
spec = _sd.get_spec(mid)
if spec is None:
raise ValueError(
@@ -1811,6 +1865,8 @@ class SherpaDictationBackend(ASRBackend):
f"{[s.id for s in _sd.list_specs()]}"
)
self._spec = spec
from services.performance_profiles import requested_tier
self.performance_tier = requested_tier("dictation")
self._rec = None # lazy OfflineRecognizer / OnlineRecognizer
# One backend is shared across live-dictation WS sessions (see
# get_sherpa_dictation_backend), so guard the one-time recognizer build
@@ -2272,7 +2328,7 @@ class OpenAICompatASRBackend(ASRBackend):
def is_available(cls) -> tuple[bool, str]:
base_url = resolve_openai_compat_asr_base_url()
if not base_url:
return False, "Configure a server endpoint in Model Catalogue → Engines"
return False, "Configure a server endpoint in Model Catalogue"
try:
normalize_openai_compat_asr_base_url(base_url)
except ValueError as exc:
@@ -2417,7 +2473,7 @@ _REGISTRY: dict[str, type[ASRBackend]] = _LazyASRRegistry({
})
# Short install hints surfaced as tooltips on the Model Catalogue → Engines UI
# Short install hints surfaced as tooltips on the Model Catalogue UI
# (parity with tts_backend._INSTALL_HINTS).
_INSTALL_HINTS: dict[str, str] = {
"whisperx": "pip install whisperx (CTranslate2 + wav2vec2 alignment; CUDA or CPU)",
@@ -2444,7 +2500,7 @@ _INSTALL_HINTS: dict[str, str] = {
"sherpa-onnx-asr": "uv add sherpa-onnx (ONNX live dictation; CPU, cross-platform)",
"openai-compat-asr": (
"No install needed — configure a server endpoint in "
"Model Catalogue → Engines. Points VoiceStudio at any OpenAI-compatible "
"Model Catalogue. Points VoiceStudio at any OpenAI-compatible "
"server (a self-hosted Qwen3-ASR/FunASR/SenseVoice server, OpenAI's "
"own Whisper API, or similar) — a path to Qwen3-ASR today, without "
"waiting on a direct transformers integration."
@@ -2770,7 +2826,7 @@ class ASRModelMissingError(RuntimeError):
super().__init__(asr_model_missing_detail(payload))
def load_active_asr_backend(*, asr_pipe=None) -> ASRBackend:
def load_active_asr_backend(*, asr_pipe=None, require_installed: bool = False) -> ASRBackend:
""":func:`get_active_asr_backend` + eager ``ensure_loaded()``, degrading
past backends whose deep import chain is broken (#1185).
@@ -2779,7 +2835,7 @@ def load_active_asr_backend(*, asr_pipe=None) -> ASRBackend:
import inside ``load_model``), so auto-detect can pick a backend that then
dies at load with ``No module named 'lightning_fabric'`` which used to
fail ASR init wholesale even though the next engine in line works fine.
Instead: record the backend as broken (Model Catalogue Engines shows why),
Instead: record the backend as broken (Model Catalogue shows why),
re-select, and load the next candidate mirroring how
:func:`_probe_available` already swallows broken natives at probe time.
@@ -2799,11 +2855,11 @@ def load_active_asr_backend(*, asr_pipe=None) -> ASRBackend:
while True:
backend = get_active_asr_backend(asr_pipe=asr_pipe)
bid = getattr(backend, "id", "?")
if tried:
if tried or require_installed:
# Preflight the SPECIFIC candidate about to load — not the global
# selection, which can disagree when an asr_pipe steers
# get_active_asr_backend (Greptile review, #1198).
missing = asr_model_missing_error(backend_id=bid)
missing = asr_model_missing_error(backend_id=bid, require_installed=require_installed)
if missing is not None:
raise ASRModelMissingError(missing)
try:
@@ -2828,7 +2884,7 @@ def load_active_asr_backend(*, asr_pipe=None) -> ASRBackend:
# ModuleNotFoundError and its ImportError parent ("cannot import
# name X" version skew) are the same env-rot class: the backend
# cannot work in this process, but siblings with independent
# import chains can. Record it either way so Model Catalogue → Engines
# import chains can. Record it either way so Model Catalogue
# reports the truth (unavailable + why + how to repair).
reason = _deep_import_reason(type(backend), e)
_DEEP_IMPORT_BROKEN[bid] = scrub_text(reason)
@@ -2878,6 +2934,109 @@ def _ref_audio_fingerprint(audio_path: str) -> str | None:
return None
def _installed_reference_fallbacks(
selected: list[ASRBackend],
) -> list[ASRBackend]:
"""Return the strongest compatible local fallbacks without changing prefs."""
fallbacks: list[ASRBackend] = []
selected_repos = {
_fw_repo(str(getattr(item, "_model_name", "")))
for item in selected
if isinstance(item, FasterWhisperBackend)
}
try:
from api.routers.setup.models import (
KNOWN_MODELS,
_model_supported,
_snapshot_dirs,
snapshot_is_complete,
)
available, _reason = FasterWhisperBackend.is_available()
if available:
compatible = sorted(
(
model
for model in KNOWN_MODELS
if str(model.get("role", "")).lower() == "asr"
and not model.get("dictation_id")
and (
str(model.get("repo_id", "")).startswith("Systran/faster-")
or model.get("repo_id")
== "deepdml/faster-whisper-large-v3-turbo-ct2"
)
and _model_supported(model)
and model.get("repo_id") not in selected_repos
),
key=lambda model: float(model.get("size_gb") or 0),
reverse=True,
)
for model in compatible:
snapshots = [
path
for path in _snapshot_dirs(str(model["repo_id"]))
if snapshot_is_complete(model, path)
]
if not snapshots:
continue
# A concrete complete snapshot cannot trigger a Hub download.
snapshot = max(snapshots, key=lambda path: os.path.getmtime(path))
backend = FasterWhisperBackend(model_name=snapshot)
setattr(backend, "_reference_ephemeral", True)
fallbacks.append(backend)
break
except Exception as exc: # noqa: BLE001 - optional local fallback
logger.warning("reference ASR cache fallback unavailable (%s)", exc)
try:
from services import sherpa_dictation
selected_sherpa = {
item.spec.id
for item in selected
if isinstance(item, SherpaDictationBackend)
}
installed = sorted(
(
spec
for spec in sherpa_dictation.list_specs()
if spec.id not in selected_sherpa
and sherpa_dictation.is_installed(spec)
),
key=lambda spec: float(spec.size_gb or 0),
reverse=True,
)
if installed:
fallbacks.append(get_sherpa_dictation_backend(installed[0].id))
except Exception as exc: # noqa: BLE001 - optional local fallback
logger.warning("reference dictation cache fallback unavailable (%s)", exc)
return fallbacks
def _transcribe_reference_candidates(
candidates: list[ASRBackend], audio_path: str,
) -> str:
for backend in candidates:
try:
result = backend.transcribe(audio_path, word_timestamps=False) or {}
candidate_text = result.get("text") or " ".join(
(seg.get("text") or "").strip()
for seg in result.get("segments", [])
)
candidate_text = (candidate_text or "").strip()
if candidate_text:
return candidate_text
except Exception as exc: # noqa: BLE001 - try the next local engine
logger.warning("transcribe_reference: %s failed (%s)", backend.id, exc)
finally:
if getattr(backend, "_reference_ephemeral", False):
try:
backend.unload()
except Exception: # noqa: BLE001 - release is best-effort
logger.warning("reference ASR fallback unload failed", exc_info=True)
return ""
def transcribe_reference(audio_path: str) -> str | None:
"""Transcribe a voice-clone reference clip with the active ASR backend.
@@ -2899,42 +3058,55 @@ def transcribe_reference(audio_path: str) -> str | None:
if cached is not None:
_ref_transcript_cache.move_to_end(fingerprint)
return cached
# No ASR model installed (TTS-only install): skip quietly instead of
# letting the backend auto-download multi-GB weights mid-/generate — this
# path is best-effort by contract (the engine's built-in fallback applies).
if asr_model_missing_error() is not None:
logger.info("transcribe_reference: no ASR model installed — skipping "
"reference auto-transcription (no silent download).")
# Prefer the selected offline ASR engine. When its selected weights are not
# installed, reuse the selected dictation engine if that model is already
# local. Short clone references need plain transcription, which dictation
# engines provide well. Falling straight through to the TTS model's bundled
# fallback produced incomplete reference conditioning for longer clips and
# introduced spurious words at the start of short generations. Neither
# branch may download weights implicitly.
candidates: list[ASRBackend] = []
offline_missing = asr_model_missing_error()
if offline_missing is None:
try:
# `load_*`, not `get_*`: a backend whose shallow probe passes but
# whose deep import chain is broken must fall through cleanly.
backend = load_active_asr_backend()
if not isinstance(backend, PyTorchWhisperBackend):
candidates.append(backend)
except Exception as e: # noqa: BLE001 — reference ASR is best-effort
logger.warning("transcribe_reference: offline ASR unavailable (%s)", e)
capture_missing = asr_model_missing_error(purpose="dictation")
if capture_missing is None:
try:
capture = get_capture_asr_backend()
if not isinstance(capture, PyTorchWhisperBackend) and not any(
type(item) is type(capture) and item.id == capture.id
for item in candidates
):
candidates.append(capture)
except Exception as e: # noqa: BLE001 — reference ASR is best-effort
logger.warning("transcribe_reference: dictation ASR unavailable (%s)", e)
text = _transcribe_reference_candidates(candidates, audio_path)
fallbacks: list[ASRBackend] = []
if not text:
fallbacks = _installed_reference_fallbacks(candidates)
text = _transcribe_reference_candidates(fallbacks, audio_path)
if not candidates and not fallbacks:
logger.info(
"transcribe_reference: no installed ASR model available — skipping "
"reference auto-transcription (no silent download)."
)
return None
try:
# `load_*`, not `get_*`: a backend whose shallow probe passes but whose
# deep import chain is broken would otherwise be handed back here and
# fail at `.transcribe()` below, costing every clone-without-transcript
# its reference text even with a healthy engine next in line (#1185).
# This path is best-effort, so a genuinely exhausted chain still just
# returns None and defers to the model's built-in fallback.
backend = load_active_asr_backend()
except Exception as e: # noqa: BLE001 — never let ASR break generation
logger.warning("transcribe_reference: no ASR backend available (%s)", e)
return None
if isinstance(backend, PyTorchWhisperBackend):
# The registry fell through to the model-attached pipeline; let the
# model load it lazily rather than constructing a second copy here.
return None
try:
result = backend.transcribe(audio_path, word_timestamps=False)
except Exception as e: # noqa: BLE001 — degrade to the model fallback
if not text:
logger.warning(
"transcribe_reference: %s failed (%s) — deferring to the model's "
"built-in ASR fallback",
backend.id, e,
"transcribe_reference: installed ASR engines returned no transcript "
"— deferring to the model's built-in ASR fallback"
)
return None
result = result or {}
text = result.get("text") or " ".join(
(seg.get("text") or "").strip() for seg in result.get("segments", [])
)
text = (text or "").strip()
if text and fingerprint is not None:
with _ref_transcript_lock:
_ref_transcript_cache[fingerprint] = text
@@ -3036,10 +3208,14 @@ def get_sherpa_dictation_backend(model_id: str) -> "SherpaDictationBackend":
:func:`get_capture_asr_backend`. Thread-safe: the recognizer is shared;
each session creates its own decode stream (see capture_ws)."""
global _capture_backend, _capture_backend_key
from services.performance_profiles import requested_tier
performance_tier = requested_tier("dictation")
_touch_capture() # any handout resets the idle clock
with _capture_backend_lock:
if (isinstance(_capture_backend, SherpaDictationBackend)
and _capture_backend_key == model_id):
and _capture_backend_key == model_id
and _capture_backend.performance_tier == performance_tier):
return _capture_backend
backend = SherpaDictationBackend(model_id=model_id)
_capture_backend = backend
@@ -3047,6 +3223,31 @@ def get_sherpa_dictation_backend(model_id: str) -> "SherpaDictationBackend":
return backend
def sherpa_engine_model_id() -> str:
"""The sherpa model the ``sherpa-onnx-asr`` engine loads when nothing pins
one explicitly: env var (power-user pin) the dictation model the user
picked in Settings / the Engines menu the catalogue default.
Unlike :func:`dictation_model_id` this ignores ``dictation.enabled`` a
user who turned the hotkey off but chose the Sherpa engine for dub/batch
transcription still means *this* model and never returns None: the
engine needs *some* model to construct. A demoted model (decoded nothing
on this host) falls through to the default rather than being re-picked.
"""
from services import sherpa_dictation as _sd
explicit = os.environ.get("OMNIVOICE_SHERPA_ASR_MODEL")
if explicit:
return explicit
try:
from core import prefs
mid = prefs.get("dictation.model_id")
except Exception: # noqa: BLE001 — prefs store unavailable → default
return _sd.DEFAULT_MODEL_ID
if _sd.is_sherpa_model(mid) and not _sd.is_demoted(mid):
return _sd.get_spec(mid).id
return _sd.DEFAULT_MODEL_ID
def dictation_model_id() -> str | None:
"""The selected sherpa dictation model id, or None when dictation is off /
no sherpa model is chosen. Env var wins (power-user pin), then prefs."""
@@ -3085,7 +3286,7 @@ def _parakeet_mlx_installed() -> bool:
trigger a surprise multi-GB download (the asr_model_missing contract).
Installed state comes from the same HF-cache helpers the model store uses
(positive results memoized see :func:`_repo_installed`), so the answer
matches the Model Catalogue Models install badges. Never raises.
matches the Model Catalogue's install badges. Never raises.
"""
try:
repo = os.environ.get("ASR_MODEL_PARAKEET_MLX", _PARAKEET_MLX_DEFAULT)
@@ -3192,8 +3393,11 @@ def get_capture_asr_backend(*, skip_sherpa: bool = False) -> ASRBackend:
if sherpa_id:
ok, _ = SherpaDictationBackend.is_available()
if ok:
from services.performance_profiles import requested_tier
performance_tier = requested_tier("dictation")
if not (isinstance(_capture_backend, SherpaDictationBackend)
and _capture_backend_key == sherpa_id):
and _capture_backend_key == sherpa_id
and _capture_backend.performance_tier == performance_tier):
try:
_capture_backend = SherpaDictationBackend(model_id=sherpa_id)
_capture_backend_key = sherpa_id
@@ -3215,7 +3419,7 @@ def get_capture_asr_backend(*, skip_sherpa: bool = False) -> ASRBackend:
# Prefer an already-installed Parakeet TDT v3 on Apple Silicon (when
# the language gate allows it — see _capture_prefers_parakeet). Gated
# on the weights being on disk so this NEVER triggers a download —
# users opt in by installing the model from Model Catalogue → Models. The
# users opt in by installing the model from the engine's Weights list in Model Catalogue. The
# gate's answer is part of the warm-singleton key so installing
# parakeet mid-session rebuilds the singleton instead of serving the
# stale whisper pick until restart (the memo in _repo_installed keeps
@@ -3268,6 +3472,34 @@ ASR_MODEL_MISSING = "asr_model_missing"
_PYTORCH_ASR_DEFAULT = "openai/whisper-large-v3-turbo"
_FASTER_WHISPER_DEFAULT = "Systran/faster-whisper-large-v3"
def faster_whisper_model_id() -> str:
"""Resolve the UI-selected CTranslate2 model, with env pins authoritative."""
from core import prefs
return str(
prefs.resolve(
"asr_model_faster",
env="ASR_MODEL_FASTER",
default=_FASTER_WHISPER_DEFAULT,
)
)
def select_faster_whisper_model(repo_id: str) -> None:
"""Persist and apply a CTranslate2 model selection for this process."""
from core import prefs
if prefs.is_env_shadowed("ASR_MODEL_FASTER"):
raise ValueError("ASR_MODEL_FASTER is set outside VoiceStudio")
prefs.set_("asr_model_faster", repo_id)
# Sidecars inherit the process environment. Updating it here makes the
# selection effective immediately as well as after the next app launch.
os.environ["ASR_MODEL_FASTER"] = repo_id
instance = _ISOLATED_INSTANCES.pop("faster-whisper-isolated", None)
if instance is not None:
instance.shutdown()
# faster-whisper / WhisperX short model aliases → the HF repo they download.
# Covers our own defaults plus the documented size aliases; an unrecognized
# alias returns None and the preflight stays out of the way (never blocks).
@@ -3300,7 +3532,7 @@ def _offline_asr_repo(backend_id: str | None = None) -> str | None:
if bid == "whisperx":
return _fw_repo(os.environ.get("ASR_MODEL_WHISPERX", "large-v3"))
if bid == "faster-whisper":
return _fw_repo(os.environ.get("ASR_MODEL_FASTER", _FASTER_WHISPER_DEFAULT))
return _fw_repo(faster_whisper_model_id())
if bid == "faster-whisper-isolated":
# Mirror the sidecar's own resolution (_asr_sidecar/main.py):
# ASR_MODEL_FW is a sidecar-only override, otherwise the shared
@@ -3308,7 +3540,7 @@ def _offline_asr_repo(backend_id: str | None = None) -> str | None:
# download a different repo than the sidecar will load.
return _fw_repo(
os.environ.get("ASR_MODEL_FW")
or os.environ.get("ASR_MODEL_FASTER")
or faster_whisper_model_id()
or _FASTER_WHISPER_DEFAULT
)
if bid == "mlx-whisper":
@@ -3321,9 +3553,7 @@ def _offline_asr_repo(backend_id: str | None = None) -> str | None:
# Unknown/none → fail open.
try:
from services import sherpa_dictation as _sd
spec = _sd.get_spec(
os.environ.get("OMNIVOICE_SHERPA_ASR_MODEL", _sd.DEFAULT_MODEL_ID)
)
spec = _sd.get_spec(sherpa_engine_model_id())
return spec.repo_id if spec is not None else None
except Exception: # noqa: BLE001 — preflight must stay best-effort
return None
@@ -3351,7 +3581,7 @@ def _capture_whisper_repo() -> str | None:
# resolve but our alias table doesn't know) yields None here — FAIL
# OPEN rather than coerce to the default repo and demand a download
# of a model the user never picked.
return _fw_repo(os.environ.get("ASR_MODEL_FASTER", _FASTER_WHISPER_DEFAULT))
return _fw_repo(faster_whisper_model_id())
return os.environ.get("OMNIVOICE_PYTORCH_ASR_MODEL", _PYTORCH_ASR_DEFAULT)
@@ -3423,12 +3653,12 @@ def _recommended_asr_model(
_INSTALLED_REPO_MEMO: set[str] = set()
def _repo_installed(repo: str) -> bool:
def _repo_installed(repo: str, *, refresh: bool = False) -> bool:
"""``is_cached`` + ``cache_is_complete`` with a positive-only session memo.
Installed state comes from the same HF-cache helpers the model store uses,
so the answer matches the Model Catalogue Models install badges."""
if repo in _INSTALLED_REPO_MEMO:
if not refresh and repo in _INSTALLED_REPO_MEMO:
return True
from api.routers.setup.models import cache_is_complete, get_model_catalog, is_cached
meta = get_model_catalog().get(repo) or {"repo_id": repo}
@@ -3453,7 +3683,7 @@ def asr_model_missing_error(*, purpose: str = "transcribe",
``sherpa_model_id`` lets the live-dictation WS pass its per-session
``?model=`` override. Installed state comes from the same HF-cache helpers
the model store uses (see :func:`_repo_installed`), so the answer matches
the Model Catalogue Models install badges.
the Model Catalogue's install badges.
``skip_sherpa`` probes only the non-Sherpa capture fallback; silent-model
recovery uses it before deciding whether persistent demotion is warranted.
``require_installed`` makes unknown/custom selections fail closed for that
@@ -3506,7 +3736,7 @@ def asr_model_missing_error(*, purpose: str = "transcribe",
return None # explicit opt-in engine — can't (and shouldn't) preflight
from api.routers.setup.models import get_model_catalog
if require_installed:
if _repo_installed(repo):
if _repo_installed(repo, refresh=True):
return None
return {
"error": ASR_MODEL_MISSING,
@@ -3531,6 +3761,14 @@ def asr_model_missing_error(*, purpose: str = "transcribe",
),
}
except Exception: # noqa: BLE001 — preflight is best-effort, never a blocker
if require_installed:
logger.warning("ASR install preflight failed; refusing implicit download", exc_info=True)
return {
"error": ASR_MODEL_MISSING,
"missing_repo_id": "unverified-local-model",
"reason": "verification_failed",
"recommended": None,
}
logger.warning("ASR install preflight failed — proceeding without it",
exc_info=True)
return None
@@ -3539,12 +3777,15 @@ def asr_model_missing_error(*, purpose: str = "transcribe",
def asr_model_missing_detail(payload: dict) -> str:
"""Human-readable (English) fallback message for the typed payload —
what legacy clients / logs see; the frontend renders its own i18n copy."""
if payload.get("reason") == "verification_failed":
return ("Could not verify the local speech-to-text model. "
"Check Settings > Logs > Backend, then retry. No model was downloaded.")
rec = payload.get("recommended") or {}
if rec.get("label"):
return (
"No speech-to-text model is installed. Download "
f"{rec['label']} ({rec['size_gb']} GB) from Model Catalogue → Models, "
f"{rec['label']} ({rec['size_gb']} GB) from the engine's Weights list in Model Catalogue, "
"then retry."
)
return ("No speech-to-text model is installed. Download one from "
"Model Catalogue → Models, then retry.")
"the engine's Weights list in Model Catalogue, then retry.")
+20
View File
@@ -186,6 +186,26 @@ def trim_trailing_silence(
return audio_tensor[..., :end]
def trim_speech_padding(audio_tensor: torch.Tensor, sample_rate: int) -> torch.Tensor:
"""Remove generated edge silence before timing, retaining 50 ms of context.
Never compress silence into the spoken slot or delete internal pauses.
Silent/invalid outputs remain intact for the generation integrity guard.
"""
if audio_tensor.numel() == 0 or sample_rate <= 0:
return audio_tensor
envelope = audio_tensor.abs()
if envelope.ndim > 1:
envelope = envelope.amax(dim=tuple(range(envelope.ndim - 1)))
voiced = torch.nonzero(envelope > 10 ** (-50 / 20))
if voiced.numel() == 0:
return audio_tensor
margin = int(sample_rate * 0.05)
start = max(0, int(voiced[0].item()) - margin)
end = min(audio_tensor.shape[-1], int(voiced[-1].item()) + 1 + margin)
return audio_tensor[..., start:end]
def apply_effects_chain(audio_tensor, sample_rate: int, chain: list[dict]) -> torch.Tensor:
"""Apply a chain of named effects to an audio tensor.
+10 -1
View File
@@ -66,6 +66,12 @@ logger = logging.getLogger("omnivoice.audio_io")
PathOrBuf = Union[str, "os.PathLike[str]", BinaryIO, io.IOBase]
def _ensure_audio_parent(path_or_buf: PathOrBuf) -> None:
"""Recover app output folders removed after backend initialization."""
if isinstance(path_or_buf, (str, os.PathLike)):
os.makedirs(os.path.dirname(os.path.abspath(path_or_buf)), exist_ok=True)
def _safe_torchaudio_save(
path_or_buf: PathOrBuf,
tensor: torch.Tensor,
@@ -156,6 +162,7 @@ def _safe_torchaudio_save(
fmt = (format or "wav").lower()
try:
_ensure_audio_parent(path_or_buf)
if fmt == "wav":
torchaudio.save(
path_or_buf,
@@ -313,6 +320,7 @@ def _safe_soundfile_write(
else:
samples = np.ascontiguousarray(samples)
_ensure_audio_parent(path)
sf.write(path, samples, sample_rate, subtype=subtype)
@@ -334,7 +342,7 @@ def atomic_save_wav(
publication AND audited tensor normalization.
Args:
target_path: Final destination. Parent directory must already exist.
target_path: Final destination. Missing parent directories are recreated.
audio: ``(channels, samples)`` or ``(samples,)`` tensor.
sample_rate: WAV sample rate in Hz.
**kwargs: Forwarded to ``_safe_torchaudio_save`` (``format``,
@@ -346,6 +354,7 @@ def atomic_save_wav(
unlinked on failure so we do not leak ``.tmp`` files in
``DUB_DIR``.
"""
_ensure_audio_parent(target_path)
target_dir = os.path.dirname(target_path) or "."
target_base = os.path.basename(target_path)
# The temp file must end in ``.wav`` even though it is conceptually a
@@ -0,0 +1,197 @@
"""Checksummed, user-triggered installer for the native audio.cpp runtime."""
from __future__ import annotations
import hashlib
import logging
import os
from pathlib import Path
import shutil
import subprocess
import tarfile
import tempfile
import threading
import urllib.request
import zipfile
from engines.audiocpp import bootstrap
logger = logging.getLogger("omnivoice.audiocpp.install")
_CHUNK = 256 * 1024
_lock = threading.Lock()
_job = {"state": "idle", "progress": 0.0, "error": None}
def _snapshot() -> dict:
with _lock:
return dict(_job)
def _update(**fields) -> None:
with _lock:
_job.update(fields)
def _runtime_paths() -> tuple[Path | None, Path | None]:
try:
server = bootstrap.resolve_server_binary()
except (OSError, RuntimeError):
return None, None
cli = server.with_name("audiocpp_cli.exe" if os.name == "nt" else "audiocpp_cli")
if not cli.is_file() or (os.name != "nt" and not os.access(cli, os.X_OK)):
return server, None
return server, cli
def status() -> dict:
server, cli = _runtime_paths()
managed = bootstrap.managed_runtime_dir()
is_managed = bool(
server
and server.resolve() == (managed / bootstrap.binary_name()).resolve()
)
return {
"supported": bootstrap.default_asset() is not None,
"installed": server is not None and cli is not None,
"managed": is_managed,
"version": bootstrap.VERSION if server and cli and is_managed else None,
"platform": bootstrap.platform_slug(),
"job": _snapshot(),
}
def _download(url: str, destination: Path, digest: str, expected_size: int) -> None:
if not url.startswith("https://github.com/"):
raise ValueError("audio.cpp downloads require the pinned GitHub release")
request = urllib.request.Request(url, headers={"User-Agent": "VoiceStudio"})
hasher = hashlib.sha256()
received = 0
with urllib.request.urlopen(request, timeout=30) as response, destination.open("wb") as out:
total = expected_size or int(response.headers.get("Content-Length") or 0)
while chunk := response.read(_CHUNK):
out.write(chunk)
hasher.update(chunk)
received += len(chunk)
if total:
_update(progress=min(received / total, 0.9))
if received != expected_size:
raise RuntimeError("The audio.cpp runtime download size did not match the release")
if hasher.hexdigest() != digest:
raise RuntimeError("The audio.cpp runtime checksum did not match the release")
def _safe_destination(root: Path, name: str) -> Path:
destination = (root / name.replace("\\", "/")).resolve()
if destination != root and root not in destination.parents:
raise RuntimeError("The audio.cpp archive contains an unsafe path")
return destination
def _extract(archive: Path, destination: Path) -> None:
destination.mkdir(parents=True)
root = destination.resolve()
if archive.suffix.lower() == ".zip":
with zipfile.ZipFile(archive) as bundle:
for member in bundle.infolist():
target = _safe_destination(root, member.filename)
if member.is_dir():
target.mkdir(parents=True, exist_ok=True)
continue
target.parent.mkdir(parents=True, exist_ok=True)
with bundle.open(member) as source, target.open("wb") as out:
shutil.copyfileobj(source, out)
return
with tarfile.open(archive, "r:*") as bundle:
for member in bundle.getmembers():
target = _safe_destination(root, member.name)
if member.isdir():
target.mkdir(parents=True, exist_ok=True)
continue
if not member.isfile():
raise RuntimeError("The audio.cpp archive contains an unsupported link")
target.parent.mkdir(parents=True, exist_ok=True)
source = bundle.extractfile(member)
if source is None:
raise RuntimeError("The audio.cpp archive contains an unreadable file")
with source, target.open("wb") as out:
shutil.copyfileobj(source, out)
target.chmod(member.mode & 0o700)
def _install() -> None:
asset = bootstrap.default_asset()
expected_size = bootstrap.default_asset_size()
if asset is None or expected_size is None:
raise RuntimeError("No audio.cpp runtime is published for this platform")
filename, digest = asset
url = (
f"https://github.com/{bootstrap.GH_REPO}/releases/download/"
f"{bootstrap.VERSION}/{filename}"
)
target = bootstrap.managed_runtime_dir()
target.parent.mkdir(parents=True, exist_ok=True)
with tempfile.TemporaryDirectory(prefix="audiocpp-install-", dir=target.parent) as temp:
temp_path = Path(temp)
archive = temp_path / filename
_download(url, archive, digest, expected_size)
extracted = temp_path / "extracted"
_extract(archive, extracted)
candidates = sorted(
extracted.rglob(bootstrap.binary_name()), key=lambda path: len(path.parts)
)
if not candidates:
raise RuntimeError("The audio.cpp release does not contain its server")
source_dir = candidates[0].parent
cli_name = "audiocpp_cli.exe" if os.name == "nt" else "audiocpp_cli"
if not (source_dir / cli_name).is_file():
raise RuntimeError("The audio.cpp release does not contain its CLI")
prepared = temp_path / "prepared"
shutil.copytree(source_dir, prepared)
for executable in (prepared / bootstrap.binary_name(), prepared / cli_name):
executable.chmod(0o700)
probe = subprocess.run( # nosec B603 -- checksummed fixed release binary
[str(prepared / bootstrap.binary_name()), "--list-devices"],
capture_output=True,
timeout=20,
check=False,
)
if probe.returncode != 0:
raise RuntimeError("The downloaded audio.cpp runtime failed its device check")
if target.exists():
shutil.rmtree(target)
os.replace(prepared, target)
bootstrap.invalidate()
def start_install(*, wait: bool = False) -> dict:
current = status()
if current["installed"]:
return {"status": "already_installed", **current}
if not current["supported"]:
raise RuntimeError("No audio.cpp runtime is published for this platform")
with _lock:
running = _job["state"] == "running"
if not running:
_job.update(state="running", progress=0.0, error=None)
if running:
return {"status": "already_running", **status()}
def worker() -> None:
try:
_install()
_update(state="done", progress=1.0, error=None)
except Exception:
logger.exception("audio.cpp runtime installation failed")
_update(
state="error",
error="The audio.cpp runtime could not be installed. Check the backend log.",
)
if wait:
worker()
else:
threading.Thread(target=worker, name="audiocpp-install", daemon=True).start()
return {"status": "started", **status()}
def reset_job_for_tests() -> None:
_update(state="idle", progress=0.0, error=None)
+37
View File
@@ -0,0 +1,37 @@
"""Resolve the installed pyannote bundle without network access at job time."""
from contextlib import contextmanager
from pathlib import Path
from tempfile import TemporaryDirectory
@contextmanager
def local_pipeline_config():
import yaml
from huggingface_hub import hf_hub_download
from huggingface_hub.constants import HF_HUB_CACHE
from services.hf_revisions import installed_revision
def cached(repo: str, filename: str) -> str:
return hf_hub_download(
repo_id=repo,
filename=filename,
revision=installed_revision(repo, HF_HUB_CACHE),
local_files_only=True,
)
config_path = cached("pyannote/speaker-diarization-3.1", "config.yaml")
config = yaml.safe_load(Path(config_path).read_text(encoding="utf-8"))
params = config["pipeline"]["params"]
# The reviewed pipeline references these two checkpoints. Local checkpoint
# paths prevent pyannote's nested Model.from_pretrained calls fetching them.
for key, repo in (
("segmentation", "pyannote/segmentation-3.0"),
("embedding", "pyannote/wespeaker-voxceleb-resnet34-LM"),
):
if params.get(key) != repo:
raise ValueError(f"Unexpected diarisation {key} repository; repair the installed pipeline")
params[key] = cached(repo, "pytorch_model.bin")
with TemporaryDirectory(prefix="voicestudio-pyannote-") as directory:
path = Path(directory) / "config.yaml"
path.write_text(yaml.safe_dump(config), encoding="utf-8")
yield str(path)
+163
View File
@@ -0,0 +1,163 @@
"""Explicit local audio.cpp Sortformer adapter for the shared diarisation flow."""
from __future__ import annotations
import json
import logging
import os
from pathlib import Path
import subprocess
import time
import threading
from tempfile import TemporaryDirectory
logger = logging.getLogger("omnivoice.diarisation.native")
_process_lock = threading.Lock()
_processes: set = set()
MAX_V1_AUDIO_SECONDS = 120.0
SORTFORMER_FRAME_SAMPLES = 1280 # 80 ms at the required 16 kHz input rate.
def _sortformer_command(binary: Path, model: Path, device, source: Path, output: Path):
return [
str(binary), "--task", "diar", "--family", "sortformer_diar",
"--model", str(model), "--backend", device.backend,
"--device", str(device.index), "--audio", str(source),
"--turns-out", str(output),
# Accelerator builds otherwise keep the default 20-second fixed graph
# and reject ordinary clips. Grow remains bounded by the v1 limit below.
"--session-option", "graph_capacity_mode=grow",
]
def _validated_turn(turn: dict, audio_frames: int) -> tuple[int, int, str]:
start, end = turn.get("start_sample"), turn.get("end_sample")
speaker = turn.get("speaker_id")
if (
type(start) is not int
or type(end) is not int
or start < 0
or start >= end
or end > audio_frames + SORTFORMER_FRAME_SAMPLES
or not isinstance(speaker, str)
or not speaker
):
raise ValueError("Invalid native speaker-turn boundaries")
# The decoder works in 80 ms frames and can pad its final turn one frame
# beyond a non-aligned WAV boundary. Keep the timeline inside the media.
return start, min(end, audio_frames), speaker
def is_running() -> bool:
with _process_lock:
return bool(_processes)
class NativeSortformer:
"""Stateless native invocation; the GGUF is never downloaded implicitly."""
def __init__(self):
from engines.audiocpp.bootstrap import resolve_server_binary
from services.diarization_runtime import sortformer_model_path
try:
self.model = sortformer_model_path()
except Exception as exc:
raise FileNotFoundError(
"Install the audio.cpp Sortformer model in Settings > Models > Diarisation"
) from exc
if not self.model.is_file() or self.model.suffix.lower() != ".gguf":
raise FileNotFoundError("The configured Sortformer GGUF is missing")
with self.model.open("rb") as model_file:
if model_file.read(4) != b"GGUF":
raise ValueError("The configured Sortformer model is not a GGUF file")
server = resolve_server_binary()
self.binary = server.with_name("audiocpp_cli.exe" if os.name == "nt" else "audiocpp_cli")
if not self.binary.is_file():
raise FileNotFoundError("The installed audio.cpp directory has no audiocpp_cli")
def __call__(self, audio_path, *, num_speakers=None, job_id=None, cancel_check=None):
if num_speakers is not None:
raise ValueError("Sortformer v1 detects up to four speakers but cannot enforce an exact speaker count")
import soundfile as sf
from pyannote.core import Annotation, Segment
from core.contained_subprocess import spawn_owned
from engines.audiocpp.bootstrap import resolve_compute_selection
from services.proc_registry import register_proc, unregister_proc
def check_cancelled():
if cancel_check is not None and cancel_check():
raise RuntimeError("Native diarisation cancelled")
check_cancelled()
audio_info = sf.info(str(audio_path))
if audio_info.duration > MAX_V1_AUDIO_SECONDS:
raise ValueError(
"Sortformer v1 supports recordings up to 120 seconds; "
"select pyannote for longer recordings"
)
device = resolve_compute_selection().device
with TemporaryDirectory(prefix="voicestudio-sortformer-") as directory:
source = Path(audio_path).resolve()
output = Path(directory) / "turns.json"
command = _sortformer_command(
self.binary, self.model, device, source, output
)
def run_owned(command, log_name):
with (Path(directory) / log_name).open("wb") as log:
check_cancelled()
process = spawn_owned(command, stdout=log, stderr=subprocess.STDOUT)
with _process_lock:
_processes.add(process)
try:
if job_id is not None:
register_proc(job_id, process)
deadline = time.monotonic() + 600
while True:
check_cancelled()
remaining = deadline - time.monotonic()
if remaining <= 0:
raise subprocess.TimeoutExpired(command, 600)
try:
code = process.wait(timeout=min(0.25, remaining))
break
except subprocess.TimeoutExpired:
continue
except BaseException:
process.kill()
process.wait()
raise
finally:
with _process_lock:
_processes.discard(process)
if job_id is not None:
unregister_proc(job_id, process)
check_cancelled()
if code != 0:
with (Path(directory) / log_name).open("rb") as diagnostic:
diagnostic.seek(0, 2)
diagnostic.seek(max(0, diagnostic.tell() - 8192))
tail = diagnostic.read().decode("utf-8", errors="replace")
logger.error("Sortformer exited with %s; native log tail:\n%s", code, tail)
raise RuntimeError(f"Native Sortformer failed (exit {code})")
if (audio_info.samplerate != 16000 or audio_info.channels != 1
or audio_info.format != "WAV" or audio_info.subtype != "PCM_16"):
from services.ffmpeg_utils import find_ffmpeg
normalized = Path(directory) / "input.wav"
run_owned([
find_ffmpeg(), "-nostdin", "-hide_banner", "-loglevel", "error", "-y",
"-i", str(source), "-vn", "-ac", "1", "-ar", "16000",
"-c:a", "pcm_s16le", str(normalized),
], "normalize.log")
source = normalized
audio_info = sf.info(str(source))
command[command.index("--audio") + 1] = str(source)
run_owned(command, "native.log")
turns = json.loads(output.read_text(encoding="utf-8"))
if not isinstance(turns, list):
raise ValueError("Invalid native speaker-turn output")
annotation = Annotation()
for index, turn in enumerate(turns):
start, end, speaker = _validated_turn(turn, audio_info.frames)
annotation[Segment(start / 16000, end / 16000), index] = speaker
return annotation
+118
View File
@@ -0,0 +1,118 @@
"""Persisted selection and installed-only resolution for diarisation runtimes."""
from __future__ import annotations
import os
from pathlib import Path
from core import prefs
PYANNOTE = "pyannote"
SORTFORMER = "audiocpp-sortformer"
SORTFORMER_REPO = "audio-cpp/audio.cpp-gguf"
SORTFORMER_FILE = "Sortformer-Diar-4spk-v1-GGUF/sortformer-diar-4spk-v1-q8_0.gguf"
_SORTFORMER_MODEL_MISSING = "Install the Sortformer model bundle"
_SORTFORMER_MODEL_BROKEN = "Repair the installed Sortformer model bundle"
_SORTFORMER_RUNTIME_MISSING = (
"The Sortformer model is installed. Install the audio.cpp runtime to use it"
)
_SORTFORMER_CLI_MISSING = (
"The installed audio.cpp runtime does not include speaker diarisation"
)
def selected_backend() -> str:
value = str(
prefs.resolve(
"diarization_backend",
env="OMNIVOICE_DIARIZATION_BACKEND",
default=PYANNOTE,
)
).strip()
return value if value in {PYANNOTE, SORTFORMER} else PYANNOTE
def sortformer_model_path() -> Path:
configured = os.environ.get("OMNIVOICE_DIARIZATION_MODEL", "").strip()
if configured:
return Path(configured).expanduser().resolve()
# An installed-only lookup never reaches the network. Installation remains
# an explicit Model Library action through the reviewed audio.cpp bundle.
from huggingface_hub import hf_hub_download
from services.hf_revisions import revision_for
return Path(
hf_hub_download(
repo_id=SORTFORMER_REPO,
filename=SORTFORMER_FILE,
revision=revision_for(SORTFORMER_REPO),
local_files_only=True,
)
).resolve()
def select_backend(backend: str) -> None:
if backend not in {PYANNOTE, SORTFORMER}:
raise ValueError("Unknown diarisation engine")
prefs.set_("diarization_backend", backend)
def sortformer_status() -> dict:
"""Return path-free readiness for the model and its native executable."""
status = {
"model": SORTFORMER_REPO,
"model_installed": False,
"runtime_installed": False,
"installed": False,
"reason": _SORTFORMER_MODEL_MISSING,
}
try:
model = sortformer_model_path()
except Exception:
return status
try:
if not model.is_file() or model.suffix.lower() != ".gguf":
return status
with model.open("rb") as model_file:
if model_file.read(4) != b"GGUF":
status["reason"] = _SORTFORMER_MODEL_BROKEN
return status
except OSError:
status["reason"] = _SORTFORMER_MODEL_BROKEN
return status
status["model_installed"] = True
status["reason"] = _SORTFORMER_RUNTIME_MISSING
try:
from engines.audiocpp.bootstrap import resolve_server_binary
server = resolve_server_binary()
except (OSError, RuntimeError):
return status
cli = server.with_name("audiocpp_cli.exe" if os.name == "nt" else "audiocpp_cli")
if not cli.is_file() or (os.name != "nt" and not os.access(cli, os.X_OK)):
status["reason"] = _SORTFORMER_CLI_MISSING
return status
status.update(runtime_installed=True, installed=True, reason=None)
return status
def installed_backends() -> set[str]:
"""Return complete local runtimes without loading weights or downloading."""
installed: set[str] = set()
from api.routers.setup.models import KNOWN_MODELS, cache_is_complete, is_cached
repo_id = "pyannote/speaker-diarization-3.1"
spec = next(model for model in KNOWN_MODELS if model["repo_id"] == repo_id)
if is_cached(repo_id) and cache_is_complete(spec):
installed.add(PYANNOTE)
try:
native = sortformer_status()
if not native["installed"]:
return installed
installed.add(SORTFORMER)
except (OSError, RuntimeError, ValueError):
pass
return installed
+130
View File
@@ -0,0 +1,130 @@
"""Dialogue-only replacement beds: original outside speech, separated bed inside."""
from __future__ import annotations
import asyncio
import hashlib
import json
import math
import os
import tempfile
from pathlib import Path
from services.ffmpeg_utils import find_ffmpeg, run_ffmpeg
from services.video_retime import expand_retime_chunks
RATE = 48000
FADE_S = .01
_locks: dict[str, asyncio.Lock] = {}
def dialogue_intervals(segments: list[dict]) -> list[tuple[float, float]]:
intervals = []
for row in segments:
a, b = float(row['start']), float(row['end'])
if not math.isfinite(a) or not math.isfinite(b) or a < 0 or b <= a:
raise ValueError('Invalid dialogue interval')
intervals.append((a, b))
merged: list[tuple[float, float]] = []
for a, b in sorted(intervals):
if merged and a <= merged[-1][1]:
merged[-1] = (merged[-1][0], max(b, merged[-1][1]))
else:
merged.append((a, b))
return merged
def splice_background(original: str, separated: str, output: str, intervals: list[tuple[float, float]]) -> None:
"""Stream in bounded memory; crossfades lie INSIDE dialogue intervals."""
import numpy as np
import soundfile as sf
with sf.SoundFile(original) as src, sf.SoundFile(separated) as bed:
if src.samplerate != bed.samplerate or src.channels != bed.channels:
raise ValueError('Background inputs must have matching sample format')
if bed.frames < src.frames - int(.1 * src.samplerate):
raise ValueError('Separated background is incomplete')
with sf.SoundFile(output, 'w', samplerate=src.samplerate, channels=src.channels, subtype='FLOAT') as out:
offset = 0
active = 0
while True:
wave = src.read(65536, dtype='float32', always_2d=True)
if not len(wave):
break
background = bed.read(len(wave), dtype='float32', always_2d=True)
if len(background) < len(wave):
background = np.pad(background, ((0, len(wave)-len(background)), (0, 0)))
times = np.arange(offset, offset + len(wave)) / src.samplerate
mask = np.zeros(len(wave), dtype='float32')
while active < len(intervals) and intervals[active][1] < times[0]:
active += 1
for a, b in intervals[active:]:
if a > times[-1]:
break
fade = min(FADE_S, (b-a)/2)
envelope = np.clip(np.minimum((times-a)/fade, (b-times)/fade), 0, 1)
mask = np.maximum(mask, envelope)
out.write(wave * (1-mask[:, None]) + background * mask[:, None])
offset += len(wave)
async def _checked(cmd: list[str]) -> None:
rc, _, error = await run_ffmpeg(cmd, timeout=1800.0)
if rc:
raise RuntimeError('Could not preserve original background audio: ' + str(error)[-500:])
async def surgical_background(source: str, separated: str, cache_dir: str, segments: list[dict], plan: list[dict], duration: float) -> str:
for chunk in plan:
ratio = float(chunk["stretch_ratio"])
if not math.isfinite(ratio) or ratio <= 0:
raise ValueError("Invalid background retiming ratio")
intervals = dialogue_intervals(segments)
if not intervals:
raise ValueError('Dialogue timing is required to preserve original background audio')
identity = [(p, os.stat(p).st_size, os.stat(p).st_mtime_ns) for p in (source, separated)]
key = hashlib.sha256(json.dumps([1, identity, intervals, plan, duration], sort_keys=True).encode()).hexdigest()[:24]
target = str(Path(cache_dir) / f'surgical_{key}.wav')
async with _locks.setdefault(target, asyncio.Lock()):
if os.path.isfile(target):
return target
ffmpeg = find_ffmpeg()
with tempfile.TemporaryDirectory(prefix='.surgical-', dir=cache_dir) as tmp:
original, bed, spliced = [str(Path(tmp)/name) for name in ('source.wav', 'bed.wav', 'spliced.wav')]
for inp, out in ((source, original), (separated, bed)):
await _checked([ffmpeg, '-y', '-i', inp, '-map', '0:a:0', '-vn', '-ar', str(RATE), '-ac', '2', '-c:a', 'pcm_f32le', out])
await asyncio.to_thread(splice_background, original, bed, spliced, intervals)
if plan and any(abs(float(p['stretch_ratio'])-1) > 1e-6 for p in plan):
chunks = expand_retime_chunks(plan, duration)
# Bound filter buffering for long projects; trim each batch's
# input before splitting it among the chunk filters.
batches = []
for batch_index in range(0, len(chunks), 16):
batch = chunks[batch_index:batch_index+16]
origin = batch[0][0]
filters = []
for i, (a, b, ratio) in enumerate(batch):
rate = 1 / ratio
tempos = []
while rate < .5:
tempos.append('atempo=0.5')
rate /= .5
while rate > 2:
tempos.append('atempo=2')
rate /= 2
tempos.append(f'atempo={rate:.9f}')
length = (b-a)*ratio
filters.append(f'[0:a]atrim=start={a-origin:.9f}:end={b-origin:.9f},asetpts=PTS-STARTPTS,' + ','.join(tempos) + f',apad,atrim=duration={length:.9f}[c{i}]')
filters.append(''.join(f'[c{i}]' for i in range(len(batch))) + f'concat=n={len(batch)}:v=0:a=1[out]')
script = Path(tmp)/'retime.txt'
script.write_text(';'.join(filters))
batch_name = f'batch{batch_index}.wav'
output = str(Path(tmp)/batch_name)
await _checked([ffmpeg, '-y', '-ss', str(origin), '-t', str(batch[-1][1]-origin), '-i', spliced, '-filter_complex_script', str(script), '-map', '[out]', '-c:a', 'pcm_f32le', output])
batches.append(batch_name)
listing = Path(tmp)/'concat.txt'
listing.write_text(''.join(f"file '{name}'\n" for name in batches))
retimed = str(Path(tmp)/'retimed.wav')
await _checked([ffmpeg, '-y', '-f', 'concat', '-safe', '1', '-i', str(listing), '-c:a', 'copy', retimed])
spliced = retimed
os.replace(spliced, target)
return target
+56
View File
@@ -0,0 +1,56 @@
"""Shared native-TTS batching policy for interactive and queued dubbing."""
from __future__ import annotations
import logging
import os
logger = logging.getLogger("omnivoice.dub_batching")
BATCH_WIDTH_ENV = "OMNIVOICE_DUB_BATCH_WIDTH"
_MAX_BATCH_WIDTH = 16
def native_batch_width(backend) -> int:
"""Return a host-safe native batch width for ``backend``."""
override = os.environ.get(BATCH_WIDTH_ENV, "").strip()
if override:
try:
return max(1, min(_MAX_BATCH_WIDTH, int(override)))
except (TypeError, ValueError):
logger.warning(
"%s=%r is not an integer; deriving the batch width from the host",
BATCH_WIDTH_ENV,
override,
)
try:
from core.device_caps import detect_host_caps
caps = detect_host_caps()
except Exception: # noqa: BLE001 - an unprobeable host takes the safe path
return 1
if caps.family == "cpu" or not caps.vram_gb:
return 1
headroom = caps.vram_gb - float(getattr(backend, "min_vram_gb", 0.0) or 0.0)
if headroom < 2.0:
return 1
if headroom < 6.0:
return 2
if headroom < 12.0:
return 4
return 8
def batch_timeout_s(texts: list[str], backend) -> float:
"""Bound one native batch without multiplying the executor base timeout."""
from services.model_manager import generate_timeout_s
floor = generate_timeout_s("", engine=backend)
overage = sum(
max(0.0, generate_timeout_s(text, engine=backend) - floor)
for text in texts
)
return floor + overage
__all__ = ["BATCH_WIDTH_ENV", "batch_timeout_s", "native_batch_width"]
+43 -3
View File
@@ -1084,9 +1084,7 @@ def yt_download_sync(
if sub_langs:
langs = list(sub_langs)
else:
orig = (info.get("language") or "").strip()
manual = list((info.get("subtitles") or {}).keys())
langs = sorted({*manual, *([orig] if orig else [])})
langs = _default_caption_languages(info)
if not langs:
logger.info("No captions available on %s (skipping subtitle pass)", log_safe(url))
else:
@@ -1115,6 +1113,48 @@ def yt_download_sync(
return video_path, title, sub_files
def _default_caption_languages(info: dict) -> list[str]:
"""Return original-language caption tracks without translated auto-captions.
Some extractors omit ``language`` even though yt-dlp exposes an original
automatic-caption track such as ``en-orig``. Treat that explicit suffix as
source metadata so caption-first ingest still works instead of needlessly
loading ASR. Manual tracks remain eligible because they are authored source
material and yt-dlp's ``skip=translated_subs`` guard still applies.
"""
original = str(info.get("language") or "").strip()
manual = {
str(language).strip()
for language in (info.get("subtitles") or {})
if str(language).strip()
}
automatic = {
str(language).strip()
for language in (info.get("automatic_captions") or {})
if str(language).strip()
}
selected = set(manual)
if original:
primary = original.split("-", 1)[0]
has_source_manual = any(
language == original or language.split("-", 1)[0] == primary
for language in manual
)
if not has_source_manual:
for candidate in (
f"{original}-orig",
f"{primary}-orig",
original,
primary,
):
if candidate in automatic:
selected.add(candidate)
break
else:
selected.update(language for language in automatic if language.endswith("-orig"))
return sorted(selected)
def parse_vtt_segments(vtt_path: str) -> list[dict]:
"""Very small WEBVTT parser → list of {start, end, text}.
+3 -3
View File
@@ -15,7 +15,7 @@ Principles (owner-set):
endpoint gets probed first never to decide. No geo-IP lookups, no
third-party calls, no telemetry.
- **Explicit choices are never auto-switched.** A user with an endpoint
configured anywhere (Model Catalogue Models, ``HF_ENDPOINT`` env, the
configured anywhere (Settings Network, ``HF_ENDPOINT`` env, the
``hf_endpoint`` pref) is in manual mode; auto applies only where nothing
was chosen. ``OMNIVOICE_HF_ENDPOINT_MODE=manual`` is a hard env opt-out.
- **Sticky, canonical-first decisions.** With both endpoints reachable the
@@ -65,7 +65,7 @@ _MODE_PREF = "hf_endpoint_mode" # "auto" | "manual"; absent → default
_DECISION_PREF = "hf_endpoint_auto" # cached decision dict (see race())
DECISION_MAX_AGE_S = 7 * 24 * 3600.0 # re-race a decision older than 7 days
PROBE_TIMEOUT_S = 3.0 # short: a probe is not a download
PROBE_TIMEOUT_S = 8.0 # high-latency / China paths often need >3s
MIRROR_SPEEDUP_FACTOR = 3.0 # mirror must be ≥3× faster to win
# Small, stable, long-lived public file for the optional ranged-GET
@@ -272,7 +272,7 @@ def explicit_endpoint():
"""The endpoint the user explicitly configured, or "".
Same resolution the download paths use: ``HF_ENDPOINT`` env (what
Model Catalogue Models persists via user_env and what main.py loads at boot)
Settings Network persists via user_env and what main.py loads at boot)
with the ``hf_endpoint`` pref as fallback. Unlike
``core.failure.configured_hf_mirror`` this does NOT filter the official
endpoint explicitly choosing huggingface.co is still an explicit
+10
View File
@@ -33,10 +33,20 @@ _ESTIMATES: dict[str, dict] = {
"destination": "hf_model_cache",
"deduplication": None,
},
"audiocpp": {
"package_download_bytes": None,
"unique_installed_bytes": None,
"potentially_shared_bytes": None,
"temporary_free_bytes": None,
"confidence": "estimated",
"destination": "hf_model_cache",
"deduplication": None,
},
}
_MODEL_REPOS = {
"omnivoice": "k2-fsa/OmniVoice",
"kittentts": "KittenML/kitten-tts-mini-0.8",
"audiocpp": "audio-cpp/audio.cpp-gguf",
}
+92 -15
View File
@@ -17,7 +17,6 @@ from __future__ import annotations
import importlib.util
import logging
import os
import sys
from typing import Optional
logger = logging.getLogger("omnivoice.engine_env")
@@ -29,6 +28,53 @@ _TORCH_COMPILE_KEY = "perf.torch_compile_disabled"
# (e.g. a brand-new architecture running through PTX forward-compat).
_FORCE_COMPILE_ENV = "OMNIVOICE_FORCE_TORCH_COMPILE"
# #2135: the environment escape hatches that torch itself honours. `main.py`
# sets TORCH_COMPILE_DISABLE/TORCHDYNAMO_DISABLE on win32, `build_engine_env`
# injects TORCH_COMPILE_DISABLE into engine subprocesses, and
# `docs/install/windows.md` tells users to export it — but the in-process gate
# below never read them, so an operator who set the documented variable still
# got a compiled model (and, on a cudagraph mode, a native crash they could not
# turn off). Reading them here makes one knob mean one thing everywhere.
_COMPILE_DISABLE_ENVS = (
"TORCH_COMPILE_DISABLE",
"TORCHDYNAMO_DISABLE",
"TORCHINDUCTOR_DISABLE",
)
_TRUTHY = frozenset({"1", "true", "yes", "on"})
def _env_compile_disabled() -> Optional[str]:
"""The name of the first set-and-truthy compile-disable env var, else None.
Mirrors torch's own reading of these variables so the app's decision and
torch's behaviour cannot disagree — the state the reporter in #2135 hit,
where the log said "torch.compile applied" while TORCH_COMPILE_DISABLE=1
was exported.
"""
for name in _COMPILE_DISABLE_ENVS:
if os.environ.get(name, "").strip().lower() in _TRUTHY:
return name
return None
def _settings_db_path() -> str:
"""The settings DB the compile toggle is actually read from (best-effort).
Logged alongside the toggle because #2135's reporter had three
`omnivoice.db` files on the box and edited one the backend never opened;
naming the path turns "the setting doesn't work" into a one-line diagnosis.
"""
try:
from core.config import DB_PATH
from core.scrub import scrub_text
return scrub_text(str(DB_PATH))
except Exception:
return "<unknown>"
# #278: set (with a reason) the first time torch.compile — or *running* the
# compiled model — fails at runtime in this process. Once set, every later
# load in the same session goes straight to eager instead of re-tripping the
@@ -251,6 +297,15 @@ def should_torch_compile(device: str) -> bool:
"""
if device != "cuda":
return False
# #2135: honoured before every other gate — an explicit env opt-out is the
# user's most direct statement of intent, and it must hold on every
# platform (the reporter was on Linux, where this used to be ignored).
disabled_by = _env_compile_disabled()
if disabled_by is not None:
logger.info(
"torch.compile skipped: %s is set — using eager mode.", disabled_by,
)
return False
if importlib.util.find_spec("triton") is None:
logger.info("torch.compile skipped: Triton unavailable — using eager mode.")
return False
@@ -258,8 +313,18 @@ def should_torch_compile(device: str) -> bool:
from services import settings_store
if settings_store.get_text(_TORCH_COMPILE_KEY, "0") == "1":
logger.info("torch.compile skipped: disabled in Settings (Performance).")
logger.info(
"torch.compile skipped: disabled in Settings (Performance) [%s].",
_settings_db_path(),
)
return False
# #2135: say which DB answered "not disabled". Without this the only
# observable outcome of a toggle that never reached the running
# backend is a log line saying compile was applied anyway.
logger.debug(
"torch.compile: %s not set in %s — compile remains eligible.",
_TORCH_COMPILE_KEY, _settings_db_path(),
)
except Exception:
logger.exception("should_torch_compile: settings read failed; proceeding")
if _compile_runtime_failure is not None:
@@ -328,19 +393,31 @@ def build_engine_env(
except Exception:
logger.exception("build_engine_env: token resolver failed (non-fatal)")
# INST-12: TORCH_COMPILE_DISABLE on Windows when the user opted in.
# The flag is a Windows-only escape hatch — torch.compile OOMs the same
# Triton kernel cache differently on macOS/Linux, so injecting on those
# platforms would just slow the engine for no gain. (The in-process
# should_torch_compile() gate handles the automatic Triton-absence case;
# the subprocess var stays user-driven by design — see test_perf_settings.)
if sys.platform.startswith("win"):
try:
from services import settings_store
# INST-12 (#65), widened to every platform by #2135: TORCH_COMPILE_DISABLE
# when the user opted in. This was win32-only on the theory that
# torch.compile only misbehaves on Windows (no Triton wheel). #2135 is the
# counter-example — a Linux/CUDA host where compile crashes the engine —
# and a Settings toggle that silently does nothing on the user's platform
# is worse than no toggle at all. Cost when enabled on Linux/macOS is a
# slower engine, which is exactly what the user asked for by enabling it.
try:
from services import settings_store
if settings_store.get_text(_TORCH_COMPILE_KEY, "0") == "1":
env["TORCH_COMPILE_DISABLE"] = "1"
except Exception:
logger.exception("build_engine_env: torch_compile_disabled read failed")
if settings_store.get_text(_TORCH_COMPILE_KEY, "0") == "1":
env["TORCH_COMPILE_DISABLE"] = "1"
except Exception:
logger.exception("build_engine_env: torch_compile_disabled read failed")
# #2135: an env opt-out on the parent must reach the child too. Without
# this a user who exported TORCH_COMPILE_DISABLE=1 got an eager parent and
# a compiled sidecar — the inconsistency that made the flag look ignored.
disabled_by = _env_compile_disabled()
if disabled_by is not None:
if env.get("TORCH_COMPILE_DISABLE") != "1":
logger.debug(
"build_engine_env: %s is set — disabling torch.compile in the "
"engine subprocess too.", disabled_by,
)
env["TORCH_COMPILE_DISABLE"] = "1"
return env
+23 -3
View File
@@ -88,14 +88,34 @@ def snapshot(
evidence_state = "loaded"
if isolated and provider is None and actual_device is None:
evidence_state = "subprocess_loaded_provider_unreported"
from core.scrub import scrub_text
runtime_device_name = routing.get("runtime_device_name")
device_name = (
getattr(caps, "device_name", "")
if runtime_device_name is None
else runtime_device_name
)
public_device_name = scrub_text(device_name)[:256]
return {
"implementation_variant": f"{engine_cls.__module__}.{engine_cls.__name__}",
"declared_device_families": list(getattr(engine_cls, "gpu_compat", ("cpu",))),
"declared_device_families": list(
routing.get("gpu_compat", getattr(engine_cls, "gpu_compat", ("cpu",)))
),
"evidence_state": evidence_state,
"actual_execution_provider": provider,
"actual_execution_device": actual_device,
"gpu_name": getattr(caps, "device_name", "") or None,
"gpu_architecture": _gpu_architecture(getattr(caps, "family", "cpu")),
"gpu_name": public_device_name or None,
"gpu_architecture": None
if (
routing.get("runtime_hardware_family")
and not routing.get("runtime_device_verified")
)
else _gpu_architecture(
routing.get("runtime_hardware_family")
or getattr(caps, "family", "cpu")
),
"runtime_vram_gb": routing.get("runtime_vram_gb"),
"precision_or_quantization": precision,
"cpu_fallback_reason": runtime_fallback_reason or (routing.get("routing_reason") if fallback else None),
"cpu_fallback_stage": runtime_fallback_stage or ("routing_preflight" if fallback else None),

Some files were not shown because too many files have changed in this diff Show More