d91beef0fd314250d8d9b94de86dfea019a8bd96
50
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4e5e8d1f89 |
Allow OmniVoice slow sidecar startup (#1743)
Fixes #1711.\n\nGives only the OmniVoice subprocess a 120-second readiness budget while retaining the shared 30-second default for all other sidecars, with regression coverage. |
||
|
|
3b6e15dad5 | fix(macos): isolate OmniVoice MPS generation | ||
|
|
ef1cb57944 | fix: route OmniVoice to ROCm GPUs (#1647) | ||
|
|
0d3c81b596 |
feat(gguf): add Linux ARM64 runtime support (#1641)
Add linux-aarch64 platform detection, Vulkan-preferred source builds with CPU fallback, ARM64-safe PyTorch dependency markers, native artifact CI, regression coverage, and synchronized architecture documentation. |
||
|
|
3f5114923b |
feat(pockettts): opt-in 24-layer checkpoints via OMNIVOICE_POCKETTTS_24L (#1613)
Adds an opt-in 24-layer PocketTTS checkpoint path, with French correctly using its required 24-layer model. |
||
|
|
43f1d46fe6 |
fix(indextts): accept the config name upstream ships, and keep long text alive (#1619)
* fix(indextts): accept the config name upstream ships, and keep long text alive Two independent defects, both reported on a working IndexTTS 2.5 install. Install always failed. IndexTeam/IndexTTS-2.5 ships the model config as config.yaml — at the pinned revision d0aa86e7 and at HEAD; config_v2_5.yaml exists in no upstream revision. VoiceStudio demanded that name, so _weights_floor_ok never found it and the install died claiming 'the download was likely interrupted' when the download had been perfect. The only way through was to hand-rename the file. Both names are accepted now, in the installer and on the load path, so installs created with the workaround keep working without a reinstall. Long text was killed at 60s. infer() is one blocking upstream call that puts nothing on the wire, and IndexTTS was the only sidecar still on the 60s recv_timeout_s class default while pockettts and omnivoice-subprocess had both raised theirs. Raising the default alone does not fix it — which is why the reporter's RECV_TIMEOUT_S=3600 edit didn't help: progress frames are also what report activity to the GPU pool's execution clock (#1367), so a silent sidecar still trips the outer generate budget. The sidecar now heartbeats every 5s while infer() runs (and during the cold model construction), _send takes a lock so the beat thread can't interleave framing, and the deadline rises to 900s via OMNIVOICE_INDEXTTS_RECV_TIMEOUT_S. test_indextts25_health_requires_25_config_name asserted the bug — that a checkout holding only config.yaml is unhealthy — so it is rewritten to the corrected contract, including that a genuinely truncated download is still caught. Fixes #1611 * test(indextts): follow the installed config name in the sidecar loader tests Two more tests encoded the config_v2_5.yaml assumption, both asserting cfg_path against a directory where no config existed at all — so they were pinning the literal name rather than the resolution. They now lay down a real checkpoints/ tree and assert the resolved path, including that a checkout carrying the pre-fix hand-renamed config still resolves. Caught by the full suite; the targeted runs during development did not reach tests/backend/services/. * test(indextts): event-driven heartbeat tests, real interleaving proof, precedence pin Review round on #1619 — all four findings taken. - The docs line naming 0.5.1 is version-neutral now ('Earlier installs') — version labels are the owner's call. - The heartbeat tests waited on wall-clock sleeps; they now block on a per-write Event with a bounded deadline, so scheduler load can't flake them. - The _send test asserted the lock EXISTS — a tautology. It now drives four concurrent writers through a stream that yields between every byte and asserts every frame decodes; verified fail-before by removing the lock (torn frame) and pass-after. - The precedence test deleted config.yaml before creating the renamed one, so reversed precedence still passed. Both files now coexist for the assertion; verified fail-before by reversing _CFG_NAMES. |
||
|
|
48c9a3b1f8 |
feat(settings): compute-device override (auto / CUDA / ROCm / XPU / MPS / CPU) (#1557)
* feat(settings): compute-device override — auto | CUDA | ROCm | XPU | MPS | CPU Auto-detect stays the default; the override kills the 'auto-detect picked wrong' issue class. Applied at the single choke point (_probe()'s family selection) so routing, get_best_device(), and every badge inherit it. Resolution: OMNIVOICE_DEVICE env > Settings pick (prefs.json) > auto (#981 pattern). An override can steer, never invent hardware: a family the host lacks is noted and ignored; cpu is always honorable. Applies at next backend start (host caps are immutable per process — same restart contract as the rest of the Performance tab, RestartBadge shown). GET/PUT /api/settings/compute-device (admin-gated) reports resolved vs applied so the panel shows restart-required truthfully and disables itself under an env pin instead of pretending. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): entry for the compute-device override (#1557) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(device-override): harvest — the override reaches CT2 ASR, full i18n, honest edge states - _ctranslate2_cuda_ok() and the ASR sidecar now gate on the probe's family, so a cpu pin (or ROCm host) can never hand CTranslate2 a CUDA device — the override reaches every CT2 loader through one shared gate - override_ignored exposed by the API and shown by the panel (env pin naming a device this machine lacks: auto is in effect, restart won't change it) - all 8 panel strings + 5 device-family labels translated into all 21 locales; failed saves keep their error visible through the re-sync - test isolation: cleanup drops OMNIVOICE_DEVICE before re-probing so no overridden caps leak into later tests; panel tests wait for loaded state - xpu/intel search keywords; oxfmt formatting Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(device-override): round 2 — fail-safe probe fallbacks, complete i18n, combined pin state - a broken capability probe now means CPU everywhere (CT2 gate + ASR sidecar) — never a torch-derived guess that would bypass a cpu pin or re-open #1529 on ROCm; regression test added - env-pinned AND not-detected shows both facts in one subtitle - device_load_failed/perf_save_failed translated into all 21 locales; CJK/th/vi/ar strings no longer say literal 'Auto' - test_ctranslate2_never_gets_cuda_on_a_rocm_build pins the probe family (it was order-dependent on the lru_cache before) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): pin the probe family in the faster-whisper OOM-fallback test Same class as the rocm-build test: it mocked torch but not the probe the new override gate consults first, so on a cpu-family CI host the CUDA fallback chain under test was unreachable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
030d5ea01f |
docs(engines): a guide for every engine + index; fix two engine-metadata bugs (#1556)
* docs(engines): a guide for every engine + index; fix two engine-metadata bugs 21 new pages under docs/engines/ (10 TTS, 10 ASR, index README) — every registered engine now has one: what it's for, platform support, model env vars, quirks with issue refs. Linked from both READMEs' engine sections. Code fixes found while verifying facts against the registries: - KittenTTS docstring claimed default voice 'Jasper'; the code default is expr-voice-2-f - the isolated-ASR sidecar read only ASR_MODEL_FW while the download preflight read ASR_MODEL_FASTER — set one and the other quietly used a different model; both now resolve ASR_MODEL_FW-override → ASR_MODEL_FASTER - moonshine's install hint named 'useful-moonshine', a package the backend never imports; now moonshine-onnx / moonshine-voice Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): entries for the engine guides + sidecar model fix (#1556) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(engines): second-harvest fixes — TLS guidance, matrix/code alignment, CN counts - README matrix aligned to gpu_compat (the code is the source of truth): CosyVoice macOS is CPU not MPS, IndexTTS and GGUF gain their real CUDA/CPU/MPS cells - gpt-sovits guide: prefer https/tunnel for non-loopback servers, plaintext warning; first-use download guidance on both OmniVoice pages - preflight empty-env fallback matches the sidecar (ASR_MODEL_FASTER='' no longer resolves a different repo) - nano installs via uv pip; kitten log level wording; index links install guides incl. the Gatekeeper step; README_CN engine counts 16/11 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme-cn): the all-engines-local claim now excludes the remote client Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b79ba9bd3b |
docs(readme): lead with download + first clone; seed benchmarks page (#1555)
* docs(readme): lead with download + first clone; seed benchmarks page Quickstart (installers, install guides, a three-step first-clone walkthrough) moves above What's-new/Features in both READMEs — visitors get the action before the pitch. New docs/benchmarks.md anchors measured per-engine/device numbers on the bench_pipeline.py harness, community-contributed, no estimates. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): entry for the README conversion restructure (#1555) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): emit RTF + CUDA peak VRAM; guard NaN RAM; define the benchmarks schema Bot harvest on #1555: the tts stage now prints RTF per warm measurement and CUDA peak VRAM (None elsewhere — no made-up zeros), the stage floor refuses unmeasurable RAM instead of sailing past a NaN comparison (FLOOR_GB=0 overrides), docs/benchmarks.md columns map 1:1 to what the harness prints, and the download badges say they open the release page. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): link palash.dev from the maker section Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): name the resolved engine, track VRAM from resolution, comment the guards Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): the quick-switch gif is the hero image The hero shows motion now; the Launchpad screenshot moves into the 0.5.0 What's-new slot so nothing appears twice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): peak VRAM is reserved memory; adapter engines name their model Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): subprocess-isolated engines report VRAM n/a, not a parent-side zero Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): out-of-process detection is declarative; sherpa rows name their model 'runs_out_of_process' is now a TTSBackend attribute set by SubprocessBackend AND omnivoice-gguf (which inherits TTSBackend directly but spawns a binary per generate — the isinstance check missed it). Duck-typed for the same module-purge reason as _is_subprocess_isolated. Sherpa-onnx identity comes from _model_dir's basename when _model_id is absent. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): backends self-report model identity via TTSBackend.model_identity() Greptile enumerated the adapter engines one at a time (mlx _model_id, sherpa _model_dir, cosyvoice env-only) — the attribute sniffing rots per engine. The hook fixes the class: each multi-model backend reports its own identity, the profiler just asks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
41722afe3b |
refactor(launchpad): quieter, borderless design refresh (#1515)
* refactor(launchpad): quieter, borderless design refresh The launchpad carried decoration from an earlier direction: icon chips, corner-hung count badges, a permanently visible filled arrow, uppercase mono card titles, and a dotted stipple divider — plus a frame that had been invisible since the app-wide border tokens were zeroed. Rework it around what the borderless direction actually implies: - Feature tiles get a whisper-faint surface instead of a dead frame, and read as three bands (bare glyph + count / title + arrow / description). `--card-hue` is spent sparingly — the glyph at rest, the surface, count and arrow only once raised. Titles move to sans sentence case; counts are plain tabular numerals. Lift softened 4px -> 2px, coloured glow -> neutral shadow, plus an explicit focus ring and a staggered entrance. - Hero drops the boxed "646" pill and the filled A/B-Compare button for quiet type, with a hairline standing in for the separation. - Section labels trade the dotted stipple for a single fading hairline; rows are transparent until hover and reveal "Open" on hover/focus (it stays in the DOM, so AT and keyboard always reach it). - Hero, tiles, recent files, callout and project lists now share one 1180px column — previously only the top half was capped, so lists ran edge-to-edge on a wide display while the deck stayed centred. Two bugs found and fixed while doing it: - Buttons that had `border border-solid border-transparent` removed fell back to the UA default border and rendered a visible 1px outline. They now carry `border-0` explicitly. - `.lp-animate` used `animation-fill-mode: both`, so after the entrance it kept owning `transform` — and animation-origin declarations outrank normal ones, which silently killed the card hover lift. Now `backwards`, which still holds the from-state through the stagger delay. Also drops CSS the page has not rendered since #904: the cursor-spotlight layer, the breath ring, and the per-card waveform strip. Verified with headless renders at 1600/1280/940 and the empty state. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(dictation): decode Wayland portal signals and show the capture pill The GlobalShortcuts portal declares Activated/Deactivated as (o session, s shortcut_id, t timestamp, a{sv} options). We decoded the timestamp as u32, so zbus rejected every signal with Signature mismatch: got `(osta{sv})`, expected `(osua{sv})` and the press was dropped as an invalid signal. Registration succeeded and the desktop even reported the bound chord back, so the hotkey looked wired up while doing nothing at all — on every Wayland compositor, for the whole life of the feature (#1490). Decode the 64-bit timestamp, and keep the 32-bit spelling as a fallback so a non-conforming portal degrades to working rather than to silence. With presses arriving, the second half of the failure showed: nothing had shown the widget window since it became a hidden recorder host, so a capture ran with no pill on screen — and a mic or Accessibility failure rendered into a window nobody could see. Add show_dictation_pill, which bottom-centres the capsule on the monitor under the pointer and shows it without taking focus (Windows keeps SW_SHOWNOACTIVATE so paste still lands in the user's document), and call it from the widget for every state but idle. Wayland denies clients their own placement, so the compositor picks the spot there; the pill still appears. dispatch_dictation_capture now logs whether a press was emitted or queued — a press that reaches Rust and produces nothing was otherwise indistinguishable from one the compositor never delivered. Tests: portal signals decode at both timestamp widths (the 64-bit case fails before this change with the exact production error); pill placement centres, respects a second monitor's origin, and clamps rather than going off-screen; the widget shows for a state needing the user, stays hidden while idle, and never shows for a press that arrives while dictation is disabled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: sync in-progress workspace changes Uncommitted work already in the tree, checkpointed so the branch matches the local machine: - Remote GPU workers: join-from-the-app flow, one-time secrets, QR join codes, a Compute control in the status bar, and the device-list Workers panel (#1516) - Model Catalogue workspace, with Settings pointing at it - Settings sidebar search and keyboard navigation - Demo assets for dubbing, dictation and voice design, plus the scripts that render them - Backend: validation-error handling, ASR request-path degradation, and the accompanying tests - CHANGELOG entries for the above and for the Wayland dictation fix Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tests): follow Engines to the Model Catalogue, and green the sweep - test_supertonic3 asserted the license gate points at "Settings" while the engine now names Model Catalogue → Engines, which is where the accept button actually lives. The assertion follows the move; what it pins is unchanged — the hint must name a place the user can reach it. - Carries the CJK allowlist entries for the rendered dub bundle (#1517) and the regenerated route snapshot for /workers/agent (#1516), both of which this branch inherits from the workspace sync. - docs/install/linux.md: the dictation capsule is bottom-anchored everywhere except Wayland, where the protocol gives applications no say in their placement. Documented rather than left as a surprise (CodeRabbit). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: stop a flaky dependency fetch from failing green runs en-core-web-sm resolves to a direct GitHub release URL, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own three retries all land within the same few seconds and fail together, so the whole job dies on a dependency that has nothing to do with the change under test — it cost #1518 and #1517 an otherwise-green run tonight. Two changes: back off between whole `uv sync` attempts, which is what actually clears it, and pass --no-sync to the pytest steps. `uv run` re-resolves the environment before running, so every test step was a fresh chance to hit the same fetch even though the install step had already synced — that is exactly how #1518 failed, in the isolated backend/tests step, with all 5467 tests already passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: one retry seam for every uv sync, not just the job that failed last en-core-web-sm resolves to a direct GitHub *release* URL rather than a package index, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own retries all land inside the same ~10 seconds and fail together, so a job dies on a dependency unrelated to the change under test. Tonight that cost four otherwise-green runs across #1515, #1517 and #1518 — and the first fix only covered the Tests job, so the next failure simply moved to Smoke (Linux), which syncs separately. The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh backs off between whole attempts (15s, 45s, 90s) and every workflow that syncs now goes through it — ci.yml (tests + the platform matrix), release.yml, security.yml, evals.yml. It still fails loudly after four attempts, so a genuinely broken lockfile is not disguised as a flake. The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke matrix has always set, which is part of why it was the one that kept dying; it has them now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ci): pin the Intel-Mac contract by intent, not by command spelling test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted the literal line `run: uv sync --extra pockettts`, so routing every sync through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The contract it exists to protect is that the pockettts extra installs ONLY on backend_supported legs — which the regex now pins, while leaving how the sync is invoked free to change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: keep every uv run out of the resolver, and bound the retry budget CodeRabbit, #1517: - `uv run` re-resolves before running, so the smoke suite, the worker-artifact tests, the release test run and the eval run were each a fresh chance to hit the flaky direct-URL fetch outside the retry loop. All of them pass --no-sync now; the environment is already synced by the step that owns the retries. security.yml's `uv run --with pip-audit` is deliberately left alone — it layers an ephemeral package rather than running the project's own tests. - The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a 120 s timeout on the smoke matrix). Three attempts and 60 s of total backoff outlast the refusals actually observed while staying well inside the jobs' timeout-minutes. - The Intel-Mac contract test pinned the smoke command literally too, so --no-sync tripped it exactly like the sync line did. Same fix: assert the contract (smoke runs only on backend_supported legs), not its spelling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
04410a458d |
Release VoiceStudio 5.0.0 (#1487)
Complete the VoiceStudio identity, release documentation, assets, version mirrors, and safe cross-platform development startup. |
||
|
|
95a35b8e07 |
feat(indextts): add native IndexTTS 2.5 support (#1485)
* feat(indextts): add native 2.5 sidecar support * fix(indextts): preserve legacy language metadata * docs(indextts): state model license terms accurately * fix: preserve IndexTTS upgrades and duration controls * fix: complete IndexTTS upgrade safeguards |
||
|
|
9f2c8ac0a4 | fix(pockettts): close license review findings | ||
|
|
77d6507318 | fix(pockettts): recheck consent after synthesis queue | ||
|
|
252c6fd149 | Merge remote-tracking branch 'origin/main' into codex/pr1442 | ||
|
|
c2da73c5c2 | fix(pockettts): enforce consent on every synthesis | ||
|
|
180d13c59e | fix(pockettts): enforce consent at construction | ||
|
|
d8268ca99c | fix(pockettts): gate unsupported Intel Mac wheels | ||
|
|
f6f2bc5dcd | fix(security): pin curated Hugging Face revisions | ||
|
|
9a770b15d1 | test(pockettts): verify pinned install on four platforms | ||
|
|
034dd2333d | fix(pockettts): complete first-use terms gate and smoke | ||
|
|
c870f794ed |
fix(engines): give every sidecar a private fd for its frames (#1428)
Protect all nine sidecar frame channels from library stdout noise and pin the complete sidecar manifest in regression coverage.\n\nCloses #1428. Thanks @1335-Group for the diagnosis and tested fix. |
||
|
|
bcb547b9f2 |
fix(engines): a slow venv probe is not a broken venv (#1414) (#1421)
Every subprocess engine confirms a candidate interpreter by spawning it and importing the engine package. For IndexTTS that is 'import indextts.infer_v2', which pulls in torch and transformers — seconds with a warm page cache, tens of seconds on a first run, a spinning disk, a network share, or Windows with real-time AV scanning every DLL. The bound was 10s (15s for three peers), and elapsing it was treated as a negative: the candidate was discarded exactly as if the import had raised. A working OMNIVOICE_INDEXTTS_DIR install was reported as 'IndexTTS-2 is not installed', or fell through into the lazy bootstrap and reinstalled over a working clone. Only successful resolution was memoised, so every retry re-ran the probe and failed identically — which is why all three reported repro paths look like one bug. A timeout is the absence of evidence, not evidence of breakage. The probe is now tri-state: yes (imported), no (ran and failed), unproven (did not finish). An unproven candidate is kept as a fallback and used only after every candidate has had its chance, so a wedged user clone cannot shadow a healthy bootstrapped venv. If an unproven venv really is broken it now fails at the sidecar handshake with a real error rather than a confident lie about the install. Fixed as a class: backend/engines/_venv_probe.py replaces the drifted copy in each of the four bootstraps, and the bound is tunable per engine, defaulting to 60s. Zero and negative values are ignored — an unbounded probe would let one wedged candidate hang engine resolution forever. Reported with a precise root cause by @OracleNightmare. (#1414) |
||
|
|
93025d9a81 |
fix(dictation): stop the widget stranding an empty square, repair the swept data dirs (#1398)
The dictation hotkey could leave a blank dark square stuck on the desktop with no way to dismiss it. Three defects compounded: the tray listener's effect depended on [state], so it detached across an await on every state change and a press landing in that gap was lost; an idle pill renders null, so the window Rust had already shown was empty; and the opaque chrome background made that empty window a hard-edged square. Nothing could hide it — dismiss() is only reachable from the X button, Esc, or a post-session timer, none of which exist for a session that never started. Fixed at the invariant rather than the call sites: the listener subscribes once for the component's lifetime, the widget window's chrome background is transparent, and an idle-but-visible window reconciles itself to hidden. The reconcile is polled (a dropped press changes no React state, so there is nothing to key an effect off) and aborts if its effect is torn down mid-check, so it can never hide a dictation that has just started. Also in scope: - The rename sweep had repointed three data-dir literals at a brand-named directory that does not exist, so smoke-test.sh verified a directory the backend never writes and desktop-prod.sh silently stopped clearing backend state on Windows. Both invisible on macOS, where they are usually run. A guard test now pins the assignments specifically. - The dictation model picker's download sizes were wrong for all seven models, in both directions — Parakeet TDT v3 (the recommended default) understated 180 MB against an actual 670 MB, while the low-RAM fallbacks were overstated threefold, discouraging exactly the choice that would have helped. Measured from the published repos and pinned by a test. - The 0.6B Parakeet models now decode on more threads, capped by host cores and still overridable. - uninstall.ps1 gained a UTF-8 BOM (Windows PowerShell 5.1 mis-decodes its non-ASCII output without one), and sponsor.yml lost its last OmniVoice references. |
||
|
|
43d3537adf |
fix(asr): stop a missing cuDNN 8 from killing the backend process (#1371)
CTranslate2 (the engine under WhisperX and faster-whisper) requires cuDNN 8 while torch ships cuDNN 9. When the side-loaded cuDNN 8 is absent it does not raise — it __fastfail()s, killing the whole backend with 0xC0000409 and no traceback, so the shell restarts it and the next attempt dies the same way. The defect was that the backend computed the answer and discarded it: the preload passed silently on a missing directory and on every OSError, then called into a library that treats the same condition as fatal. The answer is now kept, and the CTranslate2 engines report themselves unavailable so auto-detect falls through to pytorch-whisper. Two more instances of the same class went with it: the crash-isolated ASR sidecar is a child process that never preloaded at all (failing every transcribe, quietly), and the preload only searched <project root>/.venv, missing any other interpreter. Conservative by design — a false positive costs WhisperX's forced alignment, so ROCm is excluded via torch.version.hip before cuda.is_available(), which is True on HIP builds. 14 regression tests; backend suite 4374 passed. |
||
|
|
5cab8e0149 |
feat: rename the product to VoiceStudio (previously OmniVoice-Studio)
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
|
||
|
|
eeffe6c2d1 |
fix(gguf): a source-built runtime must actually run (#1348) (#1384)
A meticulous report from an LXC/CPU-only source install surfaced three real defects: the build script deleted the libggml shared libraries a dynamically-linked build needs (first spawn died with exit 127), the hardcoded 120s per-spawn kill switch reaped legitimate CPU-only renders, and OMNIVOICE_ALLOWED_ORIGINS — the only fix for cross-origin browser access — was documented nowhere. All platform branches of scripts/build-omnivoice-tts.sh now copy the shared libs next to the binary, the CI artifact glob uploads them, and the backend puts bin/ on the loader path for every spawn of the engine binary. The timeout defaults to 600s (above the pool guard's well-diagnosed 300s deadline), is tunable via OMNIVOICE_GGUF_GENERATE_TIMEOUT_S with non-finite values rejected, and the timeout error names the knob. CORS documented in api-auth.md with a pointer from remote-gpu.md. Regression tests pin the spawn-env rule, the per-branch copy rule, the artifact glob, and the timeout behavior. |
||
|
|
acd36badee |
feat(engines): PocketTTS CPU-only sidecar shape (#1306) (#1328)
* feat(engines): PocketTTS CPU-only sidecar shape (#1306) Sidecar SHAPE for review, mirroring omnivoice-subprocess: PocketTTSBackend(SubprocessBackend) (CPU-only, parent interpreter, optional-dep gate) plus a stdio sidecar (ready/ping/synthesize/shutdown, lazy TTSModel.load_model, per-ref voice cache, generate_audio to int16 PCM). Registered in services/tts_backend.py. Batch protocol; streaming raised as a follow-up. CI smoke, gated-weights preflight, 4-platform install, licence-accept gate deferred to on-top after shape review. * feat(engines): PocketTTS sidecar handles 6 languages (en/fr/de/pt/it/es) load_model(language=...) per language (cached), maps OmniVoice's language value to a pocket-tts model language, and picks the default preset voice per language when no ref clip is given. Represents PocketTTS accurately: it is multilingual, not english-only. The HF model card's 'English only' line is stale, confirmed by the GitHub README and pocket-tts 2.1.0. * fix(engines): list pockettts in docs inventory; drop unused logger docs/features.yaml tts_engines now includes pockettts, clearing the docs-drift test that failed CI (every registered engine must be in the inventory). Removed the unused logger line CodeQL flagged. No readme/doc entry yet, matching opt-in engines like supertonic3 and omnivoice-gguf; a doc page can land with the rest of the integration. * fix(engines): address PocketTTS sidecar review findings - Cold-load watchdog: heartbeat progress frames during the gated weights download so the parent does not kill a healthy sidecar mid-load, plus a 600s recv timeout on the backend. - Unsupported language: raise a clear error instead of silently falling back to English and mispronouncing. - Voice-state cache: LRU-bounded to 8 entries so a long session cannot leak memory. - ref_audio SSRF: reject URLs (local file paths only) to preserve local-first. Addresses the 3 Greptile P1 + 1 CodeRabbit Major on #1328. * fix(engines): invalidate voice cache on ref-file change; reject non-finite recv timeout - Voice-state cache key now folds the ref_audio file mtime+size, so a file replaced at the same path no longer returns a stale voice from the previous contents (Greptile P1). - recv_timeout_s rejects inf/nan env values via math.isfinite and falls back to 600s, so the deadline can't be silently disabled (CodeRabbit Major). * fix(engines): nanosecond mtime in voice cache fingerprint int(st.st_mtime) lost sub-second precision, so a file replaced at the same path within one second with the same size kept the old key and returned a stale voice. Use st.st_mtime_ns for full resolution (Greptile P1 on the follow-up fix commit). * fix(engines): raise on multi-channel audio instead of unsafe downmix The defensive mean(axis=0) assumed channels-first; on channels-last (N,2) it averaged across time, producing garbage. The engine returns mono, so the branch is unreachable in practice. Raise on ndim>1 so an upstream shape change surfaces as a loud error frame instead of silent noise. (debpalash review on #1328) * fix(engines): include import error in pockettts is_available message CodeRabbit Minor on #1328: the exception was caught as 'e' but never shown. * fix(engines): lock _send to prevent concurrent-write framing corruption Greptile P1 on #1328: the cold-load heartbeat thread and the main loop both call _send (stdout write). The stop+join serializes the normal case, but a join timeout leaves a window where both threads write length+body segments concurrently, interleaving the wire framing. Add a threading.Lock around the write so concurrent _send calls are serialized regardless. * test(engines): cover the PocketTTS sidecar's silent failure modes The four review findings fixed on this PR are all silent by construction: an unsupported language rendered fluent, confident, wrong audio; the channels-last downmix produced noise; interleaved frames desynchronized the pipe permanently; a re-recorded clip kept serving the old voice. None of them raise, and none would be caught by an end-to-end smoke test that only asserts audio came back. 49 tests over the sidecar's pure logic — language selection, PCM conversion, wire framing, the LRU voice cache — plus the backend surface (recv-timeout guards, CPU-only declaration, sample-rate lockstep with the sidecar, lazy registration). The model is mocked and the sidecar is stdlib-only at import time, so none of it needs the optional pocket-tts wheel or a child process. Verified fail-before/pass-after by reverting the lock and the multi-channel guard: the framing test fails with a length header decoded from inside another frame's body, which is the corruption itself rather than a proxy for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: debpalash <nizam4103@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3183bf5fcd |
fix(engines): downmix along the channel axis, not axis 0 (#1328) (#1366)
* fix(engines): downmix along the channel axis, not axis 0 (#1328) Found while reviewing #1328. Every subprocess sidecar guards its PCM conversion with a defensive `arr.mean(axis=0)`, which is correct only for channels-first audio. For a channels-last (N, 2) array `squeeze()` keeps both axes and the mean runs across TIME: every output sample becomes the mean of two neighbouring samples and the render collapses to 2 samples. That is not a downmix, it is a destroyed waveform played back as noise. Unreachable in all five today because every engine returns mono -- which is precisely why it could sit there being wrong. Nothing runs it, so nothing reports it, and the first engine or SDK version to emit stereo gets noise with no error anywhere. Pick the channel axis instead of assuming it, and loop so a stray extra axis reduces the whole way to mono; previously a (2, N, 2) array stayed 2-D after one mean and produced a PCM buffer whose length disagreed with the n_samples in the frame -- a desynchronized audio frame rather than a merely wrong-sounding one. Downmixing correctly rather than raising (the choice PocketTTS made on #1328): these five are shipping engines, and turning a render that works today into an error is a regression risk that the actual defect -- the wrong axis -- does not require taking. The sidecars run under different interpreters (confucius4 and dots.tts each have their own venv), so they cannot import a shared helper and the duplication cannot be refactored away. The recurrence guard is therefore a test that holds all five to the same behaviour at once, so a sixth copy pasted into a new sidecar fails there rather than shipping: 20 of its 30 cases fail before this change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(engines): let a broken sidecar fail instead of skipping CodeRabbit Major on #1366: the blanket `except Exception -> pytest.skip` turned a syntax error or an import-time regression in any of the five sidecars into a skip, so this regression suite could pass CI while running nothing. All five are stdlib-only at import (torch and the model load lazily on the first synthesize), so there is no optional dependency to tolerate -- an import failure here is a real defect in a shipping engine. Unguarded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: debpalash <nizam4103@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
47c51698c9 |
feat(engines): add omnivoice-subprocess, a crash-isolated (killable) TTS engine (#1292)
* feat(engines): add omnivoice-subprocess, a crash-isolated TTS engine The default in-process OmniVoice engine runs on the GPU ThreadPoolExecutor. When a generate or load exceeds its execution budget the pool is "reset", but the abandoned worker thread cannot be killed (Python cannot interrupt a native torch/MPS call), so it keeps holding the device until it finishes on its own and later synths queue behind it and hang. The reset restores pool capacity but not the device. This is the residual root cause behind the closed #730 and #1190: the messaging/reset mitigations address the symptom, not the device-holding zombie. Add an opt-in `omnivoice-subprocess` engine that runs the same model in a child process via SubprocessBackend. A child process can be hard-killed: on a recv-timeout the watchdog calls proc.kill(), reclaiming VRAM/device, and the next request transparently respawns a fresh sidecar. The in-process engine remains the default, so existing users see no change; this is an opt-in for unattended / scheduled / reaction-triggered synthesis where a stuck job must self-recover instead of hanging until a manual restart. Base-class and mitigation changes that ship with it: - SubprocessBackend.generate() now consumes non-terminal {"op":"progress"} frames a sidecar emits during a cold load (previously the first cold generate after spawn failed, then worked on retry). Additive: engines that reply with audio directly are unaffected. - recv_timeout_s is overridable per engine (default 60s unchanged); the new engine sets it to the generate budget so a long-but-valid synth is not falsely killed while a wedged one still is. - make_room_before_generate(): free idle GPU memory before a warm, heavy generate. The cold-load path already evicted; the warm path skipped it, so a long synth on a VRAM-tight MPS box could contend its way into the budget. Verified end-to-end against the live model (cold / warm / recovery-after-kill) and under a sustained + concurrent-pressure soak: killed-worker recovery 5/5, chunked long text 9/9, no memory leak. * Address review: install_hint + move make_room into get_model - Add `omnivoice-subprocess` to `_INSTALL_HINTS`; the test_install_hints_cover_all_registered_backends gate requires every registered backend to carry one (this was the CI failure). - Move the warm-generate VRAM eviction out of the /generate and /v1/audio/speech routes and into get_model()'s warm-return path, so EVERY native TTS generate is covered (REST, WS TTS, dub, batch, audiobook), not just the two REST routes. Drops the now-redundant per-route wiring. (Greptile P1: the per-route placement missed the other generation surfaces.) * Address review: drop dead long-text eviction path; log probe failure - _should_make_room_for_generate: the long-text headroom boost became dead code once the eviction moved into get_model() (which has no text), so the long-text branch never fired. Removed the text param, the long-text threshold/multiplier branch, and the now-unused _env_float helper. The core RAM-tight gate (the part that matters on a starved box) is unchanged. - Log the available_memory probe failure at debug instead of silently swallowing it (CodeRabbit: silent swallow breaks the debug trail). - Tests updated for the text-agnostic policy. * fix(engines): stop subprocess generate() self-deadlock on 1-worker pools SubprocessBackend.generate() acquires a GPU-pool slot for accounting, but /v1/audio/speech and /generate dispatch backend.generate() via run_on_gpu_pool_guarded, i.e. already ON a pool worker. On a 1-worker pool (MPS) the inner pool.submit queued behind the very job running it and slot_future.result(timeout=10) raised before the sidecar ever spawned, so omnivoice-subprocess (and every other subprocess engine on MPS) surfaced the in-process 300s-abandon instead of synthesizing. Skip the slot acquisition when current_thread() is already a gpu-pool worker; the outer guard already accounts for the slot. Direct callers (off the pool) still acquire one. Regression test added (generate on a pool worker). * Address review: reword slot-skip comment (fixes watermark-coverage CI) + simplify - The slot-skip comment said "dispatch backend.generate() via", and test_watermark_route_coverage's _SYNTH_CALL regex matches the literal backend.generate( anywhere in a module, so it counted subprocess_backend.py as a synthesis producer that must reference mark_synthetic (it doesn't — the routes apply mark_synthetic; the engine sits below the chokepoint, like tts_backend.py). Reworded to "dispatch generate() via". - Fold in the simplify refinement: single negated predicate, import+pool moved into the acquire branch. |
||
|
|
7f4d0ad9df | chore(audiobook): correct issue refs to #1208 + changelog entries | ||
|
|
a12a7e7ee1 |
feat(audiobook): expressive maturity — overrides, emotion, cache opt-out, discoverability (#1210)
Audiobook renders were locked to the model's most deterministic preset (32
steps / 2.0 guidance / model-default temps) with no way to change it, which is
why books sounded flatter than the same voice on the Voice page. Open that up
without changing any default byte-for-byte.
- Production Overrides in the Audiobook tab: position_temperature,
class_temperature, num_step, guidance_scale, postprocess_output (+ seed),
reusing the Voice page's panel. Unset reproduces today exactly.
- IndexTTS2 graded emotion (emo_vector / emo_text / emo_alpha) reaches the
longform path via a typed engine-options object; engines that don't
understand an option ignore it (no crash across the ~14 backends).
- Cache opt-out ("vary repeated lines") so identical lines can get distinct
takes; default off keeps the content-addressed replay.
- Markup reference now lists the reaction tags that already work in audiobooks;
docs/expressive-speech.md corrected so no recipe it names is unreachable.
- Fix AudiobookGenerateBody dropping `language`, so audiobook language
selection actually reaches the backend.
Every new param is folded into BOTH cache layers (chapter + segment) and the
per-chapter preview, so changing a knob re-renders instead of replaying stale
audio — and an all-default request keeps its old cache key, so existing books
don't re-render. Regression tests: tests/test_audiobook_expressive.py (backward
compat, cache-signature loop, preview/render parity, engine-ignores-unknown,
cache opt-out, emotion reaches engine) + audiobookOverrides.test.jsx.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
933743e336 |
fix: resolve open issue batch — #1172 #1173 #1174 #1185 #1186 #1188
- validate managed binaries before exec; 0-byte GGUF placeholders fail actionably instead of "Exec format error" (#1172) - KittenTTS: tokenizer-measured chunking to the ONNX 512-token cap; clear 400 for unspeakable input (#1173) - clean SIGTERM during weight load: shutdown-aware loader, benign cancelled-load classification, lifespan hardening, scoped log silencers (transformers load + alembic fileConfig) (#1174) - broken ASR deep-imports (lightning_fabric) mark the engine unavailable with a repair hint and fall through (#1185) - uv cache + managed Python follow the chosen install drive on Windows (cherry-picked cross-drive class fix + spaces/D: tests) (#1186) - adaptive silence-removal ladder for quiet clone references; localized actionable error for truly silent clips, all 21 locales (#1188) - CHANGELOG: consolidated Unreleased into the quiet one-liner style Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bd85bab624 |
chore: retire finished planning archives from the repo root (#1095)
Removes ~110 files of process noise (all preserved in git history): .planning/ (GSD-era phases/quick-plans/issue-clusters; workflow retired 2026-07-08), specs/ (spec-kit specs for shipped features 001-007), design/ (pre-React ASCII mockups), research/ (legacy Gradio archive), and .agents/ (rules for a third-party agent tool no longer in use). The four load-bearing decision docs move to docs/adr/ with an archival note; every live pointer follows (gguf engine module docs + quant_map, inject-apprun.sh, pyproject/test comments, fixture README + its seed script — kept byte-identical). The CJK allowlist drops the deleted legacy_gradio entries; STRUCTURE.md and ROADMAP.md document the removal instead of linking into it. Backend suite: 2891 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ff56865cf7 |
feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines (#1083)
* feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines
IndexTTS-2 required four manual terminal steps (git clone, uv venv,
uv pip install -e ., export OMNIVOICE_INDEXTTS_DIR). This turns that into
a guided in-app install:
- backend/services/sidecar_install.py — parametrized sidecar provisioner
(SidecarSpec/SPECS so future sidecar engines are one entry, not another
installer). Resumable background job with step-by-step status: disk-space
preflight (needs-X/have-Y message), source fetch (git clone --depth 1
primary, GitHub tarball fallback when git is absent/fails), dedicated
venv via uv (OMNIVOICE_BUNDLED_UV → PATH resolution; transformers<5
isolation preserved — the parent env is never touched), import-probe
verification, IndexTeam/IndexTTS-2 weights into <checkout>/checkpoints
(where the sidecar actually loads from) via snapshot_download with the
auto-selected/configured HF endpoint + token — no hardcoded
huggingface.co — and persistence of OMNIVOICE_INDEXTTS_DIR (os.environ
for immediate use, prefs.json env.* for the next launch). Idempotent:
partial installs repair, downloads resume, healthy installs (incl. a
user's own clone) report already_installed and are never touched.
- API: POST /engines/{id}/install starts the job, GET
/engines/{id}/install/status polls it, DELETE /engines/{id}/install
removes an app-managed install (loopback-gated; refuses user-managed
clones). list_backends() gains one_click_install.
- Frontend: Settings → Engines shows an Install button on the IndexTTS2
row with per-step progress, live log tail, weight-download %, and
error+remediation; the manual setup snippet is demoted to a collapsed
"Manual install" fallback. All strings via i18n (en.json).
- OMNIVOICE_INDEXTTS_DIR joins the Settings env-var allowlist
(single-sourced from the installer SPECS).
- Docs: docs/engines/indextts.md leads with the one-click flow; manual
steps become the fallback section. CHANGELOG Unreleased entry added.
- Tests: tests/test_sidecar_install.py (24 cases — happy path, disk-space
fail, git-absent/git-failing tarball fallback, partial-install repair,
already-installed/running gating, uninstall safety, spec↔bootstrap
contract, router wiring) + 6 new EngineCompatibilityMatrix RTL cases.
API route snapshot regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): harden the sidecar installer — review findings
- Route namespace: /engines/sidecar/{id}/install — a dynamic
/engines/{id}/install would shadow the literal
POST /engines/sonitranslate/install (engines router registers first);
regression-guarded by test_sidecar_routes_never_shadow_literal_engine_routes.
- Weights completion marker: a killed-mid-download multi-shard weights dir
(config.yaml + plausible shards) no longer passes for healthy; the marker
is written only after snapshot_download returns, so re-runs resume.
- _run_logged: drain thread + proc.wait(timeout) + POSIX process-group kill
— a grandchild holding the stdout pipe can no longer hang the step past
its timeout.
- Job log lock: the status poll's list(deque) copy no longer races the
worker's appends (RuntimeError under active logging).
- Self-heal: a healthy managed install whose env var was lost (prefs wiped)
is re-pointed by start_install instead of reported already_installed
while the engine stays unavailable; legacy bootstrap installs (Probe-2
venv) are trusted via the engine's own probe.
- Single-sourced uv/venv-layout resolution: engines.indextts.bootstrap now
delegates _locate_uv/_venv_python_path to services.sidecar_install.
- Frontend: stable poll interval (keyed on the running-id set, not the
status map), reload on a job that finishes before the first poll,
re-attach to an in-flight job on remount, i18n'd Install aria-label,
manual-install <details> auto-opens on failure, snippet block hoisted
out of the JSX IIFE.
- list_backends: sidecar-installable set hoisted out of the per-engine
loop; exhaustive-shape registry test updated for one_click_install.
- Tests rebind the live services.sidecar_install module per test (other
suites purge sys.modules["services"], which made router tests
order-dependent).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): fill in the PR ref (#1083)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): validated tarball fallback + scanner-clean installer
- The pre-filter= extractall fallback (Python < 3.11.4) now extracts
member-by-member behind the same guards extractall(filter="data")
enforces — regular files/dirs only, no absolute paths, no ../ escapes,
resolved-path containment. Kills the new CodeQL py/tarslip (high) and
Bandit B202 (error) alerts; regression-tested with a malicious tarball
(test_safe_extract_members_blocks_tar_slip).
- snapshot_download tracks the weights repo's default branch on purpose
(same policy as every other model download; artifacts are
checksum-verified by hf_hub) — documented + B615 waived at the call.
- Explanatory comments on the intentional empty-except blocks
(CodeQL py/empty-except notes).
Verified locally: bandit -ll -ii on the module reports 0 MEDIUM+ findings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): address Greptile review — Windows tree kill, prefs write race, poll robustness
- _kill_tree: Windows now uses taskkill /F /T so a git/uv helper spawned by
the timed-out child can't keep writing into the checkout (POSIX already
killed the process group). Unit-tested with os.name patched to nt.
- core/prefs: mutations (set_/delete) are serialized behind a module lock —
the installer worker persisting its env.* key concurrently with a Settings
write could previously drop whichever key saved first (whole-class fix:
every threaded prefs writer, not just the installer). Fail-before/
pass-after: tests/test_prefs_thread_safety.py.
- Matrix polling: at most one in-flight status request per engine (an old
'running' response can no longer land after a newer 'succeeded' and
restart the poller), and four consecutive poll failures drop the stale
snapshot instead of showing "Installing…" and hammering a dead backend
forever.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
8281b7c798 |
fix(engines): dub and batch TTS honor the active-engine selection (#987)
* fix(engines): dub and batch TTS honor the active-engine selection — with a real capability gate, not a silent OmniVoice fallback
Dub generation and batch TTS hardcoded services.model_manager.get_model()
(OmniVoice) regardless of the engine picked in Settings → Engines. A user
selecting VoxCPM2 (or any other engine) still got OmniVoice output with no
error — the silent fallback IS the bug class, not just the one report.
Root-caused and fixed for the whole class:
- New `TTSBackend.supports_cloning` capability flag (default True) marks
engines that can only offer fixed preset voices — kittentts,
supertonic3, sherpa-onnx set it False. MLXAudioBackend exposes it as an
instance property (Kokoro doesn't clone, CSM does) since the adapter
multiplexes multiple models with different capabilities.
- `cloning_capable_engine_ids()` and a shared `resolve_generation_backend()`
helper in services/tts_backend.py centralize engine resolution
(id → is_available() → routing gate → optional cloning gate), mirroring
generation.py's /generate resolution instead of inventing a third
parallel mechanism. Both routers now standardize on the existing
get_active_tts_backend() cache (unload-on-switch already handled).
- dub_generate.py's two TTS-generate call sites (main run + OOM retry) and
the /dub/preview-segment route resolve once, up front, with
require_cloning=True — dub's ref_audio is populated for essentially
every real job, so an engine that can't clone fails the whole job with
one actionable message instead of mis-cloning per segment.
- batch.py resolves once per job, require_cloning only when voice_id is
pinned — an unpinned batch job runs fine on any engine.
- Applied the three pre-existing TODO(#312) comments: mastering now skips
via `applies_own_mastering` for both pipelines, matching generation.py.
Regression tests cover the capability-id list, the fail-fast gate (proving
no OmniVoice fallback), the success path on a selected non-OmniVoice
engine, batch's pinned-vs-unpinned voice_id behavior, and the mastering
skip for both pipelines. Three existing dub tests that mocked get_model()
directly were updated to mock the new resolver instead.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(engines): exclude model-dependent adapters from cloning_capable_engine_ids()
getattr(cls, "supports_cloning", True) at the CLASS level returns a
property descriptor object (always truthy) when the flag is an instance
@property, not a plain attribute — MLXAudioBackend uses exactly this
pattern because its cloning capability depends on which of its 7+ curated
models is loaded (only CSM clones; Kokoro etc. don't). Without this fix,
the dub/batch capability-gate error message would always recommend
'switch to mlx-audio' even when the user's configured MLX model can't
clone, sending them in a circle back to the same error.
isinstance(value, bool) distinguishes a resolved boolean from a
descriptor object, so mlx-audio is excluded from the suggestion list
until its actual per-instance capability can be checked (already handled
correctly by resolve_generation_backend()'s per-call instance check).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): engine-aware dub/batch entry (#987)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
4eed552153 |
fix(engines): Confucius4-TTS validated E2E — clone sys.path import, 22.05 kHz, real install docs (#590) (#872)
Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3d0705fdb7 |
feat(engines): Confucius4-TTS — finalized (API-validated + unit-tested; opt-in, GPU run pending) (#590) (#637)
* feat(engines): Confucius4-TTS scaffold (opt-in, needs hardware validation) (#590) Plumbing for netease-youdao's Confucius4-TTS — LLM-based 14-language cross-lingual zero-shot voice cloning, Apache-2.0 — mirroring the opt-in subprocess-venv pattern of dots.tts / MOSS-TTS-v1.5: - engines/confucius4/__init__.py: Confucius4Backend(SubprocessBackend), CUDA-only (gpu_compat=("cuda",)), language passthrough, ref_audio→prompt_wav. is_available reports a clear reason and stays unavailable without a clone. - bootstrap.py: dedicated Python 3.10 venv resolution (user clone-level venv → package venv → uv bootstrap), import-probed on `confuciustts`. - main.py: sidecar speaking the same length-prefixed JSON-over-stdio protocol as the other engines, calling ConfuciusTTS(config_path, device).generate(text, lang, prompt_wav). - Registered lazily in _LAZY_REGISTRY; docs/engines/confucius4-tts.md. Gated behind OMNIVOICE_CONFUCIUS4_TTS_DIR — inert on every default install, never imports the upstream package unless opted in. The sidecar's synthesis API is derived from the upstream README and is NOT yet validated on a CUDA box; the module, docs, and CHANGELOG all flag this. 4 tests pin registration + inert-by-default. No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#590): register Confucius4 in install-hints + docs inventory (CI gates) Registering the engine tripped two completeness gates: every backend needs an install_hint (test_issue_fixes) and every registry engine must appear in the tts_engines docs inventory + README (check-docs-drift). Add the install_hint, the docs/features.yaml entry, and the README engine-table row (with the scaffold caveat). Docs-drift clean; gates pass. No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(confucius4): finalize — validate API vs upstream, add 22 sidecar unit tests, document external deps (Amphion/w2v-bert/weights) The synthesis API (ConfuciusTTS(config_path, device) → generate(text, lang, prompt_wav) → tensor, model.sample_rate) is confirmed against the netease-youdao/Confucius4-TTS repo. Added runnable unit tests for the sidecar's pure logic (language norm, tensor→PCM mono/stereo/clip, config resolution, wire framing, synthesize dispatch with the model mocked) — 22 cases, all green. Docs now list the external deps (Amphion/MaskGCT codec, facebook/w2v-bert-2.0, ~2-4GB HF checkpoint) and CUDA 12.6. Softened the scaffold warnings to reflect API-validated + unit-tested status; a one-time CUDA GPU run is still needed to confirm live inference + true sample rate. --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a7ab148483 |
fix(asr): float16-unsupported GPUs fall back to int8 instead of "no segments" (#561)
#551: both CTranslate2 ASR backends request compute_type="float16" on CUDA with NO fallback. On GPUs without efficient fp16 (older Maxwell/Pascal, GTX 16xx) or a CTranslate2/cuDNN binary mismatch, WhisperModel/whisperx.load_model raise a ValueError at construction — which escaped the existing OOM-only `except RuntimeError`, so every chunk failed and the user got "Transcription produced no segments". Add a per-device compute_type fallback chain (cuda: float16 → int8_float16 → int8; cpu: int8 → float32) to both backends + the ASR sidecar, alongside (not replacing) the existing OOM→CPU path, with an ASR_COMPUTE_TYPE override for exotic hardware (documented in README). Also in the same ASR-robustness pass: - #549: PyTorchWhisperBackend._ensure_pipe wraps the transformers pipeline load and re-raises an actionable error (reinstall transformers / use faster-whisper) instead of a bare "Could not import module 'AutoFeatureExtractor'". - #516: the /dub/transcribe SSE generator is wrapped so it can NEVER close without a terminal event — any unanticipated exception now yields a structured `error` (with build_failure's hint) + `done`, turning "stream dropped, likely ASR failed" into the real cause + Retry. - failure.py: COMPUTE_TYPE_UNSUPPORTED + TRANSFORMERS_IMPORT classes so the no-segments toast is actionable. Tests (fail-before/pass-after): float16-unsupported → int8 for both WhisperX + FasterWhisper; a generic non-OOM RuntimeError still raises; classify() maps the two new classes; the SSE stream always terminates with error→done. 7 + 1 passed, 17 in the failure suite (no regression). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3777d3a62c |
feat(tts): add MOSS-TTS-v1.5 (8B) and dots.tts (2B) as opt-in engines (#498)
Adds two zero-shot voice-cloning TTS engines requested in #498, both opt-in and subprocess-isolated with their own dedicated venv — the same pattern as IndexTTS-2. The dedicated venv is forced, not just chosen: each upstream pins a transformers version that conflicts with the parent's >=5.3 (MOSS-TTS-v1.5 ==5.0.0, dots.tts ==4.57.0), so they cannot share the parent interpreter. Because they use the clone+venv bootstrap (env var -> clone -> uv venv), this touches no pyproject.toml / uv.lock / bun.lock — `uv sync --all-extras` and Docker's `bun install --frozen-lockfile` are unchanged, so main's CI/Docker matrix stays green. Engines: - moss-tts-v15: 8B, 31 langs, ~16 GB weights, 24 kHz. AutoModel/ AutoProcessor via trust_remote_code. gpu_compat=(cuda,cpu) — MPS is undocumented/untested upstream so it is never claimed; on a Mac it runs on CPU. Apache-2.0, no license gate. - dots-tts: 2B, 24 langs, ~9 GB weights, 48 kHz. DotsTtsRuntime; continuation cloning (prompt_audio_path+prompt_text). Upstream is Linux/macOS-only, so is_available() gates it off cleanly on Windows (cross-platform parity rule — it is opt-in, never a broken default). Wiring: registered in _LAZY_REGISTRY + _INSTALL_HINTS. list_backends() surfaces both as subprocess/[cuda,cpu]/available-until-installed; the data-driven Settings engine picker needs no frontend change. Tests (19, fail-before/pass-after): registry resolution, subprocess marker, no-MPS gpu_compat, the Windows gate, not-installed honesty, and the parent-side generate() kwarg arbitration. Existing engine suite still 55 passed / 5 skipped. Sidecar inference follows the upstream-documented APIs but, like IndexTTS/Supertonic, can't be executed in CI without the multi-GB model clones. Docs (same-PR per docs-sync rule): README + README_CN engine tables, new docs/engines/moss-tts-v15.md + dots-tts.md, disk-usage.md (torch-dedup note), CHANGELOG. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2b8c8aec7c |
fix: actionable errors for non-executable engine binary (#437) + unreachable backend (#438/#454/#466) (#471)
Two reliability bugs from open issues, both first-run papercuts where the error told the user the wrong thing. #437 — `[Errno 13] Permission denied: bin/omnivoice-tts-linux-x86_64`: a git clone / zip extract on POSIX can drop the bundled binary's execute bit. It only surfaced at spawn time, and the generic synth handler then mislabeled it as "ran out of memory" and told the user to flush the model. - omnivoice_gguf.is_available() now self-heals: after the SHA check confirms the binary is the right file, it adds +x (best-effort) on POSIX; if it can't, it returns a clear "isn't executable — run chmod +x <path>" message instead of a spawn-time crash. No-op on Windows. - generation.py classifies PermissionError / EACCES / "Permission denied" as its own case ("a bundled binary lost its execute bit — reinstall or chmod +x"), so it never again masquerades as OOM. #438/#454/#466 — bare "Failed to fetch" / "NetworkError": when the local backend is still starting, crashed, or the dev server dropped, fetch() throws a TypeError that propagated raw to the user. - client.ts apiFetch now catches the thrown fetch and raises an ApiError with an actionable message ("Can't reach the local OmniVoice backend — it may still be starting up… restart the app or check Settings → Logs"), status:0 to mark a transport failure vs an HTTP error. Tests: client.test.ts +1 (thrown fetch → ApiError status 0 + actionable text); 3 pass. CJK guard green. |
||
|
|
c3b2346759 |
fix(engines): MLX platform gate (#390) + ASR gpu_compat + IndexTTS2 (#21 PR 2/5) (#431)
Builds on the device probe from PR 1. Backend-only; the routing keys are wired into /engines in PR 3. - #390 closed: MLXAudioBackend / MLXWhisperBackend now call the shared `core.device_caps.mlx_supported()` gate FIRST, before importing the package. On Linux/Windows/mac-Intel they report unavailable and never advertise a usable `mps` route, even with a stray mlx wheel installed. Replaces the ASR backend's ad-hoc inline MPS check with the one shared rule. (The Wave-4.4 OSError/RuntimeError import-guard is preserved — it now lives behind the platform gate; its test forces the gate open so the guard stays the path under test.) - `ASRBackend` ABC gains `gpu_compat: tuple[str, ...] = ("cpu",)` mirroring TTSBackend, and each subclass declares its real targets: whisperx/faster-whisper → (cuda,cpu); mlx-whisper → (mps,cpu); pytorch-whisper → (cuda,mps,cpu); nemo/funasr → (cuda,cpu); moonshine → (cpu,). Inert until PR 3 serializes them. - IndexTTS2 declares `gpu_compat = ("cuda","cpu")` so it stops advertising the inherited CPU-only default. - ROCm is deliberately NOT claimed for any ASR engine (or for IndexTTS2): CTranslate2 has no upstream HIP build, and an unverified `rocm` claim would route ROCm hosts to a broken GPU path — strictly worse than the honest `cpu_fallback` the resolver already emits ("declares CUDA only; ROCm not in its compat set"). The per-engine TTS ROCm audit is a tracked follow-up that will verify each path before claiming it. Tests: MLX gate regression (both backends, on/off Apple), ASR gpu_compat tuples + no-false-rocm invariant, IndexTTS2 override; existing MLX import-guard test updated for the new gate ordering. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dc1d36fe5f |
refactor(models): model-management v2 cleanup (mm2, all tiers) (#428)
One coherent lifecycle surface over the in-process model, diarization, and subprocess sidecars; fixes the engine-switch VRAM leak; tightens download robustness. Backend-only, response shapes preserved, no new deps. Tier 1 — correctness: - MM2-01: get_active_tts_backend() caches one instance per backend id and unload()s the outgoing engine on switch (fixes the VRAM leak behind #278); adds reset_active_backend(). - MM2-02: OmniVoiceBackend.unload() releases the shared model_manager singleton + free_vram(); SubprocessBackend.unload() -> unload_sidecar(self.id), inherited by all sidecar engines. Idempotent + preload-safe. - MM2-03: /model/loaded ASR row reports the real device + a note explaining the disabled unload button. Tier 2 — single surface: - MM2-04: new services/model_lifecycle.py owns list_loaded/unload/unload_all/ free_vram; system.py routers are thin delegations (shapes unchanged). - MM2-05: idle timeouts (in-process + sidecar) resolve via prefs.resolve (env wins, no restart); removed the duplicated _IDLE_TIMEOUT_SECONDS. Tier 3 — robustness/observability: - MM2-06: _install_cooldowns swept (1h TTL) + cleared on success — bounded. - MM2-07: per-extension weight floors (onnx 64KB, tensors 5MB) OR the original >=5MB catch — small ONNX no longer false-flagged, #352 still caught. - MM2-08: indextts GPU sidecar self-reports vram_mb in pong; parent surfaces it in list_live_sidecars (0 = CPU/unmeasured). - MM2-09: is_cached scan_cache_dir->disk fallback logs WARNING w/ exc type (#117/#118), was invisible at DEBUG. Tests: tests/test_mm2_lifecycle.py (15). Full suite: 1379 passed. Plan/summary: .planning/quick/260613-mm2-clean-model-management-v2/. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e862f0faf0 |
feat(asr): crash-isolated faster-whisper subprocess backend (Wave 4.2) (#393)
* feat(asr): crash-isolated faster-whisper subprocess backend (Wave 4.2) Native ASR engines (faster-whisper / CTranslate2) can segfault on GPU teardown — a process-level crash that kills the whole backend. Running the engine in a child process turns that into a failed job: the sidecar dies, the parent raises a decorated error (engine id + device), and the next request respawns a fresh sidecar. - services/subprocess_asr.py: SubprocessASRBackend reuses SubprocessBackend's wire protocol + lifecycle — including respawn-on-dead-process (_spawn relaunches when the child isn't alive) and GPU-slot acquire/release — adding a 'transcribe' op (the TTS 'generate' surface is stubbed). IsolatedFasterWhisperBackend wraps faster-whisper using the PARENT venv (already a dep — only the process boundary is new); opt-in via OMNIVOICE_ASR_BACKEND=faster-whisper-isolated. - engines/_asr_sidecar/main.py: the faster-whisper runner (stdlib wire protocol; torch/CT2 import lazily so the ready handshake fits the timeout). - engines/_echo/main.py: a 'transcribe' echo op so the round-trip + crash recovery are testable without a real engine. - asr_backend._REGISTRY is now a lazy dict (mirrors the TTS registry) so the isolated backend lists/resolves without importing the subprocess stack unless selected. Tests (echo sidecar, stdlib-only): round-trip, single long-lived sidecar across calls, crash-mid-transcribe → decorated error + backend healthy + next call respawns, registry exposure, generate-not-supported. Spec 7 / parity program Wave 4.2. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(asr): deterministic crash test + drift marker for lazy ASR registry (Wave 4.2 CI) CI surfaced two issues: - The echo crash test relied on the crash-AFTER-reply hook, whose reply may still reach the parent (timing-dependent) — and a leaked OMNIVOICE_ECHO_CRASH from a sibling subprocess test poisoned the non-crash tests. Fix: a deterministic OMNIVOICE_ECHO_CRASH_NO_REPLY hook that exits BEFORE replying (guaranteed dead pipe → decorated error), and the asr fixture clears both crash envs so the round-trip/two-call tests can't inherit a leak. - check-docs-drift's _ASR_MARKER didn't match the new lazy registry line (_LazyASRRegistry({); updated the marker + the self-test fixture. Verified the no-reply crash hook by driving the sidecar directly (reply=None, exit 1); drift self-test + real-repo check green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(asr): allowlist the 'segments' op so transcribe replies aren't dropped (Wave 4.2 CI) The parent's PARENT_INBOUND_OPS frozenset gated inbound sidecar frames but never included 'segments' — the ASR transcribe reply op. _recv() dropped the frame as disallowed, tail-recursed, hit EOF, and returned None, so every transcribe surfaced as a bogus 'sidecar crashed mid-transcription'. TTS ('audio') was allowlisted; ASR ('segments') was missed. Add it (and list 'transcribe' in the informational SIDECAR_INBOUND_OPS), update the exact-shape allowlist test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5ba8a5a8a0 |
fix(gguf): forward speech generation controls (#306)
Co-authored-by: openclawer <bdfzer8@gmail.com> |
||
|
|
adf486ee03 |
chore: set version to 0.3.0 across all sources (+ drop v0.4 references) (#145)
* chore: drop stray v0.4 references — everything ships on the v0.3.0 line Per the project's versioning rule (no v0.4, no unprompted version chatter): - backend/main.py + marketplace.py: the app reported version "0.4.0" (ahead of even pyproject's 0.2.7 and referencing a forbidden version). Aligned to "0.2.7" to match pyproject.toml / tauri.conf.json — a consistency fix, not a bump. - errorDocsMap.ts / indextts/bootstrap.py / _secret_key.py: reworded "v0.4" deferral comments to version-agnostic "deferred / later hardening pass". - docs/install/troubleshooting.md: the "tracked for v0.4" notarization line now matches macos.md (signing is wired; activates on the Apple cert secrets). Note: historical planning records under .planning/ still contain "defer to v0.4" notes; left as-is (a record of superseded decisions) — CLAUDE.md + the constitution are the live source of truth. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: set version to 0.3.0 across all sources (current dev line) The current/upcoming version is v0.3.0 (0.2.7 is the prior stable). Bump every version source so the codebase consistently reports 0.3.0 — the in-code dev version; the git *tag* still happens later per the release cadence. - pyproject.toml, frontend/src-tauri/Cargo.toml, tauri.conf.json, frontend/package.json: 0.2.7 → 0.3.0 - backend/main.py (FastAPI) + marketplace.py export metadata → 0.3.0 (these had drifted to a phantom "0.4.0") - CHANGELOG.md: "[0.2.7] — Unreleased" → "[0.3.0] — Unreleased" - uv.lock + Cargo.lock reconciled (1-line each) so `--frozen` installs hold. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(version): read app version from package metadata (no more drift) Greptile (#145): the FastAPI version + marketplace bundle metadata were bare string literals — they'd go stale-wrong again at the next bump (the exact class of bug this PR fixes; that's how "0.4.0" happened). Read once from importlib.metadata.version("omnivoice") via core.version.APP_VERSION, with a "0.3.0" fallback only for a non-installed source checkout. pyproject.toml is now the single source of truth for the runtime version. Tests: tests/test_app_version.py (semver + equals installed metadata). 2 pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b34dcd9e11 |
Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + integration (#100)
* Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + Wave 1 integration Integrates Serveurperso/OmniVoice-GGUF as a hardware-adaptive default voice-cloning engine, with overridable fallback to the in-process OmniVoiceBackend. Spike confirmed GO: the model is a clean quantization of k2-fsa/OmniVoice (Apache-2.0 + MIT runtime, `omnivoice-lm` custom architecture so it does NOT load in vanilla llama.cpp). Pinned SHAs: * Serveurperso/OmniVoice-GGUF revision: 361609388ae572a820d085185bbbe2a2aac4b30e * ServeurpersoCom/omnivoice.cpp master: 886fc079838ca7400cb2b42b36e2a65aa1daabe8 Implements GGUF-01 (hardware probe) through GGUF-05 (default-engine resolver with graceful fallback). The four `bin/omnivoice-tts-*` artifacts are committed as zero-byte placeholders; the new CI matrix job builds the real binaries per platform from the pinned commit SHA and appends a SHA-256 manifest used by `is_available()` for tampering detection (T-04-01). The macos-14 (Apple Silicon) slot is marked `continue-on-error: true` because omnivoice.cpp publishes no `buildmetal.sh` (Pitfall 1 / Assumption A1) — failure feeds into Task 3's GO/NO-GO call. Quant override is allow-listed against quant_map.json entries only (T-04-05). Argv is composed from typed Path objects rooted in HF_HUB_CACHE; never uses `shell=True`. HF token redaction applies to captured stderr before logging (AUTH-05 / T-04-04). Tests: 36 new (8 hardware-probe + 13 GGUF engine + 6 settings_store quant override + grep gate); 428 passed in full suite vs 402+ baseline. ADR Status stays "Proposed (research-supported)" — Task 3 (human checkpoint) flips to Accepted after CI produces real binaries and a reviewer signs off on the GGUF-06 cross-hardware smoke. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci: install libopenblas-dev on linux-x86_64 omnivoice-tts build The pinned omnivoice.cpp commit (886fc079...) ships a `buildcpu.sh` that passes `-DGGML_BLAS=ON`. ubuntu-latest has no BLAS implementation preinstalled, so the cmake configure step fails with `Could NOT find BLAS (missing: BLAS_LIBRARIES)` and the job exits in 13 s before producing the linux-x86_64 binary. macOS (Accelerate, built in) and Windows (BLAS off by default in the ggml CMakeLists for non-APPLE platforms — the build script doesn't invoke buildcpu.sh on those slots) are unaffected and stay green. Adds a Linux-gated apt step to install libopenblas-dev + pkg-config before the build, restoring cross-platform parity per the CLAUDE.md "default features must work on every platform" rule. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(gguf): constrain ref_audio to project roots — block /etc/shadow on Linux The GGUF engine's `_build_argv` previously validated ref_audio only via `ref_path.is_file()` — i.e. "does this path exist?" That check is platform-dependent: `/etc/shadow` doesn't exist on macOS (rejected naturally), but it IS a real system file on Linux, so the validation silently accepted it. CI's ubuntu-22.04 runner exposed the gap via `test_generate_blocks_freeform_ref_audio`, which exists precisely to guard the "freeform ref_audio path" attack surface. Fix: confine ref_audio to one of three allowed roots before existence checks: - VOICES_DIR (user-saved voice profiles) - DUB_DIR (per-job auto-clones extracted from source video) - tempfile.gettempdir() (browser-upload temp files; existing `cleanup_ref` flow in generation.py) Anything outside those roots → FileNotFoundError, matching the existing failure-mode contract callers handle. Existence check still runs after, so the test's mocked subprocess.run is never reached and the test passes deterministically on all three platforms. Cross-platform parity (per CLAUDE.md 2026-05-20 rule): identical behaviour on macOS / Windows / Linux — the allow-list is computed from core.config which uses platform-specific path resolution but yields the same logical "project tree" on every OS. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci(gguf): mark darwin-x86_64 binary build as experimental GitHub's macos-13 (Intel) runner pool is heavily contended — PR #100 queued for 30+ minutes waiting on darwin-x86_64 while every other platform finished in ~1m. Intel Macs are also fading hardware (Apple's platform momentum is entirely on Apple Silicon), and the GGUF engine's runtime already handles a missing binary gracefully (`is_available()` returns False on Intel Mac with a "binary not bundled for this platform" message, same path used for first-launch before any binaries build). `experimental: true` mirrors what darwin-arm64 (Metal) already has — slot still runs and uploads its binary when successful, but a failure or runner backlog no longer blocks merges. Keeps the GGUF engine shippable across the dominant arm64 / Linux / Windows surface without holding the inbox on a slow-runner queue. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
93aa66ab0a |
Phase 3 Plan 03-01: Supertonic-3 engine on SubprocessBackend (#101)
* Phase 3 Plan 03-01: Supertonic-3 engine on SubprocessBackend
Adds Supertonic-3 as a 7th opt-in TTS engine on the Phase 2
SubprocessBackend primitive. Closes TTS-01..06 (REQUIREMENTS.md):
* TTS-01 — _REGISTRY["supertonic3"] resolves to Supertonic3Backend,
a SubprocessBackend subclass.
* TTS-02 — `supertonic==1.3.1` lives under [project.optional-dependencies];
default `uv sync --no-dev` does NOT install it. Exactly one
`onnxruntime` row in `uv pip list` after `--extra supertonic`.
* TTS-03 — Model revision pinned by 40-char commit SHA
(724fb5abbf5502583fb520898d45929e62f02c0b — the "Initial
Supertonic 3 release" SHA, same as the SDK's own pin).
Resolver script for intentional bumps:
scripts/resolve_supertonic3_sha.py.
* TTS-04 — Honest CPU-only reporting. `is_available()` message says
"ready (CPU-only via onnxruntime)" and never mentions
"cuda" or "mps". `gpu_compat = ("cpu",)`.
* TTS-05 — License gate via settings_store helpers
(get/set_license_accepted) + Loopback-only
/api/settings/license endpoint + SupertonicLicenseDialog
frontend modal showing MIT (code) and OpenRAIL-M (model).
Wired into EngineCompatibilityMatrix as an "Accept license"
button on rows whose `reason` mentions "license not
accepted".
* TTS-06 — 3 langs (en/ja/ru) × 3 sec smoke test in
tests/test_supertonic3.py::test_smoke_3langs_3sec
(OMNIVOICE_SMOKE-gated; asserts no onnxruntime-gpu row
post-synthesize).
Package legitimacy gate (Task 1 in plan): supertonic on PyPI verified
to be published by Supertone Inc. (ato@supertone.ai), repo
github.com/supertone-inc/supertonic, wheel is pure-Python with no
postinstall scripts. Same publisher ships supertonic-js on npm under
the same maintainer email.
Test results:
* tests/test_supertonic3.py — 10 passed, 3 skipped (network-gated).
* tests/smoke/ — 4 passed.
* tests/ (full, --ignore=tests/manual) — 412 passed, 0 failed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(tests): uv sync --all-extras so optional-engine tests can import their package
Phase 3 added `supertonic` as an optional dependency. The CI Tests job
runs `uv sync` (no extras), so `test_cpu_only_honest` and `test_license_gate`
in tests/test_supertonic3.py hit the "supertonic package not installed"
fallback instead of the real import path, and fail.
Bare `uv sync` is the right default for users (engines are opt-in), but
the test environment should exercise the full surface. `--all-extras`
keeps the smoke job lean (still bare `uv sync`) while letting Tests
verify the integrated behavior of every optional engine.
Future-proofs against the same failure mode in Phase 4 (GGUF) and any
later optional engines.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
c3695e1668 |
Phase 2 Plan 02-03: IndexTTS on SubprocessBackend (closes #42) (#98)
Migrates IndexTTS-2 off the in-process import path and onto the SubprocessBackend primitive shipped in Plan 02-01. Closes issue #42 with a structural fix — the parent's transformers>=5.3 and IndexTTS's transformers<5 now live in separate OS processes and can never collide. * New: backend/engines/indextts/ — sidecar package (__init__.py hosts IndexTTS2Backend, main.py is the sidecar entrypoint, bootstrap.py owns the 3-step venv probe + lazy uv-based bootstrap). * services.tts_backend: IndexTTS2Backend's in-process body removed; registry resolves the class lazily via a _LazyRegistry indirection + PEP 562 __getattr__ re-export. This breaks the import cycle that arose when both subprocess_backend and tts_backend tried to import each other at module load. * docs/engines/indextts.md: install walkthrough + venv resolution order + common errors (linked from is_available()'s unavailable message). * tests: - test_indextts_backward_compat.py (8) — probe priority, no-spawn discipline, HF cache marker preservation (ENGINE-07). - test_indextts_sidecar.py (17) — subclass shape, isolation_mode, parent-side emotion arbitration (vector/audio/text/description), coexist-with-OmniVoice (headline #42 closure), env forwarding. - tests/fixtures/mock_indextts_sidecar.py — stdlib-only sidecar mimicking the production wire protocol; emits 1 s sine wave. - test_issue_fixes.py: two obsolete in-process-conflict tests rewritten to assert the new subprocess contract (no indextts.* import in the parent). Hard constraints honored: backend/services/sonitranslate.py and gpu_sandbox.py are untouched (D1 / D4). Existing v0.2.7 users with OMNIVOICE_INDEXTTS_DIR and a populated HF cache reach a working generation with zero re-download and zero re-install. 44 tests pass across the four exercised files. Full suite: 391 passed, 10 skipped, 13 xfailed, 1 xpassed in 57 s. Smoke: 4 passed. Closes #42. Requirements: ENGINE-02, ENGINE-03, ENGINE-04, ENGINE-07. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0fc5ea6cf3 |
Phase 2 Plan 02-01: SubprocessBackend primitive (Wave 1 of Phase 2) (#97)
* Phase 2 Plan 02-01: SubprocessBackend primitive + echo sidecar + ENGINE-05 wrap
Lands the durable SubprocessBackend primitive — the architectural keystone
that Plans 02-03 (IndexTTS migration), Phase 3 (Supertonic-3), and
Phase 4 (GGUF / Singing) plug into.
Files added:
- backend/services/subprocess_backend.py — base class owning spawn,
shutdown, _send/_recv (length-prefixed JSON), GPU-slot acquire-release,
atexit teardown, stderr drain, op allowlist (T-02-04), and 64 MB
frame cap (T-02-01). No multiprocessing — subprocess.Popen
exclusively so subclasses can target a *different* venv's interpreter
(Locked Decision D4 / Pitfall 1).
- backend/engines/_echo/main.py — permanent CI regression sidecar.
Stdlib-only, runs under the parent's sys.executable. Implements
ready/ping-pong/synthesize/shutdown plus test-only probe_env and
emit_unknown ops for env-forwarding and op-allowlist tests. DO NOT
DELETE — the round-trip test depends on this file.
- tests/backend/services/test_subprocess_backend.py — 13 tests:
round-trip, health_check, no-zombie, shutdown idempotency, env
forwarding (HF_TOKEN/HF_HOME/HF_ENDPOINT/HF_HUB_CACHE), oversize
frame, short read, op-allowlist drop, op-allowlist constant shape,
sidecar-crash recovery, no-multiprocessing grep gate, MAX_FRAME_BYTES.
- tests/backend/services/test_tts_backend_registry.py — 6 tests for
list_backends() resilience + shape + isolation_mode + last_error
caching + existing-engines preservation + install_hint passthrough.
Files modified:
- backend/services/tts_backend.py:
* Adds module-level _LAST_ERRORS dict for ENGINE-06.
* Rewrites list_backends() to wrap each is_available() in try/except
so one broken engine cannot blank the picker (ENGINE-05).
* Adds last_error + isolation_mode keys to each response entry
(ENGINE-06 UI in Plan 02-04 consumes via the same /engines route).
* Uses a duck-typed _is_subprocess_isolated marker rather than
issubclass(cls, SubprocessBackend) because test fixtures (token
resolver suite) purge sys.modules["services"] between tests and the
re-imported SubprocessBackend would be a different class object.
Threat-model mitigations (Plan 02-01 frontmatter):
T-02-01 DoS via length-prefix → MAX_FRAME_BYTES = 64 * 1024 * 1024
T-02-02 GPU slot leak on sidecar death → try/finally in generate
T-02-03 token bytes in stderr → drained via parent logger
(HFTokenRedactor from Phase 1 already on root)
T-02-04 unknown ops from compromised sidecar → PARENT_INBOUND_OPS
allowlist, unknown frames logged and dropped
T-02-05 Tauri group-kill scope → start_new_session=True on Unix /
CREATE_NEW_PROCESS_GROUP on Windows
Verification:
- 337 passed, 6 skipped, 12 xfailed, 1 xpassed (full suite,
`uv run pytest tests/ --ignore=tests/manual`)
- All 19 new tests pass on macOS Apple Silicon
- Smoke tests still pass: `uv run pytest tests/smoke/ -q` → 4 passed
- SoniTranslate untouched (D1 locked decision)
- Zero new Python dependencies
Closes part of ENGINE-01 + ENGINE-05.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(02-01): plan summary — public API, invariants, deviations
Documents the SubprocessBackend public API so Plan 02-03 (IndexTTS) and
Phase 3 (Supertonic-3) authors don't need to re-read the source.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|