* docs: add private production API deployment
* docs: harden private API deployment guidance
* docs: make proxy trust configuration executable
* docs: clarify private API trust boundaries
* docs: persist private service application data
Run uvicorn directly under the dev wrapper so a worker crash cannot hide behind a live reload parent. Preserve Python source reloads, restart isolated crashes with bounded diagnostics, and keep persistent crash loops loud.
Closes#1690.
Prepare the tested main branch for the v0.5.1 patch release with synchronized version sources, mirrors, lockfiles, release notes, and install guidance.
Closes#1687.
Fail before backend/window startup when Linux source hosts lack Enigo’s libxdo linker input or WebKit’s GStreamer audio sink. Print an exact distro package command, sync source-build docs, and lock the probes with deterministic tests.
Closes#1680Closes#1682
Fix fresh-clone desktop development by creating the required dist placeholder before Tauri starts, and make source setup install the selected CUDA or ROCm PyTorch stack consistently. Adds behavior-level cross-platform regression coverage.\n\nFixes #1664.\nFixes #1665.\n\nThanks @uberclokr for the contribution.
Move unbounded Stories and Audiobook project data to revisioned IndexedDB storage with bounded local fallback, durable clear tombstones, migration/recovery safeguards, and deadline-safe persistence before exits and relaunches.
Closes#1636.
* feat(install): one-command installer URL (.sh + .ps1) + 3-OS install smoke
- scripts/install.ps1: Windows source installer (winget deps, uv, bun,
clone, uv sync, frontend build); honors OMNIVOICE_PYTHON/OMNIVOICE_REGION
- infra/install-redirect: Cloudflare Worker serving /install with
User-Agent sniffing (curl -> sh, PowerShell -> ps1, browser -> landing
page); proxies live from main; /install.sh + /install.ps1 aliases
- scripts/install.sh: fix stale advertised URL (main/install.sh never
existed) and repo-root resolution so a local run no longer clones a
duplicate repo into ~/VoiceStudio (verified on macOS arm64)
- .github/workflows/install-smoke.yml: run both installers end-to-end on
ubuntu/macos/windows when they change
- docs-sync: install one-liners lead each platform guide; STRUCTURE.md
(#1626)
* fix(install): don't let a failed bun download pass silently
curl | sh runs an empty script and exits 0 when the download fails, so
a bun.sh hiccup surfaced much later as 'bun: command not found' (seen
on the macos-latest smoke runner). Fetch to a temp file, verify, fall
back to npm -g bun when node exists, and die with the manual command.
Same post-install verification for uv.
* fix(install): UTF-8 BOM for install.ps1 + quiet-style changelog entry
- tests/scripts/test_uninstall_ping.py requires shipped PowerShell
scripts with non-ASCII text to carry a UTF-8 BOM (Windows PowerShell
5.1 mis-decodes otherwise); same treatment uninstall.ps1 already gets
- test_changelog_style caps entries at ~400 chars
* fix(indextts): accept the config name upstream ships, and keep long text alive
Two independent defects, both reported on a working IndexTTS 2.5 install.
Install always failed. IndexTeam/IndexTTS-2.5 ships the model config as
config.yaml — at the pinned revision d0aa86e7 and at HEAD; config_v2_5.yaml
exists in no upstream revision. VoiceStudio demanded that name, so
_weights_floor_ok never found it and the install died claiming 'the download
was likely interrupted' when the download had been perfect. The only way
through was to hand-rename the file. Both names are accepted now, in the
installer and on the load path, so installs created with the workaround keep
working without a reinstall.
Long text was killed at 60s. infer() is one blocking upstream call that puts
nothing on the wire, and IndexTTS was the only sidecar still on the 60s
recv_timeout_s class default while pockettts and omnivoice-subprocess had both
raised theirs. Raising the default alone does not fix it — which is why the
reporter's RECV_TIMEOUT_S=3600 edit didn't help: progress frames are also what
report activity to the GPU pool's execution clock (#1367), so a silent sidecar
still trips the outer generate budget. The sidecar now heartbeats every 5s
while infer() runs (and during the cold model construction), _send takes a
lock so the beat thread can't interleave framing, and the deadline rises to
900s via OMNIVOICE_INDEXTTS_RECV_TIMEOUT_S.
test_indextts25_health_requires_25_config_name asserted the bug — that a
checkout holding only config.yaml is unhealthy — so it is rewritten to the
corrected contract, including that a genuinely truncated download is still
caught.
Fixes#1611
* test(indextts): follow the installed config name in the sidecar loader tests
Two more tests encoded the config_v2_5.yaml assumption, both asserting
cfg_path against a directory where no config existed at all — so they were
pinning the literal name rather than the resolution. They now lay down a real
checkpoints/ tree and assert the resolved path, including that a checkout
carrying the pre-fix hand-renamed config still resolves.
Caught by the full suite; the targeted runs during development did not reach
tests/backend/services/.
* test(indextts): event-driven heartbeat tests, real interleaving proof, precedence pin
Review round on #1619 — all four findings taken.
- The docs line naming 0.5.1 is version-neutral now ('Earlier installs') —
version labels are the owner's call.
- The heartbeat tests waited on wall-clock sleeps; they now block on a
per-write Event with a bounded deadline, so scheduler load can't flake them.
- The _send test asserted the lock EXISTS — a tautology. It now drives four
concurrent writers through a stream that yields between every byte and
asserts every frame decodes; verified fail-before by removing the lock
(torn frame) and pass-after.
- The precedence test deleted config.yaml before creating the renamed one, so
reversed precedence still passed. Both files now coexist for the assertion;
verified fail-before by reversing _CFG_NAMES.
* fix(setup): tolerate reserved memory in the RAM preflight, add OMNIVOICE_RAM_PREFLIGHT=0 escape hatch (#1618)
An "8 GB" machine reports ~7.8 GB usable (firmware/iGPU/kernel
reservations), so comparing OS-reported RAM against the marketing-size
8 GB threshold hard-blocked exactly the boundary hardware the minimum is
meant to admit — with no way past the wizard. Both thresholds are now
compared with a 7% reserved-memory allowance, and
OMNIVOICE_RAM_PREFLIGHT=0 downgrades a genuine fail to a warning for
users who accept the OOM risk (same opt-out shape as
OMNIVOICE_ASR_VRAM_PREFLIGHT).
Regression tests: backend/tests/test_ram_preflight_1618.py.
Docs: troubleshooting §1c.
* review: hermetic preflight stubs in tests; correct the escape-hatch doc
Greptile P1: the Settings panel can't set OMNIVOICE_RAM_PREFLIGHT (and the
blocker appears before setup completes anyway) — the doc now points at
PowerShell / shell env only.
CodeRabbit: stub _network_check and media_tools.summary so each RAM
assertion stays fast and offline (26s -> 6s locally).
* fix(dictation): refresh accessibility blocker
* docs(changelog): note accessibility refresh
* test(dictation): assert the native widget hide on accessibility grant
The recheck regression asserted only that the Accessibility pill text left
the DOM, so it still passed with hideWidgetWindow() removed and the native
capsule stranded on screen. Hold one stable getCurrentWindow().hide spy and
assert it after the poll (fails before the fix, passes after).
Refresh current v0.5.0/0.5 tag examples, document API-key and share-PIN behavior, and require encrypted private-overlay access for remote deployments. Keeps the Docker install guide and changelog synchronized.
Fix server-mode admin authentication recovery without trapping PIN-only deployments, and prevent stale 403 responses from clearing or superseding newly issued sessions. Includes backend/frontend regression coverage, docs, and changelog credit for @paoloantinori.
* feat(omnivoice): port upstream VoiceClonePrompt persistence + FlashInfer opt-in
Upstream k2-fsa teardown ports, verified with generated voice samples:
- VoiceClonePrompt.save()/.load() (upstream format v1, weights_only-safe)
on the vendored model, and a disk layer under the in-memory prompt LRU
(DATA_DIR/prompt_cache, keyed by ref path+mtime+ref_text+preprocess,
32 newest kept, OMNIVOICE_PROMPT_DISK_CACHE=0 opts out). First generation
of a session with a known voice skips the reference re-encode and any
auto-transcription pass — verified across two real processes (encodes=1
then encodes=0, same voice).
- omnivoice_flashinfer.py ported (packed CFG attention, fused kernels,
optional CUDA graphs), schedule adapted to our num_step+1 divergence.
Opt-in via OMNIVOICE_FLASHINFER=1|graph, CUDA-only, replaces
torch.compile for the session; missing package / apply failure / runtime
failure all degrade with a named reason (same #278 contract as compile:
classify → unapply → retry once, session latch). Measured 2.20x at
batch=1 on an RTX 4090 with byte-identical text and clean ASR round-trip.
- Docs: OmniVoice guide gains instruct+reference combination semantics
(consistent instruct stabilizes cloning, reference wins conflicts),
inline pronunciation control (pinyin / CMU), prompt persistence, and
corrects the 'no voice design' claim; performance.md documents both new
env knobs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: point changelog entries at the real PR number (#1565)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(pr): harden FlashInfer lifecycle + prompt-cache writes per review
Bot harvest round 1 (#1565): unapply on apply-failure (half-patched model
could crash the next render); pin eager-mode FlashInfer inference to one
thread too — the attention plan and packed position ids are per-generation
module state, so interleaved _gpu_pool workers would corrupt each other;
restore the CAPTURED pre-apply attention impl (could be flash_attention_2)
instead of assuming sdpa; unique tmp name per prompt-cache write; correct
the _forward_logits layout docstring; resolve VoiceClonePrompt at test
runtime; docs — Known limits keeps only the limitation, performance.md
states the VRAM cost and scopes the fallback claim to classified kernel
failures.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(pr): round-2 review — publish only a fully restored model, redact latch reason, tighten CPU-persistence test
Greptile: the runtime fallback now unapplies BEFORE swapping generate, so
a concurrent render keeps queuing behind the thread-affinity wrapper while
teardown mutates modules. CodeRabbit: FlashInfer failure reasons pass
through core.failure.sanitize before latching/logging (wheel paths embed
the user's home); the save-portability test now creates the tokens on CUDA
when available and asserts the persisted payload itself is CPU-resident.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(pr): fail-closed latch reason when the sanitizer itself breaks
CodeQL empty-except + CodeRabbit round 3: if core.failure.sanitize raises,
the raw reason (home paths, wheel paths) was latched anyway. Now only the
exception class survives with a fixed redaction note; two regression tests
(normal redaction + sanitizer failure).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(settings): compute-device override — auto | CUDA | ROCm | XPU | MPS | CPU
Auto-detect stays the default; the override kills the 'auto-detect picked
wrong' issue class. Applied at the single choke point (_probe()'s family
selection) so routing, get_best_device(), and every badge inherit it.
Resolution: OMNIVOICE_DEVICE env > Settings pick (prefs.json) > auto (#981
pattern). An override can steer, never invent hardware: a family the host
lacks is noted and ignored; cpu is always honorable. Applies at next
backend start (host caps are immutable per process — same restart contract
as the rest of the Performance tab, RestartBadge shown).
GET/PUT /api/settings/compute-device (admin-gated) reports resolved vs
applied so the panel shows restart-required truthfully and disables itself
under an env pin instead of pretending.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): entry for the compute-device override (#1557)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(device-override): harvest — the override reaches CT2 ASR, full i18n, honest edge states
- _ctranslate2_cuda_ok() and the ASR sidecar now gate on the probe's family,
so a cpu pin (or ROCm host) can never hand CTranslate2 a CUDA device —
the override reaches every CT2 loader through one shared gate
- override_ignored exposed by the API and shown by the panel (env pin naming
a device this machine lacks: auto is in effect, restart won't change it)
- all 8 panel strings + 5 device-family labels translated into all 21
locales; failed saves keep their error visible through the re-sync
- test isolation: cleanup drops OMNIVOICE_DEVICE before re-probing so no
overridden caps leak into later tests; panel tests wait for loaded state
- xpu/intel search keywords; oxfmt formatting
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(device-override): round 2 — fail-safe probe fallbacks, complete i18n, combined pin state
- a broken capability probe now means CPU everywhere (CT2 gate + ASR
sidecar) — never a torch-derived guess that would bypass a cpu pin or
re-open #1529 on ROCm; regression test added
- env-pinned AND not-detected shows both facts in one subtitle
- device_load_failed/perf_save_failed translated into all 21 locales;
CJK/th/vi/ar strings no longer say literal 'Auto'
- test_ctranslate2_never_gets_cuda_on_a_rocm_build pins the probe family
(it was order-dependent on the lru_cache before)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): pin the probe family in the faster-whisper OOM-fallback test
Same class as the rocm-build test: it mocked torch but not the probe the
new override gate consults first, so on a cpu-family CI host the CUDA
fallback chain under test was unreachable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* docs(engines): a guide for every engine + index; fix two engine-metadata bugs
21 new pages under docs/engines/ (10 TTS, 10 ASR, index README) — every
registered engine now has one: what it's for, platform support, model env
vars, quirks with issue refs. Linked from both READMEs' engine sections.
Code fixes found while verifying facts against the registries:
- KittenTTS docstring claimed default voice 'Jasper'; the code default is
expr-voice-2-f
- the isolated-ASR sidecar read only ASR_MODEL_FW while the download
preflight read ASR_MODEL_FASTER — set one and the other quietly used a
different model; both now resolve ASR_MODEL_FW-override → ASR_MODEL_FASTER
- moonshine's install hint named 'useful-moonshine', a package the backend
never imports; now moonshine-onnx / moonshine-voice
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): entries for the engine guides + sidecar model fix (#1556)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(engines): second-harvest fixes — TLS guidance, matrix/code alignment, CN counts
- README matrix aligned to gpu_compat (the code is the source of truth):
CosyVoice macOS is CPU not MPS, IndexTTS and GGUF gain their real
CUDA/CPU/MPS cells
- gpt-sovits guide: prefer https/tunnel for non-loopback servers, plaintext
warning; first-use download guidance on both OmniVoice pages
- preflight empty-env fallback matches the sidecar (ASR_MODEL_FASTER='' no
longer resolves a different repo)
- nano installs via uv pip; kitten log level wording; index links install
guides incl. the Gatekeeper step; README_CN engine counts 16/11
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(readme-cn): the all-engines-local claim now excludes the remote client
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* docs(readme): lead with download + first clone; seed benchmarks page
Quickstart (installers, install guides, a three-step first-clone walkthrough)
moves above What's-new/Features in both READMEs — visitors get the action
before the pitch. New docs/benchmarks.md anchors measured per-engine/device
numbers on the bench_pipeline.py harness, community-contributed, no estimates.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): entry for the README conversion restructure (#1555)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(bench): emit RTF + CUDA peak VRAM; guard NaN RAM; define the benchmarks schema
Bot harvest on #1555: the tts stage now prints RTF per warm measurement and
CUDA peak VRAM (None elsewhere — no made-up zeros), the stage floor refuses
unmeasurable RAM instead of sailing past a NaN comparison (FLOOR_GB=0
overrides), docs/benchmarks.md columns map 1:1 to what the harness prints,
and the download badges say they open the release page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(readme): link palash.dev from the maker section
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(bench): name the resolved engine, track VRAM from resolution, comment the guards
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(readme): the quick-switch gif is the hero image
The hero shows motion now; the Launchpad screenshot moves into the 0.5.0
What's-new slot so nothing appears twice.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(bench): peak VRAM is reserved memory; adapter engines name their model
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(bench): subprocess-isolated engines report VRAM n/a, not a parent-side zero
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(bench): out-of-process detection is declarative; sherpa rows name their model
'runs_out_of_process' is now a TTSBackend attribute set by SubprocessBackend
AND omnivoice-gguf (which inherits TTSBackend directly but spawns a binary
per generate — the isinstance check missed it). Duck-typed for the same
module-purge reason as _is_subprocess_isolated. Sherpa-onnx identity comes
from _model_dir's basename when _model_id is absent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(bench): backends self-report model identity via TTSBackend.model_identity()
Greptile enumerated the adapter engines one at a time (mlx _model_id,
sherpa _model_dir, cosyvoice env-only) — the attribute sniffing rots per
engine. The hook fixes the class: each multi-model backend reports its
own identity, the profiler just asks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>