Compare commits

...
Author SHA1 Message Date
velixio b6ee314ff2 fix: tolerate missing hosted system metrics 2026-08-18 13:30:50 +05:30
velixio e6f4c766f5 Merge remote-tracking branch 'origin/velixio' into velixio 2026-08-15 23:30:53 +05:30
velixio c4ea6a14b0 fix: require remote TTS render parity 2026-08-15 23:30:20 +05:30
velixio 751f04078d Revert "fix: preserve local-only voice generation"
This reverts commit 2926ce615a.
2026-08-15 23:25:11 +05:30
velixio 2926ce615a fix: preserve local-only voice generation 2026-08-15 23:15:40 +05:30
velixio 7e64d13739 Merge remote-tracking branch 'origin/main' into velixio 2026-08-15 22:28:00 +05:30
velixio 5151243ee4 fix: remove accent country flags 2026-08-15 22:28:00 +05:30
velixio eaee379dd5 fix: preserve seeded gallery renders in runtime adapter 2026-08-15 22:09:00 +05:30
velixio 0d81123954 fix: accept deterministic gallery seed 2026-08-15 21:31:02 +05:30
velixio ee35d2389e fix: require remote voice render parity 2026-08-15 18:49:08 +05:30
velixio 1fda5bdf96 fix: preserve voice identity on remote workers 2026-08-15 18:49:08 +05:30
Palash Debnath 2477dde688 docs: streamline README structure and specifications (#1560)
* docs: streamline README structure and specifications

* docs: address README review findings
2026-08-15 13:06:44 +00:00
velixio df2da4bb4d fix: attest immutable runtime model versions 2026-08-15 18:26:42 +05:30
velixio 37c8be6bfe Fix hosted job cancellation state 2026-08-15 15:13:46 +05:30
Palash DebnathandClaude Fable 5 48c9a3b1f8 feat(settings): compute-device override (auto / CUDA / ROCm / XPU / MPS / CPU) (#1557)
* feat(settings): compute-device override — auto | CUDA | ROCm | XPU | MPS | CPU

Auto-detect stays the default; the override kills the 'auto-detect picked
wrong' issue class. Applied at the single choke point (_probe()'s family
selection) so routing, get_best_device(), and every badge inherit it.
Resolution: OMNIVOICE_DEVICE env > Settings pick (prefs.json) > auto (#981
pattern). An override can steer, never invent hardware: a family the host
lacks is noted and ignored; cpu is always honorable. Applies at next
backend start (host caps are immutable per process — same restart contract
as the rest of the Performance tab, RestartBadge shown).

GET/PUT /api/settings/compute-device (admin-gated) reports resolved vs
applied so the panel shows restart-required truthfully and disables itself
under an env pin instead of pretending.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): entry for the compute-device override (#1557)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(device-override): harvest — the override reaches CT2 ASR, full i18n, honest edge states

- _ctranslate2_cuda_ok() and the ASR sidecar now gate on the probe's family,
  so a cpu pin (or ROCm host) can never hand CTranslate2 a CUDA device —
  the override reaches every CT2 loader through one shared gate
- override_ignored exposed by the API and shown by the panel (env pin naming
  a device this machine lacks: auto is in effect, restart won't change it)
- all 8 panel strings + 5 device-family labels translated into all 21
  locales; failed saves keep their error visible through the re-sync
- test isolation: cleanup drops OMNIVOICE_DEVICE before re-probing so no
  overridden caps leak into later tests; panel tests wait for loaded state
- xpu/intel search keywords; oxfmt formatting

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(device-override): round 2 — fail-safe probe fallbacks, complete i18n, combined pin state

- a broken capability probe now means CPU everywhere (CT2 gate + ASR
  sidecar) — never a torch-derived guess that would bypass a cpu pin or
  re-open #1529 on ROCm; regression test added
- env-pinned AND not-detected shows both facts in one subtitle
- device_load_failed/perf_save_failed translated into all 21 locales;
  CJK/th/vi/ar strings no longer say literal 'Auto'
- test_ctranslate2_never_gets_cuda_on_a_rocm_build pins the probe family
  (it was order-dependent on the lru_cache before)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): pin the probe family in the faster-whisper OOM-fallback test

Same class as the rocm-build test: it mocked torch but not the probe the
new override gate consults first, so on a cpu-family CI host the CUDA
fallback chain under test was unreachable.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 05:08:51 +00:00
Palash DebnathandClaude Fable 5 030d5ea01f docs(engines): a guide for every engine + index; fix two engine-metadata bugs (#1556)
* docs(engines): a guide for every engine + index; fix two engine-metadata bugs

21 new pages under docs/engines/ (10 TTS, 10 ASR, index README) — every
registered engine now has one: what it's for, platform support, model env
vars, quirks with issue refs. Linked from both READMEs' engine sections.

Code fixes found while verifying facts against the registries:
- KittenTTS docstring claimed default voice 'Jasper'; the code default is
  expr-voice-2-f
- the isolated-ASR sidecar read only ASR_MODEL_FW while the download
  preflight read ASR_MODEL_FASTER — set one and the other quietly used a
  different model; both now resolve ASR_MODEL_FW-override → ASR_MODEL_FASTER
- moonshine's install hint named 'useful-moonshine', a package the backend
  never imports; now moonshine-onnx / moonshine-voice

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): entries for the engine guides + sidecar model fix (#1556)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(engines): second-harvest fixes — TLS guidance, matrix/code alignment, CN counts

- README matrix aligned to gpu_compat (the code is the source of truth):
  CosyVoice macOS is CPU not MPS, IndexTTS and GGUF gain their real
  CUDA/CPU/MPS cells
- gpt-sovits guide: prefer https/tunnel for non-loopback servers, plaintext
  warning; first-use download guidance on both OmniVoice pages
- preflight empty-env fallback matches the sidecar (ASR_MODEL_FASTER='' no
  longer resolves a different repo)
- nano installs via uv pip; kitten log level wording; index links install
  guides incl. the Gatekeeper step; README_CN engine counts 16/11

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme-cn): the all-engines-local claim now excludes the remote client

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 03:55:59 +00:00
Palash DebnathandClaude Fable 5 b79ba9bd3b docs(readme): lead with download + first clone; seed benchmarks page (#1555)
* docs(readme): lead with download + first clone; seed benchmarks page

Quickstart (installers, install guides, a three-step first-clone walkthrough)
moves above What's-new/Features in both READMEs — visitors get the action
before the pitch. New docs/benchmarks.md anchors measured per-engine/device
numbers on the bench_pipeline.py harness, community-contributed, no estimates.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): entry for the README conversion restructure (#1555)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bench): emit RTF + CUDA peak VRAM; guard NaN RAM; define the benchmarks schema

Bot harvest on #1555: the tts stage now prints RTF per warm measurement and
CUDA peak VRAM (None elsewhere — no made-up zeros), the stage floor refuses
unmeasurable RAM instead of sailing past a NaN comparison (FLOOR_GB=0
overrides), docs/benchmarks.md columns map 1:1 to what the harness prints,
and the download badges say they open the release page.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme): link palash.dev from the maker section

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bench): name the resolved engine, track VRAM from resolution, comment the guards

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme): the quick-switch gif is the hero image

The hero shows motion now; the Launchpad screenshot moves into the 0.5.0
What's-new slot so nothing appears twice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bench): peak VRAM is reserved memory; adapter engines name their model

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bench): subprocess-isolated engines report VRAM n/a, not a parent-side zero

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bench): out-of-process detection is declarative; sherpa rows name their model

'runs_out_of_process' is now a TTSBackend attribute set by SubprocessBackend
AND omnivoice-gguf (which inherits TTSBackend directly but spawns a binary
per generate — the isinstance check missed it). Duck-typed for the same
module-purge reason as _is_subprocess_isolated. Sherpa-onnx identity comes
from _model_dir's basename when _model_id is absent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bench): backends self-report model identity via TTSBackend.model_identity()

Greptile enumerated the adapter engines one at a time (mlx _model_id,
sherpa _model_dir, cosyvoice env-only) — the attribute sniffing rots per
engine. The hook fixes the class: each multi-model backend reports its
own identity, the profiler just asks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 02:41:40 +00:00
velixio c818d235fb fix: preserve voice identity on remote workers 2026-08-15 02:42:13 +05:30
velixio 09ba4feb1c Merge remote-tracking branch 'upstream/main' 2026-08-15 01:46:29 +05:30
velixio 08791175f9 Merge branch 'feat/runtime-adapter' 2026-08-15 01:27:13 +05:30
velixio 4fa1b31eef Merge remote-tracking branch 'origin/main' into feat/runtime-adapter
# Conflicts:
#	CHANGELOG.md
#	README.md
#	backend/api/routers/archetypes.py
#	backend/tests/test_archetypes_api.py
#	frontend/src/api/types.ts
2026-08-15 00:30:48 +05:30
velixio 8654bb0225 fix: preserve gallery voice identity 2026-08-15 00:27:57 +05:30
Palash DebnathandClaude Fable 5 3d0c9605df test(shell): backend-lifecycle fault-injection harness (#1551)
* test(shell): backend-lifecycle fault-injection harness

Runs spawn_backend_and_wait/supervise_backend against REAL dying child
processes and asserts the user receives the correct NAMED diagnosis —
not merely that recovery happens. Wrong/missing explanation was 61% of
the historical "can't reach the backend" class; this rig is the
permanent regression harness for every future lifecycle fix.

Seam: OMNIVOICE_BACKEND_CMD (JSON argv or whitespace form) runs any
command as "the backend" — venv bootstrap and ffmpeg resolution are
skipped, everything else (err-log run offsets, drainer threads, env
pinning, real OS pipes, spawn-failure diagnostics) stays real. Plus
OMNIVOICE_LOG_DIR (per-test log+marker dirs, also a support tool) and
harness-only timing overrides OMNIVOICE_STARTUP_BUDGET_S /
OMNIVOICE_SUPERVISOR_POLL_MS whose production defaults are pinned by
unit tests. Lifecycle fns genericized over tauri::Runtime for the
MockRuntime app; behavior-neutral with the env unset (unit-pinned).

Scenarios (tests/backend_lifecycle.rs, scenario children = this test
binary re-invoking itself; serial by mutex + CI --test-threads=1):
- port conflict (exit 78) → the detectHints-matchable port phrasing
- generic chained traceback → root cause survives into the diagnosis
  and the crash marker
- spawn failure → spawn diagnostic reaches the user, NO bogus marker
- slow start past budget → timeout names the budget + last stderr
- post-Ready crash loop → 3 restarts announced, markers before restarts,
  "kept crashing" diagnosis naming the last exit
- SIGKILL (unix) → named as signal 9
- deliberate kill → supervisor yields silently, no marker, never Failed
- deferred-startup FATAL → the named step reaches the user, forensics,
  and the splash narration

CI: harness added to the 3-OS tauri-cross-platform matrix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(ci): temporary Windows loader bisect probe for the harness binary

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(shell): embed Common-Controls v6 manifest into Windows test binaries

Bisected on #1551: EVERY integration-test binary of this crate died at
load on Windows with STATUS_ENTRYPOINT_NOT_FOUND (0xc0000139) — cargo
gives test binaries no manifest, so the loader resolves comctl32 v5,
which lacks the TaskDialogIndirect entry point tauri's dialog/tray stack
imports. build.rs now embeds tests/windows-test.manifest via
rustc-link-arg-tests on Windows targets. Bisect probe removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): scenario gate is PID-valued — the parent can't self-inject

CodeRabbit on #1551: in a parallel local `cargo test`, the parent's own
scenario_child test could observe the armed env and start playing the
backend in-process (binding the port, idling 600s). The gate value is
now the arming process's PID; a matching PID stays inert, so only the
spawned child — a different process — runs the scenario.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 18:45:27 +00:00
Palash DebnathandClaude Fable 5 bb813ff676 feat(startup): bind the socket in ~1s and narrate startup step by step (#1550)
* feat(startup): bind the socket in ~1s and narrate startup step by step

The structural fix for the "can't reach the local backend" class (~1 in 5
of every issue ever filed): uvicorn served nothing until torch import
(10-20s cold), the 30-router fan-out, an import-time DB migration, the
cuDNN preload, and alembic all finished — every slow or fragile step
rendered as an unexplained dead backend.

main.py now keeps module scope fast and defers the heavy work:
- _phase_a_build (executor thread): prefs/env restore + #963 migration,
  yt-dlp overlay, cuDNN preload, torchaudio, model_manager, router
  imports — order preserved, literal imports so PyInstaller still traces.
- _phase_a_finalize (event loop, no awaits → atomic wrt requests):
  include_router, mounts, MCP, SPA, openapi bust.
- _phase_b: the old lifespan startup body; handles on app.state so
  shutdown survives a startup that never finished.
- Eager mode (pytest / OMNIVOICE_EAGER_INIT=1) runs everything at import
  — byte-equivalent behavior for the ~100 lifespan-less TestClient sites
  and for embedders (dump_api_routes, probe boot runner opt in).

While starting: /health answers 503 with the current step, new
/startup/progress serves the full ledger (always 200), and
StartupGateMiddleware 503s everything else with the [starting] marker
(same skip-the-Report-button convention as [shutting_down]). A deferred
failure keeps import-crash semantics: traceback to stderr → shell crash
forensics, run sentinel stays uncleared, exit 1 names the failed step.

Shell: startup_progress() probe (marker-header-gated so a foreign
responder can't narrate the splash) feeds per-step log lines into the
launch poll and the supervisor's reconnect wait. --health-check absorbs
the deferred init (60→180s); --diagnose runs Phase A up front so it
still sees restored prefs. Docker HEALTHCHECK semantics unchanged
(curl -f fails on 503 exactly as it did on connection-refused).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(startup): join the Phase A thread on shutdown; async fail-path sleep

Bot-review harvest on #1550: cancelling the deferred-startup task cannot
stop the executor thread inside Phase A's blocking imports — shutdown now
waits (bounded, only when a build started and hasn't finished) on a
thread-completion event so interpreter teardown can't race a mid-import
(#1000 class). The failure path's last-poll beat is now awaited, not
time.sleep — a blocking sleep froze the very loop that beat exists to let
serve. Also: dump_api_routes forces eager (assignment, not setdefault),
and the integration test's child gets DEVNULL instead of an undrained
pipe that could wedge a cold boot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(startup): close the Phase A submission race; CodeQL nits

Review finds on #1550: shutdown could sample _phase_a_started unset
while the executor callable was queued-but-not-running, skipping the
thread join. started is now set BEFORE submission, the submission is
shielded so a cancel can't strand a queued callable that would never set
_phase_a_finished, and the wrapper sets finished on every exit including
the already-built early return. Contract pinned by
test_phase_a_thread_join_contract. Plus explanatory comments on the new
bare excepts and a consistent return in the gate's websocket branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 15:23:02 +00:00
Palash DebnathandClaude Fable 5 bc6acec5a3 fix(shell): gate Ready on the deep health probe; pace crash-loop restarts (#1548)
* fix(shell): gate Ready on the deep health probe; pace crash-loop restarts

Two supervisor hardenings from the backend-reliability root-cause pass:

Ready now requires backend_ready() — the identity probe (/system/info
string-sniff) AND the deep probe (/profiles must 200) — at both Ready
transitions (startup poll, supervisor respawn wait). The shallow probe
alone announced a backend whose install/DB had broken underneath as up;
the UI looked alive while every real request 500'd or dead-ended on
"can't reach the backend". Death detection stays process-exit-only, so a
busy-but-alive backend is still never killed.

Supervisor respawns now back off: first respawn immediate (a one-off
crash self-heals fast), then 5s, then 15s, capped — the budget check
ends a hopeless loop, not an unbounded sleep. The pause runs behind the
already-visible "reconnecting" banner and yields within 500ms to app
quit or a deliberate retry-flow replace.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(shell): yield backoff to a tracked replacement child, not just the flag

Greptile P1 on #1548: a completed Retry/Clean&Retry sets the deliberate-
kill flag and track_backend_child clears it — possibly both between two
500ms backoff samples, so the flag alone can be missed and the old
supervisor would free_port() the retry's healthy replacement. The dead
child we observed can never read as alive again, so a live tracked child
during backoff can only be a replacement — yield to it promptly so the
retry's spawn_backend_and_wait can claim the supervisor slot at Ready.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(shell): backoff yields on spawn-generation change, not liveness

Second Greptile pass on #1548: a replacement child that itself exits
before the old supervisor's next 500ms sample read as "still dead" under
the liveness check, so ownership transfer was missed. The spawn
generation (bumped by every track_backend_child, never un-bumped) is
observable regardless of the replacement's fate — snapshot it at death
detection, yield the moment it changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(shell): snapshot spawn generation before observing the exit

Third-pass review find: sampled after try_wait, a replacement tracked in
the gap bakes its own generation into the snapshot and the ownership
transfer is missed. Snapshot first, and re-check once more before
touching the port so the zero-backoff first respawn is covered too.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 14:41:21 +00:00
Palash DebnathandClaude Fable 5 94ba362ef2 feat(triage): crash-class recurrence report — the reliability metric (#1549)
* feat(triage): crash-class recurrence report — the reliability metric

scripts/crash_class_report.py measures the "backend died / never came
up" class (the project's #1 lifetime failure, ~1 in 5 of all issues)
filtered to reports from the current version — the definition of done
for the reliability cycle. Buckets by the bug reporter's Build-status
stamp (#1547): current / outdated / unknown (pre-deflection builds), so
deflection-miss noise never pollutes the number the work is judged on.

tests/scripts/test_crash_class_report.py pins the title→sub-class
mapping against the real historical title shapes and locks the stamp
literals to frontend/src/utils/bugReport.js so a reworded marker fails
in CI instead of silently zeroing the metric.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(triage): --version is authoritative; loud fetch-cap warning

Bot-review harvest on #1549: with --version, the Environment Version
line now decides the bucket (extracted to pure classify_build + tests) —
a report stamped "current at filing time" during another version's
window no longer counts toward this version's recurrence. Hitting the
500-issue fetch cap now warns loudly instead of silently understating.
The stamp lockstep test asserts the full Build-status prefix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 14:20:35 +00:00
Palash DebnathandClaude Fable 5 aabe5783f3 fix(report): offer the latest release before filing from an outdated build (#1547)
* fix(report): offer the latest release before filing from an outdated build

6 in 10 sampled "can't reach the backend" reports came from builds that
were already obsolete when filed, and were closed with "please update" —
pure triage noise. Every Report-bug affordance now funnels through
openBugReport(): on an outdated build it offers the latest release first
(with a "File anyway" escape hatch), and the report body carries a
triage-greppable "**Build status:**" line either way, so current-version
recurrence — the reliability metric — is countable separately from
stale-build reports.

Freshness sources per deployment (behavior identical, implementation per
mode): desktop reads the Rust updater's channel-aware verdict from the
store (no new network path, no CSP widening); browser/dev/Docker make one
bounded latest-release GET, only once the user has initiated the report
flow whose destination is github.com. An 'unknown' dev build stays
silent entirely — never nudged, never claimed current.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(report): anchor version parsing; zh-TW reportBug.title in Traditional

Bot-review harvest on #1547: parseVersionTriple now rejects trailing
non-semver data (1.2.3.4, 1.2.3garbage) instead of silently reading the
leading triple into an outdated/current verdict; the pre-existing zh-TW
reportBug.title was Simplified-script — now properly Traditional.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 13:44:24 +00:00
Palash DebnathandGius 854b4852ed perf(frontend): coalesce persistence writes off input paths (#1546)
Defer and coalesce omnivoice.app and omni_ui persistence behind a 250 ms
quiet window with a 1,000 ms hard maximum, preserving storage schemas,
synchronous pending reads, legacy formats, Factory Reset semantics,
widget read-only ownership, and lifecycle (pagehide/visibilitychange)
durability. Adds scheduler, restore, reset, StrictMode, concurrent-render,
role-ownership, and migration regression tests plus an opt-in
production-bundle responsiveness harness.

Lands #1541 by @bultodepapas (maintainer landing branch; the out-of-scope
attribution-policy commit was dropped).

Co-authored-by: Gius <bultodepapas@gmail.com>
2026-08-14 13:25:36 +00:00
velixio 6111b8e4ae Merge remote-tracking branch 'origin/main' 2026-08-14 15:04:00 +05:30
Palash DebnathandClaude Fable 5 420bc73e78 fix(release): harden the AppImage repair — pinned tooling, final-writer manifest (#1545)
* fix(release): harden the AppImage repair step

All three review findings on #1544, fixed before the tag re-runs it:
- appimagetool pinned to the immutable 1.9.1 release with a verified
  SHA-256 — a mutable 'continuous' binary must not execute with the
  updater signing key and a release-write token in its environment
- the release tag reaches the script as env data, never interpolated
  into shell source (zizmor template-injection)
- a failed latest.json download now fails the step unless the asset is
  confirmed absent, and the patch refuses to upload unless at least one
  linux signature was actually replaced — a repacked AppImage can never
  ship paired with stale updater metadata

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(release): the updater manifest gets one final writer, after the matrix

CodeRabbit + Greptile on #1545: every tauri-action leg re-uploads the
shared latest.json, so patching it inside the Linux leg races the other
platforms — a later leg's upload could resurrect the stale pre-repack
signature. The manifest patch moves to a post-matrix job that runs once
after all legs: it aligns the manifest's linux entries with the .sig
asset that actually shipped (self-verifying — no cross-job state), and
no-ops when they already agree. The leg keeps asset repack/re-sign only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(release): a failed .sig download fails the manifest-align job

Same fail-closed rule as the manifest itself: absence is decided by the
asset list; any other download failure must not exit 0 with a stale
signature left in latest.json.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 01:55:10 +00:00
Palash DebnathandClaude Fable 5 fb46fa4788 fix(release): repack the AppImage with a real .DirIcon file, re-sign, re-upload (#1544)
The v0.5.0 tag build — the first real release since #1518's guard — proved
the files-map fix loses: linuxdeploy re-links .DirIcon to an ABSOLUTE
build-machine path after tauri places the real bytes, and the guard
correctly refused to publish. tauri-action's atomic build+sign+upload
leaves only a post-upload seam, so the Linux job now repairs the packed
artifact: extract, replace .DirIcon with the icon bytes as a regular file
(nothing left to dangle), repack with appimagetool, re-sign with the
updater key, clobber the draft release's asset and patch the linux
signature inside latest.json. The existing smoke then validates the
repaired AppImage. No-ops cleanly when .DirIcon already resolves.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 00:59:14 +00:00
Palash DebnathandClaude Fable 5 f83b7371c3 docs: remote-workers guide speaks VoiceStudio (#1543)
CodeRabbit's last-cycle Minor on #1540, applied as the immediate
follow-up the docs-sync rule prescribes: three OmniVoice mentions in
docs/remote-workers.md now carry the product's name.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 00:07:05 +00:00
Palash DebnathandClaude Fable 5 9184d7d625 chore: bump to 0.5.0 (#1540)
* chore: bump to 0.5.0

Owner-requested minor bump. package.json is the source of truth; the
three mirrors (Cargo.toml, pyproject.toml, _FALLBACK_VERSION), the three
lockfiles and the branding pin move in lockstep, and the accumulated
Unreleased section becomes the curated 0.5.0 release notes — quiet
Highlights first, one-liner subsections after, duplicates folded.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: fresh README for the VoiceStudio era; 0.5.0 notes wear the release

README: 551 lines from 656 — a What's-new-in-0.5.0 section with real
captures, the engine tables corrected to the actual 16 TTS registrations,
a stale Settings path and a broken Colab link fixed, roadmap/FAQ/credits
trimmed to what earns its place.

Release notes: the quick-switch GIF and catalogue/gallery screenshots,
captured from the running app during the pre-bump test pass, embedded
after the Highlights; #1542's gallery work and the ffmpeg CI fallback
recorded in their subsections.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: the feature inventory's Remote Model Downloads mention survives the README trim

check-docs-drift requires every docs/features.yaml name verbatim in the
README; the overhaul folded the phrase away. It now lives in the remote
workers feature line.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: one blockquote, and an architecture claim that survives the opt-ins

CodeRabbit on #1540: MD028 blank line inside adjacent blockquotes, and
'every layer is on your machine' contradicted the opt-in remote paths
documented two sections away — it now states local-by-default with the
opt-ins named.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 23:50:23 +00:00
Palash DebnathandClaude Fable 5 0ee62b2261 feat(gallery): save gallery voices as profiles, with validated audio references (#1542)
* feat(gallery): save gallery voices as profiles, with validated audio references

Work-in-progress lifted from the concurrent gallery session at the
owner's request (its uncommitted working tree, preserved verbatim from
base 92b1ee5d; safety snapshot remains at rescue/gallery-wip):

- gallery voices can be saved as local profiles: audio is copied into
  the profile store with content-addressed filenames, existing profiles
  are detected and refreshed only when the source clip changed
- backend/core/audio_validation.py: symlink-rejecting, root-contained
  resolution for persisted profile WAV references, with tests
- archetype/community routers and the Voice Gallery UI updated for the
  save-as-profile handoff (spec: docs/specs/longform/26-gallery-use-handoff.md)
- locale updates for the new gallery strings across all 21 files

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: drop a stray local screenshot script that rode in with the tree copy

* fix(community): explain the tolerated Content-Length parse failure; drop an unused import

CodeQL on #1542: the empty except now says why it is safe (the streamed
byte counter enforces the same cap regardless), and the test file loses
an unused Path import.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gallery): review findings — copy outside the write lock, no stale completions

CodeRabbit on #1542, all findings addressed:
- the profile-audio copy stages to a .part temp BEFORE BEGIN IMMEDIATE
  and publishes via atomic os.replace inside it — other backend writers
  no longer block for the duration of an audio copy; a mid-copy failure
  leaves no temp droppings and no profile row (both pinned by tests)
- VoiceGallery async ops carry per-operation generation tokens: a
  preview or save-as-profile that resolves after unmount (or after a
  newer operation) can no longer play audio, redirect into a workspace,
  or touch state — three fail-before regression tests
- VoiceGalleryActions imports the page at test runtime; the e2e locator
  uses a stable data-testid instead of a translated string; symlink
  tests skip cleanly where the OS can't create symlinks; the changelog
  line carries its PR ref

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: static ffmpeg fallback when the chocolatey feed is down

Third feed outage to break a PR run (2026-07-20, 2026-07-28, today —
three attempts, three 'installed 0/1'). Chocolatey is a distribution
channel, not the dependency: after the retry loop exhausts, fetch the
static gyan.dev build from its GitHub release mirror and put it on
PATH — same binary, no feed in the path. URL verified live (HTTP 200).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 22:48:05 +00:00
Palash DebnathandClaude Fable 5 77ae194f9c fix(asr): ROCm torch is not CUDA — keep CTranslate2 off the HIP GPU (#1539)
An RX 7900 XTX in the :rocm image crashed ASR init with 'CUDA driver
version is insufficient for CUDA runtime version' (#1529): ROCm torch
answers torch.cuda.is_available() and hands out 'cuda' device strings,
but whisperx/faster-whisper run on CTranslate2, whose CUDA runtime is
NVIDIA-only. Same class as the Apple/#1127 lesson, on the AMD axis.

- _ctranslate2_cuda_ok(): 'cuda' for CTranslate2 only when torch is a
  real CUDA build (torch.version.hip is the honest tell); ROCm hosts
  take CPU int8 instead of a native crash.
- _auto_detect(): on a ROCm-GPU host prefer pytorch-whisper — a pure
  transformers pipeline riding torch itself, so it actually uses the
  HIP GPU while CTranslate2 engines would idle on the CPU.

Fail-before/pass-after: 4 new tests fail on the old device pick/order.

Fixes #1529

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 21:05:53 +00:00
velixio b1f322dde2 fix(studio): restore browser interactions 2026-08-14 02:30:23 +05:30
Palash DebnathandClaude Fable 5 d3822c4976 feat(engines): open the door between the LLM family and its providers (#1538)
* feat(engines): open the door between the LLM family and its providers

The openai-compat family entry and the LLM Providers panel are one
system — llm_backend resolves every call through the active provider —
but the UI presented them as unrelated (council coherence finding). Now:

- the catalogue's openai-compat row carries a 'Provider · model' hint
  naming the endpoint that actually answers (decorative: a provider
  registry hiccup degrades to no hint, never a failed listing)
- the row offers 'Configure providers' straight into Settings → LLM
  Providers; the panel gains the backlink into catalogue → LLM family
- three new strings in all 21 locales, matching each file's provider
  terminology

Also: bugReport's encoded-ceiling test is hermetic now — it was the one
test in its file trusting ambient fetch, and hung on any machine where a
local backend holds the port without answering.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(i18n): the catalogue note names the ACTIVE provider, not the edited one

CodeRabbit on #1538: the panel can be editing a provider that is not
active, and 'this provider answers…' then points at the wrong one. The
note now says the provider MARKED ACTIVE answers, which is true under
any selection — no state-dependent copy needed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 20:46:36 +00:00
19ae20111a fix(security): replace persistent admin keys with scoped sessions (#1528)
* fix(security): replace persistent admin keys with sessions

Exchange the remote administrator key once for bounded, revocable credentials. Canonicalize backend principals, enforce cookie CSRF and exact origins, and use path-bound one-use WebSocket tickets.

Migrate the bundled UI away from durable master-key storage and credential-bearing URLs. Add unit, integration, static-hygiene, and production-browser regressions plus synchronized operator documentation.

* docs: link session hardening to PR 1528

* fix(security): key session indexes with process pepper

Use HMAC-SHA-256 instead of an unkeyed digest for in-memory session and WebSocket-ticket indexes. This preserves constant-size lookup identifiers, makes copied records unusable without the process pepper, and resolves CodeQL's weak sensitive-data hash finding.

* fix(auth): align empty bearer migration precedence

Centralize the Authorization-channel presence decision with canonical principal parsing. Bearer followed only by spaces now remains an empty channel during legacy cookie migration, while unsupported or invalid explicit credentials stay authoritative and fail closed.

* fix(security): harden admin session review boundaries

* fix(security): derive key generations with HKDF

* fix(auth): anchor the admin-session store so module reloads cannot fork it

test_master_exchange_does_not_bypass_pin_on_normal_routes failed in full-suite
runs: test_mcp_bindings' client fixture purges the services.* tree from
sys.modules and reloads main, so api.routers.auth re-imported a fresh
services.admin_sessions (new AdminSessionStore) while core.auth kept its
import-time reference to the old one — the exchange issued the cookie into
one store and the middleware resolved it against another, turning the
expected "PIN required" into "API key required".

Root cause is the class of bug, not the one test: a process-global auth
store defined as a bare module-level singleton forks under importlib.reload
or purge-and-reimport. Fix at the source: admin_session_store now resolves
through a synthetic sys.modules anchor (_omnivoice_admin_session_store_anchor)
that reloads never re-execute and package-prefix purges never match, so every
copy of the module shares the one per-process store. No consumer or behavior
changes.

Regression test reproduces both fork vectors (in-place reload and
sys.modules purge + fresh import) and asserts previously issued sessions
still resolve and the store identity is preserved; it fails before this fix
and passes after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(auth): honor X-Forwarded-Proto for CSRF origin and Secure cookies behind TLS proxies

Behind Tailscale Serve (docs/remote-gpu.md) or any TLS-terminating proxy,
the browser talks https while the backend hop stays http, so exact-origin
CSRF compared an https Origin against an http expectation and rejected
every legitimate request, and the session cookie shipped without Secure.
uvicorn's ProxyHeadersMiddleware only rewrites the scope for loopback
peers, which misses Docker and any non-loopback proxy topology.

New core.csrf.effective_scheme derives the client-facing scheme: resolved
scope first (uvicorn's trusted-proxy rewrite wins), then an upgrade-only
read of X-Forwarded-Proto's first value — https/wss promotes http to
https, everything else is ignored, and a genuine TLS hop can never be
downgraded. Used by both the destination-origin comparison and
auth._secure_cookie so the WS-ticket/logout CSRF paths and the cookie
Secure flag agree. Spoofing gains nothing: the host:port half of the
origin tuple is untouched, browsers cannot attach the header cross-site
without a preflight this API never grants, and forging it on plain http
only adds Secure (the browser then drops the cookie — self-harm only).

Regression tests: proxied https origin accepted (origin check, Secure
flag, logout), comma-separated chains, scope-fallback path, spoofed
header still rejects cross-origin, cannot downgrade real https, junk
values ignored.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(auth): consume the stored admin key only after a successful exchange

A remote-backend user upgrading with their backend unreachable lost the
only stored copy of OMNIVOICE_API_KEY: every migration path deleted the
durable ov_api_key BEFORE the session exchange settled, stranding them
until they recovered the key from the server box. Close the whole class:

- client.ts bootstrap: read the legacy key, exchange first, and remove
  the durable copy only after the exchange succeeds; on failure the key
  stays so the next launch retries the migration (auth gate still rises).
- authSession.ts exchangeApiKey: move removeLegacyMaster from before the
  fetch to the cookie/bearer success paths — the key never coexists with
  a live session, but a rejected or hung exchange no longer consumes it.
- remoteBackendProbe.ts configuredRemoteBackend: stop wiping the key on
  every app mount.
- RemoteBackendPanel: a connection test or an aborted save no longer
  wipes the pending key; only disabling the remote backend discards it.
- prefKeys.js: ov_api_key moves from PREF_KEYS to PRESERVED_KEYS —
  factory reset preserves the pending connection credential exactly like
  ov_backend_url; the successful migration is what deletes it.

Tighten the credential-hygiene static guard to match: it accepted
sessionStorage.setItem('ov_api_key', …) — the exact class it exists to
close. The guard now flags .setItem(<master key>) on any storage
receiver, quote style, or injected-store alias, with a self-test pinning
what it catches and what stays legal.

Fail-before/pass-after regression tests: backend unreachable retains the
key and the next bootstrap retries it; a successful exchange removes it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* perf(auth): make session validation occupancy-independent

* test(auth): catch optional master-key storage calls

* feat(docs): add PR control document for bultodepapas in VoiceStudio

* docs: keep the PR tracking board in the fork; credit the changelog line

The pr-control document is excellent process discipline, but it is the
contributor's own operational board (their inventory, their update
commands) — it lives naturally in their fork, and docs/agents/ here is
context every repo agent loads. Removed with appreciation; the changelog
line gains its contributor credit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 20:46:24 +00:00
Palash DebnathandClaude Fable 5 b982192011 fix(worker): reserve the capacity slot before announcing the accept (#1537)
* fix(worker): reserve the capacity slot before announcing the accept

Two changes for #1536 (flaky test_worker_at_capacity_rejects_without_penalty):

The client now inserts into _running BEFORE awaiting the accept send.
Today _send enqueues synchronously so the old order could not actually
interleave — but the reserve-then-announce order is the invariant that
stays correct if _send ever gains backpressure (a bounded outbox is the
natural evolution), instead of silently reopening an over-accept window.
If the accept send fails, the reserved task is cancelled: work the
scheduler never saw accepted must not run to double-execution.

The test now pins the real invariant — no over-concurrency — rather than
the scheduler's bookkeeping timing: on a loaded CI runner the first
attempt can die environmentally (a stream hiccup fails _run, whose
finally frees the slot), after which accepting the second task is the
CORRECT behaviour the old flat assertion punished as a failure. The
assertion now applies only while the first attempt is still running, and
names the over-accept explicitly when it fires.

Fixes #1536

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(worker): release the reserved slot on a cancelled accept-send

CodeRabbit on #1537:
- except BaseException, not Exception: a handler cancelled while the
  accept send is in flight must release the reserved slot too, or the
  unaccepted task keeps running and double-executes after reassignment.
  Fail-before/pass-after regression test included.
- the capacity test now polls the second task out of its dispatch states
  instead of sleeping 0.5s, and asserts the penalty-free invariant
  (excluded_workers empty) unconditionally — capacity rejections never
  exclude the worker regardless of the first attempt's health.
- TaskAccepted.envelope finding skipped: the field is read nowhere
  server-side (registration is the only envelope consumer) — pre-existing
  unused-field design, not introduced here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 19:19:22 +00:00
Palash DebnathandClaude Fable 5 81b146f53b fix(i18n): 53 default-value-only strings now speak all 21 languages (#1534)
* fix(i18n): 53 default-value-only strings now speak all 21 languages

Every string used via t(key, { defaultValue }) without a locale entry
rendered English for every non-English user — the worker/compute chrome
(quick settings, join panel, QR enrolment), clone/design labels, workspace
voice strip, and two crash explainers. All 53 keys now exist in en.json
and carry reviewed translations in the 20 other locales, matching each
file's established terminology (existing worker/token/engine vocabulary,
catalogue tab names for the in-text path references, registers preserved,
{{placeholders}} byte-identical, the ovw_ token prefix untranslated).

The two crash explainers are translated from their FULL concatenated
source text — the extraction initially captured only the first string
segment, which src/test/streamDropError.test.ts caught by failing on the
missing proxy/buffering guidance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(i18n): reference localized UI labels inside diagnostic strings

CodeRabbit on #1534, fixed as the class: every locale's crash_broken_env
quoted the English "Clean & Retry" although the button itself is
localized — all 19 now quote each file's own clean_retry label. Plus the
flagged singles: es unload verb disambiguated from downloading, hi unload
verb aligned with crash_oom_kill, sv kontrollplan gender agreement, de
crash_broken_env moved to the file's Sie register, ru seed_reroll_hint
mistranslation, zh-TW path label matched to the real 系統日誌 section name.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 18:27:27 +00:00
Palash DebnathandClaude Fable 5 72cabb3daf feat(launchpad): wear the signal-field waveform in the hero (#1533)
The same cover artwork the project uses on the web, bleeding in from the
hero's right — where the layout holds only air — behind a radial feather
plus right-edge fade so no box edge survives, screen-blended so its dark
field vanishes into the chrome. Decorative: aria-hidden, empty alt,
pointer-inert; 22 KB webp bundled via vite.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 17:48:38 +00:00
Palash DebnathandClaude Fable 5 4a228a00b4 fix(crash): bound crash evidence to the run that produced it (#1532)
* fix(crash): bound crash evidence to the run that produced it

backend_err.log is one file shared by every backend run, and it was
TRUNCATED on each spawn. Both properties destroyed the evidence a crash
marker exists to carry: a respawn wiped the dead process's final words,
and an unbounded tail read afterwards attached the replacement's healthy
startup to the old run's crash marker — the undiagnosable report in #1510
(startup lines, no traceback, timestamps after the recorded crash).

The file is append-only now with a run-start header; each spawn records
the byte offset where its run begins; and every death path (crash markers,
venv-heal detection, restart-budget message, the 300s startup timeout)
reads through read_error_log_tail_for_run(), which cannot see another
run's output. The log rotates to backend_err.log.1 past 1 MiB so
append-only cannot grow unbounded. Bootstrap-phase reads (uv sync) keep
the whole-file reader — no backend run exists yet there.

Fail-before/pass-after: the new tests fail under the old File::create
truncation and unbounded tail.

Fixes #1510

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(crash): flush the dying run's stderr before the next offset; redact home paths

CodeRabbit on #1532:
- the stderr drainer is tracked now and joined (2s bound) before a new
  spawn records its offset, so a dead run's buffered tail cannot be
  appended after the new run's start and misattributed. Full per-child
  offset binding is unnecessary: spawns are serialized by the #1223
  spawn-once flow; the buffered tail was the only remaining window.
- the spawn-failure diagnostic redacts the home-directory prefix — it is
  retained across runs now and lands verbatim in bug reports.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 17:41:22 +00:00
velixio 1a9a70509e feat(hosted): add opt-in voice adapters 2026-08-13 23:09:54 +05:30
velixio e877572c1a feat(runtime-adapter): load models before serving
A cold engine emits no execution evidence while it loads weights and compiles,
so the Gateway's attempt lease expires mid-load, the attempt is fenced, and the
next attempt pays the same cost — a loop that never produces audio. Loading
every READY model before the socket accepts work moves that cost to startup,
where preflight already expects to wait, so the first Execute begins inference
immediately. A prewarm failure is reported rather than fatal, and --no-prewarm
restores the previous behavior.
2026-08-13 21:13:13 +05:30
Palash DebnathandClaude Fable 5 f81ace68d1 feat(engines): frame uninstalled engines as headroom, not failures (#1531)
The catalogue's group captions read "Available" / "Not installed", which
renders a fresh install (3 of 16 engines ready) as a mostly-broken app.
The sections now say "Ready to use" / "Add more engines", and the
unavailable-row details toggle asks "What it needs" instead of "Why
unavailable?" — same information, framed as headroom to unlock.

Council outcome (first-run seat): unavailable engines must read as more
you could install, never as brokenness. Keys added to all 21 locales.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 15:38:35 +00:00
Palash DebnathandClaude Fable 5 9832fbd693 feat(engines): switch engines from anywhere — footer quick switch, workspace chips, shortcuts (#1530)
* feat(engines): add quick switching controls

* docs(changelog): note engine quick switching (#1530)

* fix(support): theme amount cards

* fix(support): restore themed amount cards

* fix(engines): address quick switch review findings

* fix(engines): green the full frontend suite around the quick switch

Three failure classes the targeted runs missed:
- the popover referenced --chrome-radius, which does not exist; it now
  wears the footer's shared MENU_SURFACE like the compute popover
- LogsFooter tests hand-wrote their api/system and api/hooks mocks, which
  drop every export the footer gains next; they are partial mocks now
- DubHeader/AudiobookHero tests rendered without a QueryClientProvider,
  which useEngines needs

Also: workspace-header chips open the popover downward (dropUp stays on
the footer instance) so it cannot clip off the top of the viewport.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(engines): one QueryClient per test module, not per render

CodeRabbit: the inline client made every wrapper render a fresh cache,
so rerender() restarted the /engines query mid-test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 15:21:05 +00:00
velixio 1f03f5632c perf(gallery): build previews concurrently and resume from disk
Rendering 1126 previews ran one clip at a time, and only the first of a
clip's five stages is on the GPU: render, then watermark embed, MP3 encode,
decode, and detection. The card idled through four CPU stages per clip.

The embed and detection are neural forward passes that ran inline on the
event loop, so they held it for the whole clip -- concurrency would have
queued behind a busy loop and bought nothing. They now go through
asyncio.to_thread, which is what makes threads the right tool here: torch
releases the GIL inside those passes, so there is no second model copy and no
IPC for the tensors. Clips then build --jobs at a time (4 by default, 1
restores the old serial behaviour) under a semaphore, because every clip in
flight holds decoded audio.

A lost watermark still stops the entire run rather than only its own clip.

--resume now also adopts MP3s already on disk. The manifest is written once,
at the end, so a run interrupted at clip 900 left 900 correct files that
--resume could not see and re-rendered every one of them. Everything an entry
needs -- sha256, byte length, duration, featured flag -- is recoverable from
the file and the catalog, so recover it.

Also add a watermark preflight. Every clip was already verified individually,
but only after the first full render, and the message blamed the bitrate when
the cause can be unrelated to audio: on a host without python3-dev, AudioSeal's
forward pass dies inside Inductor, embed_watermark catches it, and the clip is
returned unmarked. Two seconds up front, with the actual cause named. It also
warms the lazy generator/detector globals single-threaded, before --jobs fans
out.
2026-08-13 17:45:34 +05:30
velixio 4335c8c1ea test(runtime-adapter): cover the preflight contract and execute taxonomy
Fake-engine/fake-inventory tests over a real UDS gRPC server plus a fast
direct-executor path: health/capabilities shape and version identity, the
Go-preflight port passing with a READY model and failing closed without
one, only-READY-counts semantics, digest stability/sensitivity, execute
happy path (manifest checksum matches the written WAV), deadline
enforcement, cancel race with idempotent dispositions, slot exhaustion,
duplicate attempts, URL/relative handle rejection, checksum mismatch, and
the input/model-load/inference/GPU/storage failure classification.
2026-08-13 13:51:14 +05:30
velixio dcd8683f3a feat(runtime-adapter): implement the GPU-node runtime gRPC server
RuntimeAdapterService over a private Unix-domain socket (default
/run/voicestudio/runtime.sock, VOICE_STUDIO_RUNTIME_SOCKET override; no
HTTP, no TCP, no database, no outbound network):

- Health/GetCapabilities read one RuntimeContext, so runtime/adapter
  versions are identical across both calls by construction. Devices come
  from torch (CUDA per-GPU / MPS / CPU with system RAM as capacity);
  models come from the tts_backend engine registry + hf_revisions pinned
  revisions, digest-pinned via a cached sha256 snapshot digest. READY is
  explicit: probe passed, snapshot complete, digest computed — a
  loading/installed/failed model is reported truthfully, never READY.
- Execute streams started -> bounded progress -> exactly one terminal
  event, validates attempt identity, approved model digest, typed bounded
  parameters, and LOCAL absolute-path handles (URL-shaped handles are
  invalid input, never fetched), runs the engine on a worker thread,
  enforces the request deadline, and writes the output WAV atomically
  with a size/sha256/duration manifest plus raw measurements.
- Stable RTA_* failure codes map onto RuntimeFailureClass: input,
  model-load, inference, GPU-resource, local-storage, canceled, crash.
- Cancel is idempotent by attempt id (ACCEPTED / ALREADY_TERMINAL /
  NOT_FOUND) against a bounded attempt registry.
- python -m backend.runtime_adapter serves; --selfcheck starts a temp
  socket and runs a port of internal/gateway/preflight.go's checks
  against itself (verified passing on this host: 1 device, 2 ready
  digest-pinned models).
2026-08-13 13:51:03 +05:30
velixio b2f94d2bf8 feat(runtime-adapter): vendor the vssaas wire contract and committed stubs
Vendor api/proto/voicestudio/runtime/v1/runtime_adapter.proto from vssaas
byte-identically into backend/runtime_adapter/, generate the grpcio stubs
into gen/ (committed, same policy and import fixup as
backend/worker/protocol/gen/), and add the drift test that regenerates
into a tmpdir and diffs.
2026-08-13 13:37:56 +05:30
Giuseppe Rojas cd54113173 fix(security): close server-mode admin bypasses (#1525)
* fix(security): require keys for remote admin actions

* fix(frontend): guard unavailable scrollIntoView

* docs: link changelog to PR 1525

* fix(security): align PIN-only discovery policy

* fix(security): preserve strict sidecar boundary

* fix(security): normalize remote API keys

* fix(auth): normalize credential fallback order
2026-08-13 05:04:29 +00:00
Palash DebnathandClaude Opus 5 6948399e61 fix(wayland): rewrite a stale portal identity instead of trusting it (#1526)
* fix(wayland): rewrite a stale portal identity instead of trusting it

Found live: with the desktop entry's Exec pointing at a binary that had
been moved, GLib resolves the entry to NULL, the host portal answers
"Could not register app ID: App info not found", CreateSession then
fails with "An app id is required" — and the dictation shortcut is
silently dead for the entire session. Only the focused-window fallback
keeps working, which reads as "the shortcut randomly stopped".

ensure_desktop_identity() trusted any existing entry. It now validates
the USER-LOCAL entry's Exec target and rewrites the entry when the
program is gone (parsing both the current quoted spelling and the
unquoted one older builds wrote). System-dir entries stay untouched —
deb installs manage their own.

The stale-entry class is easy to hit in the wild: a dev entry pinned to
target/debug survives cargo clean; an AppImage entry survives the file
being moved or renamed.

(Rebuilt from the first push, whose `git add -A` had swept in another
working session's unrelated in-progress files; this commit carries only
the wayland fix and its changelog line.)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(wayland): read Exec from the Desktop Entry group only

CodeRabbit, #1526: find_map over every line accepted an Exec= from a
[Desktop Action …] group, so an entry with no main-group Exec — which
GLib resolves to NULL — could be retained as healthy, keeping exactly
the stale identity the rewrite exists to replace. Parsing is scoped to
[Desktop Entry] now, with an action-only regression case.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 05:00:23 +00:00
302 changed files with 28071 additions and 2943 deletions
+24
View File
@@ -271,6 +271,17 @@ jobs:
working-directory: frontend/src-tauri
run: cargo test --lib --target ${{ matrix.rust_target }} --message-format=short
# Backend-lifecycle fault-injection harness: real child processes die
# scripted deaths through the OMNIVOICE_BACKEND_CMD seam, and each
# scenario asserts the user-visible diagnosis names the actual cause
# (port conflict / traceback root cause / spawn failure / timeout /
# crash-loop exhaustion / signal 9 / deliberate replace / deferred-
# startup step). Serial: the scenarios share process-global state
# (env vars, crash store, kill-intended flag) by design.
- name: Cargo test (backend lifecycle harness)
working-directory: frontend/src-tauri
run: cargo test --test backend_lifecycle --target ${{ matrix.rust_target }} --message-format=short -- --test-threads=1
# ── Cross-platform Python runtime smoke (Phase 0 GATE-02) ───────────────
# Loads the frozen tests/fixtures/omnivoice_data/ fixture and boots the
# FastAPI app in-process via TestClient on macOS/Windows/Linux. Catches
@@ -365,6 +376,19 @@ jobs:
echo "choco attempt $i did not produce ffmpeg — retrying in $((i * 30))s"
sleep $((i * 30))
done
# Chocolatey is one distribution channel, not the dependency. When
# its feed is down across every retry (2026-08-13: three attempts,
# three 'installed 0/1'), fall back to the static gyan.dev release
# build GitHub mirror — the same binary, no feed in the path.
if ! command -v ffmpeg >/dev/null 2>&1; then
echo "::warning::choco feed down — falling back to static ffmpeg build"
curl -fsSL --retry 3 -o /tmp/ffmpeg.zip \
https://github.com/GyanD/codexffmpeg/releases/download/7.1/ffmpeg-7.1-essentials_build.zip
unzip -q /tmp/ffmpeg.zip -d /tmp/ffmpeg
bindir=$(dirname "$(find /tmp/ffmpeg -name ffmpeg.exe | head -1)")
echo "$bindir" >> "$GITHUB_PATH"
export PATH="$bindir:$PATH"
fi
ffmpeg -version
- name: System deps (Linux)
+100
View File
@@ -731,6 +731,62 @@ jobs:
find "$INSTALL" -type f -path '*backend*main.py' | grep -q . || fail "backend source main.py missing"
echo "OK — MSI installed shell + uv + backend resources"
# linuxdeploy re-links .DirIcon as an ABSOLUTE symlink into the build
# machine AFTER tauri's files-map has placed the real icon bytes — the
# exact bug #1518 guarded against, resurfacing on the first real tag
# build (v0.5.0). The seam tauri-action leaves us is post-upload: repack
# the AppImage with the icon as a REGULAR FILE, re-sign it (the updater
# signature covered the old bytes), and clobber the draft release's
# asset + the linux signature inside latest.json. The smoke below then
# validates the repaired artifact, not the broken one.
- name: Repair AppImage .DirIcon, re-sign, re-upload
if: runner.os == 'Linux'
timeout-minutes: 10
shell: bash
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
TAURI_SIGNING_PRIVATE_KEY: ${{ secrets.TAURI_SIGNING_PRIVATE_KEY }}
TAURI_SIGNING_PRIVATE_KEY_PASSWORD: ${{ secrets.TAURI_SIGNING_PRIVATE_KEY_PASSWORD }}
# Data, not shell source (zizmor template-injection): a crafted ref
# must never expand inside a script that holds the signing key.
TAG: ${{ (needs.preview-gate.outputs.is_preview == 'true') && 'preview' || github.ref_name }}
run: |
set -euo pipefail
APPIMAGE=$(find frontend/src-tauri/target/${{ matrix.rust_target }}/release/bundle/appimage -name "*.AppImage" | head -1)
APPIMAGE=$(realpath "$APPIMAGE")
WORK="$(mktemp -d)"; cd "$WORK"
"$APPIMAGE" --appimage-extract >/dev/null
ROOT="$WORK/squashfs-root"
ICON=$(readlink -f "$ROOT/.DirIcon" 2>/dev/null || true)
if [ -n "$ICON" ] && [ -f "$ICON" ] && case "$ICON" in "$ROOT"/*) true;; *) false;; esac; then
echo ".DirIcon already resolves inside the bundle — no repair needed"
exit 0
fi
# The real bytes are at the AppDir root (linuxdeploy put them there
# before mislinking). Ship a regular file: nothing left to dangle.
SRC=$(find "$ROOT" -maxdepth 1 -name "*.png" | head -1)
[ -n "$SRC" ] || SRC=$(find "$ROOT/usr/share/icons" -name "*.png" | head -1)
[ -n "$SRC" ] || { echo "no icon bytes found in bundle"; exit 1; }
rm -f "$ROOT/.DirIcon"
cp "$SRC" "$ROOT/.DirIcon"
# Pinned immutable release + checksum: this binary runs with the
# updater signing key and a release-write token in its environment,
# so a mutable 'continuous' asset is not acceptable supply chain.
AIT_URL="https://github.com/AppImage/appimagetool/releases/download/1.9.1/appimagetool-x86_64.AppImage"
AIT_SHA256="ed4ce84f0d9caff66f50bcca6ff6f35aae54ce8135408b3fa33abfc3cb384eb0"
curl -fsSL --retry 3 -o "$WORK/appimagetool" "$AIT_URL"
echo "$AIT_SHA256 $WORK/appimagetool" | sha256sum -c - || { echo "appimagetool checksum mismatch"; exit 1; }
chmod +x "$WORK/appimagetool"
# Same FUSE-less trick the build itself uses.
APPIMAGE_EXTRACT_AND_RUN=1 ARCH=x86_64 "$WORK/appimagetool" --no-appstream "$ROOT" "$APPIMAGE"
cd "$GITHUB_WORKSPACE/frontend"
bunx tauri signer sign "$APPIMAGE"
gh release upload "$TAG" "$APPIMAGE" "$APPIMAGE.sig" --clobber --repo "$GITHUB_REPOSITORY"
# latest.json is NOT patched here: every tauri-action leg re-uploads
# the shared manifest, so an in-leg patch races the other platforms —
# the repair-updater-manifest job below is the single final writer.
echo "repacked, re-signed, re-uploaded"
- name: Installer smoke (Linux)
if: runner.os == 'Linux'
timeout-minutes: 5
@@ -844,6 +900,50 @@ jobs:
# the tag (v0.3.20 shipped with only the Linux AppImage that way). `needs:
# [build]` guarantees the release already exists; `--clobber` makes a re-run
# idempotent. This can never create a second release.
# The Linux leg may repack + re-sign its AppImage (see the repair step in
# the build matrix); every tauri-action leg also re-uploads the SHARED
# latest.json, so patching the manifest inside any leg races the others.
# This job runs once after the whole matrix as the single final writer:
# it makes the manifest's linux signature agree with the .sig asset that
# actually shipped, and refuses to leave a mismatch behind.
repair-updater-manifest:
needs: [build, preview-gate]
runs-on: ubuntu-latest
timeout-minutes: 10
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
# Data, not shell source — same zizmor rule as the leg step.
TAG: ${{ (needs.preview-gate.outputs.is_preview == 'true') && 'preview' || github.ref_name }}
steps:
- name: Align latest.json's linux signature with the shipped .sig asset
shell: bash
run: |
set -euo pipefail
WORK="$(mktemp -d)"
HAS_MANIFEST=$(gh release view "$TAG" --repo "$GITHUB_REPOSITORY" --json assets --jq '[.assets[].name]|contains(["latest.json"])')
if [ "$HAS_MANIFEST" != "true" ]; then
echo "no latest.json on the release — nothing to align"; exit 0
fi
gh release download "$TAG" --pattern latest.json --output "$WORK/latest.json" --repo "$GITHUB_REPOSITORY"
# Same fail-closed rule as the manifest: absence is checked against
# the asset LIST; an actual download failure must fail the job, or
# the manifest keeps a signature nobody shipped.
HAS_SIG=$(gh release view "$TAG" --repo "$GITHUB_REPOSITORY" --json assets --jq '[.assets[].name|select(endswith(".AppImage.sig"))]|length > 0')
if [ "$HAS_SIG" != "true" ]; then
echo "no AppImage .sig asset on the release — nothing to align"; exit 0
fi
gh release download "$TAG" --pattern "*.AppImage.sig" --dir "$WORK" --repo "$GITHUB_REPOSITORY"
SIG_FILE=$(find "$WORK" -name "*.AppImage.sig" | head -1)
[ -n "$SIG_FILE" ] || { echo "sig asset listed but download produced nothing"; exit 1; }
NEW_SIG=$(cat "$SIG_FILE")
CHANGED=$(python3 -c 'import json,sys; p,sig=sys.argv[1],sys.argv[2]; d=json.load(open(p)); n=sum(1 for k,v in d.get("platforms",{}).items() if k.startswith("linux") and v.get("signature")!=sig and not v.update({"signature":sig})); json.dump(d,open(p,"w"),indent=2); print(n)' "$WORK/latest.json" "$NEW_SIG")
if [ "$CHANGED" -ge 1 ]; then
gh release upload "$TAG" "$WORK/latest.json" --clobber --repo "$GITHUB_REPOSITORY"
echo "aligned $CHANGED linux signature(s) with the shipped .sig"
else
echo "manifest already agrees with the shipped .sig — no write"
fi
uninstall-scripts:
needs: [build]
if: github.event_name == 'push' && startsWith(github.ref, 'refs/tags/v')
+61 -40
View File
@@ -10,52 +10,60 @@ the frozen-backend fallback mirror it for their toolchains.
**Highlights**
- A faster, cleaner Dub workspace for multilingual production (#1489)
- VoiceStudio now gives the app, desktop chrome, documentation, and package metadata one clear identity
- A local-first creative studio: voice cloning, design, dubbing, dictation, stories, audiobooks, and transcription without a subscription meter
- Reliability first: automatic cache repair, truthful hardware routing, safer sidecars, and actionable recovery instead of mystery failures
- Security boundaries now match the product: native file access stays native, untrusted network destinations fail closed, and public errors keep private diagnostics local
- RTX 40-series GPUs are used again instead of being sent to the CPU
- A warning before a slow generation, rather than after a five-minute wait
- The watermark can be turned off in Settings, as the docs always said
- Your other GPU can take the work now — send individual jobs to a second machine, opt-in
- More than one person can share one GPU machine, without shell access to it or taking turns
- A Model Catalogue workspace: every engine and model in one place, with the defaults set there
- Workspace tabs in the title bar, if you prefer them to the icon rail (#1412)
- macOS support now matches what the app actually delivers
- Linux AppImage: a blank white window on rolling distros (Mesa 26.1+) now starts normally
- Apple Silicon: transcription no longer needs a system ffmpeg, as the docs always said — thanks @gambletan! (#1436)
- A failed audiobook chapter says why, instead of turning red and saying nothing
- The backend now answers within a second of launch and narrates its startup step by step
- Reporting a bug from an outdated build now offers the latest release first
- The backend is only announced ready once it can actually serve, and crash-loop restarts now pace themselves
### Changed
- The backend binds its port immediately and reports startup progress live — `/health` answers 503-with-step and a new `/startup/progress` endpoint lists every step while PyTorch, API routes, and database migrations load in the background, so "starting at step X" is never mistakable for "dead"; the desktop splash narrates each step (#1550)
### Added
- The bug reporter notices when you're on an outdated build and offers the latest release before filing — with a "File anyway" escape hatch — and stamps a `Build status` line into every report so up-to-date reports are tellable from stale ones (#1547)
- Settings → Performance & Device gains a compute-device override (Auto / CUDA / ROCm / XPU / MPS / CPU, or `OMNIVOICE_DEVICE`) — pin the device when auto-detect picks wrong; only devices your machine actually has are offered (#1557)
### Docs
- The READMEs now lead with download buttons and a three-step first-clone walkthrough, and a new benchmarks page anchors measured per-engine/per-device numbers on the in-repo harness (#1555)
- Every engine now has its own guide — 21 new pages under docs/engines plus an index covering all 16 TTS and 11 ASR engines, linked from both READMEs (#1556)
### Fixed
- Hosted Studio no longer crashes when system information omits desktop-only RAM, CPU, or VRAM metrics
- The crash-isolated ASR sidecar and its download preflight now agree on which model to load — setting the shared faster-whisper model variable applies to both variants instead of the sidecar quietly using a different one (#1556)
- "Ready" now requires the deep health probe (a working database-backed route), not just the identity probe — a backend whose install broke underneath can no longer be announced up while every real request fails (#1548)
- Supervisor restarts after repeat crashes now back off (immediate, then 5s, then 15s) instead of respawning back-to-back, so a tight crash loop can't burn the whole restart budget in seconds (#1548)
- The guard that keeps transcription on the degrading ASR loader now scans the whole backend, not just the routers — a service that transcribes on a request's behalf skipped `ensure_loaded()` just as thoroughly. (#1519) — thanks @ahov520!
- The Linux app icon is no longer blank. Every AppImage since v0.4.2 shipped `.DirIcon` as an absolute symlink into the machine that built it (`/home/runner/work/…`), so the link dangled on every user's computer and file managers, app menus and desktop integration all drew nothing. The release build now verifies the icon resolves inside the bundle before publishing. (#1518)
- The Linux desktop entry no longer ships an empty `Categories=`, which `desktop-file-validate` rejects and menu builders skip. (#1518)
## [0.5.0] — 2026-08-13
### Added
**Highlights**
- The demo audio the app has always advertised now actually ships: previews for all seven voice-design presets, the three dictation replay clips, and the dubbing demo's source video plus four dubbed languages with subtitles. Every one of those was a dead link before — the tooling that renders them required macOS, so on Windows and Linux the files were never built. (#1517)
- Demo assets are rendered by VoiceStudio's own engine, so the tooling runs wherever the app does, and the demos are made by the thing they demonstrate. (#1517)
- The app is now **VoiceStudio** (previously OmniVoice-Studio) — one waveform-and-spark identity across the app, docs and installers. Your data folder, settings and Docker image paths stay put.
- **Model Catalogue** — engines and models in one workspace: every TTS, transcription and LLM engine with its device routing and install state, defaults picked there.
- Switch TTS, ASR and LLM engines from the status bar or any workspace — ready-only choices, memory status, environment-pin protection, `Ctrl/Cmd+E`. (#1530)
- Lend another machine's GPU with a join code and a QR scan — a Compute control in the status bar picks where jobs run, and several people can share one GPU box with revocable, certificate-pinned connections. (#1516, #1496)
- Server mode is locked down: admin actions require an API key (#1525), and the remote UI exchanges it for short-lived sessions that never sit in browser storage or WebSocket URLs (#1528) — thanks @bultodepapas!
- A faster, cleaner Dub workspace for multilingual production, with a production command bar and per-language cards. (#1489)
- The demo audio and video the app always advertised now actually ship, rendered by VoiceStudio's own engine. (#1517)
- Dictation works on Wayland now — the portal shortcut actually fires (#1490, #1526) — and the recording pill is back on every desktop.
- The Launchpad wears the project's signal-field waveform artwork over a quieter, borderless layout. (#1533)
- The catalogue reads as headroom, not breakage: available engines sort first, uninstalled ones say what they need (#1531), and the LLM row names the provider that actually answers (#1538).
- Gallery voices can be saved as local profiles — audio lands in your profile store with validated, content-addressed references. (#1542)
### Added
<img src="https://raw.githubusercontent.com/debpalash/VoiceStudio/main/docs/media/0.5.0/quick-switch.gif" alt="Switching TTS engines from the status bar" width="820" />
- A machine can now join a control plane from the app: Settings → System → Remote workers → **Lend this machine's GPU**, paste the join code, done — no environment variables and no restart. The address travels with the code, so the machine reconnects on its own afterwards. (#1516)
- Join codes and connection strings are shown as a **QR code** alongside the text, with a live expiry countdown — scan it from the other machine instead of retyping forty characters. (#1516)
- A **Compute** control in the status bar: pick local or a remote machine, turn remote workers on or off, and mint a join code without opening Settings. It appears only once you have opted in or enrolled a machine. (#1516)
- A worker waiting for approval can be approved from its row. The panel labelled that state before but offered no way out of it. (#1516)
- The demo audio the app has always advertised now actually ships: previews for all seven voice-design presets, the three dictation replay clips, and the dubbing demo's source video plus four dubbed languages with subtitles. Every one of those was a dead link before — the tooling that renders them required macOS, so on Windows and Linux the files were never built. (#1517)
- Demo assets are rendered by VoiceStudio's own engine, so the tooling runs wherever the app does, and the demos are made by the thing they demonstrate. (#1517)
| The Model Catalogue | The Voice Gallery |
| --- | --- |
| <img src="https://raw.githubusercontent.com/debpalash/VoiceStudio/main/docs/media/0.5.0/catalogue.png" alt="Model Catalogue — engines pane" width="420" /> | <img src="https://raw.githubusercontent.com/debpalash/VoiceStudio/main/docs/media/0.5.0/gallery-save.png" alt="Voice Gallery — save a voice as a profile" width="420" /> |
### Changed
- Gallery personas now preview through the local backend, retain their complete voice-design recipe, and open directly in Voice, Stories, or Audiobook. (#1542)
- Typing and large workspace edits no longer serialize and rewrite persisted documents on every input; writes are coalesced off the interaction path — thanks @bultodepapas! (#1541)
- Support amount choices now use every theme's shared card, accent and focus tokens. (#1530)
- Sponsoring, commercial licensing and getting in touch are one page now. They answered the same question between them and each used to live somewhere else, so they are three sections on a single scroll — the footer heart, the commercial-licence links and Contact all land on it, at the section you asked for. (#1522)
- Model Catalogue switches panes with tabs instead of a two-state toggle, and the Engine Compatibility Matrix's TTS / ASR / LLM switcher is now tabs too — arrow-key navigable, and each tab still shows the engine it would use. (#1522)
- Engines you can actually use sort to the top of the compatibility matrix, and an unavailable engine's name recedes instead of the whole row fading — the status badge and GPU chips that say *why* it is unavailable stay legible. (#1522)
- Remote workers reads as a device list: status dot, address, latency, a live task meter, resident models and last-seen per machine, with housekeeping actions revealed on hover and a three-step empty state. (#1516)
- The GPU picker and the new status-bar control paint their status dots and menu surfaces from themed tokens instead of fixed palette classes, so they stop showing Gruvbox colours on Midnight and Catppuccin. (#1516)
- Dictation shows the pill again: a capture puts a small always-on-top capsule near the bottom of the screen you are working on — listening, transcribing, the result, and any error — and takes it away when the session ends. It never takes focus, so the text still lands in the app you were typing into. On Wayland the compositor decides where it sits; everywhere else it is bottom-centred.
- Remote workers reads as a device list: status dot, address, latency, a live task meter, resident models and last-seen per machine, with housekeeping actions revealed on hover and a three-step empty state. (#1516)
- Engines and models moved out of Settings into a new Model Catalogue workspace, reachable from the icon rail (or the title-bar tabs); Settings → Engines and Settings → Models now point there, and Settings keeps the models directory and Hugging Face mirror.
- The Settings sidebar is keyboard-navigable: ⌘K / Ctrl+K jumps to the filter, ↑/↓ and Home/End move between categories, and Enter or ↓ from the filter drops into the list. Matching text in a filtered category name is highlighted, and group headers stay pinned while the list scrolls.
- The Launchpad has a quieter, more spacious look: borderless feature tiles that light up on hover or keyboard focus, plain-numeral counts, hairline section rules, and one shared page column for the hero, tiles, recent files and project lists.
@@ -75,6 +83,13 @@ the frozen-backend fallback mirror it for their toolchains.
### Added
- Gallery personas preview through the local backend, keep their full voice-design recipe, and open directly in Voice, Stories, or Audiobook — and can be saved as local profiles with validated audio references. (#1542)
- The demo audio the app has always advertised now actually ships: previews for all seven voice-design presets, the three dictation replay clips, and the dubbing demo's source video plus four dubbed languages with subtitles. Every one of those was a dead link before — the tooling that renders them required macOS, so on Windows and Linux the files were never built. (#1517)
- Demo assets are rendered by VoiceStudio's own engine, so the tooling runs wherever the app does, and the demos are made by the thing they demonstrate. (#1517)
- A machine can now join a control plane from the app: Settings → System → Remote workers → **Lend this machine's GPU**, paste the join code, done — no environment variables and no restart. The address travels with the code, so the machine reconnects on its own afterwards. (#1516)
- Join codes and connection strings are shown as a **QR code** alongside the text, with a live expiry countdown — scan it from the other machine instead of retyping forty characters. (#1516)
- A **Compute** control in the status bar: pick local or a remote machine, turn remote workers on or off, and mint a join code without opening Settings. It appears only once you have opted in or enrolled a machine. (#1516)
- A worker waiting for approval can be approved from its row. The panel labelled that state before but offered no way out of it. (#1516)
- **Model Catalogue** — a workspace of its own for engines and models: browse every TTS, transcription and LLM engine with its device routing and install state, pick the default for each, and install or remove model weights, all from one screen instead of two Settings categories.
- Remote GPU machines can now accept connections instead of dialling out, so several people can use the same box at once — each gets their own revocable connection string, with certificate-pinned TLS, a live list of who is connected, and a disconnect button. (#1496)
- Remote GPU model downloads now use the normal Models install flow and show per-worker progress. (#1478)
@@ -88,14 +103,22 @@ the frozen-backend fallback mirror it for their toolchains.
- Settings → Privacy now has an **Invisible watermark** toggle. On by default, available to everyone, and it only affects audio generated after the change. (#1308)
- A new opt-in crash-isolated TTS engine, so a native crash takes down the sidecar instead of the whole backend — thanks @paoloantinori! (#1292, #1298, #1304)
- **PocketTTS** (Kyutai), an opt-in CPU-only engine for fast, low-latency renders in six languages (en/fr/de/pt/it/es) with zero-shot cloning from a reference clip. Enable in Settings → Engines — thanks @paoloantinori! (#1306, #1328)
- A warning before a slow generation, rather than after a five-minute wait. (#1280)
### CI
### Docs
- The stdio wire protocol every engine sidecar speaks is now tested once across all nine of them, instead of against a single engine — a bug in any one sidecar's copy gets caught — thanks @paoloantinori! (#1408)
- Engine acceptance: new `docs/engine-acceptance.md` documents the job map, the bar a new engine must clear, and the out-of-tree path (#1306)
- macOS install notes and the README support table now state the real floor (#1268)
- Contact: the project X account is listed alongside Discord (#1313)
- `OMNIVOICE_ALLOWED_ORIGINS` is finally documented: a browser loading the UI from another machine's origin needs the backend's CORS allow-list, which neither server mode nor trusted networks touches — thanks @vanderlpp! (#1348)
### Fixed
- The Linux app icon is no longer blank: the AppImage shipped `.DirIcon` as a symlink into the machine that built it, so file managers and app menus drew nothing. (#1518)
- AMD/ROCm hosts no longer crash ASR with "CUDA driver version is insufficient": ROCm torch reports itself as CUDA, but whisperx/faster-whisper run on CTranslate2, which is NVIDIA-only — they now take the CPU path there, and auto-detect prefers pytorch-whisper, which genuinely uses the HIP GPU. (#1529)
- Crash reports now carry the crashed run's own stderr: the shared error log is append-only with per-run offsets, so a restart can no longer overwrite the dying process's final output with the replacement's healthy startup. (#1510)
- Wayland: a stale portal identity no longer kills the dictation shortcut for the whole session. The desktop entry the app writes for the GlobalShortcuts portal could point at a binary that has since moved (a `cargo clean`, a relocated AppImage) — GNOME then refuses the bind with "App info not found" and the hotkey silently dies. The entry is validated and rewritten at startup now. (#1526)
- The guard that keeps transcription on the degrading ASR loader now scans the whole backend, not just the routers — a service that transcribes on a request's behalf skipped `ensure_loaded()` just as thoroughly. (#1519) — thanks @ahov520!
- The Linux app icon is no longer blank. Every AppImage since v0.4.2 shipped `.DirIcon` as an absolute symlink into the machine that built it (`/home/runner/work/…`), so the link dangled on every user's computer and file managers, app menus and desktop integration all drew nothing. The release build now verifies the icon resolves inside the bundle before publishing. (#1518)
- The Linux desktop entry no longer ships an empty `Categories=`, which `desktop-file-validate` rejects and menu builders skip. (#1518)
- Wayland: the dictation shortcut now actually starts dictation. The desktop portal registered the key correctly — GNOME and KDE even showed it back — but every press was discarded while decoding the compositor's signal, so the hotkey did nothing on any Wayland session. (#1490)
- The first-run "Choose a comfortable UI size" screen no longer stutters while you sit there. Applying a scale resizes the window's own viewport, which the screen was reading back to re-pick a size — so it flipped between two sizes forever without anyone touching it. (#1514)
@@ -209,16 +232,14 @@ the frozen-backend fallback mirror it for their toolchains.
- Translation through LM Studio works. The built-in model name was the placeholder `local-model`, which LM Studio rejects because it serves whatever you have loaded — VoiceStudio now asks it, and a 404 from a local server names the models that ARE loaded instead of telling you to check a URL that was fine — thanks @biga73! (#1332)
- Generation that silently dropped the end of the input now says so. When an engine returns no audio for part of the text the result sounds clean and is simply short, so the only way to notice was to read along; the backend log now names the sentences that produced nothing. (#1330)
- Dubbing: a re-rendered line that quietly came back in a default voice instead of the cloned one now says why in the backend log — the clone clips are extracted per job and a saved dub outlives them, so regenerating after cleanup loses the reference with no error. (#1331)
### Docs
- Engine acceptance: new `docs/engine-acceptance.md` documents the job map, the bar a new engine must clear, and the out-of-tree path (#1306)
- macOS install notes and the README support table now state the real floor (#1268)
- Contact: the project X account is listed alongside Discord (#1313)
- `OMNIVOICE_ALLOWED_ORIGINS` is finally documented: a browser loading the UI from another machine's origin needs the backend's CORS allow-list, which neither server mode nor trusted networks touches — thanks @vanderlpp! (#1348)
- RTX 40-series GPUs are used again instead of being sent to the CPU. (#1289)
- Apple Silicon: transcription no longer needs a system ffmpeg, as the docs always said — thanks @gambletan! (#1436)
- A failed audiobook chapter says why, instead of turning red and saying nothing. (#1325)
### CI
- Windows CI falls back to a static ffmpeg build when the Chocolatey feed is down, instead of failing the run. (#1542)
- The stdio wire protocol every engine sidecar speaks is now tested once across all nine of them, instead of against a single engine — a bug in any one sidecar's copy gets caught — thanks @paoloantinori! (#1408)
- Windows smoke tests stopped silently passing a broken ffmpeg install, and every smoke leg is now budgeted for a cold dependency install. (#1290)
- Test suites no longer leak config paths or model-manager shutdown state into one another, which had been failing unrelated pull requests. (#1269)
- The nightly preview build stopped refusing to publish its own healthy updater manifest when the macOS legs finished a few minutes ahead of the slowest one — Preview-channel users were silently left without new builds.
+282 -553
View File
@@ -1,656 +1,385 @@
<div align="center">
<img src="docs/logo.png" alt="VoiceStudio Logo" width="120" height="120" />
<img src="docs/logo.png" alt="VoiceStudio logo" width="120" height="120" />
<h1>VoiceStudio</h1>
<p><sub><em>previously OmniVoice-Studio</em></sub></p>
<h3>Make voices. Tell stories. Keep the files. ♡</h3>
<p>Clone, design, dub, dictate, and build audiobooks in one open-source desktop studio.<br/><b>Local-first by default.</b> No subscription or usage meter. Optional online services stay opt-in.</p>
<p><sub>Previously OmniVoice-Studio</sub></p>
<h3>Local voice cloning, dubbing, dictation, and long-form audio.</h3>
<p>16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, and Linux</p>
<p><strong>Local-first.</strong> No account, API key, subscription, or usage meter for the core workflow.</p>
<p>
<a href="#quickstart">Quickstart</a> ·
<a href="#install">Install</a> ·
<a href="#features">Features</a> ·
<a href="#why-voicestudio">Why VoiceStudio</a> ·
<a href="#tts-engines">Engines</a> ·
<a href="#openai-api">API</a> ·
<a href="#sponsor--donate">Donate</a> ·
<a href="#contributing">Contributing</a> ·
<a href="https://voicestudio.sh">Website</a> ·
<a href="https://voicestudio.sh/docs">Docs</a> ·
<a href="https://status.voicestudio.sh">Status</a> ·
<a href="https://discord.gg/bzQavDfVV9">Discord</a> ·
<a href="https://x.com/idebpalash">X</a> ·
<a href="#comparison">Compare</a> ·
<a href="#requirements">Requirements</a> ·
<a href="#engines">Engines</a> ·
<a href="#architecture">Architecture</a> ·
<a href="#api">API</a> ·
<a href="#documentation">Docs</a> ·
<a href="README_CN.md"><strong>简体中文</strong></a>
</p>
<p>
<a href="https://github.com/debpalash/VoiceStudio/stargazers"><img src="https://img.shields.io/github/stars/debpalash/VoiceStudio?style=flat-square&color=f59e0b" alt="Stars" /></a>
<a href="https://github.com/debpalash/VoiceStudio/stargazers"><img src="https://img.shields.io/github/stars/debpalash/VoiceStudio?style=flat-square&color=f59e0b" alt="GitHub stars" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases"><img src="https://img.shields.io/github/downloads/debpalash/VoiceStudio/total?style=flat-square&color=8b5cf6&label=downloads" alt="Total downloads" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/github/v/release/debpalash/VoiceStudio?style=flat-square&color=10b981" alt="Release" /></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/license-AGPL--3.0-blue?style=flat-square" alt="License" /></a>
<a href="https://github.com/debpalash/VoiceStudio/issues"><img src="https://img.shields.io/github/issues/debpalash/VoiceStudio?style=flat-square&color=ef4444" alt="Issues" /></a>
<a href="https://discord.gg/bzQavDfVV9"><img src="https://img.shields.io/badge/Discord-Join_Community-5865F2?style=flat-square&logo=discord&logoColor=white" alt="Discord" /></a>
<a href="https://x.com/idebpalash"><img src="https://img.shields.io/badge/X-Follow_for_updates-000000?style=flat-square&logo=x&logoColor=white" alt="Follow on X" /></a>
<a href="https://ko-fi.com/debpalash"><img src="https://img.shields.io/badge/Ko--fi-Support_Us-FF5E5B?style=flat-square&logo=ko-fi&logoColor=white" alt="Ko-fi" /></a>
<a href="https://paypal.me/palashCoder"><img src="https://img.shields.io/badge/PayPal-Donate-00457C?style=flat-square&logo=paypal&logoColor=white" alt="PayPal" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/github/v/release/debpalash/VoiceStudio?style=flat-square&color=10b981" alt="Latest release" /></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/license-AGPL--3.0-blue?style=flat-square" alt="AGPL-3.0 license" /></a>
<a href="https://discord.gg/bzQavDfVV9"><img src="https://img.shields.io/badge/Discord-Community-5865F2?style=flat-square&logo=discord&logoColor=white" alt="Discord community" /></a>
</p>
<p>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/⬇_Download-macOS_·_Windows_·_Linux-10b981?style=for-the-badge" alt="Download the latest release" /></a>
</p>
<p>
<a href="https://trendshift.io/repositories/28176?utm_source=trendshift-badge&utm_medium=badge&utm_campaign=badge-trendshift-28176" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/trendshift/repositories/28176/daily?language=Python" alt="debpalash%2FVoiceStudio | Trendshift" width="250" height="55"/></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Download-macOS_·_Windows_·_Linux-10b981?style=for-the-badge" alt="Download VoiceStudio" /></a>
</p>
</div>
<br/>
<div align="center">
<img src="docs/screenshot-launchpad.png" alt="VoiceStudio — Launchpad" width="100%"/>
<img src="docs/media/0.5.0/quick-switch.gif" alt="Switching TTS engines from the VoiceStudio status bar" width="100%" />
</div>
> **Your voice is personal. Your studio should feel personal too.** VoiceStudio keeps its core workflow on your hardware: clone, design, dub, dictate, and publish in 646 languages without a subscription or usage meter. Network-backed engines and services are optional, visible choices—not hidden requirements.
> [!WARNING]
> **Active beta.** Things may break between releases — for the newest fixes, run from source. Bug reports and PRs are very welcome: [open an issue](https://github.com/debpalash/VoiceStudio/issues) or [join Discord](https://discord.gg/bzQavDfVV9).
> **Active beta.** Use the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) for stable work or `main` for current fixes. Report problems through [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues).
## At a glance
| | VoiceStudio |
|---|---|
| **Workflows** | Voice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation |
| **Language catalogue** | 646 TTS languages; actual coverage and quality depend on the selected engine |
| **Engines** | 16 TTS · 11 ASR · switch in Model Catalogue or with <kbd>Ctrl</kbd>/<kbd>Cmd</kbd>+<kbd>E</kbd> |
| **Platforms** | macOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+ |
| **Compute** | CUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers |
| **Interfaces** | Desktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server |
| **Storage** | Voices, projects, settings, and outputs stay on the machine by default |
| **License** | AGPL-3.0; optional engines keep their own model licenses |
<a id="install"></a>
## Install
| Platform | Package | Guide |
|---|---|---|
| macOS 13.3+ | DMG, Apple Silicon | [Install on macOS](docs/install/macos.md) |
| Windows 10/11 | MSI, x64 | [Install on Windows](docs/install/windows.md) |
| Linux | AppImage, x86_64 with glibc 2.39+ | [Install on Linux](docs/install/linux.md) |
| Docker | CUDA, ROCm, or CPU | [Run with Docker](docs/install/docker.md) |
Download packages from the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest). First launch creates a managed Python environment and downloads the default model. Later launches reuse both.
> [!NOTE]
> On macOS, first launch needs a one-time right-click → **Open** approval. Intel Macs cannot run the local Python backend; use a [remote backend](docs/install/macos.md) instead.
### First voice
1. Launch VoiceStudio and open **Voice Cloning**.
2. Add a clean voice sample. Three seconds works; 515 seconds usually gives a better prompt.
3. Enter text, choose a language, then select **Generate**.
### Run from source
Install the [development prerequisites](.github/CONTRIBUTING.md#development-setup), then:
```bash
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop
```
Use `bun run dev` for the browser UI. See [Contributing](.github/CONTRIBUTING.md) for services, tests, and platform packages.
### If setup fails
- Run **Settings → About → Run self-check** or `uv run python backend/main.py --diagnose --deep`.
- Check [install troubleshooting](docs/install/troubleshooting.md).
- Save a scrubbed diagnostic bundle from the app when opening an issue.
- For slow generation, compare [measured benchmarks](docs/benchmarks.md) and [performance settings](docs/performance.md).
<a id="features"></a>
## Features
## Features
Three flagships, five more headliners, and a dozen under the fold.
| Area | Included |
|---|---|
| **Voice Cloning** | Zero-shot synthesis from a short reference clip |
| **Voice Design** | Create a voice from age, accent, pitch, style, and delivery instructions |
| **Video Dubbing** | Transcribe, translate, preserve speakers, synthesize, and export video |
| **Stories and audiobooks** | Multi-voice scripts · EPUB/PDF import · chapter rendering · `.m4b` export |
| **Dictation Widget** | System-wide shortcut, live transcription, optional local-LLM cleanup |
| **Vocal Isolation** | Demucs speech/background separation |
| **Speaker Diarization** | Pyannote and WhisperX speaker assignment |
| **Batch Queue** | Queue large sets of audio and video jobs with per-job progress |
| **Model Catalogue** | Install, remove, select, and route TTS, ASR, and LLM models |
| **Remote Model Downloads** | Install models on enrolled remote workers with live progress |
| **GPU Auto-Detect** | CUDA, MPS, ROCm, and CPU routing with per-engine checks |
| **AI Watermark** | AudioSeal embedding and detection |
| **MCP Server** | Synthesis and transcription tools for MCP clients |
| **Diagnostics** | Self-checks, error journal, logs, and scrubbed support bundles |
| **Local-first** | Core creation stays local; network-backed features are explicit opt-ins |
| **Extensible** | Registry-based TTS, ASR, and plugin interfaces |
<table>
<tr>
<td width="33%"><img src="docs/features/clone.png" alt="Voice Cloning" width="100%"/></td>
<td width="33%"><img src="docs/features/design.png" alt="Voice Design" width="100%"/></td>
<td width="33%"><img src="docs/features/dub.png" alt="Video Dubbing" width="100%"/></td>
<td width="50%"><img src="docs/media/0.5.0/catalogue.png" alt="VoiceStudio Model Catalogue" width="100%" /></td>
<td width="50%"><img src="docs/media/0.5.0/gallery-save.png" alt="Saving a gallery voice as a local profile" width="100%" /></td>
</tr>
<tr>
<td align="center">🎙️ <b>Voice Cloning</b><br/><sub>3-sec clip → any voice · 646 languages · zero-shot</sub></td>
<td align="center">🎨 <b>Voice Design</b><br/><sub>Describe it — gender, age, accent, emotion</sub></td>
<td align="center">🎬 <b>Video Dubbing</b><br/><sub>Transcribe → translate → re-voice → MP4</sub></td>
<td align="center"><sub>Model Catalogue: engine, device, and install state</sub></td>
<td align="center"><sub>Gallery: save a shared voice as a local profile</sub></td>
</tr>
</table>
<table>
<tr>
<td align="center" width="20%">📖<br/><b>Audiobook</b><br/><sub>EPUB/PDF → .m4b, multi-voice cast</sub></td>
<td align="center" width="20%">🎭<br/><b>Stories</b><br/><sub>Multi-voice script editor</sub></td>
<td align="center" width="20%">⌨️<br/><b>Dictation Widget</b><br/><sub><kbd>⌘⇧Space</kbd> in any app</sub></td>
<td align="center" width="20%">🔐<br/><b>Local-first</b><br/><sub>Core creation stays on your machine</sub></td>
<td align="center" width="20%">🤖<br/><b>MCP Server</b><br/><sub>Use from Claude, Cursor, …</sub></td>
</tr>
</table>
<a id="comparison"></a>
<details>
<summary><b>…and 12 more</b> — isolation, diarization, batch, watermarking, diagnostics, and friends</summary>
## Comparison
<br/>
VoiceStudio trades managed cloud compute for local control. This is the practical difference:
- 🔊 **Vocal Isolation** — Demucs-powered: splits speech from music and keeps the background bed.
- 👥 **Speaker Diarization** — Pyannote + WhisperX auto-identify who said what.
- 📦 **Batch Queue** — drop 50 videos, walk away; per-job progress bars.
- 🛡️ **AI Watermark** — AudioSeal (Meta): invisible, survives compression.
- 🔬 **Diagnostics** — self-check suite, error journal, scrubbed diagnostic bundles.
- ⚡ **GPU Auto-Detect** — CUDA · MPS · ROCm (Linux, opt-in) · CPU; ≤8 GB VRAM auto-offloads.
- 📥 **Remote Model Downloads** — install pinned model weights on the selected worker with live progress.
- 🧭 **Engine routing** — preflight GPU check per engine; no silent CPU fallback.
- 📚 **Model Catalogue** — one workspace listing every TTS/ASR/LLM engine and model: set the defaults, install or remove weights.
- 🧩 **Extensible** — subclass `TTSBackend`, add any engine in ~50 lines.
- 🎒 **Portable personas** — export voices as `.ovsvoice` bundles: identity + watermark.
- ♾️ **Unlimited TTS** — sentence-chunked generation, no length cap, streaming via WebSocket.
- 🌐 **Remote backend** — point the UI at a remote server; Tailscale-friendly, bearer auth.
- 🧠 **Dictation + LLM** — local-LLM cleanup of transcripts, optional echo cancellation.
</details>
---
<a id="quickstart"></a>
## ⚡ Quickstart
<div align="center">
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/macOS-DMG_(Apple_Silicon)-000?style=for-the-badge&logo=apple&logoColor=white" alt="Download macOS DMG" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Windows-MSI_(x64)-0078D4?style=for-the-badge&logo=windows&logoColor=white" alt="Download Windows MSI" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Linux-AppImage_(x64)-FCC624?style=for-the-badge&logo=linux&logoColor=black" alt="Download Linux AppImage" /></a>
<br/>
<sub><b>macOS:</b> first launch needs a one-time approval — right-click → <b>Open</b> (or System Settings → Privacy &amp; Security → <b>"Open Anyway"</b> on macOS 15). No Terminal needed. <a href="docs/install/macos.md#gatekeeper-quarantine">Why?</a> · <b>Intel Macs:</b> local backend unsupported (<a href="https://github.com/debpalash/VoiceStudio/issues/889">#889</a>) — <a href="docs/install/macos.md">details</a>.</sub>
</div>
**Install guide:** [🍎 macOS](docs/install/macos.md) · [🪟 Windows](docs/install/windows.md) · [🐧 Linux](docs/install/linux.md) · [🐳 Docker](docs/install/docker.md)
<details>
<summary><b>🧰 Troubleshooting · slow generation · HF tokens · restricted networks</b></summary>
<br/>
- **Something broke?** Run the self-check — **Settings → About → "Run self-check"** (or `uv run python backend/main.py --diagnose --deep`) — then the [top 10 install errors](docs/install/troubleshooting.md). **"Save diagnostic bundle"** packages scrubbed logs for a bug report.
- **Feels slow?** [docs/performance.md](docs/performance.md) — where the time goes and how to tune it.
- **Want breaths, laughter, emotion?** [docs/expressive-speech.md](docs/expressive-speech.md) — what each engine can do today.
- **HF tokens · diarization · download speed / mirrors:** [tokens](docs/setup/huggingface-token.md) · [diarization](docs/features/diarization.md) · [downloads](docs/downloading-models.md).
- **Coming from [Real-Time-Voice-Cloning](https://github.com/CorentinJ/Real-Time-Voice-Cloning)?** [Migration guide](docs/migration/real-time-voice-cloning.md).
</details>
---
<a id="why-voicestudio"></a>
## ⚖️ Why VoiceStudio
Cloud voice tools are convenient, but they put your workflow behind an account, a meter, and somebody else's infrastructure. VoiceStudio gives you a capable studio that runs on your hardware, with optional integrations when you choose them.
| | **ElevenLabs** | **VoiceStudio** |
| | **VoiceStudio** | **Typical hosted voice service** |
|---|---|---|
| **Pricing** | Subscription and usage limits | Free & open-source (AGPL-3.0) · [Commercial license](#license) for proprietary use |
| **Voice Cloning** | ✅ 3s clip | ✅ 3s clip, zero-shot |
| **Voice Design** | ✅ Gender, age | ✅ Gender, age, accent, pitch, style, dialect |
| **Audiobook / Stories** | ❌ | ✅ Full audiobook editor + multi-voice stories (EPUB/PDF import, .m4b export) |
| **Languages** | Plan/model dependent | **646** |
| **Video Dubbing** | ✅ Cloud-only | ✅ Fully local |
| **Data Privacy** | Audio is processed remotely | Core workflow runs locally; online services are explicit opt-ins |
| **API Keys** | Account required | Not needed for the local workflow |
| **GPU Support** | N/A (cloud) | CUDA · Apple Silicon · ROCm (Linux) · CPU |
| **Desktop App** | ❌ | ✅ macOS · Windows · Linux |
| **TTS Engines** | 1 | **14** — [full matrix](#tts-engines) |
| **ASR Engines** | 1 | **11** — [full lineup](#asr-engines) |
| **MCP Server** | ❌ | ✅ Use from Claude, Cursor, any MCP client |
| **Self-check** | ❌ | ✅ Diagnostics suite, error journal, scrubbed debug bundles |
| **Customizable** | ❌ Closed | ✅ Fork it, extend it, ship it |
| **Best fit** | Private, offline, self-hosted, or high-volume work | Fast setup without local model management |
| **Data path** | Local by default; remote features are opt-in | Audio and text are processed by the provider |
| **Cost model** | Free software; you supply the hardware | Subscription, credits, or metered API use |
| **Setup** | Install the app and model weights | Create an account and use the web app or API |
| **Performance** | Depends on your engine and hardware | Provider manages compute and scaling |
| **Offline use** | Yes, after required models are installed | Usually requires a network connection |
| **Customization** | Source, engines, models, API, and routing are open | Limited to provider options |
| **Maintenance** | You manage updates, disk, and compute | Provider manages infrastructure |
Professional-grade voice AI, minus the subscription and the cloud.
<a id="requirements"></a>
<div align="center">
<br/>
<b>Convinced? Come build with us.</b><br/>
<a href="https://discord.gg/bzQavDfVV9"><img src="https://img.shields.io/badge/Join_Discord-5865F2?style=for-the-badge&logo=discord&logoColor=white" alt="Join Discord" /></a>
<br/><br/>
</div>
## Requirements
---
## 🖥️ System Requirements
Requirements vary by engine. These values cover the default local workflow.
| | **Minimum** | **Recommended** |
|---|---|---|
| **OS** | Windows 10, macOS 13.3+ (Apple Silicon), Ubuntu 24.04+ (glibc 2.39+) | Any modern 64-bit OS |
| **OS** | Windows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+ | Current supported OS release |
| **RAM** | 8 GB | 16 GB+ |
| **VRAM (GPU)** | 4 GB (auto-offloads TTS to CPU) | 8 GB+ (NVIDIA RTX 3060+) |
| **Disk** | 10 GB free (models + cache) | 20 GB+ SSD |
| **Python** | 3.10+ (managed by `uv`) | 3.113.12 |
| **GPU** | Optional — CPU works | NVIDIA CUDA · Apple Silicon MPS · AMD ROCm (Linux only) |
| **Disk** | 10 GB free | 20 GB+ SSD |
| **GPU** | Optional; CPU mode is supported | NVIDIA CUDA or Apple Silicon |
| **VRAM** | 4 GB when using a GPU | 8 GB+; large optional engines need more |
| **Python from source** | 3.11+ | 3.113.12 |
> [!NOTE]
> **A GPU is optional** — the whole pipeline runs on CPU (just slower), and on ≤8 GB VRAM, TTS auto-offloads to CPU. Caveats: **AMD ROCm** is Linux-only + opt-in ([Linux](docs/install/linux.md#amd-gpu-rocm)) — Windows AMD/Ryzen AI is CPU-only ([Windows](docs/install/windows.md#gpu-support)); **macOS Intel** can't run the local backend, so point it at a remote one ([#889](https://github.com/debpalash/VoiceStudio/issues/889) · [macOS](docs/install/macos.md)).
ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See [performance](docs/performance.md), [benchmarks](docs/benchmarks.md), and [engine disk usage](docs/engines/disk-usage.md).
<a id="engines"></a>
## Engines
Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: [docs/engines](docs/engines/README.md).
<a id="tts-engines"></a>
### 🗣️ TTS Engines
**14 engines, one picker.** VoiceStudio (default, 600+ languages) is always available; seven more are opt-in and auto-detected (CosyVoice 3, GPT-SoVITS, VoxCPM2, MOSS-TTS-Nano, KittenTTS, MLX-Audio, Sherpa-ONNX), plus six lazy-installed heavyweights (IndexTTS 2.5, OmniVoice GGUF, Supertonic 3, MOSS-TTS-v1.5, dots.tts, Confucius4-TTS). Switch in **Settings → TTS Engine**; the choice applies everywhere synthesis happens.
<details>
<summary><b>📊 The full matrix</b> — 14 engines × platform × clone/instruct × license</summary>
<br/>
### Text to speech
| Engine | Languages | Clone | Instruct | Linux | macOS ARM | Windows | License |
|--------|:---------:|:-----:|:--------:|:-----:|:---------:|:-------:|:-------:|
| **VoiceStudio** (default, powered by k2-fsa/OmniVoice) | 600+ | ✅ | ✅ | ✅ CUDA/CPU | MPS | CUDA/CPU | Built-in |
| **CosyVoice 3** | 9 + 18 dialects | ✅ | ✅ | CUDA/CPU | ✅ MPS | ✅ CUDA/CPU | Apache-2.0 |
| **GPT-SoVITS** | 5 | | — | CUDA/CPU | — | CUDA/CPU | MIT |
| **VoxCPM2** | 30 | ✅ | ✅ | ✅ CUDA/CPU | MPS | CUDA/CPU | Apache-2.0 |
| **MOSS-TTS-Nano** | 20 | | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| **KittenTTS** | English | — | — | CPU | CPU | CPU | MIT |
| **MLX-Audio** (Kokoro, Qwen3-TTS, CSM, Dia, …) | Multi | Varies | Varies | | ✅ Native | | Varies |
| **Sherpa-ONNX** | 20+ | — | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| **IndexTTS 2.5** ⚡ | ZH · EN · JA · ES · AR | | — | CUDA | — | CUDA | Bilibili model license¹ |
| **OmniVoice GGUF** ⚡ | 600+ | ✅ | ✅ | ✅ CPU | CPU | CPU | Built-in |
| **Supertonic 3** ⚡ | 31 | — | — | ✅ CPU | ✅ CPU | ✅ CPU | OpenRAIL-M |
| **MOSS-TTS-v1.5** ⚡ (8B) | 31 | ✅ | — | ✅ CUDA/CPU | CPU | ✅ CUDA/CPU | Apache-2.0 |
| **dots.tts** ⚡ (2B) | 24 | | — | ✅ CUDA/CPU | CPU | ❌ | Apache-2.0 |
| **Confucius4-TTS** ⚡ | 14 | | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|---|:---:|:---:|:---:|:---:|:---:|:---:|---|
| **VoiceStudio** (default, powered by k2-fsa/OmniVoice) | 600+ | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0](LICENSE-NOTICE.md) model |
| **CosyVoice 3** | 9 + 18 dialects | Yes | Yes | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| **GPT-SoVITS** | 5 | Yes | — | CUDA/CPU | — | CUDA/CPU | MIT |
| **VoxCPM2** | 30 | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | Apache-2.0 |
| **MOSS-TTS-Nano** | 20 | Yes | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| **KittenTTS** | English | — | — | CPU | CPU | CPU | MIT |
| **MLX-Audio** | Model-dependent | Varies | Varies | | MLX | | Varies |
| **Sherpa-ONNX** | 20+ | — | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| **IndexTTS 2.5** ⚡ | ZH · EN · JA · ES · AR | Yes | — | CUDA/CPU | CPU | CUDA/CPU | Bilibili model license¹ |
| **OmniVoice GGUF** ⚡ | 600+ | Yes | Yes | CUDA/CPU | MPS/CPU | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0](LICENSE-NOTICE.md) model |
| **OmniVoice (subprocess)** ⚡ | 600+ | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0](LICENSE-NOTICE.md) model |
| **PocketTTS** ⚡ | EN · FR · DE · PT · IT · ES | Yes | — | CPU | CPU | CPU | CC-BY-4.0, gated² |
| **Supertonic 3** ⚡ | 31 | | — | CPU | CPU | CPU | OpenRAIL-M |
| **MOSS-TTS-v1.5** ⚡ | 31 | Yes | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
| **dots.tts** ⚡ | 24 | Yes | — | CUDA/CPU | CPU | — | Apache-2.0 |
| **Confucius4-TTS** ⚡ | 14 | Yes | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million
monthly active users or RMB 1 billion in annual revenue. Review its
[model license](https://huggingface.co/IndexTeam/IndexTTS-2.5/blob/main/LICENSE)
before enabling the optional sidecar.
Installed or registered on demand.
GPT-SoVITS connects to `http://127.0.0.1:9880` by default. To use a server on
another machine, set `OMNIVOICE_GPTSOVITS_URL` to its credential-free
`http://` or `https://` origin and add that machine's CIDR to
`OMNIVOICE_TRUSTED_NETWORKS`; redirects and untrusted destinations are rejected.
¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the [model license](https://huggingface.co/IndexTeam/IndexTTS-2.5/blob/main/LICENSE).
> **CUDA** = GPU-accelerated · **MPS** = Apple Silicon Metal · **CPU** = runs everywhere, slower for large models · KittenTTS and MOSS-TTS-Nano run realtime on CPU · MLX-Audio is Apple Silicon only · ⚡ = lazy-registered (installed on first use)
>
> **Clone** matters beyond single-clip generation: Video Dubbing (and any Batch job with a pinned voice) needs reference-audio cloning to preserve speaker identity, so picking a Clone-less engine (KittenTTS, Sherpa-ONNX, Supertonic 3) as the active engine fails those jobs up front with an actionable message instead of silently falling back to VoiceStudio.
>
> **MOSS-TTS-v1.5** (8B, ~16 GB), **dots.tts** (2B, ~9 GB), and **Confucius4-TTS** are heavyweight opt-ins that run in their own isolated venv from a local clone. None claims Apple-Silicon MPS (CPU on Macs); dots.tts has no Windows path; Confucius4 wants CUDA (CPU works, ~17× realtime). Details: [MOSS-TTS-v1.5](docs/engines/moss-tts-v15.md) · [dots.tts](docs/engines/dots-tts.md) · [Confucius4-TTS](docs/engines/confucius4-tts.md).
² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.
</details>
Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.
<a id="asr-engines"></a>
### 🎧 ASR Engines
### Speech to text
**11 engines** — they power dictation, video dubbing, and subtitles. **WhisperX** is the cross-platform default (~100 languages, word-level timing); the rest are opt-in and auto-detected. Switch in **Model Catalogue → Engines**. Ten run fully on-device; the eleventh (OpenAI-compatible) is an optional remote client for Qwen3-ASR or any compatible server.
| Engine | ID | Languages | Best fit |
|---|---|:---:|---|
| **WhisperX** (default) | `whisperx` | ~100 | Dubbing, subtitles, word-level timing |
| **Faster-Whisper** | `faster-whisper` | ~100 | General cross-platform transcription |
| **Faster-Whisper (isolated)** | `faster-whisper-isolated` | ~100 | Crash-isolated batch transcription |
| **MLX Whisper** | `mlx-whisper` | ~100 | Apple Silicon |
| **PyTorch Whisper** | `pytorch-whisper` | ~100 | CUDA, MPS, and CPU fallback |
| **Parakeet TDT** | `nemo-parakeet` | English + 25 EU | Fast CPU/CUDA transcription |
| **Parakeet TDT v3 (MLX)** | `parakeet-mlx` | 25 EU | Apple Silicon dictation and word timestamps |
| **Moonshine** | `moonshine` | English | Low-power, low-latency ONNX |
| **FunASR** | `funasr` | 50+ | VAD and inline diarization |
| **sherpa-onnx** (live dictation) | `sherpa-onnx-asr` | Model-dependent | Streaming CPU dictation |
| **OpenAI-compatible** ⚠️ remote | `openai-compat-asr` | Server-dependent | Qwen3-ASR or another compatible endpoint; audio leaves the machine |
<details>
<summary><b>📊 The full lineup</b> — 11 engines, what each is best at, and compute-type notes</summary>
WhisperX and Faster-Whisper retry with `int8` when efficient `float16` is unavailable. Pin `ASR_COMPUTE_TYPE=int8` or `float32` only if automatic selection still fails.
<br/>
<a id="architecture"></a>
| Engine | `OMNIVOICE_ASR_BACKEND` | Languages | Best for |
|--------|-------------------------|:---------:|----------|
| **WhisperX** (default) | `whisperx` | ~100 | Dubbing & subtitles — word-level timing via wav2vec2 forced alignment |
| **Faster-Whisper** | `faster-whisper` | ~100 | Fast transcription on Linux / macOS / Windows (CTranslate2) |
| **Faster-Whisper (isolated)** | `faster-whisper-isolated` | ~100 | Same as Faster-Whisper but crash-isolated in a subprocess — an ASR crash won't take down the app |
| **MLX Whisper** | `mlx-whisper` | ~100 | Native Apple Silicon speed (Apple MLX / Metal) |
| **PyTorch Whisper** | `pytorch-whisper` | ~100 | CUDA / CPU fallback via 🤗 Transformers (no cuDNN 8 needed) |
| **Parakeet TDT** | `nemo-parakeet` | English + 25 EU | SOTA accuracy at ~10× realtime even on CPU, auto language detection (NVIDIA NeMo, CUDA/CPU) |
| **Parakeet TDT v3 (MLX)** | `parakeet-mlx` | 25 EU | The Parakeet tier for Apple Silicon — TDT word timestamps, ~2 GB unified memory, dictation-grade speed on the GPU via MLX. Install the model from **Model Catalogue → Models** and dictation prefers it automatically when your system language is one of its 25 (European) languages; other languages (CJK, Arabic, …) keep the multilingual Whisper engine so dictation coverage never regresses. |
| **Moonshine** | `moonshine` | English | Edge / low-latency, ONNX |
| **FunASR** | `funasr` | 50+ | All-in-one multilingual — built-in VAD + inline speaker diarization (SenseVoice) |
| **sherpa-onnx** (live dictation) | `sherpa-onnx-asr` | 25 EU + 90+ | Live, faster-than-real-time dictation — small streaming/offline ONNX models (Parakeet TDT v3/v2, streaming Zipformer & Paraformer, Whisper Tiny), CPU, identical on macOS / Windows / Linux. Picked per-model in **Settings → Voice**. |
| **OpenAI-compatible** ⚠️ remote | `openai-compat-asr` | Server-dependent | A path to **Qwen3-ASR** today (self-hosted server, no transformers wait), any OpenAI-compatible transcription endpoint, or OpenAI's own API — no install, configure + test the connection in **Model Catalogue → Engines** (ASR tab). Audio leaves your machine to whatever server you point it at; see [docs/engines/openai-compatible-asr.md](docs/engines/openai-compatible-asr.md). |
## Architecture
> Whisper-family engines cover ~100 languages; **FunASR / SenseVoice** adds an all-in-one multilingual path with built-in voice-activity detection and inline speaker diarization. **sherpa-onnx** powers the live dictation model picker — you talk and text appears as you speak. Every engine runs on-device — no API keys, no cloud.
> If Dubbing needs an ASR model that is not installed yet, it offers the recommended download in place, shows its progress, and retries transcription on the same job when the model is ready.
> **GPU without efficient float16?** On older NVIDIA GPUs (Maxwell/Pascal, GTX 16xx) or after a CTranslate2/cuDNN mismatch, the CTranslate2 ASR engines (WhisperX, Faster-Whisper) can't run `float16` and VoiceStudio automatically retries on `int8` — no config needed. If transcription still fails, pin the compute type with the `ASR_COMPUTE_TYPE` env var (escape hatch): `ASR_COMPUTE_TYPE=int8` (or `float32` for CPU). Set it to `int8` and restart the backend.
</details>
---
## 🏗️ Architecture
A **Tauri v2** desktop shell (Rust) wraps a **React** UI and a bundled **Python/FastAPI** backend that runs as a local sidecar on `localhost:3900`. Nothing external — every layer is on your machine.
```
┌────────────────────────────────────────────────────────────────────┐
│ Tauri v2 shell — Rust │
│ window state · global dictation hotkey · system tray · │
│ signed auto-updater (stable/preview) · single-instance · │
│ first-run bootstrap (installs uv + Python venv) · blank guard │
├────────────────────────────────────────────────────────────────────┤
│ Frontend — React + Vite │
│ Studio · Dub · Stories · Audiobook · Gallery · Dictation · │
│ Batch · Diagnostics · MCP client — Zustand store · WS bus │
│ ▲ IPC / HTTP + WS │
├──────────────────────────┼─────────────────────────────────────────┤
│ Backend — FastAPI sidecar @ localhost:3900 │
│ 100+ REST endpoints · SSE + WebSocket streaming · │
│ SQLite + Alembic (omnivoice_data/) · OpenAI-compatible API │
├───────────┬───────────┬───────────┬───────────┬────────────────────┤
│ TTS ×14 │ ASR ×11 │ Demucs │ Pyannote │ AudioSeal │
│ clone / │ WhisperX │ vocal │ speaker │ watermark │
│ design │ +10 more │ isolation│ diariz. │ embed / detect │
├───────────┴───────────┴───────────┴───────────┴────────────────────┤
│ Engine routing — per-engine GPU preflight, no silent CPU fallback │
│ Hardware: CUDA · MPS · ROCm (Linux) · CPU (auto-detected) │
└────────────────────────────────────────────────────────────────────┘
```text
Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
```
- **Shell (Rust)** — native OS integration: the system-wide dictation hotkey, tray, signed auto-updater (stable + preview channels), single-instance lock, and the first-run bootstrap that installs `uv` and a Python 3.11 venv.
- **Frontend (React)** — every workspace tab over a Zustand store, with a WebSocket event bus that live-refreshes the UI when backend data changes.
- **Backend (FastAPI)** — the bundled Python sidecar: 100+ endpoints, SSE/WSS streaming, a SQLite DB migrated by Alembic, and the OpenAI-compatible API surface.
- **Engines** — 14 TTS + 11 ASR, plus Demucs (isolation), Pyannote (diarization), and AudioSeal (watermark), all behind routing that GPU-preflights each engine and refuses to silently fall back to CPU.
| Layer | Path | Responsibility |
|---|---|---|
| Desktop shell | `frontend/src-tauri/` | Window lifecycle, tray, shortcuts, updater, sidecar bootstrap |
| Frontend | `frontend/src/` | React UI, Zustand state, API and event clients, i18n |
| API | `backend/api/` | REST routes, schemas, auth boundaries, streaming |
| Core services | `backend/services/` | Generation, dubbing, audio processing, persistence |
| Engines | `backend/engines/` | Isolated and optional engine adapters |
| Worker system | `backend/worker/` | Authenticated remote compute and job transport |
| Data | `omnivoice_data/` | Projects, voices, settings, logs, and SQLite state |
| Delivery | `scripts/`, `deploy/`, `.github/workflows/` | Development, packaging, containers, releases, CI |
<a id="openai-api"></a>
### Network boundary
## 🔌 OpenAI-compatible API
- The desktop talks to a loopback-only backend on `localhost:3900`.
- Loopback API calls need no server key. Remote access requires a share PIN or API key.
- Remote workers and OpenAI-compatible ASR are opt-in. The UI identifies when audio leaves the machine.
- Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata—not text, audio, file names, or projects.
<div align="center">
<a id="api"></a>
**Drop-in replacement for OpenAI / ElevenLabs audio.** One line — no key, no code changes:
## OpenAI-compatible API
Point an OpenAI-compatible audio client at the local backend:
```diff
- base_url="https://api.openai.com/v1"
+ base_url="http://localhost:3900/v1"
```
</div>
Your existing scripts, agents, and OpenAI/ElevenLabs SDK calls now run **locally** on whatever engine you have active. What the cloud can't do: `voice` takes **your own cloned-voice profile IDs**, and `model` can pin a **specific engine** per request.
| Endpoint | What it does |
| Endpoint | Purpose |
|---|---|
| `POST /v1/audio/speech` | TTS — text in; `mp3` / `opus` / `aac` / `flac` / `wav` / `pcm` out. `model`: `tts-1`/`tts-1-hd` (active engine) or a specific one (`voxcpm2`, `cosyvoice`, `kittentts`, …). `voice`: a cloned profile ID, `default`, or an OpenAI name (`alloy`, …). `speed` supported. |
| `POST /v1/audio/transcriptions` | STT — audio file in; `json` / `text` / `verbose_json` / `srt` / `vtt` out (`verbose_json` adds word-level timings). `whisper-1` maps to your active ASR engine. |
| `GET /v1/audio/voices` | VoiceStudio extension — lists every voice profile and engine, so clients can discover your clones. |
**Speak with your own cloned voice** — list the IDs, then pass one as `voice`:
```sh
# 1 — find a cloned voice's profile ID
curl -s http://localhost:3900/v1/audio/voices | jq '.voices[] | select(.type=="profile") | {voice_id, name}'
# 2 — synthesize with it
curl http://localhost:3900/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"model":"tts-1","voice":"<profile-id>","input":"Made on my own hardware.","response_format":"wav"}' \
--output speech.wav
```
| `POST /v1/audio/speech` | TTS to `mp3`, `opus`, `aac`, `flac`, `wav`, or `pcm`; select a profile with `voice` and an engine with `model` |
| `POST /v1/audio/transcriptions` | STT to `json`, `text`, `verbose_json`, `srt`, or `vtt` |
| `GET /v1/audio/voices` | List local voice profiles and engines |
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:3900/v1", api_key="none") # any string — nothing checks it
# TTS with your cloned voice (or "alloy" / "default"; model= can pin a specific engine)
client = OpenAI(base_url="http://localhost:3900/v1", api_key="local")
with client.audio.speech.with_streaming_response.create(
model="tts-1", voice="<profile-id>", input="Made on my own hardware.") as r:
r.stream_to_file("speech.wav")
# STT
print(client.audio.transcriptions.create(model="whisper-1", file=open("clip.wav", "rb")).text)
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) as response:
response.stream_to_file("speech.wav")
```
Want the whole surface (100+ endpoints)? The full REST API reference is embedded in the app — **Settings → OpenAPI Reference** (Scalar-powered), or the `{}` button in the footer.
The full API reference is in **Settings → OpenAPI Reference**. For LAN, Tailscale, or proxy access, read [API authentication](docs/api-auth.md) before exposing the backend.
Calling the backend from **another machine** (LAN, Tailscale, behind a proxy)? It's loopback-only and unauthenticated by default; to reach it remotely you set a share PIN or an API key. [docs/api-auth.md](docs/api-auth.md) covers the exact headers, query params, `401`/`403`/`429` meanings, and the `OMNIVOICE_TRUSTED_NETWORKS` exemption.
### Agent skills
### 📓 Run on Google Colab
Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other [skills.sh](https://skills.sh)-compatible agents:
[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/VoiceStudio_Studio_Colab.ipynb)
No local GPU? The [official notebook](notebooks/VoiceStudio_Studio_Colab.ipynb) boots the full app — web UI included — on a free Colab T4, then walks the whole feature surface (TTS, cloning, design, transcription, dubbing, audiobook, watermarking, the OpenAI-compatible API) as a guided tour with inline playback. No tunnels, no API keys.
### 🤝 Agent Skills
Teach your coding agent to speak and listen through your local VoiceStudio — one command, works with **Claude Code, Codex, Cursor, Grok, Kimi, opencode**, and any [skills.sh](https://skills.sh)-compatible agent:
```sh
```bash
npx skills add debpalash/omnivoice-studio
```
Ships two [skills](https://skills.sh):
- `omnivoice`: synthesize speech and transcribe audio through local VoiceStudio.
- `oss-maintainer`: the repository's open-source maintenance workflow.
- **`omnivoice`** — generate speech (including your cloned voices) and transcribe audio from any agent, free and fully offline via your local install.
- **`oss-maintainer`** — the maintainer methodology this project is run with, for anyone running their own OSS project with an agent.
### Google Colab
---
[![Open in Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb)
## 🗺️ Roadmap
The [notebook](notebooks/OmniVoice_Studio_Colab.ipynb) runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.
### 🔜 Up Next
<a id="documentation"></a>
- 🎬 **Lip-sync v2** — visual speech timing with wav2lip
- 🌐 **Hosted Demo** — try VoiceStudio without installing anything
- 🔌 **Plugin Marketplace** — community-contributed TTS engines and effects
- 🎵 **Real-time Voice Changer** — live microphone transformation during calls
## Documentation
| Need | Read |
|---|---|
| Install | [macOS](docs/install/macos.md) · [Windows](docs/install/windows.md) · [Linux](docs/install/linux.md) · [Docker](docs/install/docker.md) |
| Fix setup | [Troubleshooting](docs/install/troubleshooting.md) · [model downloads](docs/downloading-models.md) · [Hugging Face token](docs/setup/huggingface-token.md) |
| Choose an engine | [Engine guides](docs/engines/README.md) · [benchmarks](docs/benchmarks.md) · [expressive speech](docs/expressive-speech.md) |
| Tune hardware | [Performance](docs/performance.md) · [remote workers](docs/remote-workers.md) |
| Build integrations | [API auth](docs/api-auth.md) · [MCP](docs/mcp.md) · [examples](examples/README.md) |
| Build VoiceStudio | [Contributing](.github/CONTRIBUTING.md) · [engine acceptance](docs/engine-acceptance.md) |
| Track changes | [Changelog](CHANGELOG.md) · [roadmap](docs/ROADMAP.md) · [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) |
| Remove everything | [Uninstall guide](docs/install/uninstall.md) |
## FAQ
<details>
<summary><b>✅ Everything shipped so far</b> — the receipts, by category</summary>
<br/>
| Category | Features |
|----------|----------|
| **Longform** | Audiobook editor (text/EPUB/PDF → chaptered .m4b) with multi-voice cast, expressive controls, live per-chapter progress + Stop, and a one-click sample; Stories multi-voice editor, two-pass loudnorm mastering, crash-resume for interrupted renders, pronunciation control + SSML-lite prosody |
| **Dubbing** | Full pipeline (transcribe→translate→synthesize→mux), scene-aware splitting, lip-sync scoring, streaming TTS, per-speaker voice assignment, Smart Fit timing + second-pass QC, paste-in translations from any external tool, dedicated Dub home |
| **Voice** | Zero-shot cloning, voice design, A/B comparison, voice preview widget, gallery with favorites/tags (its voices selectable in every picker — Studio, Audiobook, Stories, Dubbing), portable persona bundles (`.ovsvoice`), voice console workspace |
| **Audio** | Demucs vocal isolation, per-segment gain, selective track export, stem/SRT/VTT/MP3 export, unlimited-length TTS via sentence-chunked generation |
| **Multi-Lang** | Translate All preserves the primary language plus every extra language chip; Generate renders and exports one retained track per language with sequential GPU execution |
| **Diarization** | Pyannote ML diarization, auto speaker clone extraction, per-speaker voice assignment |
| **ASR** | 11 engines (WhisperX, Faster-Whisper, isolated Faster-Whisper, MLX Whisper, PyTorch Whisper, Parakeet TDT, Parakeet TDT v3 MLX, Moonshine, FunASR/SenseVoice, sherpa-onnx live dictation, OpenAI-compatible remote), crash-isolated subprocess backend |
| **TTS** | 14 engines (VoiceStudio, CosyVoice 3, GPT-SoVITS, VoxCPM2, MOSS-TTS-Nano, KittenTTS, MLX-Audio, Sherpa-ONNX, + lazy: IndexTTS 2.5, OmniVoice GGUF, Supertonic 3, MOSS-TTS-v1.5, dots.tts, Confucius4-TTS), engine routing with GPU preflight |
| **Infra** | Docker deployment, CUDA/MPS/ROCm auto-detect, cuDNN 8 compat, VRAM-aware model offloading, engine routing (no silent CPU fallback), diagnostics suite & error journal, restricted-network mirror support |
| **AI Provenance** | AudioSeal invisible watermarking (SynthID-like), video logo overlay, watermark detection API |
| **UX** | Undo/redo, keyboard shortcuts, drag-and-drop, session persistence, screen-sized first-run UI scaling, and native WebKitGTK scaling |
| **Real-time Events** | WebSocket event bus — instant sidebar refresh on data mutations, exponential backoff reconnect |
| **State Management** | Zustand store migration — `uiSlice`, `pillSlice`, `dubSlice`, `generateSlice`, `prefsSlice`, `glossarySlice` |
| **Desktop** | Cross-platform Tauri installers (macOS DMG — Apple Silicon; Intel unsupported for the local backend, #889 — Windows MSI, Linux deb/AppImage), auto-update infrastructure, single-instance enforcement, close-to-tray, macOS Gatekeeper fix |
| **Dictation** | Global system-wide hotkey (`⌘+⇧+Space`), frameless floating widget, streaming ASR via WebSocket, auto-paste, customizable hotkey, local-LLM transcript refinement |
| **Batch Pipeline** | Full batch TTS: extract → transcribe → translate → generate → mix → export, with live progress tracking |
| **MCP Server** | VoiceStudio as a local TTS/STT provider for Claude, Cursor, and any MCP client |
| **Remote Backend** | Point the desktop UI at a remote backend URL with bearer auth (Tailscale-documented) |
| **Reliability** | Stall watchdog on bootstrap splash, per-engine GPU compatibility matrix, actionable errors for non-executable engine binaries, setuptools auto-repair |
<summary><strong>Does it work on Apple Silicon and Intel Macs?</strong></summary>
Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See [macOS installation](docs/install/macos.md).
</details>
---
<details>
<summary><strong>How much VRAM do I need?</strong></summary>
<a id="sponsor--donate"></a>
A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 1216 GB or more. Check the [benchmarks](docs/benchmarks.md) and engine guide.
</details>
## 💜 Sponsor / Donate
<details>
<summary><strong>Why does a longer reference clip not always improve the clone?</strong></summary>
One developer, real AI-agent bills. If VoiceStudio is useful to you, chipping in keeps development full-time — every dollar goes straight to the bills.
Cloning is zero-shot: the clip is a prompt, not training data. Use 515 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see [data preparation](docs/data_preparation.md) and [training](docs/training.md).
</details>
<details>
<summary><strong>Can I use generated audio commercially?</strong></summary>
Yes under VoiceStudio's AGPL-3.0 terms. Optional engines and model weights may use different licenses; review the selected engine's license before commercial use.
</details>
<details>
<summary><strong>Does VoiceStudio collect data?</strong></summary>
Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at **Settings → Privacy**.
</details>
<details>
<summary><strong>How do I remove VoiceStudio and its data?</strong></summary>
Use `scripts/uninstall.sh` on macOS/Linux or `scripts\uninstall.ps1` on Windows. Both show a dry run before deletion. See the [uninstall guide](docs/install/uninstall.md) for every path.
</details>
## Community and contributing
- [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues) for reproducible bugs and feature requests.
- [Discord](https://discord.gg/bzQavDfVV9) for setup help and project discussion.
- [Good first issues](https://github.com/debpalash/VoiceStudio/labels/good%20first%20issue) for a scoped starting point.
- [Contributing guide](.github/CONTRIBUTING.md) for setup, tests, and pull requests.
## Support development
VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.
[Ko-fi](https://ko-fi.com/debpalash) · [PayPal](https://paypal.me/palashCoder) · [Sponsorship details](SPONSORS.md)
## License
VoiceStudio is licensed under [AGPL-3.0](LICENSE). You may run it, modify it, use it internally, and sell generated audio. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license is available for proprietary embedding; contact **VoiceStudio@palash.dev**. See [LICENSE-NOTICE.md](LICENSE-NOTICE.md) for the plain-language scope.
Optional engines and downloaded models retain their own licenses. The bundled `omnivoice/` model remains Apache-2.0 upstream.
## Acknowledgments
VoiceStudio builds on [OmniVoice](https://github.com/k2-fsa/OmniVoice), [WhisperX](https://github.com/m-bain/whisperX), [Demucs](https://github.com/facebookresearch/demucs), [Pyannote](https://github.com/pyannote/pyannote-audio), [CTranslate2](https://github.com/OpenNMT/CTranslate2), [AudioSeal](https://github.com/facebookresearch/audioseal), [Tauri](https://tauri.app), [Supertonic](https://huggingface.co/Supertone/supertonic-3), [Sherpa-ONNX](https://github.com/k2-fsa/sherpa-onnx), [GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS), and [PocketTTS](https://kyutai.org).
<div align="center">
<img src="https://img.shields.io/badge/raised_%2410_of_%24200-5%25-EAB308?style=for-the-badge" alt="This month's agent-bill fund: $10 / $200" />
<br/><br/>
<a href="https://ko-fi.com/debpalash"><img src="https://img.shields.io/badge/Ko--fi-Support_❤️-FF5E5B?style=for-the-badge&logo=ko-fi&logoColor=white" alt="Ko-fi" /></a>
&nbsp;&nbsp;
<a href="https://paypal.me/palashCoder"><img src="https://img.shields.io/badge/PayPal-Donate-00457C?style=for-the-badge&logo=paypal&logoColor=white" alt="PayPal" /></a>
<br/><br/>
<sub>Also from the maker: <a href="https://github.com/debpalash/Opal"><b>Opal</b> 💠</a> · <a href="https://github.com/debpalash/memxt"><b>memxt</b> 🧠</a> — a ⭐ helps too.</sub>
</div>
<a id="sponsors"></a>
### 🌟 Sponsors
VoiceStudio is **free** and **AGPL-3.0** — no paid tier, no SaaS revenue. Sponsors keep development going, and in return get a logo slot here, in the app, and (for top tiers) on the project website. It's a thank-you, never a paywall. **[See tiers & become a sponsor →](SPONSORS.md)**
<div align="center">
<!-- SPONSORS:START — logo slots are filled here as sponsors come aboard; see SPONSORS.md -->
**Your logo here** — [become a sponsor](SPONSORS.md)
<!-- SPONSORS:END -->
</div>
<sub>💡 GitHub also shows a **Sponsor** button at the top of this repo, wired to the same links via <a href=".github/FUNDING.yml"><code>.github/FUNDING.yml</code></a>.</sub>
---
## 💬 Community
<div align="center">
<a href="https://discord.gg/bzQavDfVV9"><img src="https://img.shields.io/badge/💬_Discord-Join_Community-5865F2?style=for-the-badge&logo=discord&logoColor=white" alt="Join Discord" /></a>
<a href="https://x.com/idebpalash"><img src="https://img.shields.io/badge/𝕏_Follow-for_updates-000000?style=for-the-badge&logo=x&logoColor=white" alt="Follow on X" /></a>
<br/>
<sub>We respond to setup questions within hours, not days.</sub>
</div>
<details>
<summary><b>What happens in there</b></summary>
<br/>
| Channel | What happens there |
|---------|--------------------|
| `#announcements` | Release news and the big moments — new versions land here first |
| `#releases` + `#changelog` | Every build and exactly what's inside it |
| `#issues` | Bug reports as forum posts — triaged straight into GitHub issues |
| `#ideas` | Feature requests, discussed and voted on |
| `#discuss-ideas` | Design talk before things get built |
| `#general` | Setup help, GPU troubleshooting, and showing off your dubs |
</details>
---
<a id="contributing"></a>
## 🤝 Contributing
Yes please — bug fixes, new TTS engine adapters, UI improvements, docs, translations. All of it.
- 📖 Read the **[Contributing Guide](.github/CONTRIBUTING.md)** for setup, code style, and PR workflow
- 🐛 Browse [good first issues](https://github.com/debpalash/VoiceStudio/labels/good%20first%20issue)
- 💬 Join our [Discord](https://discord.gg/bzQavDfVV9) to discuss ideas or ask for help
- 𝕏 Follow [@idebpalash](https://x.com/idebpalash) for updates and what's being built next
---
## ❓ FAQ
<details>
<summary><b>Is this really as good as ElevenLabs?</b></summary>
<br/>
Honest answer: <b>it depends on what you're doing.</b>
<b>Where VoiceStudio is genuinely competitive:</b> voice cloning from a clean reference clip (state-of-the-art open diffusion TTS), language coverage (646 languages vs. their 32), and everything structural — no per-character billing, no usage caps, no audio leaving your machine, full pipeline customizability (14 TTS engines, 11 ASR engines, your choice of translation).
<b>Where ElevenLabs still wins:</b> out-of-the-box consistency and polish, especially for English TTS. Their one model is heavily tuned; our quality depends on which engine you pick, your hardware, and — for cloning — the reference audio (a dry, close-mic clip clones dramatically better than a noisy or echoey one).
<b>For dubbing specifically:</b> a dub is a chain — transcription → translation → cloning → synthesis — only as good as its weakest link on <i>your</i> source material. If parts come out incoherent, check the segment table's <i>original</i> text first: when the transcription is already wrong, switch the ASR engine or use cleaner source audio — that's usually the fix, not the voice.
Try it on your real material — it's free and takes one download. Many users replace ElevenLabs outright; some keep both. Both outcomes are fine with us.
</details>
<details>
<summary><b>Why doesn't a longer reference clip sound more like me?</b></summary>
<br/>
Because VoiceStudio's cloning is <b>zero-shot</b>: your clip is a <i>prompt</i> the model conditions on at generation time — it is never trained on. Feeding it 2 hours doesn't teach it your voice; past a short window the extra audio is simply not used. The dubbing pipeline's reference builder targets ~8 s and hard-caps at 15 s (<code>backend/services/speaker_clone.py</code>), and engines cap the prompt themselves (VoxCPM2 trims references to 30 s). This is different from ElevenLabs <i>Professional</i> Voice Cloning, which fine-tunes a model on hours of your audio — that's a training job, not a bigger prompt.
<b>What actually moves clone quality is the clip, not its length.</b> Zero-shot cloning mirrors the acoustics and delivery of the prompt, so: record 515 seconds (~8 s is the sweet spot) of continuous natural speech, close to the mic, in a quiet room with no reverb or music — an echoey clip clones echoey. One speaker only, and read in the tone and pace you want the output to have, because the clone copies your delivery, not just your timbre. Recording a few candidate clips and comparing results beats any amount of extra footage.
<b>Want audiobook-grade, trained-on-your-voice fidelity?</b> That path exists, but it's offline fine-tuning, not an in-app button: prepare a dataset of your recordings (<a href="docs/data_preparation.md">docs/data_preparation.md</a>) and fine-tune the bundled checkpoint via <code>init_from_checkpoint</code> (<a href="docs/training.md">docs/training.md</a>). Fair warning — it's a technical, command-line workflow that needs a capable GPU and hours of transcribed audio. In-app fine-tuning / long-reference "professional" cloning is on the <a href="docs/ROADMAP.md">roadmap</a> as research only; no promised date.
</details>
<details>
<summary><b>Does it work on Apple Silicon (M1/M2/M3/M4)?</b></summary>
<br/>
Yes. MPS acceleration is auto-detected. MLX-optimized Whisper models are available for faster transcription on Apple hardware. <b>Intel Macs are not supported</b>: the app UI installs, but the local Python backend cannot run because PyTorch no longer ships Intel-Mac wheels (<a href="https://github.com/debpalash/VoiceStudio/issues/889">#889</a>) — an Intel Mac can only be used with a remote backend.
</details>
<details>
<summary><b>How much VRAM do I need?</b></summary>
<br/>
<b>4 GB minimum.</b> With ≤8 GB, the TTS model is automatically offloaded to CPU during transcription. With 8+ GB, everything runs on GPU simultaneously. No GPU at all? CPU mode works — just slower (~3× for TTS).
</details>
<details>
<summary><b>Can I use this commercially?</b></summary>
<br/>
<b>Yes — commercial use is free</b> under the <a href="https://www.gnu.org/licenses/agpl-3.0.html">AGPL-3.0</a>: run it, sell the audio you make, dub client videos, deploy it across your team. One obligation: if you <b>modify</b> VoiceStudio and offer the modified version to others over a network, you must share that modified source under the same terms. Embedding it in a closed-source product instead? A commercial license is available — see <a href="#license">License</a>.
</details>
<details>
<summary><b>What languages are supported?</b></summary>
<br/>
646 languages for TTS via the VoiceStudio model. Transcription (WhisperX) supports 99 languages. Translation coverage depends on the target language pair.
</details>
<details>
<summary><b>Can I add my own TTS engine?</b></summary>
<br/>
Yes. Subclass <code>TTSBackend</code> in <code>backend/services/tts_backend.py</code> and add it to the <code>_REGISTRY</code> dictionary — ~50 lines. The fourteen built-in engines all work this way; see <a href="#tts-engines">TTS Engines</a>.
</details>
<details>
<summary><b>Does VoiceStudio collect any data about me?</b></summary>
<br/>
<b>Not unless you explicitly say yes.</b> On first run the app <i>asks</i> — one screen, two equal-weight buttons, no pre-ticked box — and until you answer yes, VoiceStudio sends nothing: no analytics, no telemetry, no accounts, no phone-home. Skipping the question means no. Your text, audio, voices, and projects never leave your machine either way.
If you do opt in (also togglable anytime under <b>Settings → Privacy → "Help improve VoiceStudio"</b>), what's sent is anonymous, content-free usage stats: generations (engine, language, generation time, character <i>count</i>, error <i>type</i>), plus app lifecycle — an install ping, updates (version-to-version), crashes (error class and a <i>bucketed</i> uptime, never logs), error <i>types</i> (capped, deduplicated), and a single uninstall ping if you remove it. Never your text, audio, file names, or anything identifying — enforced in code by a property allowlist (<code>backend/core/analytics.py</code>), not just a promise. Every build — installer, Docker, or built from source — asks the same first-run question and stays off unless you say yes (the destination is PostHog's publishable write-only client key; skipping the question means off). Your own numbers live in <b>Settings → Usage</b>, computed locally, sent nowhere.
</details>
<details>
<summary><b>How do I uninstall it / remove all its data?</b></summary>
<br/>
VoiceStudio is fully local — uninstalling is just deleting the app plus the folders it wrote (model cache, Python env, your voices/projects, config). Run <code>scripts/uninstall.sh</code> (macOS/Linux) or <code>scripts\uninstall.ps1</code> (Windows) — it prints every folder with its size as a dry-run first, then deletes on <code>--yes</code>. The full per-platform path list and app-removal steps are in <a href="docs/install/uninstall.md"><b>docs/install/uninstall.md</b></a>.
</details>
---
<a id="license"></a>
## 📜 License
VoiceStudio is free and open-source software under the [**GNU Affero General Public License v3.0 (AGPL-3.0)**](https://www.gnu.org/licenses/agpl-3.0.html).
**Free for any use — including commercial and internal business use.** Run it, sell the audio you produce with it, dub your own or clients' videos, roll it out across your team — all free, no license needed. As a **network copyleft** license, AGPL adds one obligation: if you **modify** VoiceStudio and offer that modified version to others over a network, you must make the complete corresponding source of your modified version available to them under the same AGPL-3.0 terms.
A **commercial license** is available for organizations that want to embed VoiceStudio in a **closed-source or proprietary** product or service without the AGPL-3.0 copyleft obligations. **Pricing tiers coming soon.** Inquiries: **VoiceStudio@palash.dev**.
The bundled `omnivoice/` TTS model by Han Zhu remains Apache-2.0 upstream. See [`LICENSE`](LICENSE) for the full, binding terms, and [`LICENSE-NOTICE.md`](LICENSE-NOTICE.md) for the plain-language summary and scope.
---
## 🙏 Acknowledgments
VoiceStudio is built on the shoulders of exceptional open-source work:
| Project | Role |
|---------|------|
| [**VoiceStudio (k2-fsa)**](https://github.com/k2-fsa/OmniVoice) | Zero-shot diffusion TTS engine — the core voice synthesis model |
| [**WhisperX**](https://github.com/m-bain/whisperX) | Word-level speech recognition and alignment |
| [**Demucs (Meta)**](https://github.com/facebookresearch/demucs) | Music source separation for vocal isolation |
| [**Pyannote**](https://github.com/pyannote/pyannote-audio) | Speaker diarization — who said what |
| [**CTranslate2**](https://github.com/OpenNMT/CTranslate2) | Optimized Transformer inference on CPU and GPU |
| [**AudioSeal (Meta)**](https://github.com/facebookresearch/audioseal) | Invisible neural audio watermarking for AI provenance |
| [**Tauri**](https://tauri.app) | Native desktop app framework |
| [**Supertone / Supertonic 3**](https://huggingface.co/Supertone/supertonic-3) | ONNX TTS engine — 31 languages, CPU-efficient |
| [**Sherpa-ONNX**](https://github.com/k2-fsa/sherpa-onnx) | WASM-ready universal TTS/ASR runtime |
| [**GPT-SoVITS**](https://github.com/RVC-Boss/GPT-SoVITS) | Zero-shot TTS engine — 5 languages, RTF 0.014 |
---
<a id="more-from-the-maker"></a>
## 🧰 More local open-source from the maker
Like the local-first philosophy? It runs in the family — same maker, same rule: **your data stays on your machine.**
<table>
<tr>
<td align="center" width="50%" valign="top">
<br/>
<a href="https://github.com/debpalash/Opal"><img src="https://raw.githubusercontent.com/debpalash/Opal/main/assets/opal_logo.png" width="96" alt="Opal logo"/></a>
<h3><a href="https://github.com/debpalash/Opal">Opal 💠</a></h3>
<p><b>Play everything.</b> The media player for the AI era.</p>
<p><sub>Video, anime, comics, torrents, Jellyfin & Plex — one player for all of it, with local AI memory and context built in. Written in Zig, runs on macOS & Windows.</sub></p>
<p>
<a href="https://github.com/debpalash/Opal/stargazers"><img src="https://img.shields.io/github/stars/debpalash/Opal?style=flat-square&color=f59e0b" alt="Opal stars"/></a>
<a href="https://palash.dev/opal"><img src="https://img.shields.io/badge/site-palash.dev%2Fopal-8b5cf6?style=flat-square" alt="Opal website"/></a>
</p>
</td>
<td align="center" width="50%" valign="top">
<br/>
<a href="https://github.com/debpalash/memxt"><img src="https://raw.githubusercontent.com/debpalash/memxt/main/assets/logo-mark.svg" width="96" alt="memxt logo"/></a>
<h3><a href="https://github.com/debpalash/memxt">memxt 🧠</a></h3>
<p><b>The fastest benchmarked open-source AI memory system.</b></p>
<p><sub>Local long-term memory for Claude Code and coding agents — an MCP server on SQLite + embeddings, 100% on your machine. Your agent finally remembers yesterday.</sub></p>
<p>
<a href="https://github.com/debpalash/memxt/stargazers"><img src="https://img.shields.io/github/stars/debpalash/memxt?style=flat-square&color=f59e0b" alt="memxt stars"/></a>
<a href="https://github.com/debpalash/memxt#readme"><img src="https://img.shields.io/badge/docs-README-10b981?style=flat-square" alt="memxt docs"/></a>
</p>
</td>
</tr>
</table>
---
<div align="center">
<br/>
If you read this far, you're our kind of person.<br/>
**[⭐ Star this repo](https://github.com/debpalash/VoiceStudio)** so others can find it too.<br/>
**[💬 Join the Discord](https://discord.gg/bzQavDfVV9)** to share what you build.<br/>
**[❤️ Support development](https://ko-fi.com/debpalash)** — fund the AI agent bills that keep VoiceStudio shipping.
<br/>
<a href="https://star-history.com/#debpalash/VoiceStudio&Date">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/svg?repos=debpalash/VoiceStudio&type=Date&theme=dark" />
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/svg?repos=debpalash/VoiceStudio&type=Date" />
<img alt="Star History" src="https://api.star-history.com/svg?repos=debpalash/VoiceStudio&type=Date&theme=dark" width="600" />
</picture>
</a>
<strong><a href="https://github.com/debpalash/VoiceStudio/releases/latest">Download VoiceStudio</a></strong> ·
<a href="https://github.com/debpalash/VoiceStudio">Star the project</a> ·
<a href="https://discord.gg/bzQavDfVV9">Join Discord</a>
</div>
+59 -52
View File
@@ -37,7 +37,7 @@
<br/>
<div align="center">
<img src="docs/screenshot-launchpad.png" alt="VoiceStudio — 启动台" width="100%"/>
<img src="docs/media/0.5.0/quick-switch.gif" alt="VoiceStudio — 从状态栏快速切换 TTS 引擎" width="100%"/>
</div>
> **声音很私人,创作空间也应该真正属于你。** VoiceStudio 的核心流程运行在你的硬件上:克隆、设计、配音、听写,并以 646 种语言创作,不需要订阅,也没有用量计费。联网引擎和服务始终是清晰可见的可选项,而不是隐藏依赖。
@@ -45,6 +45,56 @@
> [!WARNING]
> **活跃 Beta 阶段。** 各版本之间可能出现故障——如需最新修复,请从源码运行。非常欢迎 Bug 报告和 PR:[提交 Issue](https://github.com/debpalash/VoiceStudio/issues) 或 [加入 Discord](https://discord.gg/bzQavDfVV9)。
<a id="quickstart"></a>
## ⚡ 快速开始
<div align="center">
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/macOS-DMG_(Apple_Silicon)-000?style=for-the-badge&logo=apple&logoColor=white" alt="下载 macOS DMG" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Windows-MSI_(x64)-0078D4?style=for-the-badge&logo=windows&logoColor=white" alt="下载 Windows MSI" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Linux-AppImage_(x64)-FCC624?style=for-the-badge&logo=linux&logoColor=black" alt="下载 Linux AppImage" /></a>
<br/>
<sub>三个按钮都会打开最新发布页——在资源列表中下载对应你系统的安装包。</sub><br/>
<sub><b>macOS</b>首次启动需要一次性批准——右键点击 → <b>打开</b>macOS 15 上为 系统设置 → 隐私与安全性 → <b>“仍要打开”</b>)。无需终端。<a href="docs/install/macos.md#gatekeeper-quarantine">为什么?</a> · <b>Intel Mac</b>不支持本地后端(<a href="https://github.com/debpalash/VoiceStudio/issues/889">#889</a>)——<a href="docs/install/macos.md">详情</a>。</sub>
</div>
选择你的操作系统,按指南从头到尾操作:
- 🍎 **macOS** — [docs/install/macos.md](docs/install/macos.md)
- 🪟 **Windows** — [docs/install/windows.md](docs/install/windows.md)
- 🐧 **Linux** — [docs/install/linux.md](docs/install/linux.md)
- 🐳 **Docker** — [docs/install/docker.md](docs/install/docker.md) · [Docker Hub: `palashdeb/omnivoice-studio`](https://hub.docker.com/r/palashdeb/omnivoice-studio)
**三步克隆出你的第一个声音:**
1. **安装并启动。** 首次启动会自动搭建 Python 运行环境并下载模型权重——启动画面会逐步显示进度(仅首次,需要几分钟;之后即开即用)。
2. 从启动台打开**语音克隆**,拖入任意声音的 **3 秒音频**
3. **输入一句话,点击生成。** 音频完全属于你——在你的设备上生成和保存,支持 646 种语言。
觉得慢?[docs/performance.md](docs/performance.md) 讲清了生成时间到底花在哪里、有哪些调优开关,以及“它变慢了”的三个经典原因。各引擎/设备的实测数据见 [docs/benchmarks.md](docs/benchmarks.md)。
> 正在从 **[CorentinJ/Real-Time-Voice-Cloning](https://github.com/CorentinJ/Real-Time-Voice-Cloning)**(现已归档)迁移过来?我们有专门的迁移指南:[docs/migration/real-time-voice-cloning.md](docs/migration/real-time-voice-cloning.md)。
<details>
<summary><b>🧰 卡住了?自检、Token 与受限网络</b></summary>
<br/>
先运行内置自检——在应用中打开 **设置 → 关于 → “运行自检”**,或在源码检出目录中执行
`uv run python backend/main.py --diagnose`(加 `--deep` 还会实际加载当前引擎进行测试)。然后查看
[docs/install/troubleshooting.md](docs/install/troubleshooting.md) 中排名前
10 的安装错误。运行时出错时,应用内的错误界面会直接深链到对应条目;**设置 → 关于 →
“保存诊断包”** 会把脱敏日志与自检报告打包,方便附在 Bug 报告里。
Hugging Face Token 的配置见
[docs/setup/huggingface-token.md](docs/setup/huggingface-token.md)。说话人分离相关的模型访问门槛见
[docs/features/diarization.md](docs/features/diarization.md)。下载速度、⚡ 快速下载(Xet)状态,以及受限网络 / 镜像选项见
[docs/downloading-models.md](docs/downloading-models.md)。
</details>
---
<a id="features"></a>
## ✨ 功能
@@ -112,49 +162,6 @@
---
<a id="quickstart"></a>
## ⚡ 快速开始
<div align="center">
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/macOS-DMG_(Apple_Silicon)-000?style=for-the-badge&logo=apple&logoColor=white" alt="下载 macOS DMG" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Windows-MSI_(x64)-0078D4?style=for-the-badge&logo=windows&logoColor=white" alt="下载 Windows MSI" /></a>
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/badge/Linux-AppImage_(x64)-FCC624?style=for-the-badge&logo=linux&logoColor=black" alt="下载 Linux AppImage" /></a>
<br/>
<sub><b>macOS</b>首次启动需要一次性批准——右键点击 → <b>打开</b>macOS 15 上为 系统设置 → 隐私与安全性 → <b>“仍要打开”</b>)。无需终端。<a href="docs/install/macos.md#gatekeeper-quarantine">为什么?</a> · <b>Intel Mac</b>不支持本地后端(<a href="https://github.com/debpalash/VoiceStudio/issues/889">#889</a>)——<a href="docs/install/macos.md">详情</a>。</sub>
</div>
选择你的操作系统,按指南从头到尾操作:
- 🍎 **macOS** — [docs/install/macos.md](docs/install/macos.md)
- 🪟 **Windows** — [docs/install/windows.md](docs/install/windows.md)
- 🐧 **Linux** — [docs/install/linux.md](docs/install/linux.md)
- 🐳 **Docker** — [docs/install/docker.md](docs/install/docker.md) · [Docker Hub: `palashdeb/omnivoice-studio`](https://hub.docker.com/r/palashdeb/omnivoice-studio)
觉得慢?[docs/performance.md](docs/performance.md) 讲清了生成时间到底花在哪里、有哪些调优开关,以及“它变慢了”的三个经典原因。
> 正在从 **[CorentinJ/Real-Time-Voice-Cloning](https://github.com/CorentinJ/Real-Time-Voice-Cloning)**(现已归档)迁移过来?我们有专门的迁移指南:[docs/migration/real-time-voice-cloning.md](docs/migration/real-time-voice-cloning.md)。
<details>
<summary><b>🧰 卡住了?自检、Token 与受限网络</b></summary>
<br/>
先运行内置自检——在应用中打开 **设置 → 关于 → “运行自检”**,或在源码检出目录中执行
`uv run python backend/main.py --diagnose`(加 `--deep` 还会实际加载当前引擎进行测试)。然后查看
[docs/install/troubleshooting.md](docs/install/troubleshooting.md) 中排名前
10 的安装错误。运行时出错时,应用内的错误界面会直接深链到对应条目;**设置 → 关于 →
“保存诊断包”** 会把脱敏日志与自检报告打包,方便附在 Bug 报告里。
Hugging Face Token 的配置见
[docs/setup/huggingface-token.md](docs/setup/huggingface-token.md)。说话人分离相关的模型访问门槛见
[docs/features/diarization.md](docs/features/diarization.md)。下载速度、⚡ 快速下载(Xet)状态,以及受限网络 / 镜像选项见
[docs/downloading-models.md](docs/downloading-models.md)。
</details>
---
<a id="why-voicestudio"></a>
## 💡 为什么选择 VoiceStudio
@@ -173,8 +180,8 @@ Hugging Face Token 的配置见
| **API 密钥** | 需要账号 | 本地流程不需要 |
| **GPU 支持** | 不适用(云端) | CUDA · Apple Silicon · ROCmLinux)· CPU |
| **桌面应用** | ❌ | ✅ macOS · Windows · Linux |
| **TTS 引擎** | 1 | **14** — [完整矩阵](#tts-engines) |
| **ASR 引擎** | 1 | **10** — [完整阵容](#asr-engines) |
| **TTS 引擎** | 1 | **16** — [完整矩阵](#tts-engines) |
| **ASR 引擎** | 1 | **11** — [完整阵容](#asr-engines) |
| **MCP 服务器** | ❌ | ✅ 可从 Claude、Cursor 及任何 MCP 客户端使用 |
| **自检** | ❌ | ✅ 诊断套件、错误日志、脱敏调试包 |
| **可定制** | ❌ 闭源 | ✅ 随你 Fork、扩展、发布 |
@@ -214,10 +221,10 @@ Hugging Face Token 的配置见
### 🗣️ TTS 引擎
**14 个引擎,一个选择器。** VoiceStudio(默认,支持 600+ 语言)始终可用;另有七个引擎可选装并自动检测(CosyVoice 3、GPT-SoVITS、VoxCPM2、MOSS-TTS-Nano、KittenTTS、MLX-Audio、Sherpa-ONNX),外加个按需延迟安装的重量级引擎(IndexTTS 2.5、OmniVoice GGUF、Supertonic 3、MOSS-TTS-v1.5、dots.tts、Confucius4-TTS)。在 **设置 → TTS 引擎** 中切换;所选引擎将应用于所有语音合成场景。
**16 个引擎,一个选择器。** VoiceStudio(默认,支持 600+ 语言)始终可用;另有七个引擎可选装并自动检测(CosyVoice 3、GPT-SoVITS、VoxCPM2、MOSS-TTS-Nano、KittenTTS、MLX-Audio、Sherpa-ONNX),外加个按需延迟安装的引擎(IndexTTS 2.5、OmniVoice GGUF、OmniVoice 子进程版、PocketTTS、Supertonic 3、MOSS-TTS-v1.5、dots.tts、Confucius4-TTS)。在 **设置 → TTS 引擎** 中切换;所选引擎将应用于所有语音合成场景。**每个引擎都有独立指南:[docs/engines](docs/engines/README.md)(英文)。**
<details>
<summary><b>📊 完整矩阵</b>——14 个引擎 × 平台 × 克隆/指令 × 许可证</summary>
<summary><b>📊 完整矩阵</b>——16 个引擎 × 平台 × 克隆/指令 × 许可证</summary>
<br/>
@@ -254,10 +261,10 @@ Hugging Face Token 的配置见
### 🎧 ASR 引擎
**10 个引擎**——它们驱动听写、视频配音和字幕。**WhisperX** 是跨平台的默认引擎(约 100 种语言,词级时间对齐);其余引擎均为可选装并自动检测。在 **设置 → 引擎** 中切换。个完全在本地设备上运行;第十个(OpenAI 兼容)是可选的远程客户端,可用于 Qwen3-ASR 或任何兼容的服务器。
**11 个引擎**——它们驱动听写、视频配音和字幕。**WhisperX** 是跨平台的默认引擎(约 100 种语言,词级时间对齐);其余引擎均为可选装并自动检测。在 **设置 → 引擎** 中切换。个完全在本地设备上运行;第十个(OpenAI 兼容)是可选的远程客户端,可用于 Qwen3-ASR 或任何兼容的服务器。
<details>
<summary><b>📊 完整阵容</b>——10 个引擎、各自的强项与计算类型说明</summary>
<summary><b>📊 完整阵容</b>——11 个引擎、各自的强项与计算类型说明</summary>
<br/>
@@ -274,7 +281,7 @@ Hugging Face Token 的配置见
| **sherpa-onnx**(实时听写) | `sherpa-onnx-asr` | 25 种欧洲语言 + 90+ | 实时、快于实时的听写——小体积流式/离线 ONNX 模型(Parakeet TDT v3/v2、流式 Zipformer 与 Paraformer、Whisper Tiny),CPU 运行,macOS / Windows / Linux 表现完全一致。在 **设置 → 语音** 中按模型选择。 |
| **OpenAI 兼容** ⚠️ 远程 | `openai-compat-asr` | 取决于服务器 | 当下通往 **Qwen3-ASR** 的路径(自托管服务器,无需等 transformers 支持)、任何 OpenAI 兼容的转录端点,或 OpenAI 官方 API——无需安装,在 **设置 → 引擎**(ASR 标签页)中配置并测试连接。音频会离开你的设备,发送到你指定的任何服务器;参见 [docs/engines/openai-compatible-asr.md](docs/engines/openai-compatible-asr.md)。 |
> Whisper 系列引擎覆盖约 100 种语言;**FunASR / SenseVoice** 额外提供一条多语言一体化路径,内置语音活动检测与行内说话人分离。**sherpa-onnx** 驱动实时听写的模型选择器——你边说,文字边出现。每个引擎都在本地设备上运行——无需 API 密钥,无需云端。
> Whisper 系列引擎覆盖约 100 种语言;**FunASR / SenseVoice** 额外提供一条多语言一体化路径,内置语音活动检测与行内说话人分离。**sherpa-onnx** 驱动实时听写的模型选择器——你边说,文字边出现。除可选的 OpenAI 兼容远程客户端外,所有引擎都在本地设备上运行——无需 API 密钥,无需云端。
> **GPU 不支持高效 float16** 在较老的 NVIDIA GPUMaxwell/Pascal、GTX 16xx)上,或在 CTranslate2/cuDNN 版本不匹配之后,CTranslate2 系 ASR 引擎(WhisperX、Faster-Whisper)无法运行 `float16`VoiceStudio 会自动改用 `int8` 重试——无需配置。如果转录仍然失败,可用 `ASR_COMPUTE_TYPE` 环境变量固定计算类型(逃生舱口):`ASR_COMPUTE_TYPE=int8`CPU 用 `float32`)。将其设为 `int8` 并重启后端。
@@ -574,7 +581,7 @@ VoiceStudio 站在这些杰出开源工作的肩膀上:
## 🧰 来自同一作者的更多本地开源项目
喜欢这种本地优先的理念?它是一脉相承的——同一位作者,同一条准则:**你的数据只留在你的设备上。**
喜欢这种本地优先的理念?它是一脉相承的——同一位作者,同一条准则:**你的数据只留在你的设备上。** 全部项目见 [palash.dev](https://palash.dev)。
<table>
<tr>
+86 -118
View File
@@ -6,7 +6,10 @@ composed at the route or router level without surprises.
Currently exposed:
- `require_loopback`: 403 unless the request came from a loopback origin
(bypassed in explicit server mode see `_server_mode`).
(read-only bootstrap is allowed in explicit server mode; mutations still
require the admin API key see `_server_mode`).
- `require_admin`: method-aware admin gate for privileged routers.
- `require_admin_action`: strict admin gate for side-effectful GET actions.
- `require_native_access`: true-loopback-only access to the host filesystem;
unlike `require_loopback`, it is never bypassed by server mode.
- `ws_remote_authorized`: whether a WebSocket handshake from a non-loopback
@@ -14,64 +17,19 @@ Currently exposed:
keep their own inline loopback guards.
"""
import ipaddress
import os
import secrets
from fastapi import HTTPException, Request
# IPv4 + IPv6 loopback literals + the conventional `localhost` hostname.
# `request.client.host` carries an address, not a hostname, so the literal
# "localhost" entry is defensive — some upstream wrappers (TestClient with
# a custom client tuple, certain reverse-proxy headers) may pass strings
# rather than parsed addresses. We accept the broader set without weakening
# the guard: nothing here matches a non-loopback origin.
_LOOPBACK_HOSTS = frozenset({"127.0.0.1", "::1", "localhost"})
def _trusted_networks():
"""CIDR networks from OMNIVOICE_TRUSTED_NETWORKS (comma-separated) treated as
loopback-trusted e.g. a reverse proxy or self-hosted LAN, so the API-key /
PIN gates don't block LAN clients that can't present the credential (a proxy
that strips the Authorization header). Read at call time (matching
`_server_mode` / `remote_api_key`) so tests can monkeypatch the env; restart
to apply changes in production."""
nets = []
for cidr in os.environ.get("OMNIVOICE_TRUSTED_NETWORKS", "").split(","):
cidr = cidr.strip()
if cidr:
try:
nets.append(ipaddress.ip_network(cidr, strict=False))
except ValueError:
pass # malformed entry ignored — never wedge the auth gate
return nets
def is_loopback(host):
"""True loopback address only (127.0.0.1, ::1, localhost) — NOT a trusted
network. Admin gates (``require_loopback`` ``/system/set-env``,
``/api/settings/*``) use this so a trusted-network CIDR exempts consumption
(TTS / dictation) but never the RCE-class admin surface."""
return host in _LOOPBACK_HOSTS
def is_local_host(host):
"""Loopback address, OR on a configured trusted network. The consumption
gates (PIN/API-key middleware, WS guard) call this so a trusted LAN/proxy is
exempted. Admin gates use :func:`is_loopback` NOT this to preserve the
two-tier privilege model: consumption trust admin trust."""
if is_loopback(host):
return True
try:
ip = ipaddress.ip_address(host)
except (ValueError, TypeError):
return False
# Unwrap IPv4-mapped IPv6 (::ffff:192.168.1.5) so it matches IPv4 CIDRs —
# dual-stack proxies (Caddy, Node.js) frequently pass the mapped form.
if getattr(ip, "ipv4_mapped", None):
ip = ip.ipv4_mapped
return any(ip in net for net in _trusted_networks())
from core.auth import (
CredentialTransport,
PrincipalKind,
is_local_host,
is_loopback,
principal_for,
remote_api_key,
)
from core.csrf import SAFE_HTTP_METHODS, cookie_csrf_allowed
_TRUTHY = frozenset({"1", "true", "yes", "on"})
@@ -107,40 +65,45 @@ def _configured_pin(request) -> str | None:
def _admin_credential_configured(request) -> bool:
"""Whether the operator has set ANY credential gate — the remote API key or
a share PIN. When neither is set, server mode leaves admin open (the Docker
issue #261 flow the image depends on)."""
if os.environ.get("OMNIVOICE_API_KEY"):
"""Whether an API key or share PIN is configured.
The PIN cannot authorize admin access, but its presence means the operator
opted out of bare-server discovery. Remote admin then remains closed until
they configure and present the long API key.
"""
if remote_api_key():
return True
return bool(_configured_pin(request))
def _request_presents_admin_credential(request) -> bool:
"""Whether the request carries a valid **API key** via the channels the
middleware accepts (``Authorization: Bearer`` / ``?api_key`` / ``ov_key``
cookie).
def _request_presents_admin_credential(
request,
*,
side_effectful_get: bool = False,
) -> bool:
"""Whether the canonical principal carries remote admin capability.
Admin is RCE-class (``/system/set-env`` + ``/api/settings/*``), so only the
API key a long operator-chosen secret unlocks it. The 6-digit share PIN
is deliberately NOT accepted here: it is a *consumption* credential for LAN
playback and is short enough to brute-force (10^6, no lockout), so it must
never gate the admin surface (CodeRabbit #1213). A trusted-network CIDR
(``is_local_host`` also a consumption exemption) likewise never unlocks
admin. Net: remote admin in server mode requires the API key; a PIN-only
deployment keeps admin loopback-only. getattr-defensive so a minimal Request
stub never raises."""
api_key = os.environ.get("OMNIVOICE_API_KEY") or ""
if not api_key:
API-key and short-lived session principals may unlock server-mode admin.
PIN and trusted-network principals remain consumption-only.
"""
principal = principal_for(request)
if principal.kind not in {
PrincipalKind.API_KEY,
PrincipalKind.ADMIN_SESSION,
}:
return False
headers = getattr(request, "headers", None) or {}
query = getattr(request, "query_params", None) or {}
cookies = getattr(request, "cookies", None) or {}
auth = headers.get("authorization", "")
supplied = auth[7:].strip() if auth.lower().startswith("bearer ") else ""
if not supplied:
supplied = query.get("api_key") or cookies.get("ov_key") or ""
return bool(supplied and secrets.compare_digest(supplied, api_key))
if principal.transport not in {
CredentialTransport.COOKIE,
CredentialTransport.LEGACY_COOKIE,
}:
return True
method = str(getattr(request, "method", "GET")).upper()
if side_effectful_get or method not in SAFE_HTTP_METHODS:
return cookie_csrf_allowed(
request,
side_effectful_get=side_effectful_get,
)
return True
def require_loopback(request: Request) -> None:
@@ -163,9 +126,9 @@ def require_loopback(request: Request) -> None:
unenforceable, so the gate can't require true loopback. It then applies the
admin-credential rule instead:
- No credential configured (no API key, no PIN) open, matching the #261
Docker flow where the operator reaches ``/system/*`` off the bridge
gateway with nothing set.
- No credential configured (no API key, no PIN) read-only requests are
open, matching the #261 Docker bootstrap flow. State-changing requests
fail closed even if a route accidentally kept this legacy dependency.
- A credential IS configured the request must present the **API key**.
This keeps the two-tier privilege model intact under server mode:
``OMNIVOICE_TRUSTED_NETWORKS`` is a *consumption* exemption
@@ -173,14 +136,20 @@ def require_loopback(request: Request) -> None:
NEVER by itself unlock the admin surface (``/system/set-env`` RCE-class
and ``/api/settings/*``). The 6-digit share PIN is a consumption credential
too and does not gate admin, so a PIN-only deployment keeps admin
loopback-only; remote admin requires the (long) API key. A LAN client in a
trusted CIDR or one holding only the PIN gets 403 here even though it
sails through the consumption gates. See docs/api-auth.md (#1213).
loopback-only; remote admin requires the long API key. See
docs/api-auth.md (#1213).
"""
host = request.client.host if request.client else None
if is_loopback(host):
return
if _server_mode():
method = str(getattr(request, "method", "GET")).upper()
if method not in SAFE_HTTP_METHODS:
# Defense in depth. Privileged routers should declare
# ``require_admin`` directly, but a missed migration must not turn
# into an unauthenticated Docker write primitive.
require_admin(request)
return
if not _admin_credential_configured(request):
return
if _request_presents_admin_credential(request):
@@ -206,14 +175,32 @@ def require_admin(request: Request) -> None:
return
if _server_mode():
method = str(getattr(request, "method", "GET")).upper()
read_only = method in {"GET", "HEAD", "OPTIONS"}
if read_only and not os.environ.get("OMNIVOICE_API_KEY", "").strip():
read_only = method in SAFE_HTTP_METHODS
if read_only and not _admin_credential_configured(request):
return
if _request_presents_admin_credential(request):
return
raise HTTPException(status_code=403, detail="loopback origin or admin API key required")
def require_admin_action(request: Request) -> None:
"""Gate an administrative action even when its HTTP method is read-only.
A small number of legacy GET endpoints have real side effects. For example,
an engine health check may spawn a sidecar process. Such routes cannot use
:func:`require_admin`'s bare-server discovery exception.
"""
host = request.client.host if request.client else None
if is_loopback(host):
return
if _server_mode() and _request_presents_admin_credential(
request,
side_effectful_get=True,
):
return
raise HTTPException(status_code=403, detail="loopback origin or admin API key required")
def require_desktop(request: Request) -> None:
"""Gate capabilities that may select or execute host filesystem paths.
@@ -232,9 +219,10 @@ def require_local(request: Request) -> None:
trusted network. The consumption-tier companion to :func:`require_loopback`:
use on routes a trusted-network client (LAN/proxy) should reach without a PIN
or API key e.g. the dictation model/prefs endpoints that pair with the
dictation WebSocket. Admin routes stay on :func:`require_loopback`.
dictation WebSocket. Admin routes stay on :func:`require_admin`.
In server mode the gate is a no-op (same as :func:`require_loopback`)."""
In server mode this consumption gate is a no-op. Admin dependencies remain
method-aware and independent from this exemption."""
host = request.client.host if request.client else None
if is_local_host(host):
return
@@ -256,29 +244,9 @@ def require_native_access(request: Request) -> None:
raise HTTPException(status_code=403, detail="native filesystem access requires loopback origin")
def remote_api_key() -> str | None:
"""The remote-backend bearer key (Wave 2.3), or None when remote mode is
off. Read at call time so tests can monkeypatch the env."""
return os.environ.get("OMNIVOICE_API_KEY") or None
def ws_remote_authorized(websocket) -> bool:
"""Whether a WebSocket handshake presents the remote API key.
Browser WebSockets cannot set an Authorization header, so the key may
arrive as ``?api_key=`` or via the ``ov_key`` cookie that the bearer
middleware sets on the first authenticated HTTP request. Returns False
when remote mode is off callers keep their loopback-only behavior.
"""
key = remote_api_key()
if not key:
return False
auth = websocket.headers.get("authorization", "")
supplied = auth[7:].strip() if auth.lower().startswith("bearer ") else ""
if not supplied:
supplied = (
websocket.query_params.get("api_key")
or websocket.cookies.get("ov_key")
or ""
)
return secrets.compare_digest(supplied, key)
"""Whether the canonical WS principal has a remote admin credential."""
return principal_for(websocket).kind in {
PrincipalKind.API_KEY,
PrincipalKind.ADMIN_SESSION,
}
+252 -50
View File
@@ -26,8 +26,10 @@ Design notes
from __future__ import annotations
import hashlib
import json
import logging
import os
import re
import time
import uuid
from pathlib import Path
@@ -37,6 +39,7 @@ from fastapi import APIRouter, Body, HTTPException, Query
from fastapi.responses import FileResponse
from core import archetypes
from core.audio_validation import is_playable_wav, resolve_regular_file
from core.config import OUTPUTS_DIR, VOICES_DIR
from services import gallery
@@ -69,6 +72,153 @@ def _preview_key(a: dict) -> str:
).hexdigest()[:16]
def _design_profile_values(a: dict) -> tuple[str, str]:
"""Canonical instruct + complete picker state for a designed archetype."""
return a["instruct"], json.dumps(a["attrs"], sort_keys=True)
def _profile_audio_path(ref_audio_path: object) -> Optional[Path]:
"""Resolve only a regular, non-symlinked file inside ``VOICES_DIR``."""
return resolve_regular_file(VOICES_DIR, ref_audio_path)
def _materialized_audio_is_current(row, a: dict) -> bool:
"""Whether an existing row still has the sample described by its metadata."""
expected_filename = _profile_audio_filename(row["id"])
path = _profile_audio_path(row["ref_audio_path"])
return bool(
row["ref_audio_path"] == expected_filename
and is_playable_wav(path)
and row["instruct"] == a["instruct"]
and row["language"] == a["language"]
and row["ref_text"] == a["sample_script"]
and row["seed"] == _PREVIEW_SEED
)
def _profile_audio_filename(profile_id: str) -> str:
safe_id = (
profile_id if re.fullmatch(r"[A-Za-z0-9_-]{1,64}", profile_id or "")
else hashlib.sha256(str(profile_id).encode("utf-8")).hexdigest()[:16]
)
return f"{safe_id}.wav"
def _archetype_personality(a: dict) -> str:
return f"archetype:{a['id']}"
def _legacy_archetype_profile(conn, a: dict):
"""Adopt only a row that an older archetype materializer could have made."""
row = conn.execute(
"SELECT * FROM voice_profiles WHERE personality=? LIMIT 1",
(a["id"],),
).fetchone()
if row is None:
return None
expected_audio = _profile_audio_filename(row["id"])
try:
states_match = (
not row["vd_states"] or json.loads(row["vd_states"]) == a["attrs"]
)
except (TypeError, ValueError):
states_match = False
if (
row["ref_audio_path"] == expected_audio
and row["instruct"] == a["instruct"]
and row["language"] == a["language"]
and row["ref_text"] == a["sample_script"]
and row["seed"] == _PREVIEW_SEED
and row["kind"] in (None, "", "clone", "design")
and not row["is_locked"]
and not row["verified_own_voice"]
and states_match
):
return row
return None
def _is_materialized_archetype_row(row, a: dict) -> bool:
"""Recognize rows owned by this materializer without trusting identity text alone."""
try:
states_match = json.loads(row["vd_states"]) == a["attrs"]
except (TypeError, ValueError):
return False
return bool(
row["personality"] == _archetype_personality(a)
and row["kind"] == "design"
and row["seed"] == _PREVIEW_SEED
and row["ref_audio_path"] == _profile_audio_filename(row["id"])
and row["instruct"] == a["instruct"]
and row["language"] == a["language"]
and row["ref_text"] == a["sample_script"]
and states_match
and not row["is_locked"]
and not row["verified_own_voice"]
)
def _existing_archetype_profile(conn, a: dict):
rows = conn.execute(
"SELECT * FROM voice_profiles WHERE personality=? ORDER BY created_at, id",
(_archetype_personality(a),),
).fetchall()
owned = next((row for row in rows if _is_materialized_archetype_row(row, a)), None)
return owned if owned is not None else _legacy_archetype_profile(conn, a)
async def _render_profile_audio(
a: dict, profile_id: str, *, publish: bool = True,
) -> tuple[str, Path]:
"""Render one validated sample, optionally staging it for a later CAS."""
audio_filename = _profile_audio_filename(profile_id)
safe_id = Path(audio_filename).stem
audio_path = Path(VOICES_DIR) / audio_filename
if publish:
await _render_wav_atomic(a, audio_path, prefix=f".{safe_id}-")
else:
audio_path.parent.mkdir(parents=True, exist_ok=True)
audio_path = audio_path.parent / f".{safe_id}-{uuid.uuid4().hex}.staged.wav"
try:
await _render_archetype_wav(a, audio_path)
if not is_playable_wav(audio_path):
raise RuntimeError("the voice engine produced an invalid WAV")
except BaseException:
with __import__("contextlib").suppress(OSError):
audio_path.unlink()
raise
return audio_filename, audio_path
async def _render_wav_atomic(a: dict, out_path: Path, *, prefix: str = ".render-") -> Path:
"""Render and validate a WAV before atomically replacing *out_path*."""
audio_path = Path(out_path)
audio_path.parent.mkdir(parents=True, exist_ok=True)
tmp_path = audio_path.parent / f"{prefix}{uuid.uuid4().hex}.wav"
try:
await _render_archetype_wav(a, tmp_path)
if not is_playable_wav(tmp_path):
raise RuntimeError("the voice engine produced an invalid WAV")
os.replace(tmp_path, audio_path)
finally:
with __import__("contextlib").suppress(OSError):
tmp_path.unlink()
return audio_path
def _heal_materialized_profile(conn, row, a: dict, audio_filename: str) -> None:
"""Repair profiles created before archetype `/use` persisted design kind."""
instruct, vd_states = _design_profile_values(a)
conn.execute(
"UPDATE voice_profiles SET kind='design', instruct=?, vd_states=?, language=?, "
"ref_text=?, seed=?, ref_audio_path=?, personality=? WHERE id=?",
(
instruct, vd_states, a["language"], a["sample_script"], _PREVIEW_SEED,
audio_filename, _archetype_personality(a), row["id"],
),
)
# A non-empty script is always required — synthesizing empty text yields
# silence. Every archetype carries a use-case script, but guard the render path
# too so a malformed archetype can never drive a blank render.
@@ -255,7 +405,7 @@ def _preview_source(a: dict) -> tuple[str, str]:
"Pre-rendered preview from the voice gallery — a fixed reference "
"rendering, not a render from your current engine."
)
if (_PREVIEW_DIR / f"{key}.wav").exists():
if is_playable_wav(_PREVIEW_DIR / f"{key}.wav"):
return "cached", ""
if _no_voice_model_downloaded():
return "no_model", (
@@ -388,9 +538,9 @@ async def preview_archetype(
)
cache_path = _PREVIEW_DIR / f"{key}.wav"
if not cache_path.exists():
if not is_playable_wav(cache_path):
try:
await _render_archetype_wav(a, cache_path)
await _render_wav_atomic(a, cache_path, prefix=".preview-")
except Exception as e: # model missing / OOM / inference failure
logger.error("Archetype preview render failed", exc_info=True)
# Two different failures, two different answers. Without a model
@@ -442,70 +592,122 @@ async def use_archetype(archetype_id: str, name: Optional[str] = Query(None)):
# Idempotent (dedup): an archetype materializes to exactly ONE voice profile.
# Picking the same gallery voice again — from any picker (Gallery grid,
# VoiceSelector, …) — must reuse that one row instead of rendering + inserting
# a fresh duplicate every time. The `personality` column already carries the
# source archetype id (stamped by the INSERT below), so it's the natural
# dedup key; the expensive render + INSERT only run on first use.
# a fresh duplicate every time. Use a namespaced personality identity so an
# imported persona cannot collide with and be rewritten by an archetype id.
with db_conn() as conn:
existing = conn.execute(
"SELECT id, name FROM voice_profiles WHERE personality = ? LIMIT 1",
(a["id"],),
).fetchone()
existing = _existing_archetype_profile(conn, a)
profile_id = existing["id"] if existing is not None else str(uuid.uuid4())[:8]
audio_path: Optional[Path] = None
if existing is not None and _materialized_audio_is_current(existing, a):
audio_filename = existing["ref_audio_path"]
else:
try:
audio_filename, audio_path = await _render_profile_audio(
a, profile_id, publish=existing is None,
)
except Exception as e:
logger.error("Archetype 'use' render failed", exc_info=True)
# Same actionable/diagnostic split as /preview — minus the gallery
# suggestion, which cannot help here.
if _no_voice_model_downloaded():
detail = (
"Creating a voice needs the voice model — no voice model is "
"downloaded yet. Model Catalogue → Models → Download."
)
else:
detail = (
"Couldn't create a voice from this archetype — the voice engine "
f"reported: {e}"
)
raise HTTPException(status_code=503, detail=detail) from e
if existing is not None:
return {"profile_id": existing["id"], "name": existing["name"]}
profile_id = str(uuid.uuid4())[:8]
audio_filename = f"{profile_id}.wav"
audio_path = Path(VOICES_DIR) / audio_filename
try:
await _render_archetype_wav(a, audio_path)
except Exception as e:
logger.error("Archetype 'use' render failed", exc_info=True)
# Same actionable/diagnostic split as /preview — minus the gallery
# suggestion, which cannot help here.
if _no_voice_model_downloaded():
detail = (
"Creating a voice needs the voice model — no voice model is "
"downloaded yet. Model Catalogue → Models → Download."
with db_conn() as conn:
conn.execute("BEGIN IMMEDIATE")
current = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (existing["id"],),
).fetchone()
owned = _existing_archetype_profile(conn, a)
still_owned = current is not None and (
owned is not None and owned["id"] == current["id"]
)
if still_owned:
if audio_path is not None:
destination = Path(VOICES_DIR) / audio_filename
os.replace(audio_path, destination)
audio_path = None
_heal_materialized_profile(conn, current, a, audio_filename)
existing_result = {"profile_id": current["id"], "name": current["name"]}
else:
existing_result = None
if existing_result is not None:
event_bus.emit("profiles", {"action": "updated", "id": existing_result["profile_id"]})
return existing_result
# The row was edited/deleted while rendering. Preserve it and use the
# validated staged sample for a fresh canonical materialization.
profile_id = str(uuid.uuid4())[:8]
audio_filename = _profile_audio_filename(profile_id)
destination = Path(VOICES_DIR) / audio_filename
if audio_path is None:
try:
audio_filename, audio_path = await _render_profile_audio(a, profile_id)
except Exception as e:
raise HTTPException(
status_code=503, detail="Couldn't create a voice from this archetype.",
) from e
else:
detail = (
"Couldn't create a voice from this archetype — the voice engine "
f"reported: {e}"
)
raise HTTPException(status_code=503, detail=detail)
os.replace(audio_path, destination)
audio_path = destination
if audio_path is None: # defensive: a new profile always rendered above
raise RuntimeError("new archetype profile has no rendered audio")
profile_name = (name or a["name"]).strip() or a["name"]
try:
with db_conn() as conn:
conn.execute("BEGIN IMMEDIATE")
# Re-check under the write connection right before inserting: a
# concurrent /use for the same archetype may have inserted while we
# were rendering (the pre-render SELECT above raced). Reuse that row
# and drop our just-rendered sample instead of creating a duplicate.
# (personality is NOT globally unique — marketplace/persona imports
# reuse the column — so a UNIQUE index isn't an option; this closes
# the realistic window for the single-user desktop app.)
dup = conn.execute(
"SELECT id, name FROM voice_profiles WHERE personality = ? LIMIT 1",
(a["id"],),
).fetchone()
# `personality` is not globally UNIQUE, so serialize and re-check.
dup = _existing_archetype_profile(conn, a)
if dup is not None:
duplicate_audio = dup["ref_audio_path"]
if not _materialized_audio_is_current(dup, a):
duplicate_audio = _profile_audio_filename(dup["id"])
_duplicate_path = Path(VOICES_DIR) / duplicate_audio
_duplicate_path.parent.mkdir(parents=True, exist_ok=True)
os.replace(audio_path, _duplicate_path)
audio_path = None
_heal_materialized_profile(conn, dup, a, duplicate_audio)
with __import__("contextlib").suppress(OSError):
os.remove(audio_path)
return {"profile_id": dup["id"], "name": dup["name"]}
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, created_at) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
(
profile_id, profile_name, audio_filename, a["sample_script"],
a["instruct"], a["language"], _PREVIEW_SEED, a["id"], time.time(),
),
)
if audio_path is not None:
os.remove(audio_path)
duplicate_result = {"profile_id": dup["id"], "name": dup["name"]}
else:
duplicate_result = None
if duplicate_result is None:
instruct, vd_states = _design_profile_values(a)
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"created_at, kind, vd_states) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, 'design', ?)",
(
profile_id, profile_name, audio_filename, a["sample_script"],
instruct, a["language"], _PREVIEW_SEED,
_archetype_personality(a), time.time(), vd_states,
),
)
except Exception:
with __import__("contextlib").suppress(OSError):
os.remove(audio_path)
if audio_path is not None:
os.remove(audio_path)
raise
if duplicate_result is not None:
event_bus.emit("profiles", {"action": "updated", "id": duplicate_result["profile_id"]})
return duplicate_result
event_bus.emit("profiles", {"action": "created", "id": profile_id})
return {"profile_id": profile_id, "name": profile_name}
+231
View File
@@ -0,0 +1,231 @@
"""Short-lived credentials for the first-party remote administration UI."""
from __future__ import annotations
import math
import threading
import time
from collections import OrderedDict, deque
from collections.abc import Callable
from datetime import UTC, datetime
from typing import Literal
from fastapi import APIRouter, HTTPException, Request, Response
from fastapi.responses import JSONResponse
from pydantic import BaseModel
from core.auth import (
CredentialTransport,
PrincipalKind,
authorization_credential_present,
legacy_master_cookie_valid,
master_header_valid,
principal_for,
remote_api_key,
)
from core.csrf import cookie_csrf_allowed, effective_scheme
from services.admin_sessions import (
SESSION_TTL_SECONDS,
WS_TICKET_TTL_SECONDS,
admin_session_store,
)
router = APIRouter(prefix="/api/auth", tags=["auth"])
_FAILED_EXCHANGE_LIMIT = 10
_FAILED_EXCHANGE_WINDOW_SECONDS = 60
_MAX_TRACKED_CLIENTS = 1024
class _ExchangeAttemptLimiter:
"""Bounded per-client sliding window for failed pre-auth exchanges."""
def __init__(
self,
*,
monotonic: Callable[[], float] = time.monotonic,
limit: int = _FAILED_EXCHANGE_LIMIT,
window_seconds: int = _FAILED_EXCHANGE_WINDOW_SECONDS,
max_clients: int = _MAX_TRACKED_CLIENTS,
) -> None:
if limit <= 0 or window_seconds <= 0 or max_clients <= 0:
raise ValueError("rate-limit bounds must be positive")
self._monotonic = monotonic
self._limit = limit
self._window_seconds = window_seconds
self._max_clients = max_clients
self._attempts: OrderedDict[str, deque[float]] = OrderedDict()
self._lock = threading.Lock()
def register_failure(self, client_id: str) -> int | None:
now = self._monotonic()
cutoff = now - self._window_seconds
with self._lock:
failures = self._attempts.setdefault(client_id, deque())
while failures and failures[0] <= cutoff:
failures.popleft()
self._attempts.move_to_end(client_id)
while len(self._attempts) > self._max_clients:
self._attempts.popitem(last=False)
if len(failures) >= self._limit:
return max(
1,
math.ceil(self._window_seconds - (now - failures[0])),
)
failures.append(now)
return None
def clear(self, client_id: str) -> None:
with self._lock:
self._attempts.pop(client_id, None)
def reset(self) -> None:
with self._lock:
self._attempts.clear()
_exchange_attempt_limiter = _ExchangeAttemptLimiter()
class SessionRequest(BaseModel):
transport: Literal["cookie", "bearer"]
class WebSocketTicketRequest(BaseModel):
path: str
def _secure_cookie(request: Request) -> bool:
# Same effective-scheme logic as the exact-origin CSRF check: the resolved
# scope first (uvicorn's trusted-proxy rewrite), upgraded — never
# downgraded — by X-Forwarded-Proto for TLS-terminating proxies uvicorn
# doesn't trust (Tailscale Serve into Docker, etc.). Spoofing the header on
# a plain-http hop can only ADD the Secure flag, which fails safe: the
# browser drops such a cookie, so the spoofer only breaks their own
# session. See core.csrf.effective_scheme for the full analysis.
return effective_scheme(request) == "https"
def _set_session_cookie(response: Response, request: Request, token: str, expires_at: float) -> None:
response.set_cookie(
"ov_session",
token,
max_age=SESSION_TTL_SECONDS,
expires=datetime.fromtimestamp(expires_at, tz=UTC),
path="/",
secure=_secure_cookie(request),
httponly=True,
samesite="strict",
)
def _expire_cookie(response: Response, request: Request, name: str) -> None:
response.delete_cookie(
name,
path="/",
secure=_secure_cookie(request),
httponly=name == "ov_session",
samesite="strict",
)
def _client_id(request: Request) -> str:
host = request.client.host if request.client else "unknown"
return str(host).strip().lower()[:255] or "unknown"
def _reject_master_exchange(request: Request) -> None:
retry_after = _exchange_attempt_limiter.register_failure(_client_id(request))
if retry_after is not None:
raise HTTPException(
status_code=429,
detail="Too many authentication attempts",
headers={"Retry-After": str(retry_after)},
)
raise HTTPException(status_code=401, detail="API key required")
@router.post("/session")
def create_session(payload: SessionRequest, request: Request) -> Response:
configured = remote_api_key()
if not configured:
raise HTTPException(status_code=401, detail="API key required")
authorization_present = authorization_credential_present(request)
header_authorized = master_header_valid(request)
legacy_authorized = legacy_master_cookie_valid(request)
migrating_legacy = False
if authorization_present:
if not header_authorized:
_reject_master_exchange(request)
elif legacy_authorized:
if payload.transport != "cookie" or not cookie_csrf_allowed(request):
raise HTTPException(status_code=403, detail="browser origin rejected")
migrating_legacy = True
else:
_reject_master_exchange(request)
_exchange_attempt_limiter.clear(_client_id(request))
issued = admin_session_store.issue(configured)
if payload.transport == "bearer":
return JSONResponse(
{
"token": issued.token,
"expires_at": issued.expires_at,
"expires_in": SESSION_TTL_SECONDS,
},
status_code=201,
)
response = Response(status_code=204)
_set_session_cookie(response, request, issued.token, issued.expires_at)
if migrating_legacy or request.cookies.get("ov_key"):
_expire_cookie(response, request, "ov_key")
return response
@router.delete("/session", status_code=204)
def delete_session(request: Request) -> Response:
principal = principal_for(request)
if principal.kind is PrincipalKind.ADMIN_SESSION:
if (
principal.transport is CredentialTransport.COOKIE
and not cookie_csrf_allowed(request)
):
raise HTTPException(status_code=403, detail="browser origin rejected")
admin_session_store.revoke_by_credential(principal.credential_id)
response = Response(status_code=204)
_expire_cookie(response, request, "ov_session")
return response
@router.post("/ws-ticket")
def create_ws_ticket(payload: WebSocketTicketRequest, request: Request) -> JSONResponse:
principal = principal_for(request)
if principal.kind is not PrincipalKind.ADMIN_SESSION:
raise HTTPException(status_code=403, detail="admin session required")
if (
principal.transport is CredentialTransport.COOKIE
and not cookie_csrf_allowed(request)
):
raise HTTPException(status_code=403, detail="browser origin rejected")
try:
ticket = admin_session_store.issue_ws_ticket_for_credential(
principal.credential_id,
payload.path,
remote_api_key(),
)
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from None
except PermissionError:
raise HTTPException(status_code=401, detail="admin session required") from None
return JSONResponse(
{
"ticket": ticket.token,
"expires_at": ticket.expires_at,
"expires_in": WS_TICKET_TTL_SECONDS,
},
status_code=201,
)
+2 -2
View File
@@ -153,8 +153,8 @@ def _select_sherpa_spec(websocket: WebSocket):
async def ws_transcribe(websocket: WebSocket):
"""Stream audio in, get partial + final transcription out."""
# Loopback origin guard — refuse anything not from 127.0.0.1, ::1, or
# localhost. HTTP routers use Depends(require_loopback) at router level;
# WebSocket dependency injection differs across FastAPI versions, so we
# localhost. Privileged HTTP routers use Depends(require_admin) at router
# level; WebSocket dependency injection differs across FastAPI versions, so we
# inline the check before accept(). Without it, any local process could
# stream the user's microphone over this endpoint.
# Wave 2.3 (remote backend): a non-loopback client that presents the
+728 -99
View File
@@ -20,18 +20,26 @@ Design / safety
from __future__ import annotations
import asyncio
import contextlib
import hashlib
import json
import logging
import os
import re
import shutil
import tempfile
import time
import uuid
from pathlib import Path
from typing import Optional
from urllib.parse import urlparse
from urllib.parse import urljoin, urlparse
from fastapi import APIRouter, HTTPException, Query
from fastapi.responses import FileResponse
from core import archetypes
from core.config import DATA_DIR
from core.audio_validation import is_playable_wav, resolve_regular_file
from core.config import DATA_DIR, VOICES_DIR
logger = logging.getLogger("omnivoice.community")
router = APIRouter()
@@ -42,9 +50,32 @@ _ALLOWED_AUDIO_HOSTS = {
"cdn.jsdelivr.net", "github.com", "raw.githubusercontent.com",
"objects.githubusercontent.com", "release-assets.githubusercontent.com",
}
_ALLOWED_MANIFEST_HOSTS = {"cdn.jsdelivr.net"}
_VALID_TOKENS = set(archetypes._VD._INSTRUCT_ALL_VALID)
_USE_CASE_IDS = {c["id"] for c in archetypes.USE_CASES}
_SOURCE_RE = re.compile(r"^[A-Za-z0-9._-]+/[A-Za-z0-9._-]+$") # owner/repo only
_SOURCE_RE = re.compile(
r"^[A-Za-z0-9._-]{1,100}/[A-Za-z0-9._-]{1,100}$",
) # owner/repo only
_ITEM_ID_RE = re.compile(r"^[A-Za-z0-9_-]{1,128}$")
_SHA256_RE = re.compile(r"^[0-9a-f]{64}$")
# A gallery open may touch this loader several times (grid, preview, use). Keep
# a successful response for six hours, then revalidate it once. On a network
# failure the readable stale copy remains usable and its check time advances,
# preventing every offline gallery open from waiting through the same timeout.
_MANIFEST_MAX_AGE_S = 6 * 60 * 60
_MAX_MANIFEST_BYTES = 4 << 20
_MAX_SAMPLE_SCRIPT_CHARS = 2_000
_MAX_REF_TEXT_CHARS = 4_000
# Community voice submissions are documented as short clean WAV clips. The cap
# comfortably covers 15 s of uncompressed 96 kHz stereo PCM while preventing a
# remote manifest from turning Preview into an unbounded disk/memory download.
_MAX_VOICE_AUDIO_BYTES = 32 << 20
_ATTR_NAMES = (
"Gender", "Age", "Pitch", "Style", "EnglishAccent", "ChineseDialect",
)
# ── Config: which content repos to load ───────────────────────────────────────
@@ -52,14 +83,18 @@ def configured_sources() -> list[str]:
"""Gallery sources, in priority order. Env var > config file > default."""
env = os.environ.get("OMNIVOICE_GALLERY_SOURCES")
if env:
return [s.strip() for s in env.split(",") if s.strip()]
sources = [s.strip() for s in env.split(",")]
valid = [s for s in sources if _SOURCE_RE.fullmatch(s)]
return valid or list(_DEFAULT_SOURCES)
cfg = Path(DATA_DIR) / "gallery_sources.json"
if cfg.exists():
try:
data = json.loads(cfg.read_text(encoding="utf-8"))
srcs = data.get("sources")
if isinstance(srcs, list) and srcs:
return [str(s) for s in srcs]
valid = [s for s in srcs if isinstance(s, str) and _SOURCE_RE.fullmatch(s)]
if valid:
return valid
except Exception:
logger.warning("gallery_sources.json unreadable; using default")
return list(_DEFAULT_SOURCES)
@@ -81,9 +116,51 @@ def _safe_audio_url(url: str) -> bool:
return False
def _safe_manifest_url(url: str) -> bool:
try:
parsed = urlparse(url or "")
return parsed.scheme == "https" and parsed.hostname in _ALLOWED_MANIFEST_HOSTS
except Exception:
return False
def normalize_preset_instruct(instruct: str) -> Optional[tuple[str, dict]]:
"""Normalize one validator-safe tag per design category.
Membership in the vocabulary is not enough: ``male, female`` contains two
individually valid tokens but the engine rejects the pair as conflicting.
Build the frontend's full ``vd_states`` shape at this trust boundary too,
so Magic Wand never inherits stale sliders from the previous voice.
"""
attrs = {name: "Auto" for name in _ATTR_NAMES}
normalized: list[str] = []
seen_categories: set[int] = set()
for raw in re.split("[," + chr(0xFF0C) + "]", str(instruct or "")):
token = raw.strip().lower()
if not token or token not in _VALID_TOKENS:
return None
category = archetypes._VD._instruct_category_index(token)
if category < 0 or category in seen_categories:
return None
seen_categories.add(category)
# The picker represents the universal gender/age/pitch/style axes in
# English even for Chinese speech; dialect remains Chinese-only.
canonical = archetypes._VD._INSTRUCT_ZH_TO_EN.get(token, token)
attrs[_ATTR_NAMES[category]] = canonical
normalized.append(canonical)
if not normalized:
return None
# Accent and Chinese dialect are separate taxonomy buckets but the engine
# deliberately forbids mixing them in a single design.
if 4 in seen_categories and 5 in seen_categories:
return None
return ", ".join(normalized), attrs
def is_valid_instruct(instruct: str) -> bool:
toks = [t.strip() for t in (instruct or "").split(",") if t.strip()]
return bool(toks) and all(t in _VALID_TOKENS for t in toks)
return normalize_preset_instruct(instruct) is not None
def validate_item(raw: dict) -> Optional[dict]:
@@ -93,62 +170,203 @@ def validate_item(raw: dict) -> Optional[dict]:
it = dict(raw)
if it.get("type") not in ("preset", "voice"):
return None
if not it.get("id") or not it.get("name"):
if not isinstance(it.get("id"), str) or not _ITEM_ID_RE.fullmatch(it["id"]):
return None
if not isinstance(it.get("name"), str) or not it["name"].strip():
return None
it["name"] = it["name"].strip()[:80]
if it.get("use_case") not in _USE_CASE_IDS:
return None
if it["type"] == "preset" and not is_valid_instruct(it.get("instruct", "")):
return None # would crash synthesis — drop it
if it["type"] == "voice" and not _safe_audio_url((it.get("audio") or {}).get("url", "")):
return None
it.setdefault("facets", {})
raw_facets = it.get("facets")
if not isinstance(raw_facets, dict):
raw_facets = {}
language = it.get("language")
if not isinstance(language, str) or not language.strip():
language = raw_facets.get("lang", "English")
it["language"] = language.strip() if isinstance(language, str) and language.strip() else "English"
facets = dict(raw_facets)
if it["type"] == "preset":
normalized = normalize_preset_instruct(it.get("instruct", ""))
if normalized is None:
return None # unknown/conflicting tokens would crash synthesis
it["instruct"], it["attrs"] = normalized
attrs = it["attrs"]
facets.update({
"gender": None if attrs["Gender"] == "Auto" else attrs["Gender"],
"age": None if attrs["Age"] == "Auto" else attrs["Age"],
"pitch": None if attrs["Pitch"] == "Auto" else attrs["Pitch"],
"accent": None if attrs["EnglishAccent"] == "Auto" else attrs["EnglishAccent"],
"whisper": attrs["Style"] == "whisper",
"lang": it["language"],
})
sample_script = it.get("sample_script")
it["sample_script"] = (
sample_script.strip()[:_MAX_SAMPLE_SCRIPT_CHARS]
if isinstance(sample_script, str) else ""
)
else:
audio = it.get("audio")
if not isinstance(audio, dict) or not _safe_audio_url(audio.get("url", "")):
return None
expected = audio.get("sha256")
if expected is not None:
expected = str(expected).lower()
if not _SHA256_RE.fullmatch(expected):
return None
audio = {**audio, "sha256": expected}
ref_text = audio.get("ref_text")
audio = {
**audio,
"ref_text": (
ref_text.strip()[:_MAX_REF_TEXT_CHARS]
if isinstance(ref_text, str) else ""
),
}
it["audio"] = audio
facets.setdefault("gender", None)
facets.setdefault("age", None)
facets.setdefault("pitch", None)
facets.setdefault("accent", None)
facets.setdefault("whisper", False)
facets.setdefault("lang", it["language"])
it["facets"] = facets
it.setdefault("icon", archetypes._USE_ICON.get(it["use_case"], "Sparkles"))
it.setdefault("language", it.get("facets", {}).get("lang", "English"))
it["is_community"] = it.get("source") != "starter"
it["preview_url"] = f"/community/items/{it['id']}/preview"
return it
def _merge(manifests: list[tuple[str, Optional[dict]]]) -> tuple[list, list]:
items, packs, seen = [], [], set()
for src, m in manifests:
if not m:
if not isinstance(m, dict):
continue
for raw in (m.get("items") or []):
raw_items = m.get("items")
for raw in raw_items if isinstance(raw_items, list) else []:
v = validate_item(raw)
if v and v["id"] not in seen:
v["_source_repo"] = src
seen.add(v["id"])
items.append(v)
for p in (m.get("packs") or []):
raw_packs = m.get("packs")
for p in raw_packs if isinstance(raw_packs, list) else []:
if isinstance(p, dict):
packs.append({**p, "_source_repo": src})
return items, packs
def _fetch_manifest(source: str, refresh: bool) -> Optional[dict]:
"""Return a source's manifest from cache, or fetch + cache it. None if both fail."""
cache = _cache_path(source)
if not refresh and cache.exists():
try:
return json.loads(cache.read_text(encoding="utf-8"))
except Exception:
pass
def _read_manifest_cache(cache: Path) -> Optional[dict]:
try:
import httpx
with httpx.Client(timeout=15.0, follow_redirects=True) as client:
resp = client.get(_manifest_url(source))
resp.raise_for_status()
data = resp.json()
cache.parent.mkdir(parents=True, exist_ok=True)
cache.write_text(json.dumps(data), encoding="utf-8")
if cache.stat().st_size > _MAX_MANIFEST_BYTES:
return None
data = json.loads(cache.read_text(encoding="utf-8"))
return data if isinstance(data, dict) else None
except (OSError, ValueError, TypeError):
return None
def _write_bytes_atomic(path: Path, data: bytes) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
fd, tmp = tempfile.mkstemp(dir=str(path.parent), prefix=f".{path.name}-", suffix=".part")
try:
with os.fdopen(fd, "wb") as handle:
handle.write(data)
handle.flush()
os.fsync(handle.fileno())
os.replace(tmp, path)
except BaseException:
with contextlib.suppress(OSError):
os.unlink(tmp)
raise
def _fetch_remote_manifest(source: str, *, client=None) -> dict:
"""Fetch one bounded manifest, validating every redirect before request."""
import httpx
if not _SOURCE_RE.fullmatch(source or ""):
raise ValueError("invalid gallery source")
owned_client = client is None
http = client or httpx.Client(timeout=15.0, follow_redirects=False)
current_url = _manifest_url(source)
payload = bytearray()
try:
fetched = False
for _redirect in range(6):
if not _safe_manifest_url(current_url):
raise ValueError("gallery manifest URL is not from an allowed host")
with http.stream("GET", current_url, follow_redirects=False) as response:
if response.status_code in (301, 302, 303, 307, 308):
location = response.headers.get("location")
next_url = urljoin(current_url, location or "")
if not location or not _safe_manifest_url(next_url):
raise ValueError("gallery manifest redirected to a disallowed host")
current_url = next_url
continue
response.raise_for_status()
length = response.headers.get("content-length")
if length:
try:
declared_length = int(length)
except ValueError:
declared_length = None
if declared_length is not None and declared_length > _MAX_MANIFEST_BYTES:
raise ValueError("gallery manifest exceeded the size limit")
for chunk in response.iter_bytes():
if not chunk:
continue
if len(payload) + len(chunk) > _MAX_MANIFEST_BYTES:
raise ValueError("gallery manifest exceeded the size limit")
payload.extend(chunk)
fetched = True
break
if not fetched:
raise ValueError("gallery manifest followed too many redirects")
finally:
if owned_client:
http.close()
if not payload:
raise ValueError("gallery manifest was empty")
data = json.loads(payload)
if not isinstance(data, dict):
raise ValueError("gallery manifest is not a JSON object")
return data
def _fetch_manifest(
source: str, refresh: bool, *, now: Optional[float] = None,
) -> Optional[dict]:
"""Return a fresh manifest, with a throttled stale-cache offline fallback."""
cache = _cache_path(source)
cached = _read_manifest_cache(cache)
checked_at = time.time() if now is None else float(now)
if not refresh and cached is not None:
try:
if checked_at - cache.stat().st_mtime < _MANIFEST_MAX_AGE_S:
return cached
except OSError:
pass # treat a stat race as stale and try the source once
try:
data = _fetch_remote_manifest(source)
encoded = json.dumps(
data, ensure_ascii=False, separators=(",", ":"),
).encode("utf-8")
if len(encoded) > _MAX_MANIFEST_BYTES:
raise ValueError("gallery manifest exceeded the cache size limit")
_write_bytes_atomic(cache, encoded)
# Tests inject their own clock; production's value equals wall time.
os.utime(cache, (checked_at, checked_at))
return data
except Exception as e: # offline / 404 / bad json
logger.warning("manifest fetch failed for %s: %s", source, e)
if cache.exists():
try:
return json.loads(cache.read_text(encoding="utf-8"))
except Exception:
pass
if cached is not None:
# This mtime is a last-*check* marker. Advancing it on failure keeps
# an offline app responsive while guaranteeing another check after
# the bounded freshness interval.
with contextlib.suppress(OSError):
os.utime(cache, (checked_at, checked_at))
return cached
return None
@@ -214,6 +432,385 @@ def community_submit_url(item_type: str = Query("preset", alias="type"), source:
return {"url": f"https://github.com/{src}/issues/new?template={template}"}
def _find_item(items: list[dict], item_id: str) -> dict:
if not _ITEM_ID_RE.fullmatch(item_id or ""):
raise HTTPException(status_code=404, detail="Item not found in the gallery.")
item = next((it for it in items if it["id"] == item_id), None)
if item is None:
raise HTTPException(status_code=404, detail="Item not found in the gallery.")
return item
def _canonical_archetype(item: dict) -> Optional[dict]:
"""The built-in archetype represented exactly by a marketplace preset."""
if item.get("type") != "preset":
return None
canonical = archetypes.get_archetype(item["id"])
if canonical is None:
return None
if (canonical.get("instruct") != item.get("instruct")
or canonical.get("language") != item.get("language")):
return None
remote_script = (item.get("sample_script") or "").strip()
if remote_script and remote_script != (canonical.get("sample_script") or "").strip():
return None
return canonical
def _preset_preview_path(item: dict) -> Path:
fingerprint = hashlib.sha256(
json.dumps({
"instruct": item.get("instruct"),
"language": item.get("language"),
"sample_script": item.get("sample_script"),
}, sort_keys=True).encode("utf-8")
).hexdigest()[:16]
return _CACHE_DIR / "previews" / f"{item['id']}-{fingerprint}.wav"
def _voice_audio_fingerprint(item: dict) -> str:
audio = item.get("audio") or {}
return hashlib.sha256(
f"{audio.get('url', '')}|{audio.get('sha256', '')}".encode("utf-8")
).hexdigest()[:16]
def _voice_audio_path(item: dict) -> Path:
return _CACHE_DIR / "audio" / f"{item['id']}-{_voice_audio_fingerprint(item)}.wav"
async def _render_preset_atomic(item: dict, out_path: Path) -> Path:
if is_playable_wav(out_path):
return out_path
from api.routers.archetypes import _render_archetype_wav
out_path.parent.mkdir(parents=True, exist_ok=True)
fd, tmp_name = tempfile.mkstemp(dir=str(out_path.parent), prefix=".preview-", suffix=".wav")
os.close(fd)
tmp = Path(tmp_name)
try:
await _render_archetype_wav({
"instruct": item["instruct"],
"language": item.get("language", "English"),
"sample_script": (
(item.get("sample_script") or "").strip()
or "Hello — this is a preview of this voice."
),
}, tmp)
if not is_playable_wav(tmp):
raise RuntimeError("the voice engine produced an invalid preview WAV")
os.replace(tmp, out_path)
return out_path
finally:
with contextlib.suppress(OSError):
tmp.unlink()
def _download_voice_audio(item: dict, out_path: Path, *, client=None) -> None:
"""Stream one allow-listed voice clip into an atomic, size-bounded file."""
audio = item.get("audio") or {}
url = audio.get("url", "")
if not _safe_audio_url(url):
raise HTTPException(status_code=400, detail="Voice audio URL is not from an allowed host.")
import httpx
owned_client = client is None
http = client or httpx.Client(timeout=30.0, follow_redirects=False)
out_path.parent.mkdir(parents=True, exist_ok=True)
fd, tmp_name = tempfile.mkstemp(dir=str(out_path.parent), prefix=".voice-", suffix=".part")
total = 0
digest = hashlib.sha256()
try:
with os.fdopen(fd, "wb") as handle:
current_url = url
downloaded = False
for _redirect in range(6):
with http.stream("GET", current_url, follow_redirects=False) as response:
if response.status_code in (301, 302, 303, 307, 308):
location = response.headers.get("location")
next_url = urljoin(current_url, location or "")
if not location or not _safe_audio_url(next_url):
raise HTTPException(
status_code=502,
detail="Community voice audio redirected to a disallowed host.",
)
current_url = next_url
continue
response.raise_for_status()
length = response.headers.get("content-length")
if length:
try:
if int(length) > _MAX_VOICE_AUDIO_BYTES:
raise HTTPException(
status_code=502,
detail="Community voice audio exceeded the download size limit.",
)
except ValueError:
# A non-numeric Content-Length header is the
# server's problem, not a reason to refuse the
# download — the streamed byte counter below
# still enforces the same cap on what actually
# arrives.
pass
for chunk in response.iter_bytes():
if not chunk:
continue
total += len(chunk)
if total > _MAX_VOICE_AUDIO_BYTES:
raise HTTPException(
status_code=502,
detail="Community voice audio exceeded the download size limit.",
)
digest.update(chunk)
handle.write(chunk)
downloaded = True
break
if not downloaded:
raise HTTPException(
status_code=502,
detail="Community voice audio followed too many redirects.",
)
if total == 0:
raise HTTPException(status_code=502, detail="Community voice audio was empty.")
expected = audio.get("sha256")
if expected and digest.hexdigest() != expected:
raise HTTPException(
status_code=502,
detail="Downloaded voice failed its integrity check.",
)
handle.flush()
os.fsync(handle.fileno())
if not is_playable_wav(Path(tmp_name)):
raise HTTPException(
status_code=502, detail="Community voice audio was not a valid WAV.",
)
os.replace(tmp_name, out_path)
except BaseException:
with contextlib.suppress(OSError):
os.unlink(tmp_name)
raise
finally:
if owned_client:
http.close()
def _cached_voice_audio(item: dict) -> Path:
path = _voice_audio_path(item)
if is_playable_wav(path):
return path
with contextlib.suppress(OSError):
path.unlink()
_download_voice_audio(item, path)
return path
def _copy_atomic(source: Path, destination: Path) -> None:
destination.parent.mkdir(parents=True, exist_ok=True)
fd, tmp_name = tempfile.mkstemp(
dir=str(destination.parent), prefix=f".{destination.name}-", suffix=".part",
)
try:
with os.fdopen(fd, "wb") as out, source.open("rb") as src:
shutil.copyfileobj(src, out)
out.flush()
os.fsync(out.fileno())
os.replace(tmp_name, destination)
except BaseException:
with contextlib.suppress(OSError):
os.unlink(tmp_name)
raise
@router.get("/community/items/{item_id}/preview")
async def community_preview(
item_id: str,
local: bool = Query(False, description="Bypass canonical gallery audio after decode failure"),
):
"""Serve every community preview through the authenticated same-origin API."""
_, items, _, _ = await asyncio.to_thread(_load, False)
item = _find_item(items, item_id)
canonical = _canonical_archetype(item)
if canonical is not None:
# Reuse the signed-gallery/local-render fallback and cache owned by the
# canonical endpoint rather than synthesizing the same preset twice.
# Delegate in-process: a root-relative HTTP redirect drops supported
# reverse-proxy path prefixes such as ``https://host/api``.
from api.routers.archetypes import preview_archetype
return await preview_archetype(canonical["id"], local=local)
try:
if item["type"] == "preset":
path = await _render_preset_atomic(item, _preset_preview_path(item))
else:
path = await asyncio.to_thread(_cached_voice_audio, item)
except HTTPException:
raise
except Exception as exc:
logger.warning("Community preview unavailable (%s)", type(exc).__name__)
raise HTTPException(
status_code=503, detail="This community voice preview is unavailable right now.",
) from exc
return FileResponse(
path, media_type="audio/wav",
headers={"Cache-Control": "no-cache", "X-OmniVoice-Preview-Source": "community"},
)
def _profile_fields(item: dict) -> tuple[str, str, Optional[str], Optional[int]]:
if item["type"] == "preset":
return "design", item["instruct"], json.dumps(item["attrs"]), 42
return "clone", "", None, None
def _community_profile_audio_filename(profile_id: str, item: dict) -> str:
safe_id = (
profile_id if re.fullmatch(r"[A-Za-z0-9_-]{1,64}", profile_id or "")
else hashlib.sha256(str(profile_id).encode("utf-8")).hexdigest()[:16]
)
if item["type"] == "voice":
# The manifest URL/checksum fingerprint makes a changed submission
# invalidate its already-materialized clone without a schema change.
return f"{safe_id}-community-{_voice_audio_fingerprint(item)}.wav"
return f"{safe_id}.wav"
def _stored_profile_audio(ref_audio_path: object) -> Optional[Path]:
return resolve_regular_file(VOICES_DIR, ref_audio_path)
def _community_audio_is_current(row, item: dict, ref_text: str) -> bool:
path = _stored_profile_audio(row["ref_audio_path"])
expected_filename = _community_profile_audio_filename(row["id"], item)
if row["ref_audio_path"] != expected_filename or not is_playable_wav(path):
return False
kind, instruct, _vd_states, seed = _profile_fields(item)
inputs_match = (
row["instruct"] == instruct
and row["language"] == item.get("language", "Auto")
and row["ref_text"] == ref_text
and row["seed"] == seed
)
if not inputs_match:
return False
return True
async def _materialize_item_audio(
item: dict, profile_id: str, *, publish: bool = True,
) -> tuple[str, Path]:
"""Copy the current manifest audio, optionally staging it for a later CAS."""
audio_filename = _community_profile_audio_filename(profile_id, item)
destination = Path(VOICES_DIR) / audio_filename
audio_path = destination
if not publish:
destination.parent.mkdir(parents=True, exist_ok=True)
audio_path = destination.parent / f".{Path(audio_filename).stem}-{uuid.uuid4().hex}.staged.wav"
if item["type"] == "preset":
cached = await _render_preset_atomic(item, _preset_preview_path(item))
else:
cached = await asyncio.to_thread(_cached_voice_audio, item)
await asyncio.to_thread(_copy_atomic, cached, audio_path)
return audio_filename, audio_path
def _community_personality(item: dict) -> str:
source = item.get("_source_repo")
if not isinstance(source, str) or not _SOURCE_RE.fullmatch(source):
source = _DEFAULT_SOURCES[0]
return f"community:{source}:{item['id']}"
def _is_materialized_community_row(row, item: dict) -> bool:
if (
row["personality"] != _community_personality(item)
or row["is_locked"] or row["verified_own_voice"]
):
return False
if item["type"] == "voice":
safe_id = Path(_community_profile_audio_filename(row["id"], item)).name.split(
"-community-", 1,
)[0]
return bool(
row["kind"] == "clone"
and row["seed"] is None
and not row["vd_states"]
and row["instruct"] == ""
and row["language"] == item.get("language", "Auto")
and row["ref_text"] == (item.get("audio") or {}).get("ref_text", "")
and re.fullmatch(
rf"{re.escape(safe_id)}-community-[0-9a-f]{{16}}\.wav",
row["ref_audio_path"] or "",
)
)
try:
states = json.loads(row["vd_states"])
except (TypeError, ValueError):
return False
return bool(
row["kind"] == "design"
and row["seed"] == 42
and row["ref_audio_path"] == _community_profile_audio_filename(row["id"], item)
and row["instruct"] == item["instruct"]
and row["language"] == item.get("language", "Auto")
and row["ref_text"] == (item.get("sample_script") or "")
and states == item["attrs"]
)
def _existing_community_profile(conn, item: dict, personality: str):
candidates = conn.execute(
"SELECT * FROM voice_profiles WHERE personality=? ORDER BY created_at, id",
(personality,),
).fetchall()
existing = next(
(row for row in candidates if _is_materialized_community_row(row, item)), None,
)
if existing is not None:
return existing
# Old builds stored the bare item id. Import formats preserve arbitrary
# personality text too, so adopt only the exact shape the old materializer
# wrote; otherwise a remote item id could rewrite a user's imported voice.
if archetypes.get_archetype(item["id"]) is None:
legacy = conn.execute(
"SELECT * FROM voice_profiles WHERE personality=? LIMIT 1",
(item["id"],),
).fetchone()
if legacy is not None:
kind, instruct, _vd_states, _seed = _profile_fields(item)
ref_text = item.get("sample_script") or (item.get("audio") or {}).get(
"ref_text", "",
)
if (
legacy["ref_audio_path"] == f"{legacy['id']}.wav"
and legacy["kind"] == kind
and legacy["instruct"] == instruct
and legacy["language"] == item.get("language", "Auto")
and legacy["ref_text"] == ref_text
and legacy["seed"] is None
and not legacy["vd_states"]
and not legacy["is_locked"]
and not legacy["verified_own_voice"]
):
return legacy
return None
def _heal_existing_profile(
conn, row, item: dict, ref_text: str, personality: str, audio_filename: str,
) -> None:
kind, instruct, vd_states, seed = _profile_fields(item)
conn.execute(
"UPDATE voice_profiles SET kind=?, instruct=?, vd_states=?, language=?, "
"ref_text=?, seed=?, personality=?, ref_audio_path=? WHERE id=?",
(
kind, instruct, vd_states, item.get("language", "Auto"), ref_text,
seed, personality, audio_filename, row["id"],
),
)
@router.post("/community/items/{item_id}/use")
async def community_use(item_id: str, name: Optional[str] = Query(None)):
"""Materialize a community item into a reusable voice profile.
@@ -223,76 +820,108 @@ async def community_use(item_id: str, name: Optional[str] = Query(None)):
``voice_profiles`` row usable everywhere voices are picked.
"""
_, items, _, _ = await asyncio.to_thread(_load, False)
item = next((it for it in items if it["id"] == item_id), None)
if item is None:
raise HTTPException(status_code=404, detail="Item not found in the gallery.")
item = _find_item(items, item_id)
canonical = _canonical_archetype(item)
if canonical is not None:
from api.routers.archetypes import use_archetype
return await use_archetype(canonical["id"], name)
import time
import uuid
from core import event_bus
from core.db import db_conn
from core.config import VOICES_DIR
profile_id = str(uuid.uuid4())[:8]
audio_filename = f"{profile_id}.wav"
audio_path = Path(VOICES_DIR) / audio_filename
profile_name = (name or item["name"]).strip() or item["name"]
instruct = item.get("instruct", "") if item["type"] == "preset" else ""
ref_text = item.get("sample_script") or (item.get("audio") or {}).get("ref_text", "")
personality = _community_personality(item)
with db_conn() as conn:
existing = _existing_community_profile(conn, item, personality)
try:
if item["type"] == "preset":
from api.routers.archetypes import _render_archetype_wav
pseudo = {
"instruct": instruct,
"language": item.get("language", "English"),
"sample_script": ref_text or "Hello — this is a preview of this voice.",
}
await _render_archetype_wav(pseudo, audio_path)
else: # voice — download the reference clip (off the event loop)
await asyncio.to_thread(_download_voice_audio, item, audio_path)
except HTTPException:
raise
except Exception as e:
logger.error("Community 'use' failed", exc_info=True)
raise HTTPException(status_code=503, detail=f"Couldn't add this voice right now. Error: {e}")
try:
# A community "preset" is a synthetic designed voice (rendered from an
# instruct string) → kind='design'; a "voice" carries a real reference
# clip → kind='clone'. Setting kind makes the persona-gallery
# synthetic-only gating work (§R3) instead of defaulting all imports to
# 'clone'.
kind = "design" if item["type"] == "preset" else "clone"
with db_conn() as conn:
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, created_at, kind) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(profile_id, profile_name, audio_filename, ref_text, instruct,
item.get("language", "Auto"), None, item["id"], time.time(), kind),
profile_id = existing["id"] if existing is not None else str(uuid.uuid4())[:8]
audio_path: Optional[Path] = None
if existing is not None and _community_audio_is_current(existing, item, ref_text):
audio_filename = existing["ref_audio_path"]
else:
try:
audio_filename, audio_path = await _materialize_item_audio(
item, profile_id, publish=existing is None,
)
except HTTPException:
raise
except Exception as e:
logger.error("Community 'use' failed", exc_info=True)
raise HTTPException(
status_code=503, detail="Couldn't add this voice right now.",
) from e
if existing is not None:
with db_conn() as conn:
conn.execute("BEGIN IMMEDIATE")
current = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (existing["id"],),
).fetchone()
owned = _existing_community_profile(conn, item, personality)
still_owned = current is not None and (
_is_materialized_community_row(current, item)
or (owned is not None and owned["id"] == current["id"])
)
if still_owned:
if audio_path is not None:
destination = Path(VOICES_DIR) / audio_filename
os.replace(audio_path, destination)
audio_path = None
_heal_existing_profile(
conn, current, item, ref_text, personality, audio_filename,
)
existing_result = {"profile_id": current["id"], "name": current["name"]}
else:
existing_result = None
if existing_result is not None:
event_bus.emit("profiles", {"action": "updated", "id": existing_result["profile_id"]})
return existing_result
profile_id = str(uuid.uuid4())[:8]
audio_filename = _community_profile_audio_filename(profile_id, item)
destination = Path(VOICES_DIR) / audio_filename
if audio_path is None:
audio_filename, audio_path = await _materialize_item_audio(item, profile_id)
else:
os.replace(audio_path, destination)
audio_path = destination
if audio_path is None: # defensive: a new profile always materialized above
raise RuntimeError("new community profile has no materialized audio")
profile_name = (name or item["name"]).strip() or item["name"]
kind, instruct, vd_states, seed = _profile_fields(item)
try:
with db_conn() as conn:
conn.execute("BEGIN IMMEDIATE")
duplicate = _existing_community_profile(conn, item, personality)
if duplicate is not None:
duplicate_audio = duplicate["ref_audio_path"]
if not _community_audio_is_current(duplicate, item, ref_text):
duplicate_audio = _community_profile_audio_filename(duplicate["id"], item)
duplicate_path = Path(VOICES_DIR) / duplicate_audio
_copy_atomic(audio_path, duplicate_path)
_heal_existing_profile(
conn, duplicate, item, ref_text, personality, duplicate_audio,
)
with contextlib.suppress(OSError):
audio_path.unlink()
duplicate_result = {"profile_id": duplicate["id"], "name": duplicate["name"]}
else:
duplicate_result = None
if duplicate_result is None:
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"created_at, kind, vd_states) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(profile_id, profile_name, audio_filename, ref_text, instruct,
item.get("language", "Auto"), seed, personality, time.time(), kind, vd_states),
)
except Exception:
with __import__("contextlib").suppress(OSError):
os.remove(audio_path)
with contextlib.suppress(OSError):
audio_path.unlink()
raise
if duplicate_result is not None:
event_bus.emit("profiles", {"action": "updated", "id": duplicate_result["profile_id"]})
return duplicate_result
event_bus.emit("profiles", {"action": "created", "id": profile_id})
return {"profile_id": profile_id, "name": profile_name}
def _download_voice_audio(item: dict, out_path: Path) -> None:
import hashlib
audio = item.get("audio") or {}
url = audio.get("url", "")
if not _safe_audio_url(url):
raise HTTPException(status_code=400, detail="Voice audio URL is not from an allowed host.")
import httpx
with httpx.Client(timeout=30.0, follow_redirects=True) as client:
resp = client.get(url)
resp.raise_for_status()
data = resp.content
expected = audio.get("sha256")
if expected and hashlib.sha256(data).hexdigest() != expected:
raise HTTPException(status_code=502, detail="Downloaded voice failed its integrity check.")
out_path.parent.mkdir(parents=True, exist_ok=True)
out_path.write_bytes(data)
+42 -31
View File
@@ -25,7 +25,7 @@ from huggingface_hub import utils as hf_utils
from huggingface_hub.errors import HFValidationError
from pydantic import BaseModel
from api.dependencies import require_loopback
from api.dependencies import require_admin, require_admin_action, require_desktop
from core import prefs
from services import tts_backend, asr_backend, llm_backend, translation_engines
from services.audio_dsp import list_effect_presets
@@ -41,6 +41,15 @@ _FAMILIES = {
"llm": (llm_backend, "llm_backend"),
}
def _family_payload(family: str, module):
"""Public inventory plus whether an environment pin owns this family."""
return {
"active": module.active_backend_id(),
"env_override": bool(os.environ.get(f"OMNIVOICE_{family.upper()}_BACKEND")),
"backends": public_backends(module.list_backends()),
}
def _is_hf_repo_id(value: str) -> bool:
"""Validate the route's ``owner/repo`` contract in bounded time."""
if not isinstance(value, str) or len(value) > 96 or value.count("/") != 1:
@@ -55,34 +64,25 @@ def _is_hf_repo_id(value: str) -> bool:
@router.get("/engines")
def list_all_engines():
return {
"tts": {
"active": tts_backend.active_backend_id(),
"backends": public_backends(tts_backend.list_backends()),
},
"asr": {
"active": asr_backend.active_backend_id(),
"backends": public_backends(asr_backend.list_backends()),
},
"llm": {
"active": llm_backend.active_backend_id(),
"backends": public_backends(llm_backend.list_backends()),
},
"tts": _family_payload("tts", tts_backend),
"asr": _family_payload("asr", asr_backend),
"llm": _family_payload("llm", llm_backend),
}
@router.get("/engines/tts")
def list_tts_backends():
return {"active": tts_backend.active_backend_id(), "backends": public_backends(tts_backend.list_backends())}
return _family_payload("tts", tts_backend)
@router.get("/engines/asr")
def list_asr_backends():
return {"active": asr_backend.active_backend_id(), "backends": public_backends(asr_backend.list_backends())}
return _family_payload("asr", asr_backend)
@router.get("/engines/llm")
def list_llm_backends():
return {"active": llm_backend.active_backend_id(), "backends": public_backends(llm_backend.list_backends())}
return _family_payload("llm", llm_backend)
@router.get("/engines/effects/presets", response_model=EffectPresetsResponse)
@@ -113,7 +113,10 @@ def list_translation_engines():
}
@router.post("/engines/translation/{engine_id}/install")
@router.post(
"/engines/translation/{engine_id}/install",
dependencies=[Depends(require_admin)],
)
async def install_translation_engine(engine_id: str):
entry = translation_engines.get_engine(engine_id)
if not entry:
@@ -149,7 +152,10 @@ async def install_translation_engine(engine_id: str):
}
@router.delete("/engines/translation/{engine_id}")
@router.delete(
"/engines/translation/{engine_id}",
dependencies=[Depends(require_admin)],
)
async def uninstall_translation_engine(engine_id: str):
entry = translation_engines.get_engine(engine_id)
if not entry:
@@ -188,15 +194,16 @@ async def uninstall_translation_engine(engine_id: str):
# POST /engines/sonitranslate/install). Mirrors the
# /engines/translation/{engine_id}/install namespace pattern.
#
# Loopback-gated: installing spawns subprocesses (git/uv) and writes to the
# data directory — only the local desktop frontend may trigger it. The job
# runs fine in packaged builds: the venv lives under the user data dir, not
# inside the signed app bundle, and uv resolves via OMNIVOICE_BUNDLED_UV/PATH.
# Desktop-only: installing spawns git/uv against mutable source and writes an
# editable environment. An API key does not make that supply-chain path safe to
# trigger remotely. The job runs fine in packaged builds: the venv lives under
# the user data dir, not inside the signed app bundle, and uv resolves via
# OMNIVOICE_BUNDLED_UV/PATH.
@router.post(
"/engines/sidecar/{engine_id}/install",
dependencies=[Depends(require_loopback)],
dependencies=[Depends(require_admin), Depends(require_desktop)],
)
def install_sidecar_engine(engine_id: str):
"""Start (or report) the one-click install for a sidecar engine.
@@ -222,7 +229,7 @@ def install_sidecar_engine(engine_id: str):
@router.get(
"/engines/sidecar/{engine_id}/install/status",
dependencies=[Depends(require_loopback)],
dependencies=[Depends(require_admin)],
)
def sidecar_install_status(engine_id: str):
"""Step-by-step status of the sidecar install job (poll while running).
@@ -243,7 +250,7 @@ def sidecar_install_status(engine_id: str):
@router.delete(
"/engines/sidecar/{engine_id}/install",
dependencies=[Depends(require_loopback)],
dependencies=[Depends(require_admin)],
)
def uninstall_sidecar_engine(engine_id: str):
"""Remove an app-managed sidecar install (checkout + venv + weights) and
@@ -274,8 +281,8 @@ def uninstall_sidecar_engine(engine_id: str):
# frame. Result includes wall-clock latency so the UI can render
# "1234 ms — pong" inline next to the button.
#
# Loopback-gated (T-02-13): only the local desktop frontend may trigger
# a sidecar spawn through this endpoint.
# Admin-gated (T-02-13): only the local desktop frontend or an authenticated
# server-mode administrator may trigger a sidecar spawn through this endpoint.
# Engine instances cached for the lifetime of the FastAPI process so that
# repeated health checks don't spawn a new SubprocessBackend (each spawn
@@ -311,7 +318,7 @@ def _resolve_engine_class(engine_id: str):
@router.get(
"/engines/{engine_id}/health",
dependencies=[Depends(require_loopback)],
dependencies=[Depends(require_admin_action)],
)
def engine_health(engine_id: str):
"""Spawn-and-ping a SubprocessBackend; ``is_available()`` for the rest.
@@ -385,7 +392,7 @@ def engine_health(engine_id: str):
# hanging the Settings panel. The orphaned worker is best-effort daemon.
# * A process-wide lock serialises self-tests so a click-storm can't stack
# concurrent model loads.
# * Only ever on user click (POST) — never on Settings load. Loopback-gated.
# * Only ever on user click (POST) — never on Settings load. Admin-gated.
# Deliberately short + ASCII so the synth stays CPU-cheap and the phrase never
# trips the no-hardcoded-CJK guard.
@@ -452,7 +459,7 @@ class SelfTestResponse(BaseModel):
@router.post(
"/engines/{engine_id}/selftest",
response_model=SelfTestResponse,
dependencies=[Depends(require_loopback)],
dependencies=[Depends(require_admin)],
)
def engine_selftest(engine_id: str):
"""Run a bounded, real synthesis on an available in-process TTS engine.
@@ -551,7 +558,11 @@ class SelectEngineResponse(BaseModel):
routing_reason: str | None = None
@router.post("/engines/select", response_model=SelectEngineResponse)
@router.post(
"/engines/select",
response_model=SelectEngineResponse,
dependencies=[Depends(require_admin)],
)
def select_engine(req: SelectEngineRequest):
"""Persist a family's engine pick to prefs.json. Refuses unknown backends,
backends whose deps aren't installed, AND backends that cannot run on THIS
+229 -86
View File
@@ -1,18 +1,24 @@
import os
import json
import uuid
import time
import asyncio
import contextlib
import json
import logging
from typing import Optional, List
import os
import re
import shutil
import tempfile
import time
import uuid
from pathlib import Path
from typing import List, Optional
from fastapi import APIRouter, File, Form, UploadFile, HTTPException, Query
from fastapi.responses import FileResponse, RedirectResponse
from fastapi.responses import FileResponse
from pydantic import BaseModel
from core.db import db_conn
from core.config import VOICES_DIR, OUTPUTS_DIR
from core import event_bus
from core.audio_validation import resolve_regular_file
from core.file_cleanup import FileCleanupError, unlink_if_present
from services.ffmpeg_utils import spawn_subprocess
@@ -360,46 +366,223 @@ async def upload_voice_clip(
}
def _stage_profile_audio(source: Path, directory: Path) -> Path:
"""Copy an imported clip to a hidden temp file inside ``directory``.
The temp lives in the destination directory itself so a later
``os.replace`` to the final name is an atomic same-filesystem rename
cheap enough to run while holding a DB write lock, unlike the copy.
Callers own cleanup of the returned path if they never publish it.
"""
directory.mkdir(parents=True, exist_ok=True)
fd, tmp_name = tempfile.mkstemp(
dir=str(directory), prefix=".gallery-import-", suffix=".part",
)
os.close(fd)
try:
shutil.copy2(source, tmp_name)
except BaseException:
with contextlib.suppress(OSError):
os.unlink(tmp_name)
raise
return Path(tmp_name)
def _copy_profile_audio(source: Path, destination: Path) -> None:
"""Copy an imported clip without exposing a partial profile audio file."""
staged = _stage_profile_audio(source, destination.parent)
try:
os.replace(staged, destination)
except BaseException:
with contextlib.suppress(OSError):
os.unlink(staged)
raise
def _gallery_profile_audio_filename(profile_id: str, source: Path) -> str:
"""Return the canonical, portable filename for a My Imports profile."""
safe_id = (
profile_id if re.fullmatch(r"[A-Za-z0-9_-]{1,64}", profile_id or "")
else uuid.uuid5(uuid.NAMESPACE_URL, str(profile_id)).hex[:16]
)
suffix = source.suffix.lower()
if not re.fullmatch(r"\.[a-z0-9]{1,8}", suffix):
suffix = ".wav"
return f"{safe_id}_gallery{suffix}"
def _is_materialized_gallery_profile(row, voice: dict, audio_filename: str) -> bool:
"""Recognize only rows created by this materializer, not identity collisions."""
return bool(
row["personality"] == f"gallery:{voice['id']}"
and row["ref_audio_path"] == audio_filename
and row["ref_text"] == ""
and row["instruct"] == ""
and row["language"] == "Auto"
and row["seed"] is None
and row["kind"] == "clone"
and not row["vd_states"]
and row["description"] == (voice.get("description") or "")
and not row["is_locked"]
and not row["verified_own_voice"]
and not row["locked_audio_path"]
)
def _existing_gallery_profile(conn, voice: dict, source: Path):
personality = f"gallery:{voice['id']}"
rows = conn.execute(
"SELECT * FROM voice_profiles WHERE personality=? ORDER BY created_at, id",
(personality,),
).fetchall()
for row in rows:
expected = _gallery_profile_audio_filename(row["id"], source)
if _is_materialized_gallery_profile(row, voice, expected):
return row
return None
def _gallery_profile_audio_is_current(row, source: Path) -> bool:
"""Detect missing/replaced copies without re-hashing unchanged imports."""
destination = resolve_regular_file(VOICES_DIR, row["ref_audio_path"])
if destination is None:
return False
try:
source_stat = source.stat()
destination_stat = destination.stat()
# copy2 preserves mtime; size + nanosecond mtime catches ordinary edits
# and partial writes while keeping repeated Use clicks inexpensive.
return (
source_stat.st_size == destination_stat.st_size
and source_stat.st_mtime_ns == destination_stat.st_mtime_ns
)
except OSError:
return False
def _materialize_gallery_profile(
voice_id: str, requested_name: Optional[str] = None,
) -> dict:
"""Idempotently materialize/heal one My Imports clip as a clone profile."""
personality = f"gallery:{voice_id}"
copied_path: Optional[Path] = None
created = False
staged_path: Optional[Path] = None
staged_source: Optional[Path] = None
try:
# Stage the (potentially large) audio copy BEFORE taking SQLite's
# write lock: copying inside BEGIN IMMEDIATE would stall every other
# backend writer for the whole copy. The staged temp lives in
# VOICES_DIR itself, so publishing it inside the transaction is an
# atomic same-filesystem os.replace. This pre-read is advisory only —
# the locked transaction below re-reads and re-decides everything.
copy_needed = False
with db_conn() as conn:
pre_row = conn.execute(
"SELECT * FROM voice_gallery WHERE id = ?", (voice_id,),
).fetchone()
if pre_row is not None:
pre_source = Path(pre_row["audio_path"])
if pre_source.is_file():
pre_existing = _existing_gallery_profile(conn, dict(pre_row), pre_source)
copy_needed = pre_existing is None or not _gallery_profile_audio_is_current(
pre_existing, pre_source,
)
if copy_needed:
staged_path = _stage_profile_audio(pre_source, Path(VOICES_DIR))
staged_source = pre_source
with db_conn() as conn:
# The identity is not globally UNIQUE because personality is shared
# with other import mechanisms. Serialize this check+insert in
# SQLite so simultaneous Use clicks cannot both create a row.
conn.execute("BEGIN IMMEDIATE")
row = conn.execute(
"SELECT * FROM voice_gallery WHERE id = ?", (voice_id,),
).fetchone()
if row is None:
raise HTTPException(status_code=404, detail="Voice not found")
voice = dict(row)
source = Path(voice["audio_path"])
if not source.is_file():
raise HTTPException(status_code=404, detail="Audio file not found on disk")
def _install_audio(destination: Path) -> None:
"""Publish the staged copy under the lock via atomic rename."""
nonlocal staged_path
if staged_path is not None and staged_source == source:
os.replace(staged_path, destination)
staged_path = None
else:
# Rare race: the gallery row changed between the advisory
# pre-read and taking the lock, so any staged bytes may be
# from the wrong source. Fall back to the blocking copy
# rather than publish stale audio.
_copy_profile_audio(source, destination)
existing = _existing_gallery_profile(conn, voice, source)
if existing is not None:
ref_filename = _gallery_profile_audio_filename(existing["id"], source)
if not _gallery_profile_audio_is_current(existing, source):
ref_path = Path(VOICES_DIR) / ref_filename
_install_audio(ref_path)
copied_path = ref_path
conn.execute(
"UPDATE voice_profiles SET ref_audio_path=?, ref_text='', instruct='', "
"language='Auto', seed=NULL, description=?, kind='clone', vd_states=NULL, "
"personality=? WHERE id=?",
(
ref_filename, voice["description"] or "", personality,
existing["id"],
),
)
result = {"profile_id": existing["id"], "name": existing["name"]}
else:
profile_id = str(uuid.uuid4())[:8]
profile_name = (requested_name or voice["name"]).strip() or voice["name"]
ref_filename = _gallery_profile_audio_filename(profile_id, source)
copied_path = Path(VOICES_DIR) / ref_filename
_install_audio(copied_path)
conn.execute(
"""INSERT INTO voice_profiles
(id, name, ref_audio_path, ref_text, instruct, language, seed,
personality, is_locked, locked_audio_path, description, kind,
vd_states, created_at)
VALUES (?, ?, ?, '', '', 'Auto', NULL, ?, 0, '', ?, 'clone', NULL, ?)""",
(
profile_id, profile_name, ref_filename, personality,
voice["description"] or "", time.time(),
),
)
created = True
result = {"profile_id": profile_id, "name": profile_name}
except BaseException:
if copied_path is not None:
with contextlib.suppress(OSError):
copied_path.unlink()
raise
finally:
# Staged but never published (failure, or a concurrent request healed
# the profile first) — never leave .part droppings in VOICES_DIR.
if staged_path is not None:
with contextlib.suppress(OSError):
os.unlink(staged_path)
event_bus.emit(
"profiles", {"action": "created" if created else "updated", "id": result["profile_id"]},
)
return result
@router.post("/gallery/voices/{voice_id}/save-as-profile")
async def save_voice_as_profile(
voice_id: str,
profile_name: str = Query(..., description="Name for the voice profile"),
):
"""Save a gallery voice as a voice profile for cloning."""
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_gallery WHERE id = ?", (voice_id,)
).fetchone()
if not row:
raise HTTPException(status_code=404, detail="Voice not found")
profile_id = str(uuid.uuid4())[:8]
import shutil
ext = os.path.splitext(row["audio_path"])[1]
new_audio_path = os.path.join(VOICES_DIR, f"{profile_id}{ext}")
shutil.copy(row["audio_path"], new_audio_path)
conn.execute(
"""
INSERT INTO voice_profiles (id, name, ref_audio_path, ref_text, instruct, language, seed, created_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?)
""",
(
profile_id,
profile_name,
f"{profile_id}{ext}",
row["description"] or "",
row["character"] or "",
"Auto",
None,
time.time(),
),
)
event_bus.emit("profiles", {"action": "created", "id": profile_id})
return {"profile_id": profile_id, "name": profile_name}
result = await asyncio.to_thread(_materialize_gallery_profile, voice_id, profile_name)
return {"profile_id": result["profile_id"], "name": result["name"]}
@router.get("/gallery/voices/{voice_id}/preview")
@@ -415,22 +598,10 @@ def preview_voice(voice_id: str):
audio_path = row["audio_path"]
# Debug logging
is_absolute = os.path.isabs(audio_path)
path_exists = os.path.exists(audio_path) if audio_path else False
# If absolute path, serve directly or redirect
if is_absolute and path_exists:
# Get just the relative path from outputs dir
outputs_path = str(OUTPUTS_DIR)
if audio_path.startswith(outputs_path):
# Remove outputs_dir prefix to get relative path within outputs
rel_path = os.path.relpath(audio_path, outputs_path)
# The audio_path is like: /Users/user4/.../outputs/voice_gallery/file.wav
# rel_path becomes: voice_gallery/file.wav
# We want to serve from /audio/ so: /audio/voice_gallery/file.wav
return RedirectResponse(f"/audio/{rel_path}")
return FileResponse(audio_path, media_type="audio/wav")
if os.path.isabs(audio_path) and os.path.exists(audio_path):
# Serve the file from this API route so deployments mounted below a
# path prefix do not lose that prefix while following a redirect.
return FileResponse(audio_path)
raise HTTPException(
status_code=404,
@@ -503,33 +674,5 @@ def batch_delete_voices(body: dict):
@router.post("/gallery/voices/{voice_id}/to-profile")
def voice_to_profile(voice_id: str):
"""Create a voice profile from a gallery clip."""
with db_conn() as conn:
row = conn.execute("SELECT * FROM voice_gallery WHERE id = ?", (voice_id,)).fetchone()
if not row:
raise HTTPException(status_code=404, detail="Voice not found")
voice = dict(row)
audio_path = voice["audio_path"]
if not os.path.exists(audio_path):
raise HTTPException(status_code=404, detail="Audio file not found on disk")
import shutil
import uuid
profile_id = str(uuid.uuid4())[:8]
# Copy audio to voices dir
dest_filename = f"{profile_id}_gallery.wav"
dest_path = os.path.join(VOICES_DIR, dest_filename)
shutil.copy2(audio_path, dest_path)
import time
now = time.time()
conn.execute(
"""INSERT INTO voice_profiles
(id, name, ref_audio_path, ref_text, instruct, seed, is_locked, locked_audio_path, created_at, updated_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)""",
(profile_id, voice["name"], dest_filename, "", None, None, 0, None, now, now),
)
event_bus.emit("profiles", {"action": "created", "id": profile_id})
return {"success": True, "profile_id": profile_id, "name": voice["name"]}
result = _materialize_gallery_profile(voice_id)
return {"success": True, "profile_id": result["profile_id"], "name": result["name"]}
+33
View File
@@ -1146,6 +1146,9 @@ async def generate_speech(
# classic flow, so streaming is purely a delivery channel — engine-agnostic
# (text-level chunking, no per-engine token streaming).
stream: bool = Form(False),
# Explicit opt-in. The absence of this field preserves the local-first
# /generate contract even when an administrator configured hosted values.
hosted: bool = Form(False),
):
# #502: NFC-normalize the input text so decomposed (NFD) diacritics — common
# in pasted Vietnamese and other Latin-with-marks text — are composed to the
@@ -1156,6 +1159,36 @@ async def generate_speech(
import unicodedata
text = unicodedata.normalize("NFC", text)
if hosted:
# Hosted execution accepts only a previously, explicitly synchronized
# consent-verified profile. Never silently sync a local recording from
# a synthesis request: that would make normal offline use an upload.
if not profile_id:
raise HTTPException(status_code=422, detail="Hosted synthesis requires a synchronized voice profile.")
from services.hosted_voice_api import HostedSettings, HostedVoiceClient, HostedVoiceError
try:
settings = HostedSettings.from_environment()
except HostedVoiceError as exc:
raise HTTPException(status_code=503, detail=str(exc)) from exc
if settings is None:
raise HTTPException(status_code=409, detail="Hosted synthesis is not configured on this device.")
with db_conn() as conn:
profile = conn.execute("SELECT hosted_voice_id, language FROM voice_profiles WHERE id=?", (profile_id,)).fetchone()
if not profile:
raise HTTPException(status_code=404, detail="Voice profile not found")
if not profile["hosted_voice_id"]:
raise HTTPException(status_code=422, detail="Sync this consent-verified profile to hosted before hosted synthesis.")
client = HostedVoiceClient(settings)
try:
audio = await client.synthesize(
text=text, profile_voice_id=profile["hosted_voice_id"], language=language or profile["language"],
)
except HostedVoiceError as exc:
raise HTTPException(status_code=502, detail=str(exc)) from exc
finally:
await client.aclose()
return StreamingResponse(io.BytesIO(audio), media_type="audio/wav", headers={"X-OmniVoice-Execution": "hosted"})
# ── Engine resolution (issue #312) ──────────────────────────────────────
# The request runs on the engine selected in Settings (POST /engines/select,
# env var OMNIVOICE_TTS_BACKEND wins), or an explicit per-request `engine`
+2 -2
View File
@@ -9,13 +9,13 @@ from __future__ import annotations
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field
from api.dependencies import require_loopback
from api.dependencies import require_admin
from services import mcp_bindings
router = APIRouter(
prefix="/api/mcp",
tags=["mcp"],
dependencies=[Depends(require_loopback)],
dependencies=[Depends(require_admin)],
)
+2 -2
View File
@@ -12,10 +12,10 @@ import logging
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel
from api.dependencies import require_loopback
from api.dependencies import require_admin
logger = logging.getLogger("omnivoice.api")
router = APIRouter(dependencies=[Depends(require_loopback)])
router = APIRouter(dependencies=[Depends(require_admin)])
class CustomPathRequest(BaseModel):
+45
View File
@@ -14,6 +14,7 @@ from core import event_bus
from core.personalities import get_personalities
from omnivoice.utils.voice_design import heal_design_instruct, sanitize_instruct
from core.path_security import UnsafePath, resolve_within
from services.hosted_voice_api import HostedSettings, HostedVoiceClient, HostedVoiceError
router = APIRouter()
@@ -184,6 +185,50 @@ def get_profile(profile_id: str):
return dict(row)
@router.post("/profiles/{profile_id}/hosted-sync")
async def sync_profile_to_hosted(profile_id: str):
"""Explicitly copy a consent-verified local clone to the hosted library.
This is deliberately not part of local profile creation: merely creating a
profile must never upload biometric source audio. The hosted service records
the existing spoken-consent evidence as its versioned attestation; it does
not receive the consent recording itself.
"""
try:
settings = HostedSettings.from_environment()
except HostedVoiceError as exc:
raise HTTPException(status_code=503, detail=str(exc)) from exc
if settings is None:
raise HTTPException(status_code=409, detail="Hosted voice sync is not configured on this device.")
with db_conn() as conn:
row = conn.execute(
"SELECT id, name, description, ref_text, ref_audio_path, verified_own_voice, consent_text, hosted_voice_id "
"FROM voice_profiles WHERE id=?", (profile_id,)
).fetchone()
if not row:
raise HTTPException(status_code=404, detail="Profile not found")
if row["hosted_voice_id"]:
return {"profile_id": profile_id, "hosted_voice_id": row["hosted_voice_id"], "state": "already_synced"}
if not row["verified_own_voice"] or not row["consent_text"].strip():
raise HTTPException(status_code=422, detail="Record the voice-ownership consent statement before hosted sync.")
reference_path = _voices_path(row["ref_audio_path"] or "")
if not reference_path or not os.path.isfile(reference_path):
raise HTTPException(status_code=422, detail="This profile has no local reference recording to sync.")
client = HostedVoiceClient(settings)
try:
hosted_voice_id = await client.create_voice(
name=row["name"], description=row["description"] or row["ref_text"] or "", reference_path=reference_path,
)
except HostedVoiceError as exc:
raise HTTPException(status_code=502, detail=str(exc)) from exc
finally:
await client.aclose()
with db_conn() as conn:
conn.execute("UPDATE voice_profiles SET hosted_voice_id=? WHERE id=? AND hosted_voice_id=''", (hosted_voice_id, profile_id))
persisted = conn.execute("SELECT hosted_voice_id FROM voice_profiles WHERE id=?", (profile_id,)).fetchone()["hosted_voice_id"]
return {"profile_id": profile_id, "hosted_voice_id": persisted, "state": "synced"}
@router.put("/profiles/{profile_id}")
def update_profile(profile_id: str, patch: ProfileUpdate):
"""Partial update — only fields set on the payload are changed."""
+10 -10
View File
@@ -7,7 +7,7 @@ CRUD for the DB-backed, per-language pronunciation dictionary the
before synthesis (see ``services/pronunciation.apply_pronunciation`` and the
generate path), so a saved entry actually changes the audio on every engine.
Endpoints (loopback-only, like the dictation router):
Endpoints (admin-gated; loopback or authenticated server mode):
GET /pronunciation list every entry
POST /pronunciation create one entry
PUT /pronunciation/{entry_id} update an entry (partial)
@@ -30,12 +30,12 @@ from typing import List, Optional
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel
from api.dependencies import require_loopback
from api.dependencies import require_admin
from core.db import db_conn
from services.pronunciation import apply_pronunciation, entries_for_language
logger = logging.getLogger("omnivoice.pronunciation")
router = APIRouter()
router = APIRouter(dependencies=[Depends(require_admin)])
_VALID_TYPES = ("respelling", "ipa", "cmu")
_ALL_LANG = "*"
@@ -133,7 +133,7 @@ class PronImportRequest(BaseModel):
# ── CRUD ─────────────────────────────────────────────────────────────────────
@router.get("/pronunciation", dependencies=[Depends(require_loopback)])
@router.get("/pronunciation")
def list_entries():
with db_conn() as conn:
rows = conn.execute(
@@ -143,7 +143,7 @@ def list_entries():
return [_row_to_dict(r) for r in rows]
@router.post("/pronunciation", dependencies=[Depends(require_loopback)])
@router.post("/pronunciation")
def create_entry(entry: PronEntry):
term = entry.term.strip()
if not term:
@@ -171,7 +171,7 @@ def create_entry(entry: PronEntry):
return _row_to_dict(row)
@router.put("/pronunciation/{entry_id}", dependencies=[Depends(require_loopback)])
@router.put("/pronunciation/{entry_id}")
def update_entry(entry_id: str, patch: PronEntryUpdate):
with db_conn() as conn:
existing = conn.execute(
@@ -226,7 +226,7 @@ def update_entry(entry_id: str, patch: PronEntryUpdate):
return _row_to_dict(row)
@router.delete("/pronunciation/{entry_id}", dependencies=[Depends(require_loopback)])
@router.delete("/pronunciation/{entry_id}")
def delete_entry(entry_id: str):
with db_conn() as conn:
cur = conn.execute("DELETE FROM pronunciation_entries WHERE id = ?", (entry_id,))
@@ -236,7 +236,7 @@ def delete_entry(entry_id: str):
# ── Dry-run + import/export ───────────────────────────────────────────────────
@router.post("/pronunciation/test", dependencies=[Depends(require_loopback)])
@router.post("/pronunciation/test")
def test_substitution(req: PronTestRequest):
"""Show the post-substitution text for ``req.text`` — no model call.
@@ -258,7 +258,7 @@ def test_substitution(req: PronTestRequest):
}
@router.get("/pronunciation/export", dependencies=[Depends(require_loopback)])
@router.get("/pronunciation/export")
def export_entries():
"""Every entry as a JSON-serializable list (round-trips ``/import``)."""
with db_conn() as conn:
@@ -273,7 +273,7 @@ def export_entries():
]}
@router.post("/pronunciation/import", dependencies=[Depends(require_loopback)])
@router.post("/pronunciation/import")
def import_entries(req: PronImportRequest):
"""Bulk-add entries. ``replace=true`` clears the table first.
+84 -4
View File
@@ -20,7 +20,7 @@ from fastapi import APIRouter, Depends, HTTPException, Query
from pydantic import BaseModel, Field
from core.logging_utils import log_safe
from api.dependencies import require_admin
from api.dependencies import require_admin, require_admin_action
logger = logging.getLogger("omnivoice.api.settings")
@@ -92,8 +92,8 @@ def get_hf_token_state(fresh: bool = Query(False)):
# ── Performance settings (INST-12) ────────────────────────────────────────
# Threat T-02-04: same loopback guard as the hf-token endpoints via the
# router-level `require_loopback` dep.
# Threat T-02-04: same admin guard as the hf-token endpoints via the
# router-level `require_admin` dep.
_TORCH_COMPILE_KEY = "perf.torch_compile_disabled"
@@ -133,6 +133,83 @@ def set_torch_compile_disabled(body: _TorchCompileBody):
return _torch_compile_state()
# ── Compute-device override (Settings → Performance) ──────────────────────
class _ComputeDeviceBody(BaseModel):
value: str = Field(..., description="auto | cuda | rocm | xpu | mps | cpu")
def _compute_device_state() -> dict:
"""Everything the Performance panel needs to render the device control:
the resolved pick (env > prefs > auto), what this process actually applied
at probe time (differs after a change until restart caps are immutable
per process), what auto would pick, and which families exist here."""
from core import device_caps
caps = device_caps.detect_host_caps()
env_pin = (os.environ.get("OMNIVOICE_DEVICE") or "").strip().lower()
auto_family = next(
(f for f in ("cuda", "rocm", "xpu", "mps") if f in caps.available_families),
"cpu",
)
value = device_caps.requested_device_override()
return {
"value": value,
"applied": caps.requested_family,
"restart_required": value != caps.requested_family,
# The running process asked for a family it doesn't have (env pin on
# the wrong machine, hardware removed): auto is in effect, and a
# restart would not change that — the panel says so instead of
# pretending the pick took.
"override_ignored": (
caps.requested_family not in ("auto", caps.family)
),
"effective_family": caps.family,
"auto_family": auto_family,
"available_families": list(caps.available_families),
"env_pinned": env_pin in device_caps.DEVICE_OVERRIDE_CHOICES and env_pin != "",
"choices": list(device_caps.DEVICE_OVERRIDE_CHOICES),
}
@router.get("/compute-device")
def get_compute_device():
"""Current compute-device override state (Settings → Performance)."""
return _compute_device_state()
@router.put("/compute-device")
def set_compute_device(body: _ComputeDeviceBody):
"""Persist the compute-device pick. Applied by the capability probe at
the next backend start (host caps are immutable per process same
restart contract as the rest of the Performance tab). ``OMNIVOICE_DEVICE``
always wins over this pick; the UI shows the pin instead of pretending."""
from core import device_caps, prefs
value = (body.value or "").strip().lower()
if value not in device_caps.DEVICE_OVERRIDE_CHOICES:
raise HTTPException(
status_code=400,
detail=f"Unknown device '{value}'. Valid: {', '.join(device_caps.DEVICE_OVERRIDE_CHOICES)}",
)
caps = device_caps.detect_host_caps()
if value not in ("auto", "cpu") and value not in caps.available_families:
raise HTTPException(
status_code=400,
detail=(
f"'{value}' is not available on this host "
f"(have: {', '.join(caps.available_families)})"
),
)
try:
prefs.set_("compute_device", value)
except Exception:
logger.exception("set_compute_device failed")
raise HTTPException(status_code=500, detail="Failed to persist setting")
return _compute_device_state()
# ── Generation-history retention (Studio takes rail) ──────────────────────
@@ -481,7 +558,10 @@ def _local_models(base_url: str, api_key: str):
return None
@router.get("/llm-providers/{provider_id}/models")
@router.get(
"/llm-providers/{provider_id}/models",
dependencies=[Depends(require_admin_action)],
)
def list_llm_provider_models(provider_id: str):
"""List model ids the provider's key can access (OpenAI-compat /models).
+5 -2
View File
@@ -11,7 +11,7 @@ from core.prefs import set_ as prefs_set, delete as prefs_delete
from services import network_share
from services import tailscale as _tailscale
from api.schemas import SysinfoResponse, SystemInfoResponse, ModelStatusResponse
from api.dependencies import is_loopback, require_admin
from api.dependencies import is_loopback, require_admin, require_admin_action
from fastapi.responses import FileResponse, StreamingResponse
import torch
import shutil
@@ -1089,7 +1089,10 @@ async def diagnostic_bundle(network: bool = Query(False, description="Include th
# ── Self-check diagnostics ────────────────────────────────────────────────
@router.get("/system/diagnose")
@router.get(
"/system/diagnose",
dependencies=[Depends(require_admin_action)],
)
async def system_diagnose(
network: bool = Query(True, description="Include the HuggingFace hub reachability probe"),
deep: bool = Query(False, description="Also load the active engine and synthesize a short utterance (may cold-load the model — minutes on first run)"),
+4 -4
View File
@@ -28,7 +28,7 @@ import logging
from fastapi import APIRouter, Depends, HTTPException, Request
from pydantic import BaseModel, Field
from api.dependencies import require_loopback
from api.dependencies import require_admin
from worker import registry, routing, service
logger = logging.getLogger("omnivoice.worker")
@@ -39,9 +39,9 @@ logger = logging.getLogger("omnivoice.worker")
# the task's own deadline does.
_DISCONNECT_POLL_SECONDS = 1.0
# Management is loopback-only: these endpoints mint join tokens and revoke
# machines, so they follow the same rule as the app's other privileged routes.
router = APIRouter(prefix="/workers", tags=["workers"], dependencies=[Depends(require_loopback)])
# Management is admin-gated: these endpoints mint join tokens and revoke
# machines, so Docker writes require the API key while desktop stays loopback.
router = APIRouter(prefix="/workers", tags=["workers"], dependencies=[Depends(require_admin)])
class EnableRequest(BaseModel):
+106
View File
@@ -0,0 +1,106 @@
"""Lightweight validation for persisted profile WAV references.
This module deliberately uses only the standard library. Gallery routers import
it during startup, so pulling in torch/torchaudio merely to validate a cached
file would make every Gallery open pay the model stack's import cost.
"""
from __future__ import annotations
import os
import wave
from pathlib import Path
from typing import Optional
from core.path_security import UnsafePath, resolve_within, safe_filename
_READ_CHUNK_BYTES = 1 << 20
_MAX_CHANNELS = 64
_MAX_SAMPLE_RATE = 768_000
_MAX_SAMPLE_WIDTH = 8
def resolve_regular_file(root: os.PathLike[str] | str, value: object) -> Optional[Path]:
"""Resolve a portable bare filename inside *root*, rejecting symlinks."""
try:
name = safe_filename(value)
unresolved = Path(root).resolve(strict=False) / name
if unresolved.is_symlink():
return None
return resolve_within(root, name)
except (OSError, UnsafePath):
return None
def is_playable_wav(path: Optional[Path]) -> bool:
"""Return true only for a regular, decodable WAV with audio frames."""
if path is None:
return False
try:
if not path.is_file() or path.is_symlink():
return False
file_size = path.stat().st_size
with wave.open(str(path), "rb") as wav:
channels = wav.getnchannels()
sample_rate = wav.getframerate()
sample_width = wav.getsampwidth()
frame_count = wav.getnframes()
if (
not 0 < channels <= _MAX_CHANNELS
or not 0 < sample_rate <= _MAX_SAMPLE_RATE
or not 0 < sample_width <= _MAX_SAMPLE_WIDTH
or frame_count <= 0
):
return False
# ``wave.getnframes`` trusts the header. Read through the declared
# payload so an interrupted write with a complete header but a
# truncated data chunk cannot masquerade as playable audio.
frame_size = channels * sample_width
expected_bytes = frame_count * frame_size
# A PCM payload cannot be larger than the containing file. Check
# before calling ``readframes`` so hostile header values cannot
# turn a tiny file into a multi-gigabyte allocation request.
if expected_bytes > file_size:
return False
read_bytes = 0
chunk_frames = max(1, min(frame_count, _READ_CHUNK_BYTES // frame_size))
while read_bytes < expected_bytes:
chunk = wav.readframes(chunk_frames)
if not chunk or len(chunk) % frame_size:
return False
read_bytes += len(chunk)
return read_bytes == expected_bytes
except (MemoryError, OSError, EOFError, OverflowError, wave.Error):
# Python 3.11's wave module rejects valid IEEE-float/WAVE_EXTENSIBLE
# files. SoundFile is already a runtime dependency and recognizes those
# containers; import it only on the uncommon fallback path.
try:
import soundfile as sf
with sf.SoundFile(str(path)) as audio:
if (
audio.format != "WAV"
or not 0 < audio.channels <= _MAX_CHANNELS
or not 0 < audio.samplerate <= _MAX_SAMPLE_RATE
or len(audio) <= 0
):
return False
remaining = len(audio)
# Decode through the declared payload in byte-bounded chunks;
# ``sf.info`` alone also trusts a truncated file's header.
chunk_frames = max(
1, _READ_CHUNK_BYTES // (audio.channels * 4),
)
while remaining:
frames = audio.read(
min(remaining, chunk_frames), dtype="float32", always_2d=True,
)
count = len(frames)
if count <= 0:
return False
remaining -= count
return True
except Exception:
return False
__all__ = ["is_playable_wav", "resolve_regular_file"]
+421
View File
@@ -0,0 +1,421 @@
"""Canonical authentication identity for HTTP and WebSocket connections.
Transport parsing belongs here; authorization remains in FastAPI dependencies.
Each ASGI scope receives exactly one secret-free :class:`AuthPrincipal` so
middleware and route guards cannot disagree about credential precedence.
"""
from __future__ import annotations
import ipaddress
import importlib
import os
import secrets
from collections.abc import Mapping
from dataclasses import dataclass, field
from enum import Enum
from services.admin_sessions import (
AdminSessionStore,
)
_AUTH_STATE_KEY = "auth_principal"
_LOOPBACK_HOSTS = frozenset({"127.0.0.1", "::1", "localhost"})
CONSUME_CAPABILITIES = frozenset({"consume"})
ADMIN_CAPABILITIES = frozenset({"consume", "admin"})
LOOPBACK_CAPABILITIES = frozenset({"consume", "admin", "native"})
class PrincipalKind(str, Enum):
ANONYMOUS = "anonymous"
LOOPBACK = "loopback"
TRUSTED_NETWORK = "trusted_network"
PIN = "pin"
API_KEY = "api_key"
ADMIN_SESSION = "admin_session"
class CredentialTransport(str, Enum):
NONE = "none"
HEADER = "header"
QUERY = "query"
COOKIE = "cookie"
LEGACY_COOKIE = "legacy_cookie"
WS_TICKET = "ws_ticket"
@dataclass(frozen=True)
class AuthPrincipal:
kind: PrincipalKind
capabilities: frozenset[str]
credential_id: str | None = None
transport: CredentialTransport = CredentialTransport.NONE
def allows(self, capability: str) -> bool:
return capability in self.capabilities
@dataclass(frozen=True)
class _CredentialCandidate:
value: str = field(repr=False)
transport: CredentialTransport
allow_master: bool = False
allow_session: bool = False
allow_ticket: bool = False
def remote_api_key() -> str | None:
"""Normalized remote operator key, read dynamically for rotation support."""
return os.environ.get("OMNIVOICE_API_KEY", "").strip() or None
def credential_matches(supplied: str | None, configured: str | None) -> bool:
"""Constant-time credential comparison that accepts the full Unicode range."""
if not supplied or not configured:
return False
return secrets.compare_digest(
supplied.encode("utf-8", errors="surrogatepass"),
configured.encode("utf-8", errors="surrogatepass"),
)
def _active_admin_session_store() -> AdminSessionStore:
"""Resolve mutable process state at call time so app reloads cannot split it."""
module = importlib.import_module("services.admin_sessions")
return module.admin_session_store
def _trusted_networks() -> tuple[ipaddress.IPv4Network | ipaddress.IPv6Network, ...]:
networks = []
for value in os.environ.get("OMNIVOICE_TRUSTED_NETWORKS", "").split(","):
value = value.strip()
if not value:
continue
try:
networks.append(ipaddress.ip_network(value, strict=False))
except ValueError:
# Invalid configuration never makes the gate fail open or wedge the
# backend. It simply contributes no trusted range.
continue
return tuple(networks)
def is_loopback(host: str | None) -> bool:
return host in _LOOPBACK_HOSTS
def is_local_host(host: str | None) -> bool:
if is_loopback(host):
return True
try:
address = ipaddress.ip_address(host)
except (TypeError, ValueError):
return False
if getattr(address, "ipv4_mapped", None):
address = address.ipv4_mapped
return any(address in network for network in _trusted_networks())
def _mapping_get(mapping: Mapping[str, str] | object, name: str) -> str:
if not mapping:
return ""
getter = getattr(mapping, "get", None)
if callable(getter):
value = getter(name, "")
if value:
return str(value)
# Real Starlette Headers are case-insensitive. This small fallback keeps
# minimal request stubs and non-Starlette callers correct too.
items = getattr(mapping, "items", None)
if callable(items):
for key, value in items():
if str(key).lower() == name.lower():
return str(value or "")
return ""
def _scope_type(connection) -> str:
scope = getattr(connection, "scope", None)
return str(scope.get("type", "http")) if isinstance(scope, dict) else "http"
def _path(connection) -> str:
scope = getattr(connection, "scope", None)
if isinstance(scope, dict):
return str(scope.get("path", ""))
return str(getattr(connection, "url", "") or "")
def _canonical_websocket_path(connection) -> str:
"""Remove only the ASGI-configured deployment prefix from a WS path."""
path = _path(connection)
scope = getattr(connection, "scope", None)
if not isinstance(scope, dict):
return path
root_path = str(scope.get("root_path", "") or "").rstrip("/")
if not root_path or root_path == "/":
return path
root_path = "/" + root_path.lstrip("/")
if path.startswith(root_path + "/"):
return path[len(root_path) :]
return path
def _client_host(connection) -> str | None:
client = getattr(connection, "client", None)
if client is not None:
return getattr(client, "host", None)
scope = getattr(connection, "scope", None)
if isinstance(scope, dict) and scope.get("client"):
return scope["client"][0]
return None
def _credential_candidate(connection) -> _CredentialCandidate | None:
query = getattr(connection, "query_params", None) or {}
cookies = getattr(connection, "cookies", None) or {}
raw_authorization = authorization_header(connection)
authorization = raw_authorization.strip()
if raw_authorization.lower().startswith("bearer "):
value = raw_authorization[7:].strip()
if value:
return _CredentialCandidate(
value=value,
transport=CredentialTransport.HEADER,
allow_master=True,
allow_session=True,
)
# Preserve the legacy normalization contract: ``Bearer`` followed
# only by whitespace is equivalent to an empty credential channel.
elif authorization:
# Any non-empty explicit Authorization value is authoritative, even
# when its scheme is unsupported or its Bearer payload is missing.
# It must never fall through to a stale ambient cookie.
return _CredentialCandidate(
value=authorization,
transport=CredentialTransport.HEADER,
)
if _scope_type(connection) == "websocket":
ticket = _mapping_get(query, "ws_ticket").strip()
if ticket:
return _CredentialCandidate(
value=ticket,
transport=CredentialTransport.WS_TICKET,
allow_ticket=True,
)
query_key = _mapping_get(query, "api_key").strip()
if query_key:
return _CredentialCandidate(
value=query_key,
transport=CredentialTransport.QUERY,
allow_master=True,
)
session = _mapping_get(cookies, "ov_session").strip()
if session:
return _CredentialCandidate(
value=session,
transport=CredentialTransport.COOKIE,
allow_session=True,
)
legacy_key = _mapping_get(cookies, "ov_key").strip()
if legacy_key:
return _CredentialCandidate(
value=legacy_key,
transport=CredentialTransport.LEGACY_COOKIE,
allow_master=True,
)
return None
def presented_api_key(connection) -> str:
"""Compatibility extractor for the durable API-key transports only."""
candidate = _credential_candidate(connection)
if candidate is None or not candidate.allow_master:
return ""
return candidate.value
def authorization_header(connection) -> str:
headers = getattr(connection, "headers", None) or {}
return _mapping_get(headers, "authorization")
def authorization_credential_present(connection) -> bool:
"""Whether Authorization contains an authoritative credential channel.
This deliberately mirrors :func:`_credential_candidate`: whitespace and
``Bearer`` followed only by spaces are empty channels that may fall back to
legacy migration state. Unsupported schemes and ``Bearer`` without the
required separating space remain explicit invalid credentials.
"""
authorization = authorization_header(connection)
if authorization.lower().startswith("bearer ") and not authorization[7:].strip():
return False
return bool(authorization.strip())
def bearer_header_value(connection) -> str:
authorization = authorization_header(connection)
if not authorization.lower().startswith("bearer "):
return ""
return authorization[7:].strip()
def legacy_master_cookie_valid(connection) -> bool:
configured = remote_api_key()
cookies = getattr(connection, "cookies", None) or {}
supplied = _mapping_get(cookies, "ov_key").strip()
return credential_matches(supplied, configured)
def master_header_valid(connection) -> bool:
configured = remote_api_key()
supplied = bearer_header_value(connection)
return credential_matches(supplied, configured)
def _configured_pin(connection) -> str | None:
app = getattr(connection, "app", None)
state = getattr(app, "state", None) if app is not None else None
network_share = getattr(state, "network_share", None) if state is not None else None
pin = getattr(network_share, "pin", None) if network_share is not None else None
return str(pin) if pin else None
def _valid_pin(connection) -> bool:
configured = _configured_pin(connection)
if not configured:
return False
headers = getattr(connection, "headers", None) or {}
query = getattr(connection, "query_params", None) or {}
cookies = getattr(connection, "cookies", None) or {}
supplied = (
_mapping_get(headers, "x-omnivoice-pin").strip()
or _mapping_get(query, "pin").strip()
or _mapping_get(cookies, "ov_pin").strip()
)
return credential_matches(supplied, configured)
def _attached_principal(connection) -> AuthPrincipal | None:
scope = getattr(connection, "scope", None)
if not isinstance(scope, dict):
return None
state = scope.get("state")
if isinstance(state, dict):
principal = state.get(_AUTH_STATE_KEY)
return principal if isinstance(principal, AuthPrincipal) else None
return None
def _attach_principal(connection, principal: AuthPrincipal) -> AuthPrincipal:
scope = getattr(connection, "scope", None)
if isinstance(scope, dict):
state = scope.setdefault("state", {})
if isinstance(state, dict):
state[_AUTH_STATE_KEY] = principal
return principal
def resolve_principal(
connection,
*,
store: AdminSessionStore | None = None,
) -> AuthPrincipal:
"""Resolve and attach the single authentication decision for one scope."""
attached = _attached_principal(connection)
if attached is not None:
return attached
if store is None:
store = _active_admin_session_store()
host = _client_host(connection)
if is_loopback(host):
return _attach_principal(
connection,
AuthPrincipal(PrincipalKind.LOOPBACK, LOOPBACK_CAPABILITIES),
)
candidate = _credential_candidate(connection)
configured_key = remote_api_key()
if candidate is not None:
principal: AuthPrincipal | None = None
if (
candidate.allow_master
and credential_matches(candidate.value, configured_key)
):
principal = AuthPrincipal(
PrincipalKind.API_KEY,
ADMIN_CAPABILITIES,
credential_id="api-key",
transport=candidate.transport,
)
elif candidate.allow_session:
session = store.resolve(candidate.value, configured_key)
if session is not None:
principal = AuthPrincipal(
PrincipalKind.ADMIN_SESSION,
session.capabilities,
credential_id=session.credential_id,
transport=candidate.transport,
)
elif candidate.allow_ticket:
session = store.consume_ws_ticket(
candidate.value,
_canonical_websocket_path(connection),
configured_key,
)
if session is not None:
principal = AuthPrincipal(
PrincipalKind.ADMIN_SESSION,
session.capabilities,
credential_id=session.credential_id,
transport=candidate.transport,
)
if principal is not None:
return _attach_principal(connection, principal)
# An explicit, non-empty credential is authoritative. Do not silently
# fall back to network or PIN trust after an invalid higher-priority
# credential was presented.
return _attach_principal(
connection,
AuthPrincipal(
PrincipalKind.ANONYMOUS,
frozenset(),
transport=candidate.transport,
),
)
if is_local_host(host):
return _attach_principal(
connection,
AuthPrincipal(PrincipalKind.TRUSTED_NETWORK, CONSUME_CAPABILITIES),
)
if _valid_pin(connection):
return _attach_principal(
connection,
AuthPrincipal(
PrincipalKind.PIN,
CONSUME_CAPABILITIES,
transport=CredentialTransport.HEADER,
),
)
return _attach_principal(
connection,
AuthPrincipal(PrincipalKind.ANONYMOUS, frozenset()),
)
def principal_for(
connection,
*,
store: AdminSessionStore | None = None,
) -> AuthPrincipal:
return _attached_principal(connection) or resolve_principal(connection, store=store)
+140
View File
@@ -0,0 +1,140 @@
"""Exact-origin CSRF checks for ambient browser authentication."""
from __future__ import annotations
import os
from urllib.parse import SplitResult, urlsplit
CSRF_HEADER = "x-voicestudio-csrf"
CSRF_VALUE = "1"
SAFE_HTTP_METHODS = frozenset({"GET", "HEAD", "OPTIONS"})
_FORWARDED_PROTO_HEADER = "x-forwarded-proto"
def effective_scheme(connection) -> str:
"""Scheme of the client-facing hop: the resolved scope, TLS-upgraded by proxy evidence.
Behind a TLS-terminating proxy (Tailscale Serve the flagship remote-GPU
deployment in docs/remote-gpu.md nginx, Caddy, ...) the browser talks
``https`` while the backend hop is plain ``http``. uvicorn's
ProxyHeadersMiddleware (on by default in both launch paths: ``uvicorn.run``
in backend/main.py and the Docker ``python -m uvicorn`` entrypoint) already
rewrites the ASGI scope from ``X-Forwarded-Proto``, but only when the peer
is in ``--forwarded-allow-ips`` (default: loopback). That covers Serve on
bare metal, and we prefer that signal the scope is consulted first but
it misses Docker (the proxy connects from the bridge gateway) and any other
non-loopback proxy topology, so the header is honored here as well.
Spoofing analysis why honoring it never weakens a check: the upgrade is
one-way. ``https``/``wss`` as the first forwarded value promotes ``http``
to ``https``; every other value is ignored, so a forged header can never
downgrade a genuine TLS hop. For the exact-origin comparison the host:port
half of the tuple is untouched, a browser cannot attach X-Forwarded-Proto
cross-site without a CORS preflight this API never grants, and a
non-browser client able to forge the header can already forge Origin
itself it gains nothing. For cookies the upgrade can only ADD the Secure
flag (a Secure cookie set over plain http is simply dropped by the
browser the spoofer only breaks their own session), never strip it.
"""
url = getattr(connection, "url", None)
scheme = getattr(url, "scheme", None)
if not scheme:
scope = getattr(connection, "scope", None)
scheme = scope.get("scheme", "http") if isinstance(scope, dict) else "http"
scheme = {"ws": "http", "wss": "https"}.get(scheme, scheme)
if scheme != "https":
headers = getattr(connection, "headers", None) or {}
forwarded = (
headers.get(_FORWARDED_PROTO_HEADER, "") if hasattr(headers, "get") else ""
)
if forwarded.split(",")[0].strip().lower() in {"https", "wss"}:
scheme = "https"
return scheme
def _origin_tuple(value: str | None) -> tuple[str, str, int | None] | None:
if not value or value == "null":
return None
try:
parsed: SplitResult = urlsplit(value)
port = parsed.port
except (TypeError, ValueError):
return None
if (
not parsed.scheme
or not parsed.hostname
or parsed.username is not None
or parsed.password is not None
or parsed.path not in ("", "/")
or parsed.query
or parsed.fragment
):
return None
scheme = parsed.scheme.lower()
if scheme not in {"http", "https", "tauri"}:
return None
if port is None:
if scheme == "http":
port = 80
elif scheme == "https":
port = 443
return scheme, parsed.hostname.lower(), port
def configured_allowed_origins() -> frozenset[tuple[str, str, int | None]]:
raw_port = os.environ.get("OMNIVOICE_UI_PORT", "3901")
try:
ui_port = int(raw_port)
except (TypeError, ValueError):
ui_port = 3901
values = os.environ.get(
"OMNIVOICE_ALLOWED_ORIGINS",
f"http://localhost:{ui_port},http://127.0.0.1:{ui_port},"
"tauri://localhost,http://tauri.localhost",
).split(",")
return frozenset(
origin
for value in values
if (origin := _origin_tuple(value.strip())) is not None
)
def _destination_origin(connection) -> tuple[str, str, int | None] | None:
scheme = effective_scheme(connection)
url = getattr(connection, "url", None)
netloc = getattr(url, "netloc", None)
if netloc:
return _origin_tuple(f"{scheme}://{netloc}")
scope = getattr(connection, "scope", None)
headers = getattr(connection, "headers", None) or {}
if not isinstance(scope, dict):
return None
host = headers.get("host", "") if hasattr(headers, "get") else ""
return _origin_tuple(f"{scheme}://{host}")
def origin_allowed(connection) -> bool:
headers = getattr(connection, "headers", None) or {}
origin_value = headers.get("origin", "") if hasattr(headers, "get") else ""
presented = _origin_tuple(origin_value)
if presented is None:
return False
return presented == _destination_origin(connection) or presented in configured_allowed_origins()
def cookie_csrf_allowed(connection, *, side_effectful_get: bool = False) -> bool:
headers = getattr(connection, "headers", None) or {}
marker = headers.get(CSRF_HEADER, "") if hasattr(headers, "get") else ""
if marker != CSRF_VALUE or not origin_allowed(connection):
return False
method = getattr(connection, "method", None)
if method is None:
scope = getattr(connection, "scope", None)
method = scope.get("method", "GET") if isinstance(scope, dict) else "GET"
method = str(method).upper()
if side_effectful_get or method in SAFE_HTTP_METHODS:
fetch_site = headers.get("sec-fetch-site", "") if hasattr(headers, "get") else ""
return fetch_site == "same-origin"
return True
+3
View File
@@ -57,6 +57,9 @@ _BASE_SCHEMA = """
consent_recorded_at REAL DEFAULT NULL,
kind TEXT DEFAULT 'clone',
vd_states TEXT DEFAULT NULL,
-- Hosted Voice ID is opt-in synchronization metadata. Local synthesis
-- never depends on it, so existing offline profiles remain useful.
hosted_voice_id TEXT DEFAULT '',
created_at REAL
);
CREATE TABLE IF NOT EXISTS generation_history (
+49
View File
@@ -374,6 +374,33 @@ class HostCaps:
probe_ok: bool = True
"""``False`` only when torch could not be imported (degraded CPU-only)."""
requested_family: str = "auto"
"""The user's compute-device override as requested — ``"auto"`` when none.
``family`` reflects what was actually honored: an override that names a
family this host doesn't have is noted and ignored, never obeyed blindly."""
#: Every value the compute-device override accepts. "auto" = today's
#: priority pick; "cpu" is always honorable (invariant: cpu is always
#: available); accelerator names are honored only when detected.
DEVICE_OVERRIDE_CHOICES: tuple[str, ...] = ("auto", "cuda", "rocm", "xpu", "mps", "cpu")
def requested_device_override() -> str:
"""The user's compute-device pick: ``OMNIVOICE_DEVICE`` env > the Settings
choice (``compute_device`` in prefs.json) > ``"auto"``. Env wins so
power-users can pin a device without the UI silently undoing it (same
resolution order as engine selection, #981). Unknown values normalize to
``"auto"`` the probe must never raise."""
try:
from core import prefs
raw = prefs.resolve("compute_device", env="OMNIVOICE_DEVICE", default="auto")
except Exception:
raw = os.environ.get("OMNIVOICE_DEVICE", "auto")
val = str(raw or "auto").strip().lower()
return val if val in DEVICE_OVERRIDE_CHOICES else "auto"
def _probe() -> HostCaps:
"""Run the probe once. Enumerates every failure branch from the spec's
@@ -386,6 +413,7 @@ def _probe() -> HostCaps:
available_families=("cpu",),
notes=("torch not importable; treating host as CPU-only",),
probe_ok=False,
requested_family=requested_device_override(),
)
notes: list[str] = []
@@ -507,6 +535,26 @@ def _probe() -> HostCaps:
# available_families: every detected accelerator + cpu, deduped, cpu last.
available: tuple[DeviceFamily, ...] = tuple(dict.fromkeys([*detected, "cpu"]))
# User override (Settings → Performance, or OMNIVOICE_DEVICE): honored
# only when the named family actually exists on this host — an override
# can steer, it cannot invent hardware. Applied here, at the single
# choke point, so routing, model loads (get_best_device delegates its
# family decision here), and every badge inherit it for free.
requested = requested_device_override()
if requested != "auto":
if requested in available:
if requested != family:
notes.append(
f"compute device pinned to '{requested}' by user override "
f"(auto would pick '{family}')"
)
family = requested # type: ignore[assignment]
else:
notes.append(
f"requested compute device '{requested}' is not available on "
f"this host (have: {', '.join(available)}) — using '{family}'"
)
return HostCaps(
family=family,
available_families=available,
@@ -515,6 +563,7 @@ def _probe() -> HostCaps:
driver=driver,
notes=tuple(notes),
probe_ok=True,
requested_family=requested,
)
+125
View File
@@ -0,0 +1,125 @@
"""Startup progress ledger — what the backend is doing before it can serve.
Why this exists: the project's #1 lifetime failure class is "can't reach the
local backend", and a large slice of it was never a dead backend at all —
just one that couldn't say "I'm starting, currently loading PyTorch" because
nothing listened until every heavy import and migration finished. main.py now
binds the socket early and defers the heavy work; this module is the shared
state the early `/health` + `/startup/progress` endpoints report from while
that work runs.
Thread-safety: the deferred init runs Phase A in an executor thread while the
event loop serves probes, so every mutation and snapshot takes the lock.
"""
from __future__ import annotations
import threading
import time
# Execution order matters only for display; the ledger records whatever order
# steps actually begin in. Keep ids stable — the desktop shell field-sniffs
# them and tests pin them.
STEPS: "dict[str, str]" = {
"env_prefs": "Restoring settings…",
"native_preload": "Preparing GPU libraries…",
"ml_imports": "Loading ML runtime (PyTorch)…",
"api_routes": "Loading API routes…",
"db_migrate": "Preparing database…",
"services_start": "Starting background services…",
}
_lock = threading.Lock()
_t0 = time.monotonic()
_current: "str | None" = None
_done: "list[tuple[str, float]]" = [] # (step_id, seconds it took)
_started_at: float = 0.0
_ready = False
_error: "dict | None" = None
def begin_step(step_id: str) -> None:
global _current, _started_at
with _lock:
_finish_current_locked()
_current = step_id
_started_at = time.monotonic()
def _finish_current_locked() -> None:
global _current
if _current is not None:
_done.append((_current, round(time.monotonic() - _started_at, 2)))
_current = None
def mark_ready() -> None:
global _ready
with _lock:
_finish_current_locked()
_ready = True
def fail(message: str) -> None:
"""Record a startup failure against the step that was running."""
global _error
with _lock:
_error = {"step": _current, "message": str(message)[:500]}
def is_ready() -> bool:
with _lock:
return _ready
def current_step() -> "tuple[str | None, str | None]":
"""(step_id, human label) of the active step, or (None, None)."""
with _lock:
if _current is None:
return None, None
return _current, STEPS.get(_current, _current)
def snapshot() -> dict:
"""The `/startup/progress` body. Always safe to call, never raises."""
with _lock:
if _error is not None:
status = "failed"
elif _ready:
status = "ready"
else:
status = "starting"
states = {sid: "pending" for sid in STEPS}
for sid, _t in _done:
states[sid] = "done"
if _current is not None:
states[_current] = "active"
if _error is not None and _error.get("step"):
states[_error["step"]] = "failed"
durations = dict(_done)
return {
"status": status,
"step": _current,
"label": STEPS.get(_current, _current) if _current else None,
"steps": [
{
"id": sid,
"label": label,
"state": states.get(sid, "pending"),
**({"t": durations[sid]} if sid in durations else {}),
}
for sid, label in STEPS.items()
],
"elapsed_s": round(time.monotonic() - _t0, 2),
"error": _error,
}
def _reset_for_tests() -> None:
global _current, _ready, _error, _started_at
with _lock:
_current = None
_done.clear()
_ready = False
_error = None
_started_at = 0.0
+1 -1
View File
@@ -24,7 +24,7 @@ from pathlib import Path
# tests/test_app_version.py::test_all_version_files_in_lockstep and bumped by
# release.yml's version-bump job, so it stays equal to
# pyproject/tauri.conf/Cargo/package.json.
_FALLBACK_VERSION = "0.4.2"
_FALLBACK_VERSION = "0.5.0"
def _fallback_version() -> str:
+19 -3
View File
@@ -85,11 +85,27 @@ def _get_model():
global _model
if _model is None:
from faster_whisper import WhisperModel
name = os.environ.get("ASR_MODEL_FW", "large-v3")
# Same weights as in-process faster-whisper: ASR_MODEL_FASTER selects
# for BOTH variants, ASR_MODEL_FW stays as a sidecar-only override.
# Before this, the sidecar read only ASR_MODEL_FW while the download
# preflight read ASR_MODEL_FASTER — set one and the other variant (or
# the preflight) quietly used a different model.
name = (
os.environ.get("ASR_MODEL_FW")
or os.environ.get("ASR_MODEL_FASTER")
or "large-v3"
)
try:
import torch
device = "cuda" if torch.cuda.is_available() else "cpu"
# The probe honors the user compute-device override and the
# ROCm/CT2 incompatibility (#1529) — the child must agree with
# the parent's device decision, not re-derive its own.
from core.device_caps import detect_host_caps
device = "cuda" if detect_host_caps().family == "cuda" else "cpu"
except Exception:
# Fail SAFE: guessing "cuda" from torch here would bypass a cpu
# override and hand CTranslate2 HIP-flavoured cuda on ROCm
# (#1529). CPU always works; say why in the sidecar log.
print("asr-sidecar: device probe failed — using cpu", file=sys.stderr, flush=True)
device = "cpu"
# Degrade fp16 → int8 rather than crash on GPUs without efficient fp16
# (older Maxwell/Pascal, GTX 16xx, CTranslate2/cuDNN mismatch) (#551).
@@ -353,6 +353,9 @@ def _make_backend_class():
display_name = "OmniVoice (GGUF, hardware-adaptive)"
gpu_compat = ("cuda", "mps", "cpu")
supports_voice_design = False
# Every generate() spawns the external binary — allocations live in
# that process, invisible to parent-side accelerator counters.
runs_out_of_process = True
# 24 kHz mono Higgs Audio v2 — same as the in-process OmniVoice.
_SAMPLE_RATE = 24_000
+704 -447
View File
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,30 @@
"""Opt-in hosted Voice ID on local profiles.
Revision ID: 0011_hosted_voice_sync
Revises: 0010_remote_worker_schema
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0011_hosted_voice_sync"
down_revision: Union[str, None] = "0010_remote_worker_schema"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def _has_column(table: str, column: str) -> bool:
rows = op.get_bind().execute(sa.text(f"PRAGMA table_info({table})")).fetchall()
return any(row[1] == column for row in rows)
def upgrade() -> None:
if not _has_column("voice_profiles", "hosted_voice_id"):
op.add_column("voice_profiles", sa.Column("hosted_voice_id", sa.Text(), nullable=True, server_default=""))
def downgrade() -> None:
if _has_column("voice_profiles", "hosted_voice_id"):
op.drop_column("voice_profiles", "hosted_voice_id")
@@ -0,0 +1,54 @@
"""Mark materialized gallery archetypes as voice-design profiles.
Revision ID: 0012_mark_archetype_profiles_design
Revises: 0011_hosted_voice_sync
Create Date: 2026-08-15 00:00:00.000000
``POST /archetypes/{id}/use`` stores the archetype id in ``personality`` and
also stores a locally rendered identity WAV. That WAV must not make the
profile a clone: the archetype's instruct recipe is authoritative. Older
rows relied on the ``kind='clone'`` default and therefore selected the clone
generation path. This data-only migration fixes every row whose personality
is a current archetype id, leaving unrelated persona and marketplace imports
untouched.
"""
from typing import Sequence, Union
from alembic import op
from sqlalchemy import inspect
revision: str = "0012_mark_archetype_profiles_design"
down_revision: Union[str, None] = "0011_hosted_voice_sync"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
bind = op.get_bind()
inspector = inspect(bind)
if "voice_profiles" not in inspector.get_table_names():
return
columns = {column["name"] for column in inspector.get_columns("voice_profiles")}
if not {"kind", "personality"}.issubset(columns):
return
# The catalog is intentionally a value object, so checking an id against
# its current generated list is the precise provenance test. The
# parameterized update avoids treating any other personality string as an
# archetype.
from core import archetypes
archetype_ids = [item["id"] for item in archetypes.list_archetypes()]
for archetype_id in archetype_ids:
bind.exec_driver_sql(
"UPDATE voice_profiles SET kind = 'design' "
"WHERE personality = ? AND (kind IS NULL OR kind = '' OR kind = 'clone')",
(archetype_id,),
)
def downgrade() -> None:
# Do not silently convert voice-design profiles back to clones: that would
# reintroduce the generation mismatch for existing user data.
pass
+95
View File
@@ -0,0 +1,95 @@
# VoiceStudio runtime adapter
A local gRPC server implementing the vssaas GPU-node runtime contract
`voicestudio.runtime.v1.RuntimeAdapterService`, so a vssaas GPU Gateway can
drive this VoiceStudio backend as its inference runtime.
## Boundary (deliberate non-capabilities)
- Binds **only** a Unix-domain socket (default `/run/voicestudio/runtime.sock`,
override with `VOICE_STUDIO_RUNTIME_SOCKET`). No HTTP listener, no TCP.
- Never reaches PostgreSQL, customer credentials, or arbitrary network URLs.
`Execute` accepts **local file handles only** — absolute paths generated by
the Gateway; any URL-shaped or relative handle is rejected as invalid input.
- The Gateway owns leases, artifact transfer, retries, and billing. This
adapter owns approved model loading and inference only.
## Running
```sh
# serve (production socket):
VOICE_STUDIO_RUNTIME_SOCKET=/run/voicestudio/runtime.sock \
python -m backend.runtime_adapter
# self-check: starts the server on a private temp socket and validates the
# same expectations the Go preflight (cmd/runtime-adapter-preflight) enforces:
python -m backend.runtime_adapter --selfcheck
```
Environment:
| Variable | Default | Meaning |
| --- | --- | --- |
| `VOICE_STUDIO_RUNTIME_SOCKET` | `/run/voicestudio/runtime.sock` | Unix socket path (must be absolute; parent dir must exist and not be world-writable). |
| `VOICE_STUDIO_RUNTIME_SLOTS` | `1` | Concurrent execution slots per device. |
## Wire contract and generated stubs
`runtime_adapter.proto` is a **byte-identical vendored copy** of the vssaas
contract `api/proto/voicestudio/runtime/v1/runtime_adapter.proto`. Do not edit
it here; re-vendor from vssaas when the contract changes, then regenerate.
The `gen/` stubs are committed (same policy as `backend/worker/protocol/gen/`).
Regenerate with:
```sh
uv run python scripts/gen_runtime_adapter_protocol.py
```
`tests/test_runtime_adapter_gen.py` fails if the committed stubs drift from
the proto.
## Preflight expectations honoured
The Go preflight (`internal/gateway/preflight.go`) fails closed unless:
- the socket path is absolute, a real Unix socket (not a symlink), and its
parent directory is not world-writable — `server.prepare_socket` enforces
the same rules at bind time;
- `Health` returns `SERVING_STATE_READY` with nonempty runtime + adapter
versions, and `GetCapabilities` returns **identical** versions — both
handlers read the same constants, so they cannot disagree;
- at least one device with nonempty id/hardware class, nonzero VRAM and
slots, `free_slots <= total_slots`, unique ids;
- at least one model **explicitly READY** with `catalog_model_id`,
`model_version`, `model_digest`, and ≥1 precision. A loading, installed,
or failed model is reported with its true state and never as READY.
## Model identity
- `catalog_model_id` — the VoiceStudio TTS engine id (`omnivoice`,
`voxcpm2`, …) from `services.tts_backend`'s registry.
- `model_version` — an immutable catalog version comprising the installed
Hugging Face revision (40-char commit SHA) and the first 16 hex characters
of the attested snapshot digest. This creates a new catalog identity when
snapshot bytes change; it never rewrites an identity retained by a Job.
- `model_digest``sha256:<hex>` computed over the installed snapshot files
(sorted relative path + per-file SHA-256), cached next to the repo cache
keyed by (revision, file list, sizes, mtimes) so multi-GB weights are
hashed once. See `digest.py`.
## Failure taxonomy
Stable codes (prefix `RTA_`) map onto the proto's `RuntimeFailureClass`:
invalid input (`RTA_INPUT_*`), model load (`RTA_MODEL_LOAD_FAILED`),
inference (`RTA_INFERENCE_*`), GPU resource (`RTA_GPU_*`), local storage
(`RTA_STORAGE_*`), cancellation (terminal `ExecutionCanceled`), and adapter
crash (`RTA_RUNTIME_CRASH`). See `codes.py`.
## Tests
```sh
uv run pytest backend/tests/test_runtime_adapter_capabilities.py \
backend/tests/test_runtime_adapter_execute.py \
tests/test_runtime_adapter_gen.py
```
+18
View File
@@ -0,0 +1,18 @@
"""VoiceStudio runtime adapter — the vssaas GPU-node runtime boundary.
Implements ``voicestudio.runtime.v1.RuntimeAdapterService`` over a private
Unix-domain socket so a vssaas GPU Gateway can drive VoiceStudio's TTS
engines as its inference runtime. No HTTP listener, no database access, no
outbound network: the adapter reads and writes only the local file handles
each ``Execute`` request carries. See ``README.md`` in this directory.
"""
from __future__ import annotations
#: Version of this adapter layer (the gRPC boundary), independent of the app
#: version, which is reported as ``runtime_version``. Bump on any behavioral
#: change to the adapter itself.
ADAPTER_VERSION = "0.1.0"
DEFAULT_SOCKET_PATH = "/run/voicestudio/runtime.sock"
SOCKET_ENV = "VOICE_STUDIO_RUNTIME_SOCKET"
SLOTS_ENV = "VOICE_STUDIO_RUNTIME_SLOTS"
+65
View File
@@ -0,0 +1,65 @@
"""Entry point: ``python -m backend.runtime_adapter``.
Serves the runtime adapter on a private Unix-domain socket (default
``/run/voicestudio/runtime.sock``, override ``VOICE_STUDIO_RUNTIME_SOCKET``
or ``--socket``). ``--selfcheck`` instead starts the server on a temp socket
and validates the GPU Gateway preflight expectations against it.
"""
from __future__ import annotations
import argparse
import sys
from ._paths import ensure_backend_on_path
def main(argv: list[str] | None = None) -> int:
ensure_backend_on_path()
parser = argparse.ArgumentParser(
prog="backend.runtime_adapter",
description="VoiceStudio runtime adapter (vssaas GPU-node gRPC server)",
)
parser.add_argument(
"--socket",
default=None,
help="absolute Unix socket path (default: $VOICE_STUDIO_RUNTIME_SOCKET "
"or /run/voicestudio/runtime.sock)",
)
parser.add_argument(
"--selfcheck",
action="store_true",
help="start on a temp socket and validate the preflight expectations",
)
parser.add_argument(
"--timeout",
type=float,
default=10.0,
help="selfcheck RPC timeout in seconds (default: 10)",
)
parser.add_argument(
"--no-prewarm",
action="store_true",
help="serve immediately without loading models first (the first "
"execution then pays weight loading and compilation)",
)
args = parser.parse_args(argv)
if args.selfcheck:
from .selfcheck import selfcheck # noqa: PLC0415
return selfcheck(timeout_s=args.timeout)
from .production import build_runtime_context, prewarm_engines # noqa: PLC0415
from .server import resolve_socket_path, serve # noqa: PLC0415
context = build_runtime_context()
if not args.no_prewarm:
# Deliberately before the socket exists: the Gateway's preflight and
# first offer should both find a runtime that can start inference at
# once, rather than one that spends an attempt lease compiling.
prewarm_engines(context)
return serve(context, resolve_socket_path(args.socket))
if __name__ == "__main__":
sys.exit(main())
+20
View File
@@ -0,0 +1,20 @@
"""Import-path bootstrap for running outside the FastAPI app.
The backend is laid out to run with ``--app-dir backend`` (imports like
``services.tts_backend`` resolve against the ``backend/`` directory). When
the adapter is launched as ``python -m backend.runtime_adapter`` from the
repo root, ``backend/`` is a namespace package but not on ``sys.path`` so
call :func:`ensure_backend_on_path` before any ``services.*`` / ``core.*``
import. Idempotent; mirrors ``backend/tests/conftest.py``.
"""
from __future__ import annotations
import os
import sys
def ensure_backend_on_path() -> str:
backend_dir = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
if backend_dir not in sys.path:
sys.path.insert(0, backend_dir)
return backend_dir
+163
View File
@@ -0,0 +1,163 @@
"""Stable failure codes and exception classification for Execute.
The vssaas API Gateway keys retry and customer-charge policy off these codes,
so they are a wire contract: never rename an existing code, only add. Every
code maps to exactly one proto ``RuntimeFailureClass``.
"""
from __future__ import annotations
import re
from .gen import runtime_adapter_pb2 as pb2
# ── invalid approved input ────────────────────────────────────────────────
INPUT_ATTEMPT_IDENTITY = "RTA_INPUT_ATTEMPT_IDENTITY"
INPUT_ATTEMPT_DUPLICATE = "RTA_INPUT_ATTEMPT_DUPLICATE"
INPUT_MODEL_UNKNOWN = "RTA_INPUT_MODEL_UNKNOWN"
INPUT_MODEL_NOT_READY = "RTA_INPUT_MODEL_NOT_READY"
INPUT_MODEL_DIGEST_MISMATCH = "RTA_INPUT_MODEL_DIGEST_MISMATCH"
INPUT_MODEL_PRECISION = "RTA_INPUT_MODEL_PRECISION_UNSUPPORTED"
INPUT_DEVICE_UNKNOWN = "RTA_INPUT_DEVICE_UNKNOWN"
INPUT_HANDLE_INVALID = "RTA_INPUT_HANDLE_INVALID"
INPUT_ARTIFACTS_INVALID = "RTA_INPUT_ARTIFACTS_INVALID"
INPUT_CHECKSUM_MISMATCH = "RTA_INPUT_CHECKSUM_MISMATCH"
INPUT_TEXT_EMPTY = "RTA_INPUT_TEXT_EMPTY"
INPUT_TEXT_TOO_LARGE = "RTA_INPUT_TEXT_TOO_LARGE"
INPUT_TEXT_ENCODING = "RTA_INPUT_TEXT_ENCODING"
INPUT_PARAMETER_UNKNOWN = "RTA_INPUT_PARAMETER_UNKNOWN"
INPUT_PARAMETER_TYPE = "RTA_INPUT_PARAMETER_TYPE"
INPUT_PARAMETER_RANGE = "RTA_INPUT_PARAMETER_RANGE"
INPUT_DEADLINE_INVALID = "RTA_INPUT_DEADLINE_INVALID"
INPUT_REJECTED = "RTA_INPUT_REJECTED" # engine-level TTSInputError
# ── model load / inference ────────────────────────────────────────────────
MODEL_LOAD_FAILED = "RTA_MODEL_LOAD_FAILED"
MODEL_LOAD_DEADLINE = "RTA_MODEL_LOAD_DEADLINE_EXCEEDED"
INFERENCE_FAILED = "RTA_INFERENCE_FAILED"
INFERENCE_BAD_OUTPUT = "RTA_INFERENCE_BAD_OUTPUT"
INFERENCE_DEADLINE = "RTA_INFERENCE_DEADLINE_EXCEEDED"
# ── GPU resource ──────────────────────────────────────────────────────────
GPU_OUT_OF_MEMORY = "RTA_GPU_OUT_OF_MEMORY"
GPU_SLOTS_EXHAUSTED = "RTA_GPU_SLOTS_EXHAUSTED"
# ── local storage ─────────────────────────────────────────────────────────
STORAGE_READ_FAILED = "RTA_STORAGE_READ_FAILED"
STORAGE_WRITE_FAILED = "RTA_STORAGE_WRITE_FAILED"
# ── adapter crash ─────────────────────────────────────────────────────────
RUNTIME_CRASH = "RTA_RUNTIME_CRASH"
_INPUT = pb2.RUNTIME_FAILURE_CLASS_INPUT
_MODEL_LOAD = pb2.RUNTIME_FAILURE_CLASS_MODEL_LOAD
_INFERENCE = pb2.RUNTIME_FAILURE_CLASS_INFERENCE
_GPU = pb2.RUNTIME_FAILURE_CLASS_GPU_RESOURCE
_STORAGE = pb2.RUNTIME_FAILURE_CLASS_LOCAL_STORAGE
_RUNTIME = pb2.RUNTIME_FAILURE_CLASS_RUNTIME
CODE_CLASS: dict[str, int] = {
INPUT_ATTEMPT_IDENTITY: _INPUT,
INPUT_ATTEMPT_DUPLICATE: _INPUT,
INPUT_MODEL_UNKNOWN: _INPUT,
INPUT_MODEL_NOT_READY: _INPUT,
INPUT_MODEL_DIGEST_MISMATCH: _INPUT,
INPUT_MODEL_PRECISION: _INPUT,
INPUT_DEVICE_UNKNOWN: _INPUT,
INPUT_HANDLE_INVALID: _INPUT,
INPUT_ARTIFACTS_INVALID: _INPUT,
INPUT_CHECKSUM_MISMATCH: _INPUT,
INPUT_TEXT_EMPTY: _INPUT,
INPUT_TEXT_TOO_LARGE: _INPUT,
INPUT_TEXT_ENCODING: _INPUT,
INPUT_PARAMETER_UNKNOWN: _INPUT,
INPUT_PARAMETER_TYPE: _INPUT,
INPUT_PARAMETER_RANGE: _INPUT,
INPUT_DEADLINE_INVALID: _INPUT,
INPUT_REJECTED: _INPUT,
MODEL_LOAD_FAILED: _MODEL_LOAD,
MODEL_LOAD_DEADLINE: _MODEL_LOAD,
INFERENCE_FAILED: _INFERENCE,
INFERENCE_BAD_OUTPUT: _INFERENCE,
INFERENCE_DEADLINE: _INFERENCE,
GPU_OUT_OF_MEMORY: _GPU,
GPU_SLOTS_EXHAUSTED: _GPU,
STORAGE_READ_FAILED: _STORAGE,
STORAGE_WRITE_FAILED: _STORAGE,
RUNTIME_CRASH: _RUNTIME,
}
class ExecutionFailure(Exception):
"""A classified, wire-safe execution failure."""
def __init__(self, stable_code: str, safe_detail: str = ""):
if stable_code not in CODE_CLASS: # programming error, not a wire case
raise ValueError(f"unknown stable code {stable_code!r}")
super().__init__(stable_code)
self.stable_code = stable_code
self.failure_class = CODE_CLASS[stable_code]
self.safe_detail = scrub_detail(safe_detail)
_PATHISH = re.compile(r"(?:[A-Za-z]:)?[/\\][^\s'\"]+")
_MAX_DETAIL = 240
def scrub_detail(detail: str) -> str:
"""Bound and de-path a detail string before it crosses the wire.
Local handles are server-generated, but engine exceptions routinely embed
checkpoint paths, cache dirs, and home directories. None of that belongs
in an event the Gateway relays upstream.
"""
scrubbed = _PATHISH.sub("<path>", detail or "").strip()
return scrubbed[:_MAX_DETAIL]
_OOM_MARKERS = (
"out of memory",
"cuda error: out of memory",
"mps backend out of memory",
"hip out of memory",
"cublas_status_alloc_failed",
)
def _is_oom(exc: BaseException) -> bool:
if type(exc).__name__ == "OutOfMemoryError": # torch.cuda.OutOfMemoryError
return True
message = str(exc).lower()
return any(marker in message for marker in _OOM_MARKERS)
def _is_engine_input_error(exc: BaseException) -> bool:
try:
from services.tts_backend import TTSInputError # noqa: PLC0415
except Exception:
return False
return isinstance(exc, TTSInputError)
def classify_engine_error(exc: BaseException, phase: str) -> ExecutionFailure:
"""Map an engine exception to a stable failure code.
``phase`` is ``"model_load"`` or ``"synthesis"`` the phase the engine
thread was in when it raised.
"""
if isinstance(exc, ExecutionFailure):
return exc
detail = f"{type(exc).__name__}: {exc}"
if _is_oom(exc):
return ExecutionFailure(GPU_OUT_OF_MEMORY, detail)
if _is_engine_input_error(exc):
return ExecutionFailure(INPUT_REJECTED, detail)
if isinstance(exc, OSError):
return ExecutionFailure(STORAGE_READ_FAILED, detail)
if phase == "model_load":
return ExecutionFailure(MODEL_LOAD_FAILED, detail)
return ExecutionFailure(INFERENCE_FAILED, detail)
def deadline_failure(phase: str) -> ExecutionFailure:
code = MODEL_LOAD_DEADLINE if phase == "model_load" else INFERENCE_DEADLINE
return ExecutionFailure(code, "attempt deadline exceeded")
+112
View File
@@ -0,0 +1,112 @@
"""Stable digests for locally installed model snapshots.
``model_digest`` in the wire contract pins the exact bytes a READY model will
execute with. Hugging Face snapshots are symlink farms into ``blobs/``, so the
digest is computed over the *resolved* file contents: SHA-256 of the sorted
sequence ``<posix relpath>\\n<file sha256>\\n``. That is stable across hosts,
cache locations, and symlink layout, and changes whenever any weight byte or
the file set changes.
Hashing multi-GB weights on every ``GetCapabilities`` call would be absurd, so
the result is cached in a JSON sidecar keyed by a cheap fingerprint of the
file list (relpath, size, mtime_ns). Any file change invalidates the cache and
forces a full re-hash.
"""
from __future__ import annotations
import hashlib
import json
import os
from pathlib import Path
DIGEST_PREFIX = "sha256:"
_CHUNK = 1024 * 1024
def file_sha256(path: str | os.PathLike[str]) -> str:
hasher = hashlib.sha256()
with open(path, "rb") as fh:
while True:
chunk = fh.read(_CHUNK)
if not chunk:
break
hasher.update(chunk)
return hasher.hexdigest()
def _manifest(root: Path) -> list[tuple[str, int, int]]:
"""Sorted (relpath, size, mtime_ns) for every regular file under root.
Follows symlinks (HF snapshot layout); a dangling symlink raises
``FileNotFoundError`` callers treat that as an incomplete install.
"""
entries: list[tuple[str, int, int]] = []
for current, dirs, files in os.walk(root, followlinks=True):
dirs.sort()
for name in sorted(files):
path = Path(current) / name
stat = path.stat() # resolves symlinks; raises if dangling
rel = path.relative_to(root).as_posix()
entries.append((rel, stat.st_size, stat.st_mtime_ns))
entries.sort()
return entries
def _fingerprint(entries: list[tuple[str, int, int]]) -> str:
return hashlib.sha256(
json.dumps(entries, separators=(",", ":")).encode("utf-8")
).hexdigest()
def snapshot_digest(root: str | os.PathLike[str], cache_path: str | os.PathLike[str] | None = None) -> str:
"""``sha256:<hex>`` digest of the snapshot at ``root``.
Raises ``FileNotFoundError`` for a missing/empty snapshot or dangling
symlink and ``OSError`` for unreadable files callers classify those as
not-READY rather than fabricating a digest.
"""
root = Path(root)
entries = _manifest(root)
if not entries:
raise FileNotFoundError(f"empty model snapshot: {root}")
fingerprint = _fingerprint(entries)
if cache_path is not None:
cached = _read_cache(cache_path)
if cached is not None and cached.get("fingerprint") == fingerprint:
digest = cached.get("digest", "")
if isinstance(digest, str) and digest.startswith(DIGEST_PREFIX):
return digest
hasher = hashlib.sha256()
for rel, _size, _mtime in entries:
hasher.update(rel.encode("utf-8"))
hasher.update(b"\n")
hasher.update(file_sha256(root / rel).encode("ascii"))
hasher.update(b"\n")
digest = DIGEST_PREFIX + hasher.hexdigest()
if cache_path is not None:
_write_cache(cache_path, fingerprint, digest)
return digest
def _read_cache(cache_path: str | os.PathLike[str]) -> dict | None:
try:
with open(cache_path, encoding="utf-8") as fh:
data = json.load(fh)
return data if isinstance(data, dict) else None
except (OSError, ValueError):
return None
def _write_cache(cache_path: str | os.PathLike[str], fingerprint: str, digest: str) -> None:
cache_path = Path(cache_path)
payload = json.dumps({"fingerprint": fingerprint, "digest": digest})
try:
cache_path.parent.mkdir(parents=True, exist_ok=True)
temporary = cache_path.with_suffix(f".tmp-{os.getpid()}")
temporary.write_text(payload, encoding="utf-8")
os.replace(temporary, cache_path)
except OSError:
pass # cache is an optimization; the digest itself is already computed
+664
View File
@@ -0,0 +1,664 @@
"""Execute/Cancel: attempt registry, validation, and the event stream.
One ``Execute`` call is one *attempt*. The generator emits::
started progress* exactly one of completed | failed | canceled
The engine call itself (``ensure_ready`` + ``generate``) runs on a daemon
worker thread; the streaming generator polls it, emitting bounded heartbeat
progress and enforcing the request deadline and cancellation. A blocking
engine cannot be interrupted mid-kernel, so on cancel/deadline the thread is
abandoned and its result discarded the terminal event is what the Gateway
acts on, and slot accounting is released only when the thread actually exits.
The adapter never turns a customer string into a filesystem path: it touches
exactly the local handles the request carries, after validation.
"""
from __future__ import annotations
import os
import threading
import time
from collections import OrderedDict
from dataclasses import dataclass, field
from . import codes
from ._paths import ensure_backend_on_path
from .digest import file_sha256
from .gen import runtime_adapter_pb2 as pb2
from .inventory import STATE_READY
_MAX_TEXT_BYTES = 512_000
_MAX_REF_AUDIO_BYTES = 100 * 1024 * 1024
_MAX_DEADLINE_S = 24 * 3600.0
_MAX_PROGRESS_EVENTS = 512
#: Typed, bounded Execute parameters → the engine ``generate()`` kwarg of the
#: same name. Kinds: ("string", max_len) / ("integer", lo, hi) /
#: ("number", lo, hi) / ("boolean",).
PARAMETER_SPECS: dict[str, tuple] = {
"language": ("string", 32),
"ref_text": ("string", 4096),
"instruct": ("string", 2048),
"description": ("string", 2048),
"speed": ("number", 0.25, 4.0),
"guidance_scale": ("number", 0.0, 16.0),
"num_step": ("integer", 1, 128),
# Gallery reference voices persist their OSS design seed. Accept it at
# the hosted runtime boundary so a selected voice produces the same take.
"seed": ("integer", 0, 4_294_967_295),
}
# ── attempt registry ──────────────────────────────────────────────────────
@dataclass
class AttemptRecord:
job_id: str
attempt_id: str
cancel: threading.Event = field(default_factory=threading.Event)
terminal: str | None = None # "completed" | "failed" | "canceled"
class AttemptRegistry:
"""Attempt bookkeeping: admission, idempotent cancel, bounded history."""
def __init__(self, max_terminal: int = 4096):
self._lock = threading.Lock()
self._active: dict[str, AttemptRecord] = {}
self._terminal: OrderedDict[str, AttemptRecord] = OrderedDict()
self._max_terminal = max_terminal
def begin(self, job_id: str, attempt_id: str, slot_limit: int) -> AttemptRecord:
with self._lock:
if attempt_id in self._active or attempt_id in self._terminal:
raise codes.ExecutionFailure(
codes.INPUT_ATTEMPT_DUPLICATE, "attempt id already used"
)
if len(self._active) >= max(1, slot_limit):
raise codes.ExecutionFailure(
codes.GPU_SLOTS_EXHAUSTED, "no free execution slot"
)
record = AttemptRecord(job_id=job_id, attempt_id=attempt_id)
self._active[attempt_id] = record
return record
def finish(self, attempt_id: str, terminal: str) -> None:
with self._lock:
record = self._active.pop(attempt_id, None)
if record is None:
return
record.terminal = terminal
self._terminal[attempt_id] = record
while len(self._terminal) > self._max_terminal:
self._terminal.popitem(last=False)
def active_count(self) -> int:
with self._lock:
return len(self._active)
def cancel(self, job_id: str, attempt_id: str) -> int:
"""Idempotent by attempt id; returns a proto CancelDisposition."""
with self._lock:
record = self._active.get(attempt_id)
if record is not None:
if job_id and record.job_id and job_id != record.job_id:
return pb2.CANCEL_DISPOSITION_NOT_FOUND
record.cancel.set()
return pb2.CANCEL_DISPOSITION_ACCEPTED
record = self._terminal.get(attempt_id)
if record is not None:
if job_id and record.job_id and job_id != record.job_id:
return pb2.CANCEL_DISPOSITION_NOT_FOUND
return pb2.CANCEL_DISPOSITION_ALREADY_TERMINAL
return pb2.CANCEL_DISPOSITION_NOT_FOUND
# ── request validation ────────────────────────────────────────────────────
@dataclass
class ValidatedRequest:
text: str
output_handle: str
output_media_type: str
output_size_bound: int
engine_kwargs: dict
deadline_monotonic: float
catalog_model_id: str
def _validate_handle(handle: str, code: str = codes.INPUT_HANDLE_INVALID) -> str:
cleaned = (handle or "").strip()
if (
not cleaned
or "\x00" in cleaned
or "://" in cleaned
or not os.path.isabs(cleaned)
or os.path.normpath(cleaned) != cleaned
):
raise codes.ExecutionFailure(code, "local handle must be an absolute path")
return cleaned
def _read_input_file(artifact, max_bytes: int) -> bytes:
path = _validate_handle(artifact.local_handle)
try:
stat = os.lstat(path)
except OSError as exc:
raise codes.ExecutionFailure(
codes.STORAGE_READ_FAILED, f"input handle unreadable: {type(exc).__name__}"
)
import stat as stat_module # noqa: PLC0415
if not stat_module.S_ISREG(stat.st_mode):
raise codes.ExecutionFailure(
codes.INPUT_HANDLE_INVALID, "input handle must be a regular file"
)
bound = max_bytes
if 0 < artifact.expected_size_bytes <= max_bytes:
bound = artifact.expected_size_bytes
if stat.st_size > bound:
raise codes.ExecutionFailure(
codes.INPUT_TEXT_TOO_LARGE, "input exceeds its size bound"
)
try:
with open(path, "rb") as fh:
data = fh.read(bound + 1)
except OSError as exc:
raise codes.ExecutionFailure(
codes.STORAGE_READ_FAILED, f"input read failed: {type(exc).__name__}"
)
if len(data) > bound:
raise codes.ExecutionFailure(
codes.INPUT_TEXT_TOO_LARGE, "input exceeds its size bound"
)
expected = (artifact.expected_sha256 or "").strip().lower().removeprefix("sha256:")
if expected:
import hashlib # noqa: PLC0415
if hashlib.sha256(data).hexdigest() != expected:
raise codes.ExecutionFailure(
codes.INPUT_CHECKSUM_MISMATCH, "input checksum mismatch"
)
return data
def _typed_parameter(name: str, value) -> object:
spec = PARAMETER_SPECS.get(name)
if spec is None:
raise codes.ExecutionFailure(
codes.INPUT_PARAMETER_UNKNOWN, f"unknown parameter {name!r}"
)
kind = spec[0]
which = value.WhichOneof("value")
if kind == "string":
if which != "string_value":
raise codes.ExecutionFailure(
codes.INPUT_PARAMETER_TYPE, f"parameter {name!r} must be a string"
)
text = value.string_value
if len(text) > spec[1]:
raise codes.ExecutionFailure(
codes.INPUT_PARAMETER_RANGE, f"parameter {name!r} too long"
)
return text
if kind == "integer":
if which != "integer_value":
raise codes.ExecutionFailure(
codes.INPUT_PARAMETER_TYPE, f"parameter {name!r} must be an integer"
)
number = value.integer_value
if not spec[1] <= number <= spec[2]:
raise codes.ExecutionFailure(
codes.INPUT_PARAMETER_RANGE, f"parameter {name!r} out of range"
)
return int(number)
if kind == "number":
if which == "number_value":
number = value.number_value
elif which == "integer_value":
number = float(value.integer_value)
else:
raise codes.ExecutionFailure(
codes.INPUT_PARAMETER_TYPE, f"parameter {name!r} must be a number"
)
if not spec[1] <= number <= spec[2]:
raise codes.ExecutionFailure(
codes.INPUT_PARAMETER_RANGE, f"parameter {name!r} out of range"
)
return float(number)
if which != "boolean_value":
raise codes.ExecutionFailure(
codes.INPUT_PARAMETER_TYPE, f"parameter {name!r} must be a boolean"
)
return bool(value.boolean_value)
# ── the executor ──────────────────────────────────────────────────────────
class Executor:
"""Validates and runs attempts against an inventory + engine provider."""
def __init__(
self,
inventory,
engine_provider,
registry: AttemptRegistry,
*,
slot_limit: int = 1,
progress_interval: float = 0.5,
poll_interval: float = 0.02,
clock=time.monotonic,
):
self._inventory = inventory
self._engine_provider = engine_provider
self._registry = registry
self._slot_limit = max(1, slot_limit)
self._progress_interval = progress_interval
self._poll_interval = poll_interval
self._clock = clock
# -- validation ----------------------------------------------------
def _validate(self, request) -> ValidatedRequest:
now_ms = int(time.time() * 1000)
if request.deadline_unix_ms <= now_ms:
raise codes.ExecutionFailure(
codes.INPUT_DEADLINE_INVALID, "deadline is not in the future"
)
budget_s = min((request.deadline_unix_ms - now_ms) / 1000.0, _MAX_DEADLINE_S)
model = self._validate_model(request.model)
self._validate_device(request.device_id)
text_artifact, ref_artifact = self._split_inputs(request.inputs)
output = self._single_output(request.outputs)
output_handle = _validate_handle(output.local_handle)
parent = os.path.dirname(output_handle)
if not os.path.isdir(parent):
raise codes.ExecutionFailure(
codes.INPUT_HANDLE_INVALID, "output handle directory does not exist"
)
raw = _read_input_file(text_artifact, _MAX_TEXT_BYTES)
try:
text = raw.decode("utf-8").strip()
except UnicodeDecodeError:
raise codes.ExecutionFailure(
codes.INPUT_TEXT_ENCODING, "input text is not valid UTF-8"
)
if not text:
raise codes.ExecutionFailure(codes.INPUT_TEXT_EMPTY, "input text is empty")
engine_kwargs: dict = {}
for name in sorted(request.parameters):
engine_kwargs[name] = _typed_parameter(name, request.parameters[name])
if ref_artifact is not None:
_read_input_file(ref_artifact, _MAX_REF_AUDIO_BYTES) # existence/bounds/checksum
engine_kwargs["ref_audio"] = _validate_handle(ref_artifact.local_handle)
return ValidatedRequest(
text=text,
output_handle=output_handle,
output_media_type=output.media_type or "audio/wav",
output_size_bound=int(output.expected_size_bytes),
engine_kwargs=engine_kwargs,
deadline_monotonic=self._clock() + budget_s,
catalog_model_id=request.model.catalog_model_id,
)
def _validate_model(self, spec):
wanted = (spec.catalog_model_id or "").strip()
if not wanted:
raise codes.ExecutionFailure(
codes.INPUT_MODEL_UNKNOWN, "catalog model id is required"
)
matches = [
model
for model in self._inventory.models()
if model.catalog_model_id == wanted
]
if not matches:
raise codes.ExecutionFailure(codes.INPUT_MODEL_UNKNOWN, "model not present")
model = matches[0]
if model.state != STATE_READY:
raise codes.ExecutionFailure(
codes.INPUT_MODEL_NOT_READY, "model is not READY"
)
if spec.model_version and spec.model_version != model.model_version:
raise codes.ExecutionFailure(
codes.INPUT_MODEL_UNKNOWN, "model version mismatch"
)
if not spec.model_digest or spec.model_digest != model.model_digest:
raise codes.ExecutionFailure(
codes.INPUT_MODEL_DIGEST_MISMATCH, "approved model digest mismatch"
)
if spec.precision and spec.precision not in model.precisions:
raise codes.ExecutionFailure(
codes.INPUT_MODEL_PRECISION, "precision not offered by this model"
)
return model
def _validate_device(self, device_id: str) -> None:
wanted = (device_id or "").strip()
if not wanted:
raise codes.ExecutionFailure(
codes.INPUT_DEVICE_UNKNOWN, "device id is required"
)
known = {device.device_id for device in self._inventory.devices()}
if wanted not in known:
raise codes.ExecutionFailure(
codes.INPUT_DEVICE_UNKNOWN, "device id not in inventory"
)
@staticmethod
def _split_inputs(inputs):
text_artifacts, audio_artifacts = [], []
for artifact in inputs:
if artifact.operation != pb2.LOCAL_ARTIFACT_OPERATION_READ:
raise codes.ExecutionFailure(
codes.INPUT_ARTIFACTS_INVALID, "inputs must be READ artifacts"
)
media = artifact.media_type or ""
if media.startswith("audio/"):
audio_artifacts.append(artifact)
elif media == "" or media.startswith("text/"):
text_artifacts.append(artifact)
else:
raise codes.ExecutionFailure(
codes.INPUT_ARTIFACTS_INVALID, f"unsupported input media {media!r}"
)
if len(text_artifacts) != 1 or len(audio_artifacts) > 1:
raise codes.ExecutionFailure(
codes.INPUT_ARTIFACTS_INVALID,
"tts needs exactly one text input and at most one reference audio",
)
return text_artifacts[0], (audio_artifacts[0] if audio_artifacts else None)
@staticmethod
def _single_output(outputs):
if len(outputs) != 1:
raise codes.ExecutionFailure(
codes.INPUT_ARTIFACTS_INVALID, "tts needs exactly one output artifact"
)
output = outputs[0]
if output.operation != pb2.LOCAL_ARTIFACT_OPERATION_WRITE:
raise codes.ExecutionFailure(
codes.INPUT_ARTIFACTS_INVALID, "output must be a WRITE artifact"
)
media = output.media_type or ""
if media and not media.startswith("audio/"):
raise codes.ExecutionFailure(
codes.INPUT_ARTIFACTS_INVALID, f"unsupported output media {media!r}"
)
return output
# -- execution -----------------------------------------------------
def execute(self, request, grpc_context=None):
"""Generator of ``pb2.ExecuteResponse``. Never raises for a
classified failure failures become terminal events."""
session = _Session(self, request)
return session.run(grpc_context)
class _Session:
def __init__(self, executor: Executor, request):
self._x = executor
self.request = request
self.job_id = request.job_id
self.attempt_id = request.attempt_id
self.sequence = 0
self.phase = "model_load"
self.terminal_sent = False
self.chars = 0
self.gpu_ms = 0
self.cpu_ms = 0
self.output_audio_ms = 0
# event builders ---------------------------------------------------
def _event(self, **payload):
self.sequence += 1
return pb2.ExecuteResponse(
event=pb2.ExecutionEvent(
job_id=self.job_id,
attempt_id=self.attempt_id,
sequence=self.sequence,
observed_at_unix_ms=int(time.time() * 1000),
**payload,
)
)
def _measurements(self):
return pb2.RuntimeMeasurements(
normalized_input_characters=self.chars,
output_audio_ms=self.output_audio_ms,
gpu_execution_ms=self.gpu_ms,
cpu_execution_ms=self.cpu_ms,
)
def _failed(self, failure: codes.ExecutionFailure):
self.terminal_sent = True
return self._event(
failed=pb2.ExecutionFailed(
failure_class=failure.failure_class,
stable_code=failure.stable_code,
safe_detail=failure.safe_detail,
measurements=self._measurements(),
)
)
def _canceled(self):
self.terminal_sent = True
return self._event(
canceled=pb2.ExecutionCanceled(measurements=self._measurements())
)
# main flow --------------------------------------------------------
def run(self, grpc_context):
if not self.attempt_id.strip() or not self.job_id.strip():
yield self._failed(
codes.ExecutionFailure(
codes.INPUT_ATTEMPT_IDENTITY, "job and attempt ids are required"
)
)
return
registry = self._x._registry
try:
record = registry.begin(self.job_id, self.attempt_id, self._x._slot_limit)
except codes.ExecutionFailure as failure:
yield self._failed(failure)
return
try:
yield from self._run_admitted(record, grpc_context)
finally:
terminal = "canceled"
if self.terminal_sent:
terminal = self._terminal_kind or "failed"
registry.finish(self.attempt_id, terminal)
_terminal_kind: str | None = None
def _run_admitted(self, record, grpc_context):
try:
validated = self._x._validate(self.request)
except codes.ExecutionFailure as failure:
self._terminal_kind = "failed"
yield self._failed(failure)
return
except Exception as exc: # adapter bug — still a classified event
self._terminal_kind = "failed"
yield self._failed(
codes.ExecutionFailure(codes.RUNTIME_CRASH, f"{type(exc).__name__}")
)
return
self.chars = len(validated.text)
yield self._event(started=pb2.ExecutionStarted())
worker = _EngineWorker(self._x._engine_provider, validated, self)
worker.start()
clock = self._x._clock
next_progress = clock() + self._x._progress_interval
progress_events = 0
while not worker.done.wait(self._x._poll_interval):
if record.cancel.is_set() or (
grpc_context is not None and not grpc_context.is_active()
):
self._terminal_kind = "canceled"
yield self._canceled()
return
now = clock()
if now >= validated.deadline_monotonic:
self._terminal_kind = "failed"
yield self._failed(codes.deadline_failure(self.phase))
return
if now >= next_progress and progress_events < _MAX_PROGRESS_EVENTS:
progress_events += 1
next_progress = now + self._x._progress_interval
permille = 100 if self.phase == "model_load" else 550
yield self._event(
progress=pb2.ExecutionProgress(
progress_permille=permille, stage_code=self.phase
)
)
if record.cancel.is_set():
self._terminal_kind = "canceled"
yield self._canceled()
return
if worker.error is not None:
self._terminal_kind = "failed"
yield self._failed(codes.classify_engine_error(worker.error, worker.phase))
return
try:
manifest = self._write_output(worker, validated)
except codes.ExecutionFailure as failure:
self._terminal_kind = "failed"
yield self._failed(failure)
return
self._terminal_kind = "completed"
self.terminal_sent = True
yield self._event(
completed=pb2.ExecutionCompleted(
outputs=[manifest], measurements=self._measurements()
)
)
def _write_output(self, worker, validated: ValidatedRequest):
ensure_backend_on_path()
tensor = worker.result
sample_rate = worker.sample_rate
if tensor is None or not hasattr(tensor, "numel") or tensor.numel() == 0:
raise codes.ExecutionFailure(
codes.INFERENCE_BAD_OUTPUT, "engine returned no audio"
)
if not isinstance(sample_rate, int) or sample_rate <= 0:
raise codes.ExecutionFailure(
codes.INFERENCE_BAD_OUTPUT, "engine reported no sample rate"
)
try:
from services.audio_io import atomic_save_wav # noqa: PLC0415
atomic_save_wav(validated.output_handle, tensor.detach().cpu(), sample_rate)
except codes.ExecutionFailure:
raise
except Exception as exc:
raise codes.ExecutionFailure(
codes.STORAGE_WRITE_FAILED, f"{type(exc).__name__}: {exc}"
)
try:
size = os.stat(validated.output_handle).st_size
sha = file_sha256(validated.output_handle)
except OSError as exc:
raise codes.ExecutionFailure(
codes.STORAGE_WRITE_FAILED, f"{type(exc).__name__}"
)
if 0 < validated.output_size_bound < size:
raise codes.ExecutionFailure(
codes.STORAGE_WRITE_FAILED, "output exceeds its size bound"
)
samples = tensor.numel() if tensor.dim() == 1 else tensor.shape[-1]
self.output_audio_ms = int(samples * 1000 / sample_rate)
return pb2.LocalArtifactManifest(
artifact_id=self.request.outputs[0].artifact_id,
local_handle=validated.output_handle,
size_bytes=size,
sha256=sha,
media_type=validated.output_media_type,
duration_ms=self.output_audio_ms,
)
class _EngineWorker:
"""Runs the engine on a daemon thread, recording phase and timings."""
def __init__(self, engine_provider, validated: ValidatedRequest, session: _Session):
self._engine_provider = engine_provider
self._validated = validated
self._session = session
self.done = threading.Event()
self.error: BaseException | None = None
self.result = None
self.sample_rate: int | None = None
self.phase = "model_load"
def start(self) -> None:
thread = threading.Thread(
target=self._run,
name=f"runtime-adapter-attempt-{self._session.attempt_id}",
daemon=True,
)
thread.start()
@staticmethod
def _synthesize(engine, text: str, params: dict):
"""Use the same seeded native path as OSS Gallery and ovnode workers."""
from services import tts_backend # noqa: PLC0415
if isinstance(engine, tts_backend.OmniVoiceBackend):
from api.routers.generation import _run_inference # noqa: PLC0415
with tts_backend.engine_in_use(engine):
return _run_inference(
engine._model, text, params.get("language"),
params.get("ref_audio"), params.get("ref_text"),
params.get("instruct"), params.get("duration"),
params.get("num_step", 16), params.get("guidance_scale", 2.0),
params.get("speed", 1.0), params.get("t_shift"),
params.get("denoise", True), params.get("postprocess_output", True),
params.get("layer_penalty_factor"),
params.get("position_temperature"),
params.get("class_temperature"), params.get("seed"),
)
return engine.generate(text, **params)
def _run(self) -> None:
wall_start = time.monotonic()
cpu_start = time.process_time()
try:
engine = self._engine_provider(self._validated.catalog_model_id)
ensure_ready = getattr(engine, "ensure_ready", None)
if callable(ensure_ready):
ensure_ready()
self.phase = "synthesis"
self._session.phase = "synthesis"
synth_start = time.monotonic()
self.result = self._synthesize(engine, self._validated.text, self._validated.engine_kwargs)
rate = getattr(engine, "sample_rate", None)
self.sample_rate = int(rate) if isinstance(rate, (int, float)) and rate else None
self._session.gpu_ms = int((time.monotonic() - synth_start) * 1000)
except BaseException as exc: # classified later, never lost
self.error = exc
finally:
self._session.cpu_ms = int((time.process_time() - cpu_start) * 1000)
if self._session.gpu_ms == 0 and self.error is None:
self._session.gpu_ms = int((time.monotonic() - wall_start) * 1000)
self.done.set()
+5
View File
@@ -0,0 +1,5 @@
"""Generated protocol stubs — DO NOT EDIT.
Regenerate with ``uv run python scripts/gen_runtime_adapter_protocol.py``
after any change to ``../runtime_adapter.proto``.
"""
File diff suppressed because one or more lines are too long
@@ -0,0 +1,330 @@
from google.protobuf.internal import containers as _containers
from google.protobuf.internal import enum_type_wrapper as _enum_type_wrapper
from google.protobuf import descriptor as _descriptor
from google.protobuf import message as _message
from collections.abc import Iterable as _Iterable, Mapping as _Mapping
from typing import ClassVar as _ClassVar, Optional as _Optional, Union as _Union
DESCRIPTOR: _descriptor.FileDescriptor
class ServingState(int, metaclass=_enum_type_wrapper.EnumTypeWrapper):
__slots__ = ()
SERVING_STATE_UNSPECIFIED: _ClassVar[ServingState]
SERVING_STATE_READY: _ClassVar[ServingState]
SERVING_STATE_DEGRADED: _ClassVar[ServingState]
SERVING_STATE_UNHEALTHY: _ClassVar[ServingState]
class RuntimeModelState(int, metaclass=_enum_type_wrapper.EnumTypeWrapper):
__slots__ = ()
RUNTIME_MODEL_STATE_UNSPECIFIED: _ClassVar[RuntimeModelState]
RUNTIME_MODEL_STATE_INSTALLED: _ClassVar[RuntimeModelState]
RUNTIME_MODEL_STATE_LOADING: _ClassVar[RuntimeModelState]
RUNTIME_MODEL_STATE_READY: _ClassVar[RuntimeModelState]
RUNTIME_MODEL_STATE_FAILED: _ClassVar[RuntimeModelState]
class LocalArtifactOperation(int, metaclass=_enum_type_wrapper.EnumTypeWrapper):
__slots__ = ()
LOCAL_ARTIFACT_OPERATION_UNSPECIFIED: _ClassVar[LocalArtifactOperation]
LOCAL_ARTIFACT_OPERATION_READ: _ClassVar[LocalArtifactOperation]
LOCAL_ARTIFACT_OPERATION_WRITE: _ClassVar[LocalArtifactOperation]
class RuntimeFailureClass(int, metaclass=_enum_type_wrapper.EnumTypeWrapper):
__slots__ = ()
RUNTIME_FAILURE_CLASS_UNSPECIFIED: _ClassVar[RuntimeFailureClass]
RUNTIME_FAILURE_CLASS_INPUT: _ClassVar[RuntimeFailureClass]
RUNTIME_FAILURE_CLASS_MODEL_LOAD: _ClassVar[RuntimeFailureClass]
RUNTIME_FAILURE_CLASS_INFERENCE: _ClassVar[RuntimeFailureClass]
RUNTIME_FAILURE_CLASS_GPU_RESOURCE: _ClassVar[RuntimeFailureClass]
RUNTIME_FAILURE_CLASS_LOCAL_STORAGE: _ClassVar[RuntimeFailureClass]
RUNTIME_FAILURE_CLASS_RUNTIME: _ClassVar[RuntimeFailureClass]
RUNTIME_FAILURE_CLASS_CANCELED: _ClassVar[RuntimeFailureClass]
class CancelDisposition(int, metaclass=_enum_type_wrapper.EnumTypeWrapper):
__slots__ = ()
CANCEL_DISPOSITION_UNSPECIFIED: _ClassVar[CancelDisposition]
CANCEL_DISPOSITION_ACCEPTED: _ClassVar[CancelDisposition]
CANCEL_DISPOSITION_ALREADY_TERMINAL: _ClassVar[CancelDisposition]
CANCEL_DISPOSITION_NOT_FOUND: _ClassVar[CancelDisposition]
SERVING_STATE_UNSPECIFIED: ServingState
SERVING_STATE_READY: ServingState
SERVING_STATE_DEGRADED: ServingState
SERVING_STATE_UNHEALTHY: ServingState
RUNTIME_MODEL_STATE_UNSPECIFIED: RuntimeModelState
RUNTIME_MODEL_STATE_INSTALLED: RuntimeModelState
RUNTIME_MODEL_STATE_LOADING: RuntimeModelState
RUNTIME_MODEL_STATE_READY: RuntimeModelState
RUNTIME_MODEL_STATE_FAILED: RuntimeModelState
LOCAL_ARTIFACT_OPERATION_UNSPECIFIED: LocalArtifactOperation
LOCAL_ARTIFACT_OPERATION_READ: LocalArtifactOperation
LOCAL_ARTIFACT_OPERATION_WRITE: LocalArtifactOperation
RUNTIME_FAILURE_CLASS_UNSPECIFIED: RuntimeFailureClass
RUNTIME_FAILURE_CLASS_INPUT: RuntimeFailureClass
RUNTIME_FAILURE_CLASS_MODEL_LOAD: RuntimeFailureClass
RUNTIME_FAILURE_CLASS_INFERENCE: RuntimeFailureClass
RUNTIME_FAILURE_CLASS_GPU_RESOURCE: RuntimeFailureClass
RUNTIME_FAILURE_CLASS_LOCAL_STORAGE: RuntimeFailureClass
RUNTIME_FAILURE_CLASS_RUNTIME: RuntimeFailureClass
RUNTIME_FAILURE_CLASS_CANCELED: RuntimeFailureClass
CANCEL_DISPOSITION_UNSPECIFIED: CancelDisposition
CANCEL_DISPOSITION_ACCEPTED: CancelDisposition
CANCEL_DISPOSITION_ALREADY_TERMINAL: CancelDisposition
CANCEL_DISPOSITION_NOT_FOUND: CancelDisposition
class ExecuteResponse(_message.Message):
__slots__ = ("event",)
EVENT_FIELD_NUMBER: _ClassVar[int]
event: ExecutionEvent
def __init__(self, event: _Optional[_Union[ExecutionEvent, _Mapping]] = ...) -> None: ...
class HealthRequest(_message.Message):
__slots__ = ()
def __init__(self) -> None: ...
class HealthResponse(_message.Message):
__slots__ = ("state", "runtime_version", "adapter_version", "health_flags")
STATE_FIELD_NUMBER: _ClassVar[int]
RUNTIME_VERSION_FIELD_NUMBER: _ClassVar[int]
ADAPTER_VERSION_FIELD_NUMBER: _ClassVar[int]
HEALTH_FLAGS_FIELD_NUMBER: _ClassVar[int]
state: ServingState
runtime_version: str
adapter_version: str
health_flags: _containers.RepeatedScalarFieldContainer[str]
def __init__(self, state: _Optional[_Union[ServingState, str]] = ..., runtime_version: _Optional[str] = ..., adapter_version: _Optional[str] = ..., health_flags: _Optional[_Iterable[str]] = ...) -> None: ...
class GetCapabilitiesRequest(_message.Message):
__slots__ = ()
def __init__(self) -> None: ...
class GetCapabilitiesResponse(_message.Message):
__slots__ = ("runtime_version", "adapter_version", "devices", "models")
RUNTIME_VERSION_FIELD_NUMBER: _ClassVar[int]
ADAPTER_VERSION_FIELD_NUMBER: _ClassVar[int]
DEVICES_FIELD_NUMBER: _ClassVar[int]
MODELS_FIELD_NUMBER: _ClassVar[int]
runtime_version: str
adapter_version: str
devices: _containers.RepeatedCompositeFieldContainer[RuntimeDevice]
models: _containers.RepeatedCompositeFieldContainer[RuntimeModel]
def __init__(self, runtime_version: _Optional[str] = ..., adapter_version: _Optional[str] = ..., devices: _Optional[_Iterable[_Union[RuntimeDevice, _Mapping]]] = ..., models: _Optional[_Iterable[_Union[RuntimeModel, _Mapping]]] = ...) -> None: ...
class RuntimeDevice(_message.Message):
__slots__ = ("device_id", "hardware_class", "total_vram_bytes", "total_slots", "free_slots")
DEVICE_ID_FIELD_NUMBER: _ClassVar[int]
HARDWARE_CLASS_FIELD_NUMBER: _ClassVar[int]
TOTAL_VRAM_BYTES_FIELD_NUMBER: _ClassVar[int]
TOTAL_SLOTS_FIELD_NUMBER: _ClassVar[int]
FREE_SLOTS_FIELD_NUMBER: _ClassVar[int]
device_id: str
hardware_class: str
total_vram_bytes: int
total_slots: int
free_slots: int
def __init__(self, device_id: _Optional[str] = ..., hardware_class: _Optional[str] = ..., total_vram_bytes: _Optional[int] = ..., total_slots: _Optional[int] = ..., free_slots: _Optional[int] = ...) -> None: ...
class RuntimeModel(_message.Message):
__slots__ = ("catalog_model_id", "model_version", "model_digest", "precisions", "features", "state")
CATALOG_MODEL_ID_FIELD_NUMBER: _ClassVar[int]
MODEL_VERSION_FIELD_NUMBER: _ClassVar[int]
MODEL_DIGEST_FIELD_NUMBER: _ClassVar[int]
PRECISIONS_FIELD_NUMBER: _ClassVar[int]
FEATURES_FIELD_NUMBER: _ClassVar[int]
STATE_FIELD_NUMBER: _ClassVar[int]
catalog_model_id: str
model_version: str
model_digest: str
precisions: _containers.RepeatedScalarFieldContainer[str]
features: _containers.RepeatedScalarFieldContainer[str]
state: RuntimeModelState
def __init__(self, catalog_model_id: _Optional[str] = ..., model_version: _Optional[str] = ..., model_digest: _Optional[str] = ..., precisions: _Optional[_Iterable[str]] = ..., features: _Optional[_Iterable[str]] = ..., state: _Optional[_Union[RuntimeModelState, str]] = ...) -> None: ...
class ExecuteRequest(_message.Message):
__slots__ = ("job_id", "attempt_id", "device_id", "slot_id", "model", "parameters", "inputs", "outputs", "deadline_unix_ms", "maximum_preview_bytes")
class ParametersEntry(_message.Message):
__slots__ = ("key", "value")
KEY_FIELD_NUMBER: _ClassVar[int]
VALUE_FIELD_NUMBER: _ClassVar[int]
key: str
value: ParameterValue
def __init__(self, key: _Optional[str] = ..., value: _Optional[_Union[ParameterValue, _Mapping]] = ...) -> None: ...
JOB_ID_FIELD_NUMBER: _ClassVar[int]
ATTEMPT_ID_FIELD_NUMBER: _ClassVar[int]
DEVICE_ID_FIELD_NUMBER: _ClassVar[int]
SLOT_ID_FIELD_NUMBER: _ClassVar[int]
MODEL_FIELD_NUMBER: _ClassVar[int]
PARAMETERS_FIELD_NUMBER: _ClassVar[int]
INPUTS_FIELD_NUMBER: _ClassVar[int]
OUTPUTS_FIELD_NUMBER: _ClassVar[int]
DEADLINE_UNIX_MS_FIELD_NUMBER: _ClassVar[int]
MAXIMUM_PREVIEW_BYTES_FIELD_NUMBER: _ClassVar[int]
job_id: str
attempt_id: str
device_id: str
slot_id: str
model: ModelSpec
parameters: _containers.MessageMap[str, ParameterValue]
inputs: _containers.RepeatedCompositeFieldContainer[LocalArtifact]
outputs: _containers.RepeatedCompositeFieldContainer[LocalArtifact]
deadline_unix_ms: int
maximum_preview_bytes: int
def __init__(self, job_id: _Optional[str] = ..., attempt_id: _Optional[str] = ..., device_id: _Optional[str] = ..., slot_id: _Optional[str] = ..., model: _Optional[_Union[ModelSpec, _Mapping]] = ..., parameters: _Optional[_Mapping[str, ParameterValue]] = ..., inputs: _Optional[_Iterable[_Union[LocalArtifact, _Mapping]]] = ..., outputs: _Optional[_Iterable[_Union[LocalArtifact, _Mapping]]] = ..., deadline_unix_ms: _Optional[int] = ..., maximum_preview_bytes: _Optional[int] = ...) -> None: ...
class ModelSpec(_message.Message):
__slots__ = ("catalog_model_id", "model_version", "model_digest", "precision")
CATALOG_MODEL_ID_FIELD_NUMBER: _ClassVar[int]
MODEL_VERSION_FIELD_NUMBER: _ClassVar[int]
MODEL_DIGEST_FIELD_NUMBER: _ClassVar[int]
PRECISION_FIELD_NUMBER: _ClassVar[int]
catalog_model_id: str
model_version: str
model_digest: str
precision: str
def __init__(self, catalog_model_id: _Optional[str] = ..., model_version: _Optional[str] = ..., model_digest: _Optional[str] = ..., precision: _Optional[str] = ...) -> None: ...
class ParameterValue(_message.Message):
__slots__ = ("string_value", "integer_value", "number_value", "boolean_value")
STRING_VALUE_FIELD_NUMBER: _ClassVar[int]
INTEGER_VALUE_FIELD_NUMBER: _ClassVar[int]
NUMBER_VALUE_FIELD_NUMBER: _ClassVar[int]
BOOLEAN_VALUE_FIELD_NUMBER: _ClassVar[int]
string_value: str
integer_value: int
number_value: float
boolean_value: bool
def __init__(self, string_value: _Optional[str] = ..., integer_value: _Optional[int] = ..., number_value: _Optional[float] = ..., boolean_value: _Optional[bool] = ...) -> None: ...
class LocalArtifact(_message.Message):
__slots__ = ("artifact_id", "local_handle", "operation", "expected_size_bytes", "expected_sha256", "media_type")
ARTIFACT_ID_FIELD_NUMBER: _ClassVar[int]
LOCAL_HANDLE_FIELD_NUMBER: _ClassVar[int]
OPERATION_FIELD_NUMBER: _ClassVar[int]
EXPECTED_SIZE_BYTES_FIELD_NUMBER: _ClassVar[int]
EXPECTED_SHA256_FIELD_NUMBER: _ClassVar[int]
MEDIA_TYPE_FIELD_NUMBER: _ClassVar[int]
artifact_id: str
local_handle: str
operation: LocalArtifactOperation
expected_size_bytes: int
expected_sha256: str
media_type: str
def __init__(self, artifact_id: _Optional[str] = ..., local_handle: _Optional[str] = ..., operation: _Optional[_Union[LocalArtifactOperation, str]] = ..., expected_size_bytes: _Optional[int] = ..., expected_sha256: _Optional[str] = ..., media_type: _Optional[str] = ...) -> None: ...
class ExecutionEvent(_message.Message):
__slots__ = ("job_id", "attempt_id", "sequence", "observed_at_unix_ms", "started", "progress", "preview", "completed", "failed", "canceled")
JOB_ID_FIELD_NUMBER: _ClassVar[int]
ATTEMPT_ID_FIELD_NUMBER: _ClassVar[int]
SEQUENCE_FIELD_NUMBER: _ClassVar[int]
OBSERVED_AT_UNIX_MS_FIELD_NUMBER: _ClassVar[int]
STARTED_FIELD_NUMBER: _ClassVar[int]
PROGRESS_FIELD_NUMBER: _ClassVar[int]
PREVIEW_FIELD_NUMBER: _ClassVar[int]
COMPLETED_FIELD_NUMBER: _ClassVar[int]
FAILED_FIELD_NUMBER: _ClassVar[int]
CANCELED_FIELD_NUMBER: _ClassVar[int]
job_id: str
attempt_id: str
sequence: int
observed_at_unix_ms: int
started: ExecutionStarted
progress: ExecutionProgress
preview: PreviewChunk
completed: ExecutionCompleted
failed: ExecutionFailed
canceled: ExecutionCanceled
def __init__(self, job_id: _Optional[str] = ..., attempt_id: _Optional[str] = ..., sequence: _Optional[int] = ..., observed_at_unix_ms: _Optional[int] = ..., started: _Optional[_Union[ExecutionStarted, _Mapping]] = ..., progress: _Optional[_Union[ExecutionProgress, _Mapping]] = ..., preview: _Optional[_Union[PreviewChunk, _Mapping]] = ..., completed: _Optional[_Union[ExecutionCompleted, _Mapping]] = ..., failed: _Optional[_Union[ExecutionFailed, _Mapping]] = ..., canceled: _Optional[_Union[ExecutionCanceled, _Mapping]] = ...) -> None: ...
class ExecutionStarted(_message.Message):
__slots__ = ()
def __init__(self) -> None: ...
class ExecutionProgress(_message.Message):
__slots__ = ("progress_permille", "stage_code")
PROGRESS_PERMILLE_FIELD_NUMBER: _ClassVar[int]
STAGE_CODE_FIELD_NUMBER: _ClassVar[int]
progress_permille: int
stage_code: str
def __init__(self, progress_permille: _Optional[int] = ..., stage_code: _Optional[str] = ...) -> None: ...
class PreviewChunk(_message.Message):
__slots__ = ("sequence", "media_type", "data")
SEQUENCE_FIELD_NUMBER: _ClassVar[int]
MEDIA_TYPE_FIELD_NUMBER: _ClassVar[int]
DATA_FIELD_NUMBER: _ClassVar[int]
sequence: int
media_type: str
data: bytes
def __init__(self, sequence: _Optional[int] = ..., media_type: _Optional[str] = ..., data: _Optional[bytes] = ...) -> None: ...
class ExecutionCompleted(_message.Message):
__slots__ = ("outputs", "measurements")
OUTPUTS_FIELD_NUMBER: _ClassVar[int]
MEASUREMENTS_FIELD_NUMBER: _ClassVar[int]
outputs: _containers.RepeatedCompositeFieldContainer[LocalArtifactManifest]
measurements: RuntimeMeasurements
def __init__(self, outputs: _Optional[_Iterable[_Union[LocalArtifactManifest, _Mapping]]] = ..., measurements: _Optional[_Union[RuntimeMeasurements, _Mapping]] = ...) -> None: ...
class LocalArtifactManifest(_message.Message):
__slots__ = ("artifact_id", "local_handle", "size_bytes", "sha256", "media_type", "duration_ms")
ARTIFACT_ID_FIELD_NUMBER: _ClassVar[int]
LOCAL_HANDLE_FIELD_NUMBER: _ClassVar[int]
SIZE_BYTES_FIELD_NUMBER: _ClassVar[int]
SHA256_FIELD_NUMBER: _ClassVar[int]
MEDIA_TYPE_FIELD_NUMBER: _ClassVar[int]
DURATION_MS_FIELD_NUMBER: _ClassVar[int]
artifact_id: str
local_handle: str
size_bytes: int
sha256: str
media_type: str
duration_ms: int
def __init__(self, artifact_id: _Optional[str] = ..., local_handle: _Optional[str] = ..., size_bytes: _Optional[int] = ..., sha256: _Optional[str] = ..., media_type: _Optional[str] = ..., duration_ms: _Optional[int] = ...) -> None: ...
class ExecutionFailed(_message.Message):
__slots__ = ("failure_class", "stable_code", "safe_detail", "measurements")
FAILURE_CLASS_FIELD_NUMBER: _ClassVar[int]
STABLE_CODE_FIELD_NUMBER: _ClassVar[int]
SAFE_DETAIL_FIELD_NUMBER: _ClassVar[int]
MEASUREMENTS_FIELD_NUMBER: _ClassVar[int]
failure_class: RuntimeFailureClass
stable_code: str
safe_detail: str
measurements: RuntimeMeasurements
def __init__(self, failure_class: _Optional[_Union[RuntimeFailureClass, str]] = ..., stable_code: _Optional[str] = ..., safe_detail: _Optional[str] = ..., measurements: _Optional[_Union[RuntimeMeasurements, _Mapping]] = ...) -> None: ...
class ExecutionCanceled(_message.Message):
__slots__ = ("measurements",)
MEASUREMENTS_FIELD_NUMBER: _ClassVar[int]
measurements: RuntimeMeasurements
def __init__(self, measurements: _Optional[_Union[RuntimeMeasurements, _Mapping]] = ...) -> None: ...
class RuntimeMeasurements(_message.Message):
__slots__ = ("normalized_input_characters", "input_audio_ms", "output_audio_ms", "gpu_execution_ms", "cpu_execution_ms")
NORMALIZED_INPUT_CHARACTERS_FIELD_NUMBER: _ClassVar[int]
INPUT_AUDIO_MS_FIELD_NUMBER: _ClassVar[int]
OUTPUT_AUDIO_MS_FIELD_NUMBER: _ClassVar[int]
GPU_EXECUTION_MS_FIELD_NUMBER: _ClassVar[int]
CPU_EXECUTION_MS_FIELD_NUMBER: _ClassVar[int]
normalized_input_characters: int
input_audio_ms: int
output_audio_ms: int
gpu_execution_ms: int
cpu_execution_ms: int
def __init__(self, normalized_input_characters: _Optional[int] = ..., input_audio_ms: _Optional[int] = ..., output_audio_ms: _Optional[int] = ..., gpu_execution_ms: _Optional[int] = ..., cpu_execution_ms: _Optional[int] = ...) -> None: ...
class CancelRequest(_message.Message):
__slots__ = ("job_id", "attempt_id", "reason_code", "deadline_unix_ms")
JOB_ID_FIELD_NUMBER: _ClassVar[int]
ATTEMPT_ID_FIELD_NUMBER: _ClassVar[int]
REASON_CODE_FIELD_NUMBER: _ClassVar[int]
DEADLINE_UNIX_MS_FIELD_NUMBER: _ClassVar[int]
job_id: str
attempt_id: str
reason_code: str
deadline_unix_ms: int
def __init__(self, job_id: _Optional[str] = ..., attempt_id: _Optional[str] = ..., reason_code: _Optional[str] = ..., deadline_unix_ms: _Optional[int] = ...) -> None: ...
class CancelResponse(_message.Message):
__slots__ = ("disposition",)
DISPOSITION_FIELD_NUMBER: _ClassVar[int]
disposition: CancelDisposition
def __init__(self, disposition: _Optional[_Union[CancelDisposition, str]] = ...) -> None: ...
@@ -0,0 +1,229 @@
# Generated by the gRPC Python protocol compiler plugin. DO NOT EDIT!
"""Client and server classes corresponding to protobuf-defined services."""
import grpc
import warnings
from . import runtime_adapter_pb2 as runtime__adapter__pb2
GRPC_GENERATED_VERSION = '1.81.1'
GRPC_VERSION = grpc.__version__
_version_not_supported = False
try:
from grpc._utilities import first_version_is_lower
_version_not_supported = first_version_is_lower(GRPC_VERSION, GRPC_GENERATED_VERSION)
except ImportError:
_version_not_supported = True
if _version_not_supported:
raise RuntimeError(
f'The grpc package installed is at version {GRPC_VERSION},'
+ ' but the generated code in runtime_adapter_pb2_grpc.py depends on'
+ f' grpcio>={GRPC_GENERATED_VERSION}.'
+ f' Please upgrade your grpc module to grpcio>={GRPC_GENERATED_VERSION}'
+ f' or downgrade your generated code using grpcio-tools<={GRPC_VERSION}.'
)
class RuntimeAdapterServiceStub:
"""RuntimeAdapterService is local to a GPU Node and is never publicly exposed.
"""
def __init__(self, channel):
"""Constructor.
Args:
channel: A grpc.Channel.
"""
self.Health = channel.unary_unary(
'/voicestudio.runtime.v1.RuntimeAdapterService/Health',
request_serializer=runtime__adapter__pb2.HealthRequest.SerializeToString,
response_deserializer=runtime__adapter__pb2.HealthResponse.FromString,
_registered_method=True)
self.GetCapabilities = channel.unary_unary(
'/voicestudio.runtime.v1.RuntimeAdapterService/GetCapabilities',
request_serializer=runtime__adapter__pb2.GetCapabilitiesRequest.SerializeToString,
response_deserializer=runtime__adapter__pb2.GetCapabilitiesResponse.FromString,
_registered_method=True)
self.Execute = channel.unary_stream(
'/voicestudio.runtime.v1.RuntimeAdapterService/Execute',
request_serializer=runtime__adapter__pb2.ExecuteRequest.SerializeToString,
response_deserializer=runtime__adapter__pb2.ExecuteResponse.FromString,
_registered_method=True)
self.Cancel = channel.unary_unary(
'/voicestudio.runtime.v1.RuntimeAdapterService/Cancel',
request_serializer=runtime__adapter__pb2.CancelRequest.SerializeToString,
response_deserializer=runtime__adapter__pb2.CancelResponse.FromString,
_registered_method=True)
class RuntimeAdapterServiceServicer:
"""RuntimeAdapterService is local to a GPU Node and is never publicly exposed.
"""
def Health(self, request, context):
"""Missing associated documentation comment in .proto file."""
context.set_code(grpc.StatusCode.UNIMPLEMENTED)
context.set_details('Method not implemented!')
raise NotImplementedError('Method not implemented!')
def GetCapabilities(self, request, context):
"""Missing associated documentation comment in .proto file."""
context.set_code(grpc.StatusCode.UNIMPLEMENTED)
context.set_details('Method not implemented!')
raise NotImplementedError('Method not implemented!')
def Execute(self, request, context):
"""Missing associated documentation comment in .proto file."""
context.set_code(grpc.StatusCode.UNIMPLEMENTED)
context.set_details('Method not implemented!')
raise NotImplementedError('Method not implemented!')
def Cancel(self, request, context):
"""Missing associated documentation comment in .proto file."""
context.set_code(grpc.StatusCode.UNIMPLEMENTED)
context.set_details('Method not implemented!')
raise NotImplementedError('Method not implemented!')
def add_RuntimeAdapterServiceServicer_to_server(servicer, server):
rpc_method_handlers = {
'Health': grpc.unary_unary_rpc_method_handler(
servicer.Health,
request_deserializer=runtime__adapter__pb2.HealthRequest.FromString,
response_serializer=runtime__adapter__pb2.HealthResponse.SerializeToString,
),
'GetCapabilities': grpc.unary_unary_rpc_method_handler(
servicer.GetCapabilities,
request_deserializer=runtime__adapter__pb2.GetCapabilitiesRequest.FromString,
response_serializer=runtime__adapter__pb2.GetCapabilitiesResponse.SerializeToString,
),
'Execute': grpc.unary_stream_rpc_method_handler(
servicer.Execute,
request_deserializer=runtime__adapter__pb2.ExecuteRequest.FromString,
response_serializer=runtime__adapter__pb2.ExecuteResponse.SerializeToString,
),
'Cancel': grpc.unary_unary_rpc_method_handler(
servicer.Cancel,
request_deserializer=runtime__adapter__pb2.CancelRequest.FromString,
response_serializer=runtime__adapter__pb2.CancelResponse.SerializeToString,
),
}
generic_handler = grpc.method_handlers_generic_handler(
'voicestudio.runtime.v1.RuntimeAdapterService', rpc_method_handlers)
server.add_generic_rpc_handlers((generic_handler,))
server.add_registered_method_handlers('voicestudio.runtime.v1.RuntimeAdapterService', rpc_method_handlers)
# This class is part of an EXPERIMENTAL API.
class RuntimeAdapterService:
"""RuntimeAdapterService is local to a GPU Node and is never publicly exposed.
"""
@staticmethod
def Health(request,
target,
options=(),
channel_credentials=None,
call_credentials=None,
insecure=False,
compression=None,
wait_for_ready=None,
timeout=None,
metadata=None):
return grpc.experimental.unary_unary(
request,
target,
'/voicestudio.runtime.v1.RuntimeAdapterService/Health',
runtime__adapter__pb2.HealthRequest.SerializeToString,
runtime__adapter__pb2.HealthResponse.FromString,
options,
channel_credentials,
insecure,
call_credentials,
compression,
wait_for_ready,
timeout,
metadata,
_registered_method=True)
@staticmethod
def GetCapabilities(request,
target,
options=(),
channel_credentials=None,
call_credentials=None,
insecure=False,
compression=None,
wait_for_ready=None,
timeout=None,
metadata=None):
return grpc.experimental.unary_unary(
request,
target,
'/voicestudio.runtime.v1.RuntimeAdapterService/GetCapabilities',
runtime__adapter__pb2.GetCapabilitiesRequest.SerializeToString,
runtime__adapter__pb2.GetCapabilitiesResponse.FromString,
options,
channel_credentials,
insecure,
call_credentials,
compression,
wait_for_ready,
timeout,
metadata,
_registered_method=True)
@staticmethod
def Execute(request,
target,
options=(),
channel_credentials=None,
call_credentials=None,
insecure=False,
compression=None,
wait_for_ready=None,
timeout=None,
metadata=None):
return grpc.experimental.unary_stream(
request,
target,
'/voicestudio.runtime.v1.RuntimeAdapterService/Execute',
runtime__adapter__pb2.ExecuteRequest.SerializeToString,
runtime__adapter__pb2.ExecuteResponse.FromString,
options,
channel_credentials,
insecure,
call_credentials,
compression,
wait_for_ready,
timeout,
metadata,
_registered_method=True)
@staticmethod
def Cancel(request,
target,
options=(),
channel_credentials=None,
call_credentials=None,
insecure=False,
compression=None,
wait_for_ready=None,
timeout=None,
metadata=None):
return grpc.experimental.unary_unary(
request,
target,
'/voicestudio.runtime.v1.RuntimeAdapterService/Cancel',
runtime__adapter__pb2.CancelRequest.SerializeToString,
runtime__adapter__pb2.CancelResponse.FromString,
options,
channel_credentials,
insecure,
call_credentials,
compression,
wait_for_ready,
timeout,
metadata,
_registered_method=True)
+318
View File
@@ -0,0 +1,318 @@
"""Device and model inventory reported through Health/GetCapabilities.
The server is written against the small protocol at the top of this module so
tests can substitute fakes; :class:`ProductionInventory` is the real thing,
wired to ``services.tts_backend``'s engine registry, ``services.hf_revisions``
pinned revisions, and :mod:`runtime_adapter.digest`.
State rules (mirrors the Go preflight's expectations):
- READY is **explicit**: engine registered, availability probe passed, the
pinned snapshot fully present on disk, and a digest computed. Anything
less is INSTALLED / LOADING / FAILED never READY.
- A loading or failed model is still listed (with its true state) so the
Gateway can observe it; only READY models are schedulable.
"""
from __future__ import annotations
import os
import threading
import time
from dataclasses import dataclass, field
from . import SLOTS_ENV
from ._paths import ensure_backend_on_path
from .digest import snapshot_digest
STATE_INSTALLED = "installed"
STATE_LOADING = "loading"
STATE_READY = "ready"
STATE_FAILED = "failed"
@dataclass(frozen=True)
class DeviceInfo:
device_id: str
hardware_class: str
total_vram_bytes: int
total_slots: int
free_slots: int
@dataclass(frozen=True)
class ModelInfo:
catalog_model_id: str
model_version: str
model_digest: str
precisions: tuple[str, ...] = ()
features: tuple[str, ...] = ()
state: str = STATE_INSTALLED
#: Engines this adapter can attest as digest-pinned models: TTS engine id →
#: curated Hugging Face repo (must be pinned in ``services.hf_revisions``).
#: Engines without a single pinned weights repo (external API servers,
#: multi-model muxes) are deliberately absent — they cannot be digest-pinned.
ENGINE_MODEL_REPOS: dict[str, str] = {
"omnivoice": "k2-fsa/OmniVoice",
"voxcpm2": "openbmb/VoxCPM2",
"moss-tts-nano": "OpenMOSS-Team/MOSS-TTS-Nano-100M",
"kittentts": "KittenML/kitten-tts-mini-0.8",
"cosyvoice": "FunAudioLLM/Fun-CosyVoice3-0.5B-2512",
"moss-tts-v15": "OpenMOSS-Team/MOSS-TTS-v1.5",
}
def catalog_model_version(revision: str, model_digest: str) -> str:
"""Return the immutable catalog version for an attested model snapshot.
A Hugging Face revision names source history, not necessarily the exact
snapshot bytes installed on a node. The catalog version therefore carries
a short, deterministic digest suffix. A changed snapshot becomes a new
catalog identity instead of mutating an identity retained by Jobs.
"""
digest = model_digest.removeprefix("sha256:")
if len(revision) != 40 or len(digest) != 64:
raise ValueError("model identity requires a SHA revision and SHA-256 digest")
return f"{revision}+sha256-{digest[:16]}"
def slots_per_device(default: int = 1) -> int:
raw = os.environ.get(SLOTS_ENV, "").strip()
try:
value = int(raw) if raw else default
except ValueError:
return default
return max(1, min(value, 64))
@dataclass
class ProductionInventory:
"""Real host inventory. All heavy imports happen inside methods.
``models()`` is memoized for ``model_ttl_s`` under a lock: the first call
hashes every installed snapshot (minutes for multi-GB weights, then cached
in the on-disk digest sidecar), and Health + GetCapabilities arrive
back-to-back. Call :meth:`warm` before serving so the first RPC never
pays the hashing cost inside its deadline.
"""
slots: int = field(default_factory=slots_per_device)
model_ttl_s: float = 15.0
def __post_init__(self):
self._model_lock = threading.Lock()
self._model_cache: list[ModelInfo] | None = None
self._model_cache_at = 0.0
def warm(self) -> None:
self.models()
def devices(self, busy_slots: int = 0) -> list[DeviceInfo]:
ensure_backend_on_path()
devices = self._accelerators() or [self._cpu_device()]
return [self._with_slots(device, busy_slots) for device in devices]
def _with_slots(self, device: DeviceInfo, busy_slots: int) -> DeviceInfo:
free = max(0, min(device.total_slots - busy_slots, device.total_slots))
return DeviceInfo(
device_id=device.device_id,
hardware_class=device.hardware_class,
total_vram_bytes=device.total_vram_bytes,
total_slots=device.total_slots,
free_slots=free,
)
def _accelerators(self) -> list[DeviceInfo]:
try:
import torch # noqa: PLC0415
except Exception:
return []
found: list[DeviceInfo] = []
try:
if torch.cuda.is_available():
for index in range(torch.cuda.device_count()):
props = torch.cuda.get_device_properties(index)
found.append(
DeviceInfo(
device_id=f"cuda:{index}",
hardware_class=torch.cuda.get_device_name(index),
total_vram_bytes=int(props.total_memory),
total_slots=self.slots,
free_slots=self.slots,
)
)
return found
except Exception:
pass
try:
if getattr(torch.backends, "mps", None) and torch.backends.mps.is_available():
vram = 0
recommended = getattr(torch.mps, "recommended_max_memory", None)
if callable(recommended):
try:
vram = int(recommended())
except Exception:
vram = 0
if vram <= 0:
vram = _system_memory_bytes()
return [
DeviceInfo(
device_id="mps:0",
hardware_class="apple-silicon-mps",
total_vram_bytes=vram,
total_slots=self.slots,
free_slots=self.slots,
)
]
except Exception:
pass
return []
def _cpu_device(self) -> DeviceInfo:
# A CPU-only node is a valid (slow) execution device. total_vram_bytes
# carries system memory so the Gateway's ">0" validity check reflects
# real capacity rather than a made-up constant.
import platform # noqa: PLC0415
return DeviceInfo(
device_id="cpu:0",
hardware_class=platform.processor() or platform.machine() or "cpu",
total_vram_bytes=_system_memory_bytes(),
total_slots=self.slots,
free_slots=self.slots,
)
def models(self) -> list[ModelInfo]:
with self._model_lock:
now = time.monotonic()
if (
self._model_cache is not None
and now - self._model_cache_at < self.model_ttl_s
):
return list(self._model_cache)
self._model_cache = self._scan_models()
self._model_cache_at = time.monotonic()
return list(self._model_cache)
def _scan_models(self) -> list[ModelInfo]:
ensure_backend_on_path()
from services.hf_cache_repair import repo_cache_dir # noqa: PLC0415
from services.hf_revisions import installed_revision # noqa: PLC0415
from services.tts_backend import get_backend_class # noqa: PLC0415
models: list[ModelInfo] = []
for engine_id, repo_id in sorted(ENGINE_MODEL_REPOS.items()):
try:
backend_cls = get_backend_class(engine_id)
except Exception:
continue # engine not registered in this build
repo_dir = repo_cache_dir(repo_id)
try:
revision = installed_revision(repo_id, os.path.dirname(repo_dir))
except ValueError:
continue # repo not in the curated catalog — cannot attest
snapshot = os.path.join(repo_dir, "snapshots", revision)
if not os.path.isdir(snapshot):
continue # weights not installed at the pinned revision
models.append(
self._model_state(engine_id, backend_cls, repo_dir, revision, snapshot)
)
return models
def _model_state(
self, engine_id: str, backend_cls, repo_dir: str, revision: str, snapshot: str
) -> ModelInfo:
base = ModelInfo(
catalog_model_id=engine_id,
model_version=revision,
model_digest="",
precisions=self._precisions(backend_cls),
features=self._features(backend_cls),
)
try:
ok, _message = backend_cls.is_available()
except Exception:
return _replace_state(base, STATE_FAILED)
if not ok:
return _replace_state(base, STATE_INSTALLED)
if _snapshot_incomplete(repo_dir, snapshot):
return _replace_state(base, STATE_LOADING)
try:
model_digest = snapshot_digest(
snapshot,
cache_path=os.path.join(repo_dir, f"voicestudio-digest-{revision}.json"),
)
except OSError:
return _replace_state(base, STATE_LOADING)
return ModelInfo(
catalog_model_id=base.catalog_model_id,
model_version=catalog_model_version(base.model_version, model_digest),
model_digest=model_digest,
precisions=base.precisions,
features=base.features,
state=STATE_READY,
)
def _precisions(self, backend_cls) -> tuple[str, ...]:
# Advisory execution precisions. fp32 always works; fp16 is offered
# when the engine targets an accelerator this host actually has.
compat = tuple(getattr(backend_cls, "gpu_compat", ("cpu",)))
try:
from core.device_caps import detect_host_caps # noqa: PLC0415
family = detect_host_caps().family
except Exception:
family = "cpu"
if family != "cpu" and family in compat:
return ("fp16", "fp32")
return ("fp32",)
def _features(self, backend_cls) -> tuple[str, ...]:
features = ["tts"]
if getattr(backend_cls, "supports_cloning", False) is True:
features.append("voice_clone")
if getattr(backend_cls, "supports_voice_design", False):
features.append("voice_design")
if getattr(backend_cls, "supports_emotion", False):
features.append("emotion")
return tuple(features)
def _replace_state(model: ModelInfo, state: str) -> ModelInfo:
return ModelInfo(
catalog_model_id=model.catalog_model_id,
model_version=model.model_version,
model_digest=model.model_digest,
precisions=model.precisions,
features=model.features,
state=state,
)
def _snapshot_incomplete(repo_dir: str, snapshot: str) -> bool:
"""A download in flight leaves ``*.incomplete`` blobs or dangling links."""
blobs = os.path.join(repo_dir, "blobs")
try:
if any(name.endswith(".incomplete") for name in os.listdir(blobs)):
return True
except OSError:
pass
for current, _dirs, files in os.walk(snapshot):
for name in files:
path = os.path.join(current, name)
if not os.path.exists(path): # dangling symlink
return True
return False
def _system_memory_bytes() -> int:
try:
import psutil # noqa: PLC0415
return int(psutil.virtual_memory().total)
except Exception:
try:
return os.sysconf("SC_PAGE_SIZE") * os.sysconf("SC_PHYS_PAGES")
except (ValueError, OSError, AttributeError):
return 1 # still nonzero: the preflight requires > 0
+63
View File
@@ -0,0 +1,63 @@
"""Wires the adapter to the real VoiceStudio backend.
Kept separate from ``server.py`` so tests can build a
:class:`~runtime_adapter.server.RuntimeContext` from fakes without importing
torch or the engine registry.
"""
from __future__ import annotations
import sys
from . import ADAPTER_VERSION
from ._paths import ensure_backend_on_path
from .inventory import ProductionInventory, slots_per_device
from .server import RuntimeContext
def production_engine_provider(catalog_model_id: str):
"""Resolve a READY catalog model id to its cached engine instance."""
ensure_backend_on_path()
from services.tts_backend import get_engine_instance_for # noqa: PLC0415
return get_engine_instance_for(catalog_model_id)
def build_runtime_context() -> RuntimeContext:
ensure_backend_on_path()
from core.version import APP_VERSION # noqa: PLC0415
slots = slots_per_device()
return RuntimeContext(
runtime_version=APP_VERSION,
adapter_version=ADAPTER_VERSION,
inventory=ProductionInventory(slots=slots),
engine_provider=production_engine_provider,
slot_limit=slots,
)
def prewarm_engines(context: RuntimeContext) -> None:
"""Load and compile every READY model before the socket accepts work.
The GPU Gateway leases an attempt for a bounded window and renews it from
execution evidence. A cold engine produces no evidence: weight loading and
torch compilation can run for minutes emitting nothing, so the lease
expires mid-load, the attempt is fenced, the Job requeues, and the next
attempt pays the same cost a loop that never yields audio.
Paying that cost once at startup, before the adapter is reachable, means
the first real Execute begins inference immediately. Preflight already
refuses a runtime with no READY model, so a failure here is reported and
the model is dropped from the advertised set rather than being offered as
schedulable capacity the node cannot actually serve promptly.
"""
ensure_backend_on_path()
for model in context.inventory.models():
if model.state != "ready":
continue
try:
context.engine_provider(model.catalog_model_id)
except Exception as error: # noqa: BLE001 - reported, never fatal
print(
f"runtime adapter: prewarm of {model.catalog_model_id} failed: {error}",
file=sys.stderr,
)
@@ -0,0 +1,199 @@
syntax = "proto3";
package voicestudio.runtime.v1;
option go_package = "github.com/velixio/vssaas/api/gen/runtime/v1;runtimev1";
// RuntimeAdapterService is local to a GPU Node and is never publicly exposed.
service RuntimeAdapterService {
rpc Health(HealthRequest) returns (HealthResponse);
rpc GetCapabilities(GetCapabilitiesRequest) returns (GetCapabilitiesResponse);
rpc Execute(ExecuteRequest) returns (stream ExecuteResponse);
rpc Cancel(CancelRequest) returns (CancelResponse);
}
message ExecuteResponse { ExecutionEvent event = 1; }
message HealthRequest {}
message HealthResponse {
ServingState state = 1;
string runtime_version = 2;
string adapter_version = 3;
repeated string health_flags = 4;
}
enum ServingState {
SERVING_STATE_UNSPECIFIED = 0;
SERVING_STATE_READY = 1;
SERVING_STATE_DEGRADED = 2;
SERVING_STATE_UNHEALTHY = 3;
}
message GetCapabilitiesRequest {}
message GetCapabilitiesResponse {
string runtime_version = 1;
string adapter_version = 2;
repeated RuntimeDevice devices = 3;
repeated RuntimeModel models = 4;
}
message RuntimeDevice {
string device_id = 1;
string hardware_class = 2;
uint64 total_vram_bytes = 3;
uint32 total_slots = 4;
uint32 free_slots = 5;
}
message RuntimeModel {
string catalog_model_id = 1;
string model_version = 2;
string model_digest = 3;
repeated string precisions = 4;
repeated string features = 5;
RuntimeModelState state = 6;
}
enum RuntimeModelState {
RUNTIME_MODEL_STATE_UNSPECIFIED = 0;
RUNTIME_MODEL_STATE_INSTALLED = 1;
RUNTIME_MODEL_STATE_LOADING = 2;
RUNTIME_MODEL_STATE_READY = 3;
RUNTIME_MODEL_STATE_FAILED = 4;
}
message ExecuteRequest {
string job_id = 1;
string attempt_id = 2;
string device_id = 3;
string slot_id = 4;
ModelSpec model = 5;
map<string, ParameterValue> parameters = 6;
repeated LocalArtifact inputs = 7;
repeated LocalArtifact outputs = 8;
int64 deadline_unix_ms = 9;
uint32 maximum_preview_bytes = 10;
}
message ModelSpec {
string catalog_model_id = 1;
string model_version = 2;
string model_digest = 3;
string precision = 4;
}
message ParameterValue {
oneof value {
string string_value = 1;
int64 integer_value = 2;
double number_value = 3;
bool boolean_value = 4;
}
}
message LocalArtifact {
string artifact_id = 1;
string local_handle = 2;
LocalArtifactOperation operation = 3;
uint64 expected_size_bytes = 4;
string expected_sha256 = 5;
string media_type = 6;
}
enum LocalArtifactOperation {
LOCAL_ARTIFACT_OPERATION_UNSPECIFIED = 0;
LOCAL_ARTIFACT_OPERATION_READ = 1;
LOCAL_ARTIFACT_OPERATION_WRITE = 2;
}
message ExecutionEvent {
string job_id = 1;
string attempt_id = 2;
uint64 sequence = 3;
int64 observed_at_unix_ms = 4;
oneof payload {
ExecutionStarted started = 10;
ExecutionProgress progress = 11;
PreviewChunk preview = 12;
ExecutionCompleted completed = 13;
ExecutionFailed failed = 14;
ExecutionCanceled canceled = 15;
}
}
message ExecutionStarted {}
message ExecutionProgress {
uint32 progress_permille = 1;
string stage_code = 2;
}
message PreviewChunk {
uint64 sequence = 1;
string media_type = 2;
bytes data = 3;
}
message ExecutionCompleted {
repeated LocalArtifactManifest outputs = 1;
RuntimeMeasurements measurements = 2;
}
message LocalArtifactManifest {
string artifact_id = 1;
string local_handle = 2;
uint64 size_bytes = 3;
string sha256 = 4;
string media_type = 5;
uint64 duration_ms = 6;
}
message ExecutionFailed {
RuntimeFailureClass failure_class = 1;
string stable_code = 2;
string safe_detail = 3;
RuntimeMeasurements measurements = 4;
}
message ExecutionCanceled {
RuntimeMeasurements measurements = 1;
}
enum RuntimeFailureClass {
RUNTIME_FAILURE_CLASS_UNSPECIFIED = 0;
RUNTIME_FAILURE_CLASS_INPUT = 1;
RUNTIME_FAILURE_CLASS_MODEL_LOAD = 2;
RUNTIME_FAILURE_CLASS_INFERENCE = 3;
RUNTIME_FAILURE_CLASS_GPU_RESOURCE = 4;
RUNTIME_FAILURE_CLASS_LOCAL_STORAGE = 5;
RUNTIME_FAILURE_CLASS_RUNTIME = 6;
RUNTIME_FAILURE_CLASS_CANCELED = 7;
}
message RuntimeMeasurements {
uint64 normalized_input_characters = 1;
uint64 input_audio_ms = 2;
uint64 output_audio_ms = 3;
uint64 gpu_execution_ms = 4;
uint64 cpu_execution_ms = 5;
}
message CancelRequest {
string job_id = 1;
string attempt_id = 2;
string reason_code = 3;
int64 deadline_unix_ms = 4;
}
message CancelResponse {
CancelDisposition disposition = 1;
}
enum CancelDisposition {
CANCEL_DISPOSITION_UNSPECIFIED = 0;
CANCEL_DISPOSITION_ACCEPTED = 1;
CANCEL_DISPOSITION_ALREADY_TERMINAL = 2;
CANCEL_DISPOSITION_NOT_FOUND = 3;
}
+168
View File
@@ -0,0 +1,168 @@
"""``--selfcheck``: validate the Go preflight's expectations against ourselves.
Starts the server on a private temp socket, then runs a Python port of
``internal/gateway/preflight.go``'s checks over the wire: socket-path safety,
READY health with version evidence, identical versions across Health and
GetCapabilities, valid unique devices, and at least one explicitly READY,
digest-pinned model with a version and precisions. Prints only a bounded
readiness summary (never handles, paths, or credentials) and exits nonzero on
any failed expectation the same fail-closed behavior a node deployment gets
from ``cmd/runtime-adapter-preflight``.
"""
from __future__ import annotations
import os
import stat as stat_module
import tempfile
from dataclasses import dataclass
import grpc
from .gen import runtime_adapter_pb2 as pb2
from .gen import runtime_adapter_pb2_grpc as pb2_grpc
_MAX_UINT32 = 2**32 - 1
class PreflightError(Exception):
"""One failed preflight expectation, with a bounded message."""
@dataclass(frozen=True)
class PreflightSummary:
socket_path: str
runtime_version: str
adapter_version: str
device_count: int
ready_model_count: int
total_slots: int
free_slots: int
def render(self) -> str:
return (
f"runtime={self.runtime_version} adapter={self.adapter_version} "
f"devices={self.device_count} ready_models={self.ready_model_count} "
f"slots={self.free_slots}/{self.total_slots}"
)
def validate_socket_file(socket_path: str) -> None:
if not socket_path or not os.path.isabs(socket_path):
raise PreflightError("socket path must be absolute")
info = os.lstat(socket_path)
if stat_module.S_ISLNK(info.st_mode) or not stat_module.S_ISSOCK(info.st_mode):
raise PreflightError("endpoint must be a local Unix socket")
parent = os.stat(os.path.dirname(socket_path))
if not stat_module.S_ISDIR(parent.st_mode) or parent.st_mode & 0o002:
raise PreflightError("socket directory is unsafe")
def run_preflight(socket_path: str, timeout_s: float = 10.0) -> PreflightSummary:
"""Port of ``PreflightRuntime`` + ``validateRuntimeCapabilities``."""
validate_socket_file(socket_path)
with grpc.insecure_channel(f"unix:{socket_path}") as channel:
stub = pb2_grpc.RuntimeAdapterServiceStub(channel)
try:
health = stub.Health(pb2.HealthRequest(), timeout=timeout_s)
except grpc.RpcError as exc:
raise PreflightError(f"health call failed: {exc.code().name}")
if (
health.state != pb2.SERVING_STATE_READY
or not health.runtime_version.strip()
or not health.adapter_version.strip()
):
raise PreflightError("runtime is not ready with versioned adapter evidence")
try:
caps = stub.GetCapabilities(pb2.GetCapabilitiesRequest(), timeout=timeout_s)
except grpc.RpcError as exc:
raise PreflightError(f"capabilities call failed: {exc.code().name}")
return _validate_capabilities(socket_path, health, caps)
def _validate_capabilities(socket_path, health, caps) -> PreflightSummary:
if not caps.runtime_version.strip() or not caps.adapter_version.strip():
raise PreflightError("capabilities lack version evidence")
if (
caps.runtime_version != health.runtime_version
or caps.adapter_version != health.adapter_version
):
raise PreflightError("health and capabilities versions disagree")
if not caps.devices:
raise PreflightError("no execution devices reported")
total_slots = free_slots = 0
seen_devices: set[str] = set()
for device in caps.devices:
if (
not device.device_id.strip()
or not device.hardware_class.strip()
or device.total_vram_bytes == 0
or device.total_slots == 0
or device.free_slots > device.total_slots
):
raise PreflightError("invalid execution device reported")
if device.device_id in seen_devices:
raise PreflightError("duplicate execution device reported")
seen_devices.add(device.device_id)
if (
total_slots + device.total_slots > _MAX_UINT32
or free_slots + device.free_slots > _MAX_UINT32
):
raise PreflightError("slot total overflows protocol limit")
total_slots += device.total_slots
free_slots += device.free_slots
ready = 0
seen_models: set[tuple[str, str, str]] = set()
for model in caps.models:
if model.state != pb2.RUNTIME_MODEL_STATE_READY:
continue
if (
not model.catalog_model_id.strip()
or not model.model_version.strip()
or not model.model_digest.strip()
or not model.precisions
):
raise PreflightError("invalid ready model reported")
identity = (model.catalog_model_id, model.model_version, model.model_digest)
if identity in seen_models:
raise PreflightError("duplicate ready model reported")
seen_models.add(identity)
ready += 1
if ready == 0:
raise PreflightError("no ready model reported")
return PreflightSummary(
socket_path=socket_path,
runtime_version=health.runtime_version,
adapter_version=health.adapter_version,
device_count=len(caps.devices),
ready_model_count=ready,
total_slots=total_slots,
free_slots=free_slots,
)
def selfcheck(timeout_s: float = 10.0) -> int:
"""Start the production server on a temp socket and preflight it."""
from .production import build_runtime_context # noqa: PLC0415
from .server import create_server # noqa: PLC0415
context = build_runtime_context()
warm = getattr(context.inventory, "warm", None)
if callable(warm):
print("selfcheck: warming model inventory (first run hashes weights)…")
warm()
# Short prefix: macOS caps Unix-socket paths at 103 characters and the
# default macOS tempdir is already ~60 characters deep.
with tempfile.TemporaryDirectory(prefix="vs-rta-") as tmp:
os.chmod(tmp, 0o700)
socket_path = os.path.join(tmp, "runtime.sock")
server = create_server(context, socket_path)
server.start()
try:
summary = run_preflight(socket_path, timeout_s=timeout_s)
except PreflightError as failure:
print(f"selfcheck: FAIL: {failure}")
return 1
finally:
server.stop(grace=2).wait()
print(f"selfcheck: OK: {summary.render()}")
return 0
+208
View File
@@ -0,0 +1,208 @@
"""The gRPC server: Unix-domain socket only, no HTTP, no TCP.
``Health`` and ``GetCapabilities`` read the same version constants from one
:class:`RuntimeContext`, so the "identical versions" preflight expectation
holds by construction. Socket-path safety mirrors the Go preflight's checks
(absolute path, no symlink, parent directory not world-writable) at bind time
so an unsafe deployment fails closed on our side too.
"""
from __future__ import annotations
import os
import stat as stat_module
import threading
from concurrent import futures
from dataclasses import dataclass, field
import grpc
from . import ADAPTER_VERSION, DEFAULT_SOCKET_PATH, SOCKET_ENV
from .executor import AttemptRegistry, Executor
from .gen import runtime_adapter_pb2 as pb2
from .gen import runtime_adapter_pb2_grpc as pb2_grpc
from .inventory import (
STATE_FAILED,
STATE_INSTALLED,
STATE_LOADING,
STATE_READY,
)
_MODEL_STATE_TO_PB = {
STATE_INSTALLED: pb2.RUNTIME_MODEL_STATE_INSTALLED,
STATE_LOADING: pb2.RUNTIME_MODEL_STATE_LOADING,
STATE_READY: pb2.RUNTIME_MODEL_STATE_READY,
STATE_FAILED: pb2.RUNTIME_MODEL_STATE_FAILED,
}
@dataclass
class RuntimeContext:
"""Everything the servicer needs; tests build it from fakes."""
runtime_version: str
inventory: object
engine_provider: object
adapter_version: str = ADAPTER_VERSION
slot_limit: int = 1
progress_interval: float = 0.5
poll_interval: float = 0.02
registry: AttemptRegistry = field(default_factory=AttemptRegistry)
def executor(self) -> Executor:
return Executor(
self.inventory,
self.engine_provider,
self.registry,
slot_limit=self.slot_limit,
progress_interval=self.progress_interval,
poll_interval=self.poll_interval,
)
class RuntimeAdapterServicer(pb2_grpc.RuntimeAdapterServiceServicer):
def __init__(self, context: RuntimeContext):
self._context = context
self._executor = context.executor()
def Health(self, request, grpc_context):
flags: list[str] = []
state = pb2.SERVING_STATE_READY
try:
devices = self._context.inventory.devices(
busy_slots=self._context.registry.active_count()
)
models = self._context.inventory.models()
except Exception:
return pb2.HealthResponse(
state=pb2.SERVING_STATE_UNHEALTHY,
runtime_version=self._context.runtime_version,
adapter_version=self._context.adapter_version,
health_flags=["inventory-error"],
)
if not devices:
state = pb2.SERVING_STATE_UNHEALTHY
flags.append("no-device")
if not any(model.state == STATE_READY for model in models):
state = max(state, pb2.SERVING_STATE_DEGRADED)
flags.append("no-ready-model")
return pb2.HealthResponse(
state=state,
runtime_version=self._context.runtime_version,
adapter_version=self._context.adapter_version,
health_flags=flags,
)
def GetCapabilities(self, request, grpc_context):
busy = self._context.registry.active_count()
response = pb2.GetCapabilitiesResponse(
runtime_version=self._context.runtime_version,
adapter_version=self._context.adapter_version,
)
for device in self._context.inventory.devices(busy_slots=busy):
response.devices.append(
pb2.RuntimeDevice(
device_id=device.device_id,
hardware_class=device.hardware_class,
total_vram_bytes=device.total_vram_bytes,
total_slots=device.total_slots,
free_slots=device.free_slots,
)
)
for model in self._context.inventory.models():
response.models.append(
pb2.RuntimeModel(
catalog_model_id=model.catalog_model_id,
model_version=model.model_version,
model_digest=model.model_digest,
precisions=list(model.precisions),
features=list(model.features),
state=_MODEL_STATE_TO_PB.get(
model.state, pb2.RUNTIME_MODEL_STATE_UNSPECIFIED
),
)
)
return response
def Execute(self, request, grpc_context):
yield from self._executor.execute(request, grpc_context)
def Cancel(self, request, grpc_context):
disposition = self._context.registry.cancel(request.job_id, request.attempt_id)
return pb2.CancelResponse(disposition=disposition)
def resolve_socket_path(explicit: str | None = None) -> str:
return (
(explicit or "").strip()
or os.environ.get(SOCKET_ENV, "").strip()
or DEFAULT_SOCKET_PATH
)
def prepare_socket(socket_path: str) -> str:
"""Fail closed on any unsafe socket placement; remove only a stale socket."""
if not socket_path or not os.path.isabs(socket_path):
raise ValueError("runtime socket path must be absolute")
parent = os.path.dirname(socket_path)
try:
parent_stat = os.stat(parent)
except OSError as exc:
raise ValueError(f"runtime socket directory is missing: {exc}") from exc
if not stat_module.S_ISDIR(parent_stat.st_mode) or parent_stat.st_mode & 0o002:
raise ValueError("runtime socket directory is unsafe (world-writable?)")
try:
existing = os.lstat(socket_path)
except FileNotFoundError:
return socket_path
if stat_module.S_ISSOCK(existing.st_mode):
os.unlink(socket_path) # stale socket from a previous run
return socket_path
raise ValueError("runtime socket path exists and is not a socket")
def create_server(
context: RuntimeContext, socket_path: str, *, max_workers: int | None = None
) -> grpc.Server:
prepare_socket(socket_path)
workers = max_workers or max(8, context.slot_limit * 2 + 4)
server = grpc.server(
futures.ThreadPoolExecutor(
max_workers=workers, thread_name_prefix="runtime-adapter"
)
)
pb2_grpc.add_RuntimeAdapterServiceServicer_to_server(
RuntimeAdapterServicer(context), server
)
bound = server.add_insecure_port(f"unix:{socket_path}")
if bound == 0:
raise RuntimeError("failed to bind the runtime adapter socket")
return server
def serve(context: RuntimeContext, socket_path: str) -> int:
"""Run until SIGINT/SIGTERM. Returns a process exit code."""
import signal # noqa: PLC0415
warm = getattr(context.inventory, "warm", None)
if callable(warm):
warm() # hash installed snapshots before the socket exists
server = create_server(context, socket_path)
server.start()
try:
os.chmod(socket_path, 0o660) # gateway runs under the same service identity
except OSError:
pass
stop = threading.Event()
def _stop(_signum, _frame):
stop.set()
signal.signal(signal.SIGTERM, _stop)
signal.signal(signal.SIGINT, _stop)
stop.wait()
server.stop(grace=10).wait()
try:
os.unlink(socket_path)
except OSError:
pass
return 0
+402
View File
@@ -0,0 +1,402 @@
"""Process-bound credentials for the first-party remote administration UI.
The durable ``OMNIVOICE_API_KEY`` is an operator secret, not a browser session.
This module exchanges it for opaque, bounded-lifetime credentials without
depending on FastAPI or persisting a verifier to disk.
"""
from __future__ import annotations
import hmac
import re
import secrets
import sys
import threading
import time
from types import ModuleType
from base64 import urlsafe_b64encode
from collections import OrderedDict
from collections.abc import Callable
from dataclasses import dataclass, field
from cryptography.hazmat.primitives import hashes
from cryptography.hazmat.primitives.kdf.hkdf import HKDF
SESSION_TTL_SECONDS = 8 * 60 * 60
WS_TICKET_TTL_SECONDS = 30
MAX_ADMIN_SESSIONS = 256
MAX_WS_TICKETS = 512
ADMIN_SESSION_PREFIX = "ovs_admin_session_"
WS_TICKET_PREFIX = "ovs_ws_ticket_"
_TOKEN_BYTES = 32
_ENCODED_TOKEN_LENGTH = 43
_TOKEN_BODY_RE = re.compile(rf"^[A-Za-z0-9_-]{{{_ENCODED_TOKEN_LENGTH}}}$")
_ALLOWED_WS_PATHS = frozenset({"/ws/events", "/ws/transcribe"})
_ADMIN_CAPABILITIES = frozenset({"consume", "admin"})
_KEY_GENERATION_INFO = b"omnivoice-admin-key-generation-v1"
def _hash_token(token: str, pepper: bytes) -> str:
# These are 256-bit random values, not user-chosen passwords. A keyed,
# process-local index is the right primitive: there is no feasible password
# dictionary to slow down, and a copied record is unusable without the
# store's independently generated pepper.
return hmac.digest(pepper, token.encode("utf-8"), "sha256").hex()
def _encode_token(raw: bytes) -> str:
return urlsafe_b64encode(raw).rstrip(b"=").decode("ascii")
@dataclass(frozen=True)
class IssuedSession:
token: str = field(repr=False)
expires_at: float
@dataclass(frozen=True)
class IssuedTicket:
token: str = field(repr=False)
expires_at: float
@dataclass(frozen=True)
class SessionRecord:
credential_id: str
capabilities: frozenset[str]
issued_at: float
expires_at: float
@dataclass(frozen=True)
class _StoredSession:
credential_id: str
issued_monotonic: float
expires_monotonic: float
issued_at: float
expires_at: float
def public(self) -> SessionRecord:
return SessionRecord(
credential_id=self.credential_id,
capabilities=_ADMIN_CAPABILITIES,
issued_at=self.issued_at,
expires_at=self.expires_at,
)
@dataclass(frozen=True)
class _StoredTicket:
session_hash: str
path: str
issued_monotonic: float
expires_monotonic: float
class AdminSessionStore:
"""Thread-safe, process-local store for admin sessions and WS tickets."""
def __init__(
self,
*,
monotonic: Callable[[], float] = time.monotonic,
wall_time: Callable[[], float] = time.time,
token_bytes: Callable[[int], bytes] = secrets.token_bytes,
pepper: bytes | None = None,
session_ttl_seconds: int = SESSION_TTL_SECONDS,
ws_ticket_ttl_seconds: int = WS_TICKET_TTL_SECONDS,
max_sessions: int = MAX_ADMIN_SESSIONS,
max_tickets: int = MAX_WS_TICKETS,
) -> None:
if session_ttl_seconds <= 0 or ws_ticket_ttl_seconds <= 0:
raise ValueError("credential TTLs must be positive")
if max_sessions <= 0 or max_tickets <= 0:
raise ValueError("credential store capacities must be positive")
self._monotonic = monotonic
self._wall_time = wall_time
self._token_bytes = token_bytes
self._pepper = pepper if pepper is not None else secrets.token_bytes(32)
if len(self._pepper) < 32:
raise ValueError("session-store pepper must contain at least 256 bits")
self._session_ttl = session_ttl_seconds
self._ticket_ttl = ws_ticket_ttl_seconds
self._max_sessions = max_sessions
self._max_tickets = max_tickets
self._sessions: OrderedDict[str, _StoredSession] = OrderedDict()
self._tickets: OrderedDict[str, _StoredTicket] = OrderedDict()
self._ticket_hashes_by_session: dict[str, set[str]] = {}
self._key_generation: bytes | None = None
self._lock = threading.RLock()
def __repr__(self) -> str:
snapshot = self.debug_snapshot()
return (
"AdminSessionStore("
f"sessions={snapshot['sessions']}, ws_tickets={snapshot['ws_tickets']})"
)
@staticmethod
def _normalize_master(api_key: str | None) -> str:
return api_key.strip() if isinstance(api_key, str) else ""
def _generation(self, api_key: str) -> bytes:
return HKDF(
algorithm=hashes.SHA256(),
length=32,
salt=self._pepper,
info=_KEY_GENERATION_INFO,
).derive(api_key.encode("utf-8", errors="surrogatepass"))
def _sync_key_locked(self, api_key: str | None) -> bool:
normalized = self._normalize_master(api_key)
if not normalized:
self._clear_credentials_locked()
self._key_generation = None
return False
generation = self._generation(normalized)
if self._key_generation is None:
self._key_generation = generation
return True
if not hmac.compare_digest(self._key_generation, generation):
self._clear_credentials_locked()
self._key_generation = generation
return True
@staticmethod
def _valid_token(token: str | None, prefix: str) -> bool:
if not isinstance(token, str) or not token.startswith(prefix):
return False
return bool(_TOKEN_BODY_RE.fullmatch(token.removeprefix(prefix)))
def _new_token_locked(self, prefix: str, existing: object) -> tuple[str, str]:
for _attempt in range(8):
raw = self._token_bytes(_TOKEN_BYTES)
if not isinstance(raw, bytes) or len(raw) != _TOKEN_BYTES:
raise RuntimeError("token source must return exactly 32 bytes")
token = prefix + _encode_token(raw)
token_hash = _hash_token(token, self._pepper)
if token_hash not in existing:
return token, token_hash
raise RuntimeError("credential token source produced repeated collisions")
def _clear_credentials_locked(self) -> None:
self._sessions.clear()
self._tickets.clear()
self._ticket_hashes_by_session.clear()
def _remove_ticket_locked(self, ticket_hash: str) -> _StoredTicket | None:
ticket = self._tickets.pop(ticket_hash, None)
if ticket is None:
return None
session_tickets = self._ticket_hashes_by_session.get(ticket.session_hash)
if session_tickets is not None:
session_tickets.discard(ticket_hash)
if not session_tickets:
self._ticket_hashes_by_session.pop(ticket.session_hash, None)
return ticket
def _remove_session_locked(self, session_hash: str) -> _StoredSession | None:
record = self._sessions.pop(session_hash, None)
for ticket_hash in tuple(self._ticket_hashes_by_session.get(session_hash, ())):
self._remove_ticket_locked(ticket_hash)
# Defensive cleanup keeps a prior partial mutation from preserving a
# dangling reverse-index bucket even when the session was already gone.
self._ticket_hashes_by_session.pop(session_hash, None)
return record
def _purge_locked(self, now: float) -> None:
# TTLs are fixed per store and monotonic issue times never decrease, so
# insertion order is expiry order. Only the expired prefix can require
# work; the common request path examines at most one record per type.
while self._sessions:
session_hash = next(iter(self._sessions))
if now < self._sessions[session_hash].expires_monotonic:
break
self._remove_session_locked(session_hash)
while self._tickets:
ticket_hash = next(iter(self._tickets))
if now < self._tickets[ticket_hash].expires_monotonic:
break
self._remove_ticket_locked(ticket_hash)
def _evict_sessions_locked(self) -> None:
while len(self._sessions) >= self._max_sessions:
self._remove_session_locked(next(iter(self._sessions)))
def _evict_tickets_locked(self) -> None:
while len(self._tickets) >= self._max_tickets:
self._remove_ticket_locked(next(iter(self._tickets)))
def issue(self, api_key: str) -> IssuedSession:
normalized = self._normalize_master(api_key)
if not normalized:
raise ValueError("configured API key required")
with self._lock:
self._sync_key_locked(normalized)
now = self._monotonic()
wall_now = self._wall_time()
self._purge_locked(now)
self._evict_sessions_locked()
token, token_hash = self._new_token_locked(ADMIN_SESSION_PREFIX, self._sessions)
expires_monotonic = now + self._session_ttl
expires_at = wall_now + self._session_ttl
self._sessions[token_hash] = _StoredSession(
credential_id=token_hash,
issued_monotonic=now,
expires_monotonic=expires_monotonic,
issued_at=wall_now,
expires_at=expires_at,
)
return IssuedSession(token=token, expires_at=expires_at)
def resolve(self, token: str | None, api_key: str | None) -> SessionRecord | None:
if not self._valid_token(token, ADMIN_SESSION_PREFIX):
return None
assert isinstance(token, str)
with self._lock:
if not self._sync_key_locked(api_key):
return None
now = self._monotonic()
self._purge_locked(now)
record = self._sessions.get(_hash_token(token, self._pepper))
if record is None or now >= record.expires_monotonic:
return None
return record.public()
def revoke(self, token: str | None) -> bool:
if not self._valid_token(token, ADMIN_SESSION_PREFIX):
return False
assert isinstance(token, str)
token_hash = _hash_token(token, self._pepper)
with self._lock:
return self._remove_session_locked(token_hash) is not None
def revoke_by_credential(self, credential_id: str | None) -> bool:
if not isinstance(credential_id, str) or len(credential_id) != 64:
return False
with self._lock:
return self._remove_session_locked(credential_id) is not None
def issue_ws_ticket(
self,
session_token: str | None,
path: str,
api_key: str | None,
) -> IssuedTicket:
if path not in _ALLOWED_WS_PATHS:
raise ValueError("WebSocket path is not allowed")
if not self._valid_token(session_token, ADMIN_SESSION_PREFIX):
raise PermissionError("valid admin session required")
assert isinstance(session_token, str)
session_hash = _hash_token(session_token, self._pepper)
return self.issue_ws_ticket_for_credential(session_hash, path, api_key)
def issue_ws_ticket_for_credential(
self,
credential_id: str | None,
path: str,
api_key: str | None,
) -> IssuedTicket:
if path not in _ALLOWED_WS_PATHS:
raise ValueError("WebSocket path is not allowed")
with self._lock:
if not isinstance(credential_id, str) or len(credential_id) != 64:
raise PermissionError("valid admin session required")
if not self._sync_key_locked(api_key):
raise PermissionError("valid admin session required")
now = self._monotonic()
self._purge_locked(now)
session = self._sessions.get(credential_id)
if session is None or now >= session.expires_monotonic:
raise PermissionError("valid admin session required")
self._evict_tickets_locked()
token, token_hash = self._new_token_locked(WS_TICKET_PREFIX, self._tickets)
expires_at = self._wall_time() + self._ticket_ttl
self._tickets[token_hash] = _StoredTicket(
session_hash=credential_id,
path=path,
issued_monotonic=now,
expires_monotonic=now + self._ticket_ttl,
)
self._ticket_hashes_by_session.setdefault(credential_id, set()).add(
token_hash
)
return IssuedTicket(token=token, expires_at=expires_at)
def consume_ws_ticket(
self,
ticket_token: str | None,
path: str,
api_key: str | None,
) -> SessionRecord | None:
if not self._valid_token(ticket_token, WS_TICKET_PREFIX):
return None
assert isinstance(ticket_token, str)
with self._lock:
if not self._sync_key_locked(api_key):
return None
now = self._monotonic()
self._purge_locked(now)
ticket = self._remove_ticket_locked(
_hash_token(ticket_token, self._pepper)
)
if ticket is None or now >= ticket.expires_monotonic or ticket.path != path:
return None
session = self._sessions.get(ticket.session_hash)
if session is None or now >= session.expires_monotonic:
return None
return session.public()
def clear(self) -> None:
with self._lock:
self._clear_credentials_locked()
self._key_generation = None
@property
def active_session_count(self) -> int:
with self._lock:
self._purge_locked(self._monotonic())
return len(self._sessions)
def debug_snapshot(self) -> dict[str, int]:
with self._lock:
self._purge_locked(self._monotonic())
return {"sessions": len(self._sessions), "ws_tickets": len(self._tickets)}
#: Synthetic ``sys.modules`` key holding the one per-process store. A module
#: object in ``sys.modules`` is the only namespace that survives everything
#: test suites do to this package: ``importlib.reload`` re-executes module
#: code but never touches unrelated ``sys.modules`` entries, and the purges
#: that pop whole ``services.*`` / ``api.*`` trees match package prefixes this
#: underscore-prefixed top-level name is outside of.
_ANCHOR_MODULE_NAME = "_omnivoice_admin_session_store_anchor"
def _process_store() -> AdminSessionStore:
"""Return THE per-process store, however this module was (re)imported.
Auth is process-global state: the copy of this module that issues a
credential and the copy that later resolves it must always be looking at
the same store. A bare module-level ``AdminSessionStore()`` breaks that
the moment anything reloads or re-imports this module (fresh module dict
fresh store freshly issued sessions vanish for holders of the old
reference, and vice versa). Anchoring the instance outside the module's
own namespace makes every copy of this module share one store.
"""
anchor = sys.modules.get(_ANCHOR_MODULE_NAME)
if not isinstance(anchor, ModuleType):
anchor = ModuleType(_ANCHOR_MODULE_NAME)
anchor.__doc__ = "Process-global anchor for the VoiceStudio admin-session store."
sys.modules[_ANCHOR_MODULE_NAME] = anchor
store = getattr(anchor, "admin_session_store", None)
if store is None:
store = AdminSessionStore()
anchor.admin_session_store = store
return store
admin_session_store = _process_store()
+83 -16
View File
@@ -520,12 +520,10 @@ class WhisperXBackend(ASRBackend):
def _pick_device() -> tuple[str, str]:
# CUDA fp16 when available; otherwise CPU int8 (fastest CPU path,
# negligible WER regression vs fp32 for whisper-large-v3).
try:
import torch
if torch.cuda.is_available():
return "cuda", "float16"
except Exception:
pass
# _ctranslate2_cuda_ok, not torch.cuda.is_available: ROCm torch also
# answers True there, and CTranslate2 has no HIP backend (#1529).
if _ctranslate2_cuda_ok():
return "cuda", "float16"
return "cpu", "int8"
# Peak VRAM (GB) to load *and transcribe* whisper large-v3 per CTranslate2
@@ -981,12 +979,10 @@ class FasterWhisperBackend(ASRBackend):
# - Apple Silicon / CPU → CPU int8 (fastest on CPU, negligible
# WER regression vs fp32 for whisper-large-v3)
device, compute_type = "cpu", "int8"
try:
import torch
if torch.cuda.is_available():
device, compute_type = "cuda", "float16"
except Exception:
pass
# _ctranslate2_cuda_ok, not torch.cuda.is_available: ROCm torch also
# answers True there, and CTranslate2 has no HIP backend (#1529).
if _ctranslate2_cuda_ok():
device, compute_type = "cuda", "float16"
logger.info(
"faster-whisper loading %s on %s (%s)",
self._model_name, device, compute_type,
@@ -2329,7 +2325,7 @@ _INSTALL_HINTS: dict[str, str] = {
"mac-ARM source installs since 0.3.22. Parakeet TDT v3 on the GPU via "
"MLX: 25 European languages, word timestamps, ~2 GB unified memory.)"
),
"moonshine": "pip install useful-moonshine (edge/CPU-optimized ASR)",
"moonshine": "uv pip install moonshine-onnx (or moonshine-voice; edge/CPU-optimized ASR)",
"funasr": "pip install funasr (SenseVoiceSmall + FSMN-VAD; CUDA or CPU)",
"sherpa-onnx-asr": "uv add sherpa-onnx (ONNX live dictation; CPU, cross-platform)",
"openai-compat-asr": (
@@ -2464,6 +2460,61 @@ def _mps_available() -> bool:
return False
def _cuda_reported_available() -> bool:
"""``torch.cuda.is_available()`` verbatim — True on real CUDA *and* HIP."""
try:
import torch
return bool(torch.cuda.is_available())
except Exception: # noqa: BLE001 — no torch
return False
def _rocm_torch() -> bool:
"""True when torch is the ROCm (HIP) build.
ROCm torch masquerades as CUDA: ``torch.cuda.is_available()`` answers True
and tensors live on ``"cuda"`` devices, but the CUDA *runtime libraries*
other packages ship are still NVIDIA-only. ``torch.version.hip`` is the
one honest tell.
"""
try:
import torch
return getattr(torch.version, "hip", None) is not None
except Exception: # noqa: BLE001 — no torch
return False
def _ctranslate2_cuda_ok() -> bool:
"""Whether CTranslate2 (whisperx / faster-whisper) may use ``"cuda"``.
CTranslate2 has NO HIP backend. On a ROCm host torch says cuda is
available (HIP), the device string is handed to CTranslate2, and its
NVIDIA CUDA runtime dies with "CUDA driver version is insufficient for
CUDA runtime version" — the #1529 report, an AMD RX 7900 XTX in the
:rocm Docker image. Real CUDA only; ROCm hosts take the CPU path here
(auto-detect prefers pytorch-whisper there, which does use HIP).
Also honors the user compute-device override (Settings Performance /
``OMNIVOICE_DEVICE``): a host pinned to cpu (or any non-cuda family)
must not hand CTranslate2 a CUDA device the probe applies the
override, so gating on its family covers every CT2 loader at once.
"""
try:
from core.device_caps import detect_host_caps
if detect_host_caps().family != "cuda":
return False
except Exception: # noqa: BLE001 — fail SAFE, not fast
# Without a working probe we can't know whether an override or a
# ROCm build is in play — guessing "cuda" from torch here is exactly
# the #1529 crash. CPU always works.
logger.warning("device probe failed — CTranslate2 taking the CPU path", exc_info=True)
return False
return _cuda_reported_available() and not _rocm_torch()
def _auto_detect() -> str:
"""Pick the best available ASR engine **for this hardware**.
@@ -2495,6 +2546,14 @@ def _auto_detect() -> str:
"""
if _mps_available() and _probe_available(MLXWhisperBackend):
return "mlx-whisper"
# Same class as the Apple case, on the ROCm axis (#1529): whisperx and
# faster-whisper are CTranslate2, which has no HIP backend — on a ROCm
# host they run on the CPU while the GPU sits idle (and before
# _ctranslate2_cuda_ok they died outright trying NVIDIA's runtime).
# pytorch-whisper is a pure transformers pipeline riding torch itself,
# so it genuinely uses the HIP GPU there.
if _rocm_torch() and _cuda_reported_available() and _probe_available(PyTorchWhisperBackend):
return "pytorch-whisper"
if _probe_available(WhisperXBackend):
return "whisperx"
if _probe_available(FasterWhisperBackend):
@@ -3077,10 +3136,18 @@ def _offline_asr_repo(backend_id: str | None = None) -> str | None:
bid = backend_id or active_backend_id()
if bid == "whisperx":
return _fw_repo(os.environ.get("ASR_MODEL_WHISPERX", "large-v3"))
if bid in ("faster-whisper", "faster-whisper-isolated"):
# The crash-isolated sidecar loads the SAME CT2 weights as in-process
# faster-whisper (it reuses the ASR_MODEL_FASTER selection).
if bid == "faster-whisper":
return _fw_repo(os.environ.get("ASR_MODEL_FASTER", _FASTER_WHISPER_DEFAULT))
if bid == "faster-whisper-isolated":
# Mirror the sidecar's own resolution (_asr_sidecar/main.py):
# ASR_MODEL_FW is a sidecar-only override, otherwise the shared
# ASR_MODEL_FASTER selection applies — so the preflight can never
# download a different repo than the sidecar will load.
return _fw_repo(
os.environ.get("ASR_MODEL_FW")
or os.environ.get("ASR_MODEL_FASTER")
or _FASTER_WHISPER_DEFAULT
)
if bid == "mlx-whisper":
return os.environ.get("ASR_MODEL", _MLX_MODEL_DEFAULT)
if bid == "parakeet-mlx":
+145
View File
@@ -0,0 +1,145 @@
"""Opt-in adapter from OSS profiles/generation to the hosted v1 contract.
Local VoiceStudio never calls this module unless the caller explicitly requests
``hosted`` execution *and* all VSS_HOSTED_* settings are present. It stages
text/reference bytes as hosted Artifacts, creates a consent-backed Voice, and
uses durable Jobs; no local path, source recording URL, or plaintext text is
sent in a Job snapshot.
"""
from __future__ import annotations
import asyncio
import hashlib
import os
import time
import uuid
from dataclasses import dataclass
from pathlib import Path
import httpx
class HostedVoiceError(RuntimeError):
"""A safe, user-actionable hosted adapter failure."""
@dataclass(frozen=True)
class HostedSettings:
base_url: str
token: str
project_id: str
model_id: str
model_version: str
base_voice_id: str
consent_text_version: str
@classmethod
def from_environment(cls) -> "HostedSettings | None":
values = {
name: os.environ.get(name, "").strip()
for name in (
"VSS_HOSTED_API_BASE", "VSS_HOSTED_API_TOKEN",
"VSS_HOSTED_PROJECT_ID", "VSS_HOSTED_MODEL_ID",
"VSS_HOSTED_MODEL_VERSION", "VSS_HOSTED_BASE_VOICE_ID",
)
}
if not any(values.values()):
return None
missing = [name for name, value in values.items() if not value]
if missing:
raise HostedVoiceError("Hosted execution is incomplete; configure " + ", ".join(missing) + ".")
base_url = values["VSS_HOSTED_API_BASE"].rstrip("/")
if not base_url.startswith(("http://", "https://")):
raise HostedVoiceError("VSS_HOSTED_API_BASE must be an http(s) URL.")
return cls(
base_url=base_url, token=values["VSS_HOSTED_API_TOKEN"],
project_id=values["VSS_HOSTED_PROJECT_ID"], model_id=values["VSS_HOSTED_MODEL_ID"],
model_version=values["VSS_HOSTED_MODEL_VERSION"], base_voice_id=values["VSS_HOSTED_BASE_VOICE_ID"],
consent_text_version=os.environ.get("VSS_HOSTED_CONSENT_TEXT_VERSION", "oss-spoken-consent-v1").strip() or "oss-spoken-consent-v1",
)
class HostedVoiceClient:
def __init__(self, settings: HostedSettings, client: httpx.AsyncClient | None = None):
self.settings = settings
self.client = client or httpx.AsyncClient(base_url=settings.base_url, timeout=60)
self._owns_client = client is None
async def aclose(self) -> None:
if self._owns_client:
await self.client.aclose()
def _headers(self, *, idempotency: bool = False) -> dict[str, str]:
headers = {"Authorization": f"Bearer {self.settings.token}"}
if idempotency:
headers["Idempotency-Key"] = str(uuid.uuid4())
return headers
async def _request(self, method: str, path: str, *, json: dict | None = None, headers: dict | None = None) -> httpx.Response:
response = await self.client.request(method, path, json=json, headers=headers)
if response.is_error:
detail = "hosted service rejected the request"
try:
body = response.json()
detail = body.get("error", {}).get("message") or body.get("detail") or detail
except ValueError:
pass
raise HostedVoiceError(f"Hosted request failed ({response.status_code}): {detail}")
return response
async def upload_artifact(self, *, purpose: str, media_type: str, payload: bytes) -> str:
digest = hashlib.sha256(payload).hexdigest()
grant = (await self._request("POST", "/v1/artifacts/upload-authorizations", json={
"project_id": self.settings.project_id, "purpose": purpose, "media_type": media_type,
"size_bytes": len(payload), "sha256": digest,
}, headers=self._headers())).json()
put_headers = {k: v for k, v in (grant.get("required_headers") or {}).items() if k.lower() not in {"host", "content-length"}}
put_headers.setdefault("Content-Type", media_type)
response = await self.client.request(grant.get("method", "PUT"), grant["url"], content=payload, headers=put_headers)
if response.is_error:
raise HostedVoiceError(f"Hosted Artifact upload failed ({response.status_code}).")
await self._request("POST", f"/v1/artifacts/{grant['artifact_id']}/complete", json={"size_bytes": len(payload), "sha256": digest}, headers=self._headers())
return grant["artifact_id"]
async def create_voice(self, *, name: str, description: str, reference_path: str) -> str:
payload = Path(reference_path).read_bytes()
if not payload:
raise HostedVoiceError("The reference recording is empty.")
suffix = Path(reference_path).suffix.lower()
media_type = {".wav": "audio/wav", ".mp3": "audio/mpeg", ".flac": "audio/flac"}.get(suffix, "audio/wav")
reference_id = await self.upload_artifact(purpose="reference_audio", media_type=media_type, payload=payload)
voice = await self._request("POST", "/v1/voices", json={
"project_id": self.settings.project_id, "display_name": name, "description": description[:1024],
"reference_audio_artifact_id": reference_id,
"consent": {"attestation_text_version": self.settings.consent_text_version},
}, headers=self._headers(idempotency=True))
return voice.json()["id"]
async def synthesize(self, *, text: str, profile_voice_id: str, language: str | None = None) -> bytes:
text_artifact = await self.upload_artifact(purpose="input", media_type="text/plain", payload=text.encode("utf-8"))
configuration = {"voice_id": self.settings.base_voice_id, "voice_reference_id": profile_voice_id, "output_format": "wav"}
if language and language != "Auto":
configuration["language"] = language
job = await self._request("POST", "/v1/jobs", json={
"project_id": self.settings.project_id, "workflow": "tts",
"model": {"id": self.settings.model_id, "version": self.settings.model_version},
"input": {"text_artifact_id": text_artifact}, "configuration": configuration,
}, headers=self._headers(idempotency=True))
job_id = job.json()["job_id"]
deadline = time.monotonic() + 15 * 60
while time.monotonic() < deadline:
view = (await self._request("GET", f"/v1/jobs/{job_id}", headers=self._headers())).json()
if view.get("state") == "succeeded":
outputs = view.get("output_artifact_ids") or []
if not outputs:
raise HostedVoiceError("Hosted synthesis completed without audio output.")
grant = (await self._request("POST", f"/v1/artifacts/{outputs[0]}/download-authorization", headers=self._headers())).json()
audio = await self.client.request(grant.get("method", "GET"), grant["url"])
if audio.is_error:
raise HostedVoiceError("Hosted synthesis output could not be downloaded.")
return audio.content
if view.get("state") in {"failed", "canceled"}:
raise HostedVoiceError("Hosted synthesis did not complete successfully.")
await asyncio.sleep(0.5)
raise HostedVoiceError("Hosted synthesis timed out waiting for its durable Job.")
+23
View File
@@ -251,10 +251,33 @@ def list_backends() -> list[dict]:
"effective_device": "network",
"routing_status": "n/a",
"routing_reason": None,
# The openai-compat family entry and the LLM Providers panel are
# ONE system (this backend resolves through the active provider),
# but the UI presented them as unrelated. Naming the resolved
# provider + model here lets the catalogue row say which endpoint
# actually answers, instead of a generic family label.
"hint": _provider_hint(bid) if ok else None,
})
return out
def _provider_hint(bid: str) -> str | None:
"""``Provider · model`` for the openai-compat row, None for everything else."""
if bid != "openai-compat":
return None
try:
from services import llm_providers
p = llm_providers.active_provider()
if p is None:
return None
model = llm_providers.resolve_model(p)
return f"{p.display_name} · {model}" if model else p.display_name
except Exception:
# The hint is decoration; a provider-registry hiccup must not take
# down the whole engines listing.
return None
def active_backend_id() -> str:
explicit = os.environ.get("OMNIVOICE_LLM_BACKEND")
if explicit:
+4
View File
@@ -363,6 +363,10 @@ class SubprocessBackend(TTSBackend):
# A duck-typed marker survives that.
_is_subprocess_isolated: bool = True
# Generation happens in the sidecar: parent-side accelerator counters
# can't see its allocations (see TTSBackend.runs_out_of_process).
runs_out_of_process: bool = True
# Default sample rate; subclasses override.
_DEFAULT_SAMPLE_RATE = 24000
+39 -5
View File
@@ -300,6 +300,23 @@ class TTSBackend(ABC):
#: 0 means "no meaningful floor" (CPU-class engines) and never warns.
min_vram_gb: float = 0.0
#: True when generation allocates in ANOTHER process — a dedicated-venv
#: sidecar (SubprocessBackend) or a spawned binary (omnivoice-gguf).
#: Parent-process accelerator counters cannot see those allocations, so
#: profilers/diagnostics must not attribute the parent's VRAM numbers to
#: the engine. Duck-typed (attribute, not issubclass) for the same
#: module-purge reason as `_is_subprocess_isolated`.
runs_out_of_process: bool = False
def model_identity(self) -> Optional[str]:
"""Which concrete model this backend would run, for adapter engines
that host several very different models behind one backend id
(mlx-audio, sherpa-onnx, cosyvoice). None means the engine id
already names the model. Profilers and diagnostics use this to
label results without it, Kokoro-under-mlx and Dia-under-mlx
rows are indistinguishable."""
return None
@abstractmethod
def generate(
self,
@@ -1075,7 +1092,8 @@ class KittenTTSBackend(TTSBackend):
- English only
- Much faster + much smaller install
Preset voice is chosen via `extras["voice"]` (defaults to "Jasper"). Any
Preset voice is chosen via `extras["voice"]` (defaults to DEFAULT_VOICE,
"expr-voice-2-f"). Any
`ref_audio` / `instruct` / `language` arg is ignored with a log line so
the common call-site doesn't need to know which engine it's talking to.
"""
@@ -1384,6 +1402,9 @@ class MLXAudioBackend(TTSBackend):
def sample_rate(self) -> int:
return self._sr
def model_identity(self) -> Optional[str]:
return self._model_id
@property
def supported_languages(self) -> list[str]:
# Per-model; Kokoro supports 8, Qwen3 ~4, Kugel 24. Return "multi"
@@ -1571,6 +1592,18 @@ class CosyVoiceBackend(TTSBackend):
def supported_languages(self) -> list[str]:
return ["zh", "en", "ja", "ko", "yue", "de", "es", "fr", "it", "ru"]
@staticmethod
def _resolved_model_dir() -> str:
return os.environ.get(
"OMNIVOICE_COSYVOICE_MODEL",
"pretrained_models/Fun-CosyVoice3-0.5B",
)
def model_identity(self) -> Optional[str]:
# v1/v2/v3 all live behind the one "cosyvoice" id — the directory
# basename is the only thing that tells the models apart.
return os.path.basename(os.path.normpath(self._resolved_model_dir()))
def _ensure_loaded(self):
if self._model is not None:
return
@@ -1578,10 +1611,7 @@ class CosyVoiceBackend(TTSBackend):
if not ok:
raise RuntimeError(f"CosyVoice unavailable: {msg}")
from cosyvoice.cli.cosyvoice import AutoModel # type: ignore[import-not-found]
model_dir = os.environ.get(
"OMNIVOICE_COSYVOICE_MODEL",
"pretrained_models/Fun-CosyVoice3-0.5B",
)
model_dir = self._resolved_model_dir()
logger.info("Loading CosyVoice from %s", model_dir)
self._model = AutoModel(model_dir=model_dir)
@@ -1814,6 +1844,10 @@ class SherpaOnnxBackend(TTSBackend):
self._tts = None
self._model_dir = os.environ.get("OMNIVOICE_SHERPA_MODEL", "")
def model_identity(self) -> Optional[str]:
model_dir = (self._model_dir or "").strip()
return os.path.basename(os.path.normpath(model_dir)) if model_dir else None
@classmethod
def is_available(cls) -> tuple[bool, str]:
try:
+205
View File
@@ -0,0 +1,205 @@
"""Shared fakes and harness for the runtime-adapter tests.
Not a test module (no ``test_`` prefix): imported by
``test_runtime_adapter_capabilities.py`` and
``test_runtime_adapter_execute.py``.
"""
from __future__ import annotations
import contextlib
import hashlib
import os
import shutil
import tempfile
import threading
import time
import grpc
from runtime_adapter.gen import runtime_adapter_pb2 as pb2
from runtime_adapter.gen import runtime_adapter_pb2_grpc as pb2_grpc
from runtime_adapter.inventory import (
STATE_READY,
DeviceInfo,
ModelInfo,
)
from runtime_adapter.server import RuntimeContext, create_server
READY_MODEL = ModelInfo(
catalog_model_id="fake-tts",
model_version="a" * 40,
model_digest="sha256:" + "b" * 64,
precisions=("fp32",),
features=("tts",),
state=STATE_READY,
)
DEVICE = DeviceInfo(
device_id="cpu:0",
hardware_class="test-cpu",
total_vram_bytes=8 * 1024**3,
total_slots=1,
free_slots=1,
)
class FakeInventory:
def __init__(self, models=None, devices=None):
self._models = list(models) if models is not None else [READY_MODEL]
self._devices = list(devices) if devices is not None else [DEVICE]
def devices(self, busy_slots: int = 0):
return [
DeviceInfo(
device_id=d.device_id,
hardware_class=d.hardware_class,
total_vram_bytes=d.total_vram_bytes,
total_slots=d.total_slots,
free_slots=max(0, d.total_slots - busy_slots),
)
for d in self._devices
]
def models(self):
return list(self._models)
class FakeEngine:
"""Half a second of silence at 24 kHz, instantly."""
sample_rate = 24000
def __init__(self):
self.generate_calls = []
def ensure_ready(self):
pass
def generate(self, text, **kw):
import torch
self.generate_calls.append((text, kw))
return torch.zeros(1, 12000)
class SlowEngine(FakeEngine):
"""Sleeps through generate in small slices so tests stay responsive."""
def __init__(self, seconds: float = 10.0):
super().__init__()
self.seconds = seconds
self.started = threading.Event()
def generate(self, text, **kw):
self.started.set()
deadline = time.monotonic() + self.seconds
while time.monotonic() < deadline:
time.sleep(0.01)
return super().generate(text, **kw)
class FailingEngine(FakeEngine):
def __init__(self, exc: BaseException, phase: str = "synthesis"):
super().__init__()
self._exc = exc
self._phase = phase
def ensure_ready(self):
if self._phase == "model_load":
raise self._exc
def generate(self, text, **kw):
raise self._exc
def make_context(engine=None, inventory=None, **kw) -> RuntimeContext:
engine = engine if engine is not None else FakeEngine()
engines = {READY_MODEL.catalog_model_id: engine}
kw.setdefault("progress_interval", 0.05)
kw.setdefault("poll_interval", 0.005)
return RuntimeContext(
runtime_version="1.2.3-test",
inventory=inventory if inventory is not None else FakeInventory(),
engine_provider=lambda model_id: engines[model_id],
**kw,
)
@contextlib.contextmanager
def serve_over_socket(context: RuntimeContext, tmp_path=None):
# A pytest tmp_path routinely exceeds the 103-character Unix-socket path
# limit on macOS, so the socket gets its own short private tempdir.
socket_dir = tempfile.mkdtemp(prefix="vs-rta-")
socket_path = os.path.join(socket_dir, "runtime.sock")
server = create_server(context, socket_path)
server.start()
channel = grpc.insecure_channel(f"unix:{socket_path}")
try:
yield pb2_grpc.RuntimeAdapterServiceStub(channel), socket_path
finally:
channel.close()
server.stop(grace=0).wait()
shutil.rmtree(socket_dir, ignore_errors=True)
def make_execute_request(
tmp_path,
text: str = "hello runtime",
*,
attempt_id: str = "attempt-1",
job_id: str = "job-1",
model: ModelInfo = READY_MODEL,
device_id: str = "cpu:0",
deadline_in_s: float = 30.0,
parameters: dict | None = None,
input_sha256: str | None = None,
input_handle: str | None = None,
output_handle: str | None = None,
) -> pb2.ExecuteRequest:
if input_handle is None:
input_path = tmp_path / "input.txt"
input_path.write_text(text, encoding="utf-8")
input_handle = str(input_path)
if input_sha256 is None and text is not None:
input_sha256 = hashlib.sha256(text.encode("utf-8")).hexdigest()
if output_handle is None:
output_handle = str(tmp_path / "output.wav")
return pb2.ExecuteRequest(
job_id=job_id,
attempt_id=attempt_id,
device_id=device_id,
slot_id="slot-0",
model=pb2.ModelSpec(
catalog_model_id=model.catalog_model_id,
model_version=model.model_version,
model_digest=model.model_digest,
precision="fp32",
),
parameters=parameters or {},
inputs=[
pb2.LocalArtifact(
artifact_id="in-1",
local_handle=input_handle,
operation=pb2.LOCAL_ARTIFACT_OPERATION_READ,
expected_sha256=input_sha256 or "",
media_type="text/plain",
)
],
outputs=[
pb2.LocalArtifact(
artifact_id="out-1",
local_handle=output_handle,
operation=pb2.LOCAL_ARTIFACT_OPERATION_WRITE,
media_type="audio/wav",
)
],
deadline_unix_ms=int((time.time() + deadline_in_s) * 1000),
maximum_preview_bytes=0,
)
def terminal_of(events):
last = events[-1].event
kind = last.WhichOneof("payload")
assert kind in ("completed", "failed", "canceled"), kind
return kind, last
+29
View File
@@ -49,9 +49,38 @@ if not os.environ.get("OMNIVOICE_ENV_FILE"):
os.environ["OMNIVOICE_MODEL"] = "test"
import functools
import shutil
import pytest
@functools.lru_cache(maxsize=1)
def supports_symlinks() -> bool:
"""True when this process may create symlinks. On Windows,
``os.symlink`` raises OSError without Developer Mode or admin rights, so
symlink-dependent assertions must be skipped there rather than fail."""
probe_dir = tempfile.mkdtemp(prefix="omnivoice-symlink-probe-")
try:
target = os.path.join(probe_dir, "target")
with open(target, "w", encoding="utf-8"):
pass
try:
os.symlink(target, os.path.join(probe_dir, "link"))
except (OSError, NotImplementedError):
return False
return True
finally:
shutil.rmtree(probe_dir, ignore_errors=True)
@pytest.fixture(scope="session")
def symlinks_supported() -> bool:
"""Bool fixture over :func:`supports_symlinks` for guarding the
symlink-only assertions of a test while its other assertions still run."""
return supports_symlinks()
@pytest.fixture
def asr_model_installed(monkeypatch, request):
"""Neutralize the no-ASR-installed preflight (asr_model_missing_error →
+271 -6
View File
@@ -9,7 +9,10 @@ generation.py's proven ``_run_inference`` rather than re-implementing it.
"""
from __future__ import annotations
import io
import json
from pathlib import Path
import wave
import pytest
@@ -23,6 +26,23 @@ from core import archetypes # noqa: E402
from api.routers import archetypes as arch_router # noqa: E402
def _wav_bytes() -> bytes:
buf = io.BytesIO()
with wave.open(buf, "wb") as wav:
wav.setnchannels(1)
wav.setsampwidth(2)
wav.setframerate(24_000)
wav.writeframes(b"\x00\x01" * 64)
return buf.getvalue()
def _write_wav(path: Path) -> bytes:
data = _wav_bytes()
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(data)
return data
@pytest.fixture(scope="module")
def client():
app = FastAPI()
@@ -133,8 +153,7 @@ def test_preview_serves_cached_wav_without_model(client):
key = arch_router._preview_key(sample)
cache_dir = Path(arch_router._PREVIEW_DIR)
cache_dir.mkdir(parents=True, exist_ok=True)
dummy = b"RIFF\x24\x00\x00\x00WAVEfmt cached-archetype-preview"
(cache_dir / f"{key}.wav").write_bytes(dummy)
dummy = _write_wav(cache_dir / f"{key}.wav")
r = client.get(f"/archetypes/{sample['id']}/preview")
assert r.status_code == 200
@@ -143,7 +162,7 @@ def test_preview_serves_cached_wav_without_model(client):
# ── Materialize-on-use idempotency (dedup, no re-render) ───────────────────────
def test_use_is_idempotent_dedup(client, monkeypatch):
def test_use_is_idempotent_dedup(client, tmp_path, monkeypatch, symlinks_supported):
"""The 2nd `/use` of the same archetype reuses its one materialized profile
and does NOT render again the guarantee that materialize-on-select in any
voice picker can't spawn duplicate rows on repeated picks.
@@ -151,6 +170,7 @@ def test_use_is_idempotent_dedup(client, monkeypatch):
The render boundary (``_render_archetype_wav``) is mocked so no model/GPU is
needed: it just drops a stub WAV where the row expects one.
"""
from core import event_bus
from core.db import init_db
init_db() # ensure the voice_profiles table exists in the hermetic tmp DB
@@ -159,10 +179,13 @@ def test_use_is_idempotent_dedup(client, monkeypatch):
async def _fake_render(a, out_path):
render_calls["n"] += 1
Path(out_path).parent.mkdir(parents=True, exist_ok=True)
Path(out_path).write_bytes(b"RIFF\x24\x00\x00\x00WAVEfmt stub")
_write_wav(Path(out_path))
monkeypatch.setattr(arch_router, "_render_archetype_wav", _fake_render)
emitted = []
monkeypatch.setattr(
event_bus, "emit", lambda topic, payload: emitted.append((topic, payload)),
)
sample = archetypes.list_archetypes(featured=True)[0]
@@ -181,6 +204,248 @@ def test_use_is_idempotent_dedup(client, monkeypatch):
from core.db import db_conn
with db_conn() as conn:
rows = conn.execute(
"SELECT id FROM voice_profiles WHERE personality = ?", (sample["id"],)
"SELECT * FROM voice_profiles WHERE personality = ?",
(arch_router._archetype_personality(sample),),
).fetchall()
assert len(rows) == 1
assert rows[0]["kind"] == "design"
assert json.loads(rows[0]["vd_states"]) == sample["attrs"]
with db_conn() as conn:
row = conn.execute("SELECT * FROM voice_profiles WHERE id=?", (pid,)).fetchone()
assert row["kind"] == "design"
assert row["instruct"] == sample["instruct"]
assert json.loads(row["vd_states"]) == sample["attrs"]
# A missing sample or synthesis-input drift must be repaired before the
# existing profile is returned; Preview and Use must describe one voice.
audio_path = arch_router._profile_audio_path(row["ref_audio_path"])
assert audio_path is not None
audio_path.unlink()
repaired = client.post(f"/archetypes/{sample['id']}/use")
assert repaired.status_code == 200 and repaired.json()["profile_id"] == pid
assert render_calls["n"] == 2
assert audio_path.read_bytes().startswith(b"RIFF")
with db_conn() as conn:
conn.execute("UPDATE voice_profiles SET instruct='male' WHERE id=?", (pid,))
refreshed = client.post(f"/archetypes/{sample['id']}/use")
assert refreshed.status_code == 200
assert refreshed.json()["profile_id"] != pid
assert render_calls["n"] == 3
with db_conn() as conn:
edited = conn.execute("SELECT instruct FROM voice_profiles WHERE id=?", (pid,)).fetchone()
assert edited["instruct"] == "male"
# Continue corruption checks against the new canonical materialization.
pid = refreshed.json()["profile_id"]
with db_conn() as conn:
row = conn.execute("SELECT * FROM voice_profiles WHERE id=?", (pid,)).fetchone()
audio_path = arch_router._profile_audio_path(row["ref_audio_path"])
assert audio_path is not None
audio_path.write_bytes(b"not a WAV")
repaired_corrupt = client.post(f"/archetypes/{sample['id']}/use")
assert repaired_corrupt.status_code == 200
assert render_calls["n"] == 4
if symlinks_supported: # Windows needs Developer Mode to create symlinks
outside = tmp_path / "outside.wav"
outside_bytes = _write_wav(outside)
audio_path.unlink()
audio_path.symlink_to(outside)
repaired_symlink = client.post(f"/archetypes/{sample['id']}/use")
assert repaired_symlink.status_code == 200
assert render_calls["n"] == 5
assert not audio_path.is_symlink()
assert outside.read_bytes() == outside_bytes
# A valid header with a missing payload is not playable and must self-heal.
renders_before = render_calls["n"]
truncated = _wav_bytes()[:44]
audio_path.write_bytes(truncated)
repaired_truncated = client.post(f"/archetypes/{sample['id']}/use")
assert repaired_truncated.status_code == 200
assert render_calls["n"] == renders_before + 1
assert audio_path.read_bytes() != truncated
def test_archetype_staged_repair_preserves_concurrently_edited_profile(
client, monkeypatch,
):
"""A repair may publish only if the row still belongs to the archetype."""
from core.config import VOICES_DIR
from core.db import db_conn, init_db
init_db()
sample = archetypes.list_archetypes(featured=True)[3]
personality = arch_router._archetype_personality(sample)
edited_personality = f"user-edited:{sample['id']}"
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?, ?)",
(sample["id"], personality, edited_personality),
)
original_id = {"value": None}
mutation_seen = {"value": False}
async def racing_render(_item, path):
destination = Path(path)
if destination.name.endswith(".staged.wav"):
assert original_id["value"] is not None
with db_conn() as conn:
conn.execute(
"UPDATE voice_profiles SET name='User edit', personality=? WHERE id=?",
(edited_personality, original_id["value"]),
)
mutation_seen["value"] = True
_write_wav(destination)
monkeypatch.setattr(arch_router, "_render_archetype_wav", racing_render)
first = client.post(f"/archetypes/{sample['id']}/use")
assert first.status_code == 200
original_id["value"] = first.json()["profile_id"]
with db_conn() as conn:
original = conn.execute(
"SELECT ref_audio_path FROM voice_profiles WHERE id=?",
(original_id["value"],),
).fetchone()
original_audio = arch_router._profile_audio_path(original["ref_audio_path"])
assert original_audio is not None
corrupt_bytes = b"corrupt user-owned sample"
original_audio.write_bytes(corrupt_bytes)
repaired = client.post(f"/archetypes/{sample['id']}/use")
assert repaired.status_code == 200
repaired_id = repaired.json()["profile_id"]
assert mutation_seen["value"]
assert repaired_id != original_id["value"]
with db_conn() as conn:
edited = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (original_id["value"],),
).fetchone()
canonical = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (repaired_id,),
).fetchone()
canonical_count = conn.execute(
"SELECT count(*) FROM voice_profiles WHERE personality=?", (personality,),
).fetchone()[0]
assert edited["name"] == "User edit"
assert edited["personality"] == edited_personality
assert edited["instruct"] == sample["instruct"]
assert original_audio.read_bytes() == corrupt_bytes
assert canonical["personality"] == personality
assert canonical["ref_audio_path"] == arch_router._profile_audio_filename(repaired_id)
assert canonical_count == 1
assert (Path(VOICES_DIR) / canonical["ref_audio_path"]).read_bytes() == _wav_bytes()
assert not list(Path(VOICES_DIR).glob(f".{original_id['value']}-*.staged.wav"))
def test_archetype_use_adopts_only_a_compatible_legacy_row(client, monkeypatch):
from core.config import VOICES_DIR
from core.db import db_conn, init_db
init_db()
sample = archetypes.list_archetypes(featured=True)[1]
legacy_id = "legacyarch"
legacy_audio = Path(VOICES_DIR) / f"{legacy_id}.wav"
_write_wav(legacy_audio)
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?)",
(sample["id"], arch_router._archetype_personality(sample)),
)
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"kind, vd_states, created_at) VALUES (?, 'Legacy archetype', ?, ?, ?, ?, 42, ?, "
"'clone', NULL, 1)",
(
legacy_id, legacy_audio.name, sample["sample_script"], sample["instruct"],
sample["language"], sample["id"],
),
)
async def unexpected_render(*_args):
raise AssertionError("a valid legacy archetype sample must be reused")
monkeypatch.setattr(arch_router, "_render_archetype_wav", unexpected_render)
response = client.post(f"/archetypes/{sample['id']}/use")
assert response.status_code == 200
assert response.json()["profile_id"] == legacy_id
with db_conn() as conn:
row = conn.execute("SELECT * FROM voice_profiles WHERE id=?", (legacy_id,)).fetchone()
assert row["personality"] == arch_router._archetype_personality(sample)
assert row["kind"] == "design"
assert json.loads(row["vd_states"]) == sample["attrs"]
def test_archetype_use_does_not_rewrite_an_imported_personality_collision(
client, monkeypatch,
):
from core.config import VOICES_DIR
from core.db import db_conn, init_db
init_db()
sample = archetypes.list_archetypes(featured=True)[2]
imported_id = "importedarch"
imported_ns_id = "importedarchns"
imported_audio = Path(VOICES_DIR) / f"{imported_id}.wav"
imported_ns_audio = Path(VOICES_DIR) / f"{imported_ns_id}.wav"
original_audio = _write_wav(imported_audio)
original_ns_audio = _write_wav(imported_ns_audio)
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?)",
(sample["id"], arch_router._archetype_personality(sample)),
)
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"kind, is_locked, verified_own_voice, created_at) VALUES "
"(?, 'Imported collision', ?, 'user transcript', 'male', 'Auto', NULL, ?, "
"'clone', 1, 1, 1)",
(imported_id, imported_audio.name, sample["id"]),
)
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"kind, vd_states, is_locked, verified_own_voice, created_at) VALUES "
"(?, 'Imported namespaced collision', ?, ?, ?, ?, 42, ?, "
"'design', NULL, 0, 0, 2)",
(
imported_ns_id, imported_ns_audio.name, sample["sample_script"],
sample["instruct"], sample["language"],
arch_router._archetype_personality(sample),
),
)
async def render(_item, path):
_write_wav(Path(path))
monkeypatch.setattr(arch_router, "_render_archetype_wav", render)
response = client.post(f"/archetypes/{sample['id']}/use")
assert response.status_code == 200
assert response.json()["profile_id"] != imported_id
with db_conn() as conn:
imported = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (imported_id,),
).fetchone()
imported_ns = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (imported_ns_id,),
).fetchone()
created = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (response.json()["profile_id"],),
).fetchone()
assert imported["personality"] == sample["id"]
assert imported["instruct"] == "male"
assert imported["ref_text"] == "user transcript"
assert imported_audio.read_bytes() == original_audio
assert imported_ns["instruct"] == sample["instruct"]
assert imported_ns["ref_text"] == sample["sample_script"]
assert imported_ns["vd_states"] is None
assert imported_ns_audio.read_bytes() == original_ns_audio
assert created["personality"] == arch_router._archetype_personality(sample)
+10
View File
@@ -125,6 +125,16 @@ def test_faster_whisper_float16_unsupported_falls_back_to_int8(monkeypatch):
)
monkeypatch.setitem(sys.modules, "torch", fake_torch)
# The compute-device override gate consults the capability probe before
# the torch mock above — pin it to a CUDA family so the fallback chain
# under test is reachable on a cpu-only CI host.
from core.device_caps import HostCaps
monkeypatch.setattr(
"core.device_caps.detect_host_caps",
lambda: HostCaps(family="cuda", available_families=("cuda", "cpu")),
)
be = FasterWhisperBackend()
be._ensure_model()
+44
View File
@@ -0,0 +1,44 @@
"""Regression tests for the lightweight persisted-WAV trust boundary."""
from __future__ import annotations
import struct
from core.audio_validation import is_playable_wav, resolve_regular_file
def test_oversized_declared_wav_payload_is_not_treated_as_playable(tmp_path):
"""A hostile frame count must be bounded and backed by real payload bytes."""
path = tmp_path / "oversized.wav"
declared_size = 0xFFFF_FFF0
header = struct.pack(
"<4sI4s4sIHHIIHH4sI",
b"RIFF",
0xFFFF_FFFF,
b"WAVE",
b"fmt ",
16,
1,
1,
24_000,
48_000,
2,
16,
b"data",
declared_size,
)
path.write_bytes(header + b"\x00\x01")
assert not is_playable_wav(path)
def test_profile_wav_resolution_rejects_escape_and_symlink(tmp_path, symlinks_supported):
root = tmp_path / "voices"
root.mkdir()
outside = tmp_path / "outside.wav"
outside.write_bytes(b"outside")
assert resolve_regular_file(root, "../outside.wav") is None
assert resolve_regular_file(root, str(outside)) is None
if symlinks_supported: # Windows needs Developer Mode to create symlinks
(root / "linked.wav").symlink_to(outside)
assert resolve_regular_file(root, "linked.wav") is None
+680 -6
View File
@@ -1,25 +1,43 @@
"""Tests for the community gallery (marketplace) loader.
Covers the no-network surface: strict item validation (invalid presets and
unsafe audio URLs are dropped so they can never crash synthesis or fetch from
an arbitrary host), manifest merge/dedup, offline cache reads, filtering, and
the prefilled submit URL. The render/download paths need the model/network and
are exercised at runtime.
Covers strict item validation, manifest/cache boundaries, same-origin preview,
and idempotent profile materialization without a model or network dependency.
"""
from __future__ import annotations
import io
import json
import os
from pathlib import Path
import wave
import pytest
# conftest.py puts `backend/` on sys.path and points OMNIVOICE_DATA_DIR at a
# throwaway tmpdir before this module imports the REAL core.config (the old
# sys.modules stub leaked at collection time and broke mixed runs).
from fastapi import FastAPI # noqa: E402
from fastapi import FastAPI, HTTPException, Response # noqa: E402
from fastapi.testclient import TestClient # noqa: E402
from api.routers import community # noqa: E402
def _wav_bytes() -> bytes:
buf = io.BytesIO()
with wave.open(buf, "wb") as wav:
wav.setnchannels(1)
wav.setsampwidth(2)
wav.setframerate(24_000)
wav.writeframes(b"\x00\x01" * 64)
return buf.getvalue()
def _write_wav(path: Path) -> bytes:
data = _wav_bytes()
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(data)
return data
_FIXTURE = {
"schema_version": 1,
"items": [
@@ -75,12 +93,51 @@ def test_unknown_use_case_dropped():
assert community.validate_item(_FIXTURE["items"][4]) is None
def test_malformed_manifest_entries_do_not_break_other_sources():
valid = _FIXTURE["items"][0]
items, packs = community._merge([
("bad/repo", {"items": 42, "packs": "not-a-list"}),
("good/repo", {"items": [None, "not-an-item", valid], "packs": [None]}),
])
assert [item["id"] for item in items] == [valid["id"]]
assert packs == []
def test_is_valid_instruct():
assert community.is_valid_instruct("male, elderly, very low pitch")
assert not community.is_valid_instruct("male, sultry")
assert not community.is_valid_instruct("male, female")
assert not community.is_valid_instruct("british accent, 四川话")
assert not community.is_valid_instruct("")
def test_preset_attrs_are_normalized_and_complete():
item = community.validate_item(_FIXTURE["items"][0])
assert item["instruct"] == "female, middle-aged, low pitch"
assert item["attrs"] == {
"Gender": "female", "Age": "middle-aged", "Pitch": "low pitch",
"Style": "Auto", "EnglishAccent": "Auto", "ChineseDialect": "Auto",
}
assert item["preview_url"] == "/community/items/p1/preview"
def test_remote_transcript_fields_are_bounded():
preset = community.validate_item({
**_FIXTURE["items"][0],
"sample_script": " x " * (community._MAX_SAMPLE_SCRIPT_CHARS + 10),
})
voice = community.validate_item({
**_FIXTURE["items"][3],
"audio": {
**_FIXTURE["items"][3]["audio"],
"ref_text": " y " * (community._MAX_REF_TEXT_CHARS + 10),
},
})
assert len(preset["sample_script"]) == community._MAX_SAMPLE_SCRIPT_CHARS
assert len(voice["audio"]["ref_text"]) == community._MAX_REF_TEXT_CHARS
# ── merge keeps only valid items ──────────────────────────────────────────────
def test_merge_drops_invalid_and_dedups():
items, packs = community._merge([("debpalash/omnivoice-gallery", _FIXTURE)])
@@ -116,3 +173,620 @@ def test_submit_url(client):
voice = client.get("/community/submit-url", params={"type": "voice"}).json()["url"]
assert "preset-submission.yml" in preset and "omnivoice-gallery" in preset
assert "voice-submission.yml" in voice
# ── bounded cache freshness + stale offline fallback ─────────────────────────
def test_stale_manifest_refreshes_then_stays_fresh(tmp_path, monkeypatch):
monkeypatch.setattr(community, "_CACHE_DIR", tmp_path)
source = "debpalash/omnivoice-gallery"
cache = community._cache_path(source)
cache.parent.mkdir(parents=True)
cache.write_text(json.dumps(_FIXTURE), encoding="utf-8")
os.utime(cache, (100.0, 100.0))
fresh = {**_FIXTURE, "updated_at": "new"}
calls = []
monkeypatch.setattr(
community, "_fetch_remote_manifest",
lambda src: calls.append(src) or fresh,
)
now = 100.0 + community._MANIFEST_MAX_AGE_S + 1
assert community._fetch_manifest(source, False, now=now)["updated_at"] == "new"
assert community._fetch_manifest(source, False, now=now + 1)["updated_at"] == "new"
assert calls == [source]
def test_stale_manifest_falls_back_and_throttles_offline_retry(tmp_path, monkeypatch):
monkeypatch.setattr(community, "_CACHE_DIR", tmp_path)
source = "debpalash/omnivoice-gallery"
cache = community._cache_path(source)
cache.parent.mkdir(parents=True)
cache.write_text(json.dumps(_FIXTURE), encoding="utf-8")
os.utime(cache, (100.0, 100.0))
calls = []
def offline(src):
calls.append(src)
raise OSError("offline")
monkeypatch.setattr(community, "_fetch_remote_manifest", offline)
now = 100.0 + community._MANIFEST_MAX_AGE_S + 1
assert community._fetch_manifest(source, False, now=now) == _FIXTURE
assert community._fetch_manifest(source, False, now=now + 1) == _FIXTURE
assert calls == [source]
def test_manifest_fetch_is_bounded(monkeypatch):
monkeypatch.setattr(community, "_MAX_MANIFEST_BYTES", 8)
class Response:
status_code = 200
headers = {}
def __enter__(self): return self
def __exit__(self, *_args): return False
def raise_for_status(self): return None
def iter_bytes(self): yield b'{"items":[]}'
class Client:
def stream(self, method, url, **kwargs):
assert method == "GET"
assert url.startswith("https://cdn.jsdelivr.net/")
assert kwargs == {"follow_redirects": False}
return Response()
with pytest.raises(ValueError, match="size limit"):
community._fetch_remote_manifest("test/source", client=Client())
def test_manifest_fetch_rejects_redirect_before_external_request():
requested = []
class Response:
status_code = 302
headers = {"location": "https://evil.example/manifest.json"}
def __enter__(self): return self
def __exit__(self, *_args): return False
class Client:
def stream(self, _method, url, **_kwargs):
requested.append(url)
return Response()
with pytest.raises(ValueError, match="disallowed host"):
community._fetch_remote_manifest("test/source", client=Client())
assert requested == [community._manifest_url("test/source")]
# ── Preview proxy ─────────────────────────────────────────────────────────────
def test_canonical_preset_preview_delegates_same_origin(client, monkeypatch):
from core import archetypes
from api.routers import archetypes as arch_router
canonical = archetypes.list_archetypes(featured=True)[0]
item = community.validate_item({
**canonical, "type": "preset", "source": "starter",
})
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
delegated = []
async def preview(archetype_id, local=False):
delegated.append((archetype_id, local))
return Response(_wav_bytes(), media_type="audio/wav")
monkeypatch.setattr(arch_router, "preview_archetype", preview)
response = client.get(f"/community/items/{item['id']}/preview")
local = client.get(f"/community/items/{item['id']}/preview?local=true")
assert response.status_code == local.status_code == 200
assert "location" not in response.headers
assert delegated == [(item["id"], False), (item["id"], True)]
def test_noncanonical_preset_preview_renders_once(client, tmp_path, monkeypatch):
item = community.validate_item(_FIXTURE["items"][0])
monkeypatch.setattr(community, "_CACHE_DIR", tmp_path)
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
from api.routers import archetypes as arch_router
calls = []
async def render(_item, path):
calls.append(path)
_write_wav(Path(path))
monkeypatch.setattr(arch_router, "_render_archetype_wav", render)
first = client.get("/community/items/p1/preview")
second = client.get("/community/items/p1/preview")
assert first.status_code == second.status_code == 200
assert first.content == _wav_bytes()
assert first.headers["x-omnivoice-preview-source"] == "community"
assert len(calls) == 1
community._preset_preview_path(item).write_bytes(b"not audio")
repaired = client.get("/community/items/p1/preview")
assert repaired.status_code == 200
assert repaired.content == _wav_bytes()
assert len(calls) == 2
def test_recorded_preview_is_served_from_same_origin(client, tmp_path, monkeypatch):
item = community.validate_item(_FIXTURE["items"][3])
clip = tmp_path / "voice.wav"
expected = _write_wav(clip)
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
monkeypatch.setattr(community, "_cached_voice_audio", lambda _item: clip)
response = client.get("/community/items/v1/preview")
assert response.status_code == 200
assert response.content == expected
def test_recorded_download_cap_is_atomic(tmp_path, monkeypatch):
item = community.validate_item(_FIXTURE["items"][3])
destination = tmp_path / "voice.wav"
destination.write_bytes(b"existing-good-audio")
monkeypatch.setattr(community, "_MAX_VOICE_AUDIO_BYTES", 8)
class Response:
status_code = 200
headers = {}
def __enter__(self): return self
def __exit__(self, *_args): return False
def raise_for_status(self): return None
def iter_bytes(self): yield b"123456789"
class Client:
def stream(self, method, url, **kwargs):
assert method == "GET" and url.startswith("https://github.com/")
assert kwargs == {"follow_redirects": False}
return Response()
with pytest.raises(HTTPException) as exc:
community._download_voice_audio(item, destination, client=Client())
assert getattr(exc.value, "status_code", None) == 502
assert destination.read_bytes() == b"existing-good-audio"
assert not list(tmp_path.glob(".*.part"))
def test_recorded_download_rejects_redirect_before_external_request(tmp_path):
item = community.validate_item(_FIXTURE["items"][3])
requested = []
class Response:
status_code = 302
headers = {"location": "https://evil.example/private.wav"}
def __enter__(self): return self
def __exit__(self, *_args): return False
class Client:
def stream(self, _method, url, **_kwargs):
requested.append(url)
return Response()
with pytest.raises(HTTPException) as exc:
community._download_voice_audio(item, tmp_path / "voice.wav", client=Client())
assert getattr(exc.value, "status_code", None) == 502
assert requested == [item["audio"]["url"]]
def test_recorded_download_follows_allowlisted_redirect(tmp_path):
item = community.validate_item(_FIXTURE["items"][3])
destination = tmp_path / "voice.wav"
requested = []
expected = _wav_bytes()
class Response:
def __init__(self, status, headers, body=b""):
self.status_code, self.headers, self.body = status, headers, body
def __enter__(self): return self
def __exit__(self, *_args): return False
def raise_for_status(self): return None
def iter_bytes(self): yield self.body
class Client:
def stream(self, _method, url, **_kwargs):
requested.append(url)
if len(requested) == 1:
return Response(302, {"location": "https://objects.githubusercontent.com/v1.wav"})
return Response(200, {}, expected)
community._download_voice_audio(item, destination, client=Client())
assert destination.read_bytes() == expected
assert requested == [item["audio"]["url"], "https://objects.githubusercontent.com/v1.wav"]
def test_recorded_download_rejects_non_audio_bytes(tmp_path):
item = community.validate_item(_FIXTURE["items"][3])
destination = tmp_path / "voice.wav"
class Response:
status_code = 200
headers = {}
def __enter__(self): return self
def __exit__(self, *_args): return False
def raise_for_status(self): return None
def iter_bytes(self): yield b"this is not audio"
class Client:
def stream(self, _method, _url, **_kwargs): return Response()
with pytest.raises(HTTPException, match="valid WAV"):
community._download_voice_audio(item, destination, client=Client())
assert not destination.exists()
assert not list(tmp_path.glob(".*.part"))
# ── Materialization ───────────────────────────────────────────────────────────
def test_community_use_is_idempotent_design_profile(
client, tmp_path, monkeypatch, symlinks_supported,
):
from core import event_bus
from core.db import db_conn, init_db
from api.routers import archetypes as arch_router
init_db()
item = community.validate_item(_FIXTURE["items"][0])
item["_source_repo"] = "test/source"
personality = community._community_personality(item)
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
calls = []
emitted = []
async def render(_item, path):
calls.append(path)
_write_wav(Path(path))
monkeypatch.setattr(arch_router, "_render_archetype_wav", render)
monkeypatch.setattr(
event_bus, "emit", lambda topic, payload: emitted.append((topic, payload)),
)
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?)",
(item["id"], personality),
)
first = client.post("/community/items/p1/use")
second = client.post("/community/items/p1/use")
assert first.status_code == second.status_code == 200
assert second.json()["profile_id"] == first.json()["profile_id"]
assert len(calls) == 1
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (first.json()["profile_id"],),
).fetchone()
assert row["kind"] == "design"
assert row["personality"] == personality
assert json.loads(row["vd_states"])["Gender"] == "female"
assert row["instruct"] == item["instruct"]
assert emitted[-1] == (
"profiles", {"action": "updated", "id": first.json()["profile_id"]},
)
profile_audio = community._stored_profile_audio(row["ref_audio_path"])
assert profile_audio is not None
profile_audio.unlink()
repaired = client.post("/community/items/p1/use")
assert repaired.status_code == 200
assert repaired.json()["profile_id"] == first.json()["profile_id"]
assert profile_audio.read_bytes() == _wav_bytes()
# The current preset preview cache repairs the profile without another
# model render.
assert len(calls) == 1
profile_audio.write_bytes(b"not a WAV")
repaired_corrupt = client.post("/community/items/p1/use")
assert repaired_corrupt.status_code == 200
assert profile_audio.read_bytes() == _wav_bytes()
if symlinks_supported: # Windows needs Developer Mode to create symlinks
outside = tmp_path / "outside.wav"
outside_bytes = _write_wav(outside)
profile_audio.unlink()
profile_audio.symlink_to(outside)
repaired_symlink = client.post("/community/items/p1/use")
assert repaired_symlink.status_code == 200
assert not profile_audio.is_symlink()
assert outside.read_bytes() == outside_bytes
def test_community_staged_repair_preserves_concurrently_edited_profile(
client, monkeypatch,
):
"""A staged community repair must not reclaim a row edited mid-copy."""
from core.config import VOICES_DIR
from core.db import db_conn, init_db
init_db()
item = community.validate_item(_FIXTURE["items"][0])
item["_source_repo"] = "test/source"
personality = community._community_personality(item)
edited_personality = f"user-edited:{personality}"
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?, ?)",
(item["id"], personality, edited_personality),
)
_write_wav(community._preset_preview_path(item))
original_id = {"value": None}
mutation_seen = {"value": False}
real_copy_atomic = community._copy_atomic
def racing_copy(source, destination):
destination = Path(destination)
if destination.name.endswith(".staged.wav"):
assert original_id["value"] is not None
with db_conn() as conn:
conn.execute(
"UPDATE voice_profiles SET name='User edit', personality=? WHERE id=?",
(edited_personality, original_id["value"]),
)
mutation_seen["value"] = True
real_copy_atomic(Path(source), destination)
monkeypatch.setattr(community, "_copy_atomic", racing_copy)
first = client.post(f"/community/items/{item['id']}/use")
assert first.status_code == 200
original_id["value"] = first.json()["profile_id"]
with db_conn() as conn:
original = conn.execute(
"SELECT ref_audio_path FROM voice_profiles WHERE id=?",
(original_id["value"],),
).fetchone()
original_audio = community._stored_profile_audio(original["ref_audio_path"])
assert original_audio is not None
corrupt_bytes = b"corrupt user-owned sample"
original_audio.write_bytes(corrupt_bytes)
repaired = client.post(f"/community/items/{item['id']}/use")
assert repaired.status_code == 200
repaired_id = repaired.json()["profile_id"]
assert mutation_seen["value"]
assert repaired_id != original_id["value"]
with db_conn() as conn:
edited = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (original_id["value"],),
).fetchone()
canonical = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (repaired_id,),
).fetchone()
canonical_count = conn.execute(
"SELECT count(*) FROM voice_profiles WHERE personality=?", (personality,),
).fetchone()[0]
assert edited["name"] == "User edit"
assert edited["personality"] == edited_personality
assert edited["instruct"] == item["instruct"]
assert original_audio.read_bytes() == corrupt_bytes
assert canonical["personality"] == personality
assert canonical["ref_audio_path"] == community._community_profile_audio_filename(
repaired_id, item,
)
assert canonical_count == 1
assert (Path(VOICES_DIR) / canonical["ref_audio_path"]).read_bytes() == _wav_bytes()
assert not list(Path(VOICES_DIR).glob(f".{original_id['value']}-*.staged.wav"))
def test_recorded_community_use_is_idempotent_clone_profile(client, tmp_path, monkeypatch):
from core.db import db_conn, init_db
init_db()
item = community.validate_item(_FIXTURE["items"][3])
item["_source_repo"] = "test/source"
personality = community._community_personality(item)
clip = tmp_path / "recorded.wav"
_write_wav(clip)
cache_calls = []
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
monkeypatch.setattr(
community, "_cached_voice_audio", lambda _item: cache_calls.append(_item["id"]) or clip,
)
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?)",
(item["id"], personality),
)
first = client.post("/community/items/v1/use")
second = client.post("/community/items/v1/use")
assert first.status_code == second.status_code == 200
assert second.json()["profile_id"] == first.json()["profile_id"]
assert cache_calls == ["v1"]
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (first.json()["profile_id"],),
).fetchone()
assert row["kind"] == "clone"
assert row["personality"] == personality
assert row["vd_states"] is None and row["instruct"] == ""
assert row["ref_text"] == ""
old_audio_filename = row["ref_audio_path"]
item["audio"]["url"] = "https://raw.githubusercontent.com/test/source/main/v2.wav"
refreshed = client.post("/community/items/v1/use")
assert refreshed.status_code == 200
assert refreshed.json()["profile_id"] == first.json()["profile_id"]
assert cache_calls == ["v1", "v1"]
with db_conn() as conn:
refreshed_row = conn.execute(
"SELECT ref_audio_path FROM voice_profiles WHERE id=?",
(first.json()["profile_id"],),
).fetchone()
assert refreshed_row["ref_audio_path"] != old_audio_filename
def test_noncanonical_builtin_id_cannot_heal_archetype_profile(client, monkeypatch):
from core import archetypes
from core.db import db_conn, init_db
from api.routers import archetypes as arch_router
init_db()
canonical = archetypes.list_archetypes(featured=True)[0]
changed_instruct = "female" if canonical["instruct"] != "female" else "male"
item = community.validate_item({
**canonical,
"type": "preset",
"source": "community",
"instruct": changed_instruct,
})
item["_source_repo"] = "test/source"
personality = community._community_personality(item)
builtin_profile_id = f"b{os.urandom(4).hex()[:7]}"
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?)",
(canonical["id"], personality),
)
conn.execute(
"INSERT INTO voice_profiles (id, name, personality, instruct, kind, created_at) "
"VALUES (?, 'Built-in profile', ?, 'sentinel', 'design', 1)",
(builtin_profile_id, canonical["id"]),
)
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
async def render(_item, path):
_write_wav(Path(path))
monkeypatch.setattr(arch_router, "_render_archetype_wav", render)
response = client.post(f"/community/items/{canonical['id']}/use")
assert response.status_code == 200
assert response.json()["profile_id"] != builtin_profile_id
with db_conn() as conn:
builtin = conn.execute(
"SELECT instruct FROM voice_profiles WHERE id=?", (builtin_profile_id,),
).fetchone()
community_row = conn.execute(
"SELECT personality FROM voice_profiles WHERE id=?",
(response.json()["profile_id"],),
).fetchone()
conn.execute(
"DELETE FROM voice_profiles WHERE id IN (?, ?)",
(builtin_profile_id, response.json()["profile_id"]),
)
assert builtin["instruct"] == "sentinel"
assert community_row["personality"] == personality
def test_community_use_does_not_rewrite_an_imported_bare_id_collision(
client, monkeypatch,
):
from core.config import VOICES_DIR
from core.db import db_conn, init_db
from api.routers import archetypes as arch_router
init_db()
item = community.validate_item(_FIXTURE["items"][0])
item["_source_repo"] = "test/source"
personality = community._community_personality(item)
imported_id = "importedcomm"
imported_ns_id = "importedcommns"
imported_audio = Path(VOICES_DIR) / f"{imported_id}.wav"
imported_ns_audio = Path(VOICES_DIR) / f"{imported_ns_id}.wav"
original_audio = _write_wav(imported_audio)
original_ns_audio = _write_wav(imported_ns_audio)
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?)",
(item["id"], personality),
)
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"kind, is_locked, verified_own_voice, created_at) VALUES "
"(?, 'Imported collision', ?, 'user transcript', 'male', 'Auto', NULL, ?, "
"'clone', 1, 1, 1)",
(imported_id, imported_audio.name, item["id"]),
)
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"kind, vd_states, is_locked, verified_own_voice, created_at) VALUES "
"(?, 'Imported namespaced collision', ?, ?, ?, ?, 42, ?, "
"'design', NULL, 0, 0, 2)",
(
imported_ns_id, imported_ns_audio.name, item["sample_script"],
item["instruct"], item["language"], personality,
),
)
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
async def render(_item, path):
_write_wav(Path(path))
monkeypatch.setattr(arch_router, "_render_archetype_wav", render)
response = client.post(f"/community/items/{item['id']}/use")
assert response.status_code == 200
assert response.json()["profile_id"] != imported_id
with db_conn() as conn:
imported = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (imported_id,),
).fetchone()
imported_ns = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (imported_ns_id,),
).fetchone()
created = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (response.json()["profile_id"],),
).fetchone()
assert imported["personality"] == item["id"]
assert imported["instruct"] == "male"
assert imported["ref_text"] == "user transcript"
assert imported_audio.read_bytes() == original_audio
assert imported_ns["instruct"] == item["instruct"]
assert imported_ns["ref_text"] == item["sample_script"]
assert imported_ns["vd_states"] is None
assert imported_ns_audio.read_bytes() == original_ns_audio
assert created["personality"] == personality
def test_noncolliding_legacy_community_profile_is_adopted(client, monkeypatch):
from core.config import VOICES_DIR
from core.db import db_conn, init_db
from api.routers import archetypes as arch_router
init_db()
item = community.validate_item(_FIXTURE["items"][0])
item["_source_repo"] = "test/source"
personality = community._community_personality(item)
legacy_id = f"l{os.urandom(4).hex()[:7]}"
with db_conn() as conn:
conn.execute(
"DELETE FROM voice_profiles WHERE personality IN (?, ?)",
(item["id"], personality),
)
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"kind, vd_states, created_at) VALUES "
"(?, 'Legacy community profile', ?, '', ?, ?, NULL, ?, 'design', NULL, 1)",
(legacy_id, f"{legacy_id}.wav", item["instruct"], item["language"], item["id"]),
)
_write_wav(Path(VOICES_DIR) / f"{legacy_id}.wav")
monkeypatch.setattr(
community, "_load", lambda _refresh: (["test/source"], [item], [], False),
)
community._preset_preview_path(item).unlink(missing_ok=True)
rendered = []
async def render(_item, path):
rendered.append(path)
_write_wav(Path(path))
monkeypatch.setattr(arch_router, "_render_archetype_wav", render)
response = client.post(f"/community/items/{item['id']}/use")
assert response.status_code == 200
assert response.json()["profile_id"] == legacy_id
assert len(rendered) == 1
with db_conn() as conn:
adopted = conn.execute(
"SELECT personality, kind, ref_audio_path FROM voice_profiles WHERE id=?",
(legacy_id,),
).fetchone()
assert adopted["personality"] == personality
assert adopted["kind"] == "design"
adopted_audio = community._stored_profile_audio(adopted["ref_audio_path"])
assert adopted_audio is not None and adopted_audio.is_file()
+250
View File
@@ -0,0 +1,250 @@
"""Gallery-import profile materialization contracts."""
from __future__ import annotations
import shutil
import sqlite3
import time
import uuid
from pathlib import Path
import pytest
from fastapi import FastAPI
from fastapi.testclient import TestClient
from api.routers import gallery
from core.db import db_conn, init_db
@pytest.fixture(scope="module")
def client():
init_db()
gallery._init_gallery_db()
app = FastAPI()
app.include_router(gallery.router)
return TestClient(app)
def _gallery_voice(
suffix: str = ".wav", content: bytes = b"RIFF imported voice",
) -> tuple[str, Path]:
voice_id = f"g{uuid.uuid4().hex[:7]}"
path = gallery.VOICE_GALLERY_DIR / f"{voice_id}{suffix}"
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(content)
with db_conn() as conn:
conn.execute(
"""INSERT INTO voice_gallery
(id, name, character, category, source_type, source_url, audio_path,
duration, description, tags, created_at)
VALUES (?, ?, ?, 'import', 'youtube', ?, ?, 5.0, ?, '[]', ?)""",
(
voice_id, "Imported narrator", "Video title is not an instruct",
"https://example.invalid/source", str(path),
"Source URL/notes are not a spoken transcript", time.time(),
),
)
return voice_id, path
def test_save_as_profile_keeps_import_metadata_out_of_tts_fields(client):
voice_id, _ = _gallery_voice()
response = client.post(
f"/gallery/voices/{voice_id}/save-as-profile",
params={"profile_name": "Reusable import"},
)
assert response.status_code == 200
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (response.json()["profile_id"],),
).fetchone()
assert row["kind"] == "clone"
assert row["personality"] == f"gallery:{voice_id}"
assert row["ref_text"] == ""
assert row["instruct"] == ""
assert row["description"] == "Source URL/notes are not a spoken transcript"
def test_to_profile_uses_live_schema_and_clone_metadata(client):
voice_id, _ = _gallery_voice()
response = client.post(f"/gallery/voices/{voice_id}/to-profile")
assert response.status_code == 200
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (response.json()["profile_id"],),
).fetchone()
assert row["kind"] == "clone"
assert row["personality"] == f"gallery:{voice_id}"
assert row["ref_text"] == row["instruct"] == ""
assert row["description"] == "Source URL/notes are not a spoken transcript"
def test_both_import_routes_share_one_idempotent_profile(client, monkeypatch):
emitted = []
monkeypatch.setattr(
gallery.event_bus, "emit", lambda topic, payload: emitted.append((topic, payload)),
)
voice_id, _ = _gallery_voice()
first = client.post(
f"/gallery/voices/{voice_id}/save-as-profile",
params={"profile_name": "One reusable profile"},
)
repeated = client.post(
f"/gallery/voices/{voice_id}/save-as-profile",
params={"profile_name": "Ignored duplicate name"},
)
alternate = client.post(f"/gallery/voices/{voice_id}/to-profile")
assert first.status_code == repeated.status_code == alternate.status_code == 200
assert {
first.json()["profile_id"],
repeated.json()["profile_id"],
alternate.json()["profile_id"],
} == {first.json()["profile_id"]}
with db_conn() as conn:
rows = conn.execute(
"SELECT * FROM voice_profiles WHERE personality=?",
(f"gallery:{voice_id}",),
).fetchall()
assert len(rows) == 1
assert rows[0]["name"] == "One reusable profile"
assert rows[0]["kind"] == "clone" and rows[0]["vd_states"] is None
assert emitted[-1] == (
"profiles", {"action": "updated", "id": first.json()["profile_id"]},
)
def test_gallery_profile_repairs_a_missing_copy_without_duplication(client):
voice_id, source = _gallery_voice()
first = client.post(f"/gallery/voices/{voice_id}/to-profile")
assert first.status_code == 200
with db_conn() as conn:
row = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (first.json()["profile_id"],),
).fetchone()
copied = Path(gallery.VOICES_DIR) / row["ref_audio_path"]
copied.unlink()
repaired = client.post(f"/gallery/voices/{voice_id}/to-profile")
assert repaired.status_code == 200
assert repaired.json()["profile_id"] == first.json()["profile_id"]
assert copied.read_bytes() == source.read_bytes()
def test_gallery_profile_does_not_rewrite_a_namespaced_import_collision(client):
voice_id, source = _gallery_voice()
collision_id = f"c{uuid.uuid4().hex[:7]}"
personality = f"gallery:{voice_id}"
collision_name = gallery._gallery_profile_audio_filename(collision_id, source)
collision_audio = Path(gallery.VOICES_DIR) / collision_name
collision_audio.parent.mkdir(parents=True, exist_ok=True)
collision_audio.write_bytes(b"user-owned audio")
with db_conn() as conn:
conn.execute(
"INSERT INTO voice_profiles "
"(id, name, ref_audio_path, ref_text, instruct, language, seed, personality, "
"description, kind, vd_states, is_locked, verified_own_voice, created_at) "
"VALUES (?, 'User profile', ?, '', '', 'Auto', NULL, ?, "
"'user-owned metadata', 'clone', NULL, 0, 0, ?)",
(collision_id, collision_name, personality, time.time()),
)
response = client.post(f"/gallery/voices/{voice_id}/to-profile")
assert response.status_code == 200
assert response.json()["profile_id"] != collision_id
with db_conn() as conn:
collision = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (collision_id,),
).fetchone()
created = conn.execute(
"SELECT * FROM voice_profiles WHERE id=?", (response.json()["profile_id"],),
).fetchone()
assert collision["description"] == "user-owned metadata"
assert collision_audio.read_bytes() == b"user-owned audio"
assert created["personality"] == personality
def _part_files() -> set[Path]:
return set(Path(gallery.VOICES_DIR).glob("*.part")) | set(
Path(gallery.VOICES_DIR).glob(".*.part")
)
def test_audio_copy_never_holds_the_db_write_lock(client, monkeypatch):
"""The bulk file copy must happen BEFORE the BEGIN IMMEDIATE transaction.
While the copy runs, another backend writer takes (and releases) SQLite's
write lock. If materialization copied inside its own write transaction,
this concurrent writer would hit `database is locked` and the test fails.
"""
from core.config import DB_PATH
voice_id, _ = _gallery_voice()
real_copy2 = shutil.copy2
concurrent_writes = []
def copy_and_probe(src, dst, **kwargs):
probe = sqlite3.connect(DB_PATH, timeout=0.5)
try:
probe.execute("BEGIN IMMEDIATE")
probe.execute(
"UPDATE voice_gallery SET category = category WHERE id = ?",
(voice_id,),
)
probe.commit()
concurrent_writes.append(True)
finally:
probe.close()
return real_copy2(src, dst, **kwargs)
monkeypatch.setattr(gallery.shutil, "copy2", copy_and_probe)
response = client.post(f"/gallery/voices/{voice_id}/to-profile")
assert response.status_code == 200
assert concurrent_writes == [True]
assert _part_files() == set()
def test_failed_copy_leaves_no_temp_droppings_or_profile_row(client, monkeypatch):
"""A copy that dies mid-write must not leave .part files or a DB row."""
voice_id, _ = _gallery_voice()
def exploding_copy(src, dst, **kwargs):
Path(dst).write_bytes(b"partial bytes")
raise OSError("disk full mid-copy")
monkeypatch.setattr(gallery.shutil, "copy2", exploding_copy)
with pytest.raises(OSError, match="disk full mid-copy"):
client.post(f"/gallery/voices/{voice_id}/to-profile")
assert _part_files() == set()
with db_conn() as conn:
rows = conn.execute(
"SELECT * FROM voice_profiles WHERE personality = ?",
(f"gallery:{voice_id}",),
).fetchall()
assert rows == []
def test_gallery_preview_serves_outputs_file_without_root_relative_redirect(client):
voice_id, source = _gallery_voice()
response = client.get(
f"/gallery/voices/{voice_id}/preview", follow_redirects=False,
)
assert response.status_code == 200
assert "location" not in response.headers
assert response.content == source.read_bytes()
def test_gallery_preview_preserves_non_wav_content_type(client):
voice_id, _ = _gallery_voice(".mp3", b"ID3 imported voice")
response = client.get(f"/gallery/voices/{voice_id}/preview")
assert response.status_code == 200
assert response.headers["content-type"] == "audio/mpeg"
+71
View File
@@ -0,0 +1,71 @@
import asyncio
import httpx
import pytest
from services.hosted_voice_api import HostedSettings, HostedVoiceClient, HostedVoiceError
_NAMES = (
"VSS_HOSTED_API_BASE", "VSS_HOSTED_API_TOKEN", "VSS_HOSTED_PROJECT_ID",
"VSS_HOSTED_MODEL_ID", "VSS_HOSTED_MODEL_VERSION", "VSS_HOSTED_BASE_VOICE_ID",
)
def test_hosted_adapter_is_disabled_without_configuration(monkeypatch):
for name in _NAMES:
monkeypatch.delenv(name, raising=False)
assert HostedSettings.from_environment() is None
def test_hosted_adapter_refuses_partial_configuration(monkeypatch):
for name in _NAMES:
monkeypatch.delenv(name, raising=False)
monkeypatch.setenv("VSS_HOSTED_API_BASE", "http://127.0.0.1:8080")
with pytest.raises(HostedVoiceError, match="VSS_HOSTED_API_TOKEN"):
HostedSettings.from_environment()
def test_hosted_adapter_requires_http_endpoint(monkeypatch):
values = {
"VSS_HOSTED_API_BASE": "not-a-url",
"VSS_HOSTED_API_TOKEN": "token",
"VSS_HOSTED_PROJECT_ID": "project",
"VSS_HOSTED_MODEL_ID": "model",
"VSS_HOSTED_MODEL_VERSION": "v1",
"VSS_HOSTED_BASE_VOICE_ID": "base",
}
for name, value in values.items():
monkeypatch.setenv(name, value)
with pytest.raises(HostedVoiceError, match="http"):
HostedSettings.from_environment()
def test_create_voice_uses_artifact_grants_then_canonical_voice_resource(tmp_path):
reference = tmp_path / "reference.wav"
reference.write_bytes(b"reference-audio")
settings = HostedSettings("https://api.test", "token", "project", "model", "v1", "base", "oss-spoken-consent-v1")
requests = []
def handler(request):
requests.append(request)
if request.url.path == "/v1/artifacts/upload-authorizations":
return httpx.Response(200, json={"artifact_id": "artifact-ref", "method": "PUT", "url": "https://objects.test/ref", "required_headers": {}})
if request.url.host == "objects.test":
return httpx.Response(200)
if request.url.path == "/v1/artifacts/artifact-ref/complete":
return httpx.Response(200, json={})
if request.url.path == "/v1/voices":
return httpx.Response(201, json={"id": "hosted-voice"})
return httpx.Response(404)
async def create():
client = HostedVoiceClient(settings, httpx.AsyncClient(base_url=settings.base_url, transport=httpx.MockTransport(handler)))
return await client.create_voice(name="Local profile", description="description", reference_path=str(reference))
assert asyncio.run(create()) == "hosted-voice"
voice_request = next(request for request in requests if request.url.path == "/v1/voices")
body = __import__("json").loads(voice_request.content)
assert body["project_id"] == "project"
assert body["reference_audio_artifact_id"] == "artifact-ref"
assert body["consent"]["attestation_text_version"] == "oss-spoken-consent-v1"
@@ -0,0 +1,171 @@
"""Health/GetCapabilities shape, preflight parity, digest stability.
Mirrors what ``internal/gateway/preflight.go`` in vssaas enforces: READY
health with version evidence, identical versions across both calls, valid
unique devices, and only explicitly-READY models counting as schedulable.
"""
from __future__ import annotations
import os
import sys
import pytest
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from _runtime_adapter_helpers import ( # noqa: E402
DEVICE,
READY_MODEL,
FakeInventory,
make_context,
serve_over_socket,
)
from runtime_adapter.digest import file_sha256, snapshot_digest
from runtime_adapter.gen import runtime_adapter_pb2 as pb2
from runtime_adapter.inventory import (
STATE_FAILED,
STATE_INSTALLED,
STATE_LOADING,
ModelInfo,
catalog_model_version,
)
from runtime_adapter.selfcheck import PreflightError, run_preflight
from runtime_adapter.server import prepare_socket
def _model(state, model_id="other-model", digest="sha256:" + "c" * 64):
return ModelInfo(
catalog_model_id=model_id,
model_version="d" * 40,
model_digest=digest,
precisions=("fp32",),
features=("tts",),
state=state,
)
def test_health_and_capabilities_versions_are_identical_and_ready(tmp_path):
with serve_over_socket(make_context(), tmp_path) as (stub, _):
health = stub.Health(pb2.HealthRequest(), timeout=5)
caps = stub.GetCapabilities(pb2.GetCapabilitiesRequest(), timeout=5)
assert health.state == pb2.SERVING_STATE_READY
assert health.runtime_version == "1.2.3-test"
assert health.adapter_version.strip()
assert caps.runtime_version == health.runtime_version
assert caps.adapter_version == health.adapter_version
def test_capabilities_report_device_and_ready_model_evidence(tmp_path):
inventory = FakeInventory(
models=[
READY_MODEL,
_model(STATE_LOADING, "loading-model"),
_model(STATE_FAILED, "failed-model"),
_model(STATE_INSTALLED, "installed-model"),
]
)
with serve_over_socket(make_context(inventory=inventory), tmp_path) as (stub, _):
caps = stub.GetCapabilities(pb2.GetCapabilitiesRequest(), timeout=5)
[device] = caps.devices
assert device.device_id == DEVICE.device_id
assert device.hardware_class == DEVICE.hardware_class
assert device.total_vram_bytes > 0
assert 0 < device.free_slots <= device.total_slots
by_id = {model.catalog_model_id: model for model in caps.models}
ready = by_id[READY_MODEL.catalog_model_id]
assert ready.state == pb2.RUNTIME_MODEL_STATE_READY
assert ready.model_version.startswith("d" * 40 + "+sha256-")
assert ready.model_digest.startswith("sha256:")
assert list(ready.precisions)
# A loading/failed/installed model is reported truthfully, never READY.
assert by_id["loading-model"].state == pb2.RUNTIME_MODEL_STATE_LOADING
assert by_id["failed-model"].state == pb2.RUNTIME_MODEL_STATE_FAILED
assert by_id["installed-model"].state == pb2.RUNTIME_MODEL_STATE_INSTALLED
def test_preflight_port_passes_against_a_ready_server(tmp_path):
inventory = FakeInventory(models=[READY_MODEL, _model(STATE_LOADING)])
with serve_over_socket(make_context(inventory=inventory), tmp_path) as (
stub,
socket_path,
):
summary = run_preflight(socket_path, timeout_s=5)
assert summary.ready_model_count == 1 # the loading model must not count
assert summary.device_count == 1
assert summary.runtime_version == "1.2.3-test"
assert summary.total_slots == 1
def test_preflight_fails_closed_without_a_ready_model(tmp_path):
inventory = FakeInventory(models=[_model(STATE_LOADING)])
with serve_over_socket(make_context(inventory=inventory), tmp_path) as (
stub,
socket_path,
):
health = stub.Health(pb2.HealthRequest(), timeout=5)
assert health.state == pb2.SERVING_STATE_DEGRADED
assert "no-ready-model" in health.health_flags
with pytest.raises(PreflightError):
run_preflight(socket_path, timeout_s=5)
def test_prepare_socket_rejects_unsafe_paths(tmp_path):
with pytest.raises(ValueError):
prepare_socket("relative/socket.sock")
regular = tmp_path / "not-a-socket"
regular.write_text("x")
with pytest.raises(ValueError):
prepare_socket(str(regular))
missing_parent = tmp_path / "nope" / "runtime.sock"
with pytest.raises(ValueError):
prepare_socket(str(missing_parent))
def test_snapshot_digest_is_stable_and_content_sensitive(tmp_path):
snapshot = tmp_path / "snapshots" / "rev"
snapshot.mkdir(parents=True)
(snapshot / "weights.bin").write_bytes(b"\x01\x02\x03")
(snapshot / "config.json").write_text("{}")
cache = tmp_path / "digest-cache.json"
first = snapshot_digest(snapshot, cache_path=cache)
second = snapshot_digest(snapshot, cache_path=cache) # served from cache
assert first == second
assert first.startswith("sha256:")
assert cache.exists()
# Any byte change must change the digest (cache invalidated by mtime/size).
(snapshot / "weights.bin").write_bytes(b"\x01\x02\x04")
assert snapshot_digest(snapshot, cache_path=cache) != first
with pytest.raises(FileNotFoundError):
snapshot_digest(tmp_path / "empty-none")
def test_catalog_model_version_changes_when_attested_snapshot_changes():
revision = "d" * 40
first = catalog_model_version(revision, "sha256:" + "a" * 64)
second = catalog_model_version(revision, "sha256:" + "b" * 64)
assert first.startswith(revision + "+sha256-")
assert first != second
def test_file_sha256_matches_hashlib(tmp_path):
import hashlib
payload = b"runtime adapter"
path = tmp_path / "f.bin"
path.write_bytes(payload)
assert file_sha256(path) == hashlib.sha256(payload).hexdigest()
def test_socket_file_is_private_to_the_node(tmp_path):
import stat
with serve_over_socket(make_context(), tmp_path) as (_stub, socket_path):
mode = os.lstat(socket_path).st_mode
assert stat.S_ISSOCK(mode)
@@ -0,0 +1,347 @@
"""Execute/Cancel: happy path, deadline, cancel race, failure taxonomy."""
from __future__ import annotations
import hashlib
import os
import sys
import threading
import time
import pytest
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from _runtime_adapter_helpers import ( # noqa: E402
READY_MODEL,
FailingEngine,
FakeInventory,
SlowEngine,
make_context,
make_execute_request,
serve_over_socket,
terminal_of,
)
from runtime_adapter import codes
from runtime_adapter.gen import runtime_adapter_pb2 as pb2
from runtime_adapter.inventory import STATE_INSTALLED, ModelInfo
def _run_direct(context, request):
"""Drive the executor without a live gRPC server (fast path for taxonomy)."""
return list(context.executor().execute(request, None))
def _failure(events):
kind, last = terminal_of(events)
assert kind == "failed", f"expected failed terminal, got {kind}"
return last.failed
# ── happy path ────────────────────────────────────────────────────────────
def test_execute_happy_path_streams_and_writes_the_manifest(tmp_path):
context = make_context()
request = make_execute_request(tmp_path, text="hello runtime")
with serve_over_socket(context, tmp_path) as (stub, _):
events = list(stub.Execute(request, timeout=30))
payloads = [event.event.WhichOneof("payload") for event in events]
assert payloads[0] == "started"
assert payloads[-1] == "completed"
assert all(kind == "progress" for kind in payloads[1:-1])
sequences = [event.event.sequence for event in events]
assert sequences == sorted(sequences)
assert all(event.event.attempt_id == "attempt-1" for event in events)
completed = events[-1].event.completed
[manifest] = completed.outputs
output_path = tmp_path / "output.wav"
assert manifest.local_handle == str(output_path)
assert output_path.stat().st_size == manifest.size_bytes > 0
assert manifest.sha256 == hashlib.sha256(output_path.read_bytes()).hexdigest()
assert manifest.media_type == "audio/wav"
assert manifest.duration_ms == 500 # 12000 samples at 24 kHz
measurements = completed.measurements
assert measurements.normalized_input_characters == len("hello runtime")
assert measurements.output_audio_ms == 500
def test_execute_passes_typed_parameters_to_the_engine(tmp_path):
from _runtime_adapter_helpers import FakeEngine
engine = FakeEngine()
context = make_context(engine=engine)
request = make_execute_request(
tmp_path,
parameters={
"speed": pb2.ParameterValue(number_value=1.5),
"language": pb2.ParameterValue(string_value="en"),
"num_step": pb2.ParameterValue(integer_value=8),
},
)
events = _run_direct(context, request)
assert terminal_of(events)[0] == "completed"
[(text, kwargs)] = engine.generate_calls
assert text == "hello runtime"
assert kwargs == {"speed": 1.5, "language": "en", "num_step": 8}
def test_execute_passes_seed_to_the_engine(tmp_path):
"""Hosted Gallery defaults must retain the OSS deterministic seed."""
from _runtime_adapter_helpers import FakeEngine
engine = FakeEngine()
context = make_context(engine=engine)
request = make_execute_request(
tmp_path,
parameters={"seed": pb2.ParameterValue(integer_value=42)},
)
events = _run_direct(context, request)
assert terminal_of(events)[0] == "completed"
[(_, kwargs)] = engine.generate_calls
assert kwargs == {"seed": 42}
# ── deadline ──────────────────────────────────────────────────────────────
def test_deadline_is_enforced_with_a_stable_code(tmp_path):
context = make_context(engine=SlowEngine(seconds=30))
request = make_execute_request(tmp_path, deadline_in_s=0.4)
start = time.monotonic()
events = _run_direct(context, request)
elapsed = time.monotonic() - start
failed = _failure(events)
assert failed.stable_code in (codes.INFERENCE_DEADLINE, codes.MODEL_LOAD_DEADLINE)
assert failed.failure_class in (
pb2.RUNTIME_FAILURE_CLASS_INFERENCE,
pb2.RUNTIME_FAILURE_CLASS_MODEL_LOAD,
)
assert elapsed < 5, "terminal event must arrive promptly after the deadline"
def test_deadline_in_the_past_is_invalid_input(tmp_path):
context = make_context()
request = make_execute_request(tmp_path)
request.deadline_unix_ms = int(time.time() * 1000) - 1000
failed = _failure(_run_direct(context, request))
assert failed.stable_code == codes.INPUT_DEADLINE_INVALID
assert failed.failure_class == pb2.RUNTIME_FAILURE_CLASS_INPUT
# ── cancel ────────────────────────────────────────────────────────────────
def test_cancel_race_yields_canceled_terminal_and_idempotent_dispositions(tmp_path):
engine = SlowEngine(seconds=30)
context = make_context(engine=engine)
request = make_execute_request(tmp_path)
with serve_over_socket(context, tmp_path) as (stub, _):
stream = stub.Execute(request, timeout=30)
first = next(stream)
assert first.event.WhichOneof("payload") == "started"
assert engine.started.wait(5), "engine must be mid-generate for the race"
cancel = pb2.CancelRequest(job_id="job-1", attempt_id="attempt-1")
assert stub.Cancel(cancel, timeout=5).disposition == (
pb2.CANCEL_DISPOSITION_ACCEPTED
)
# Idempotent while still running.
assert stub.Cancel(cancel, timeout=5).disposition == (
pb2.CANCEL_DISPOSITION_ACCEPTED
)
events = [first, *stream]
kind, last = terminal_of(events)
assert kind == "canceled"
assert last.canceled.HasField("measurements")
# After the terminal event the same cancel is ALREADY_TERMINAL …
assert stub.Cancel(cancel, timeout=5).disposition == (
pb2.CANCEL_DISPOSITION_ALREADY_TERMINAL
)
# … and an unknown attempt is NOT_FOUND.
unknown = pb2.CancelRequest(job_id="job-1", attempt_id="nope")
assert stub.Cancel(unknown, timeout=5).disposition == (
pb2.CANCEL_DISPOSITION_NOT_FOUND
)
def test_cancel_before_any_execute_is_not_found(tmp_path):
with serve_over_socket(make_context(), tmp_path) as (stub, _):
response = stub.Cancel(
pb2.CancelRequest(job_id="j", attempt_id="never-ran"), timeout=5
)
assert response.disposition == pb2.CANCEL_DISPOSITION_NOT_FOUND
# ── failure classification ────────────────────────────────────────────────
def test_model_load_failure_is_classified(tmp_path):
engine = FailingEngine(RuntimeError("weights corrupted"), phase="model_load")
failed = _failure(
_run_direct(make_context(engine=engine), make_execute_request(tmp_path))
)
assert failed.stable_code == codes.MODEL_LOAD_FAILED
assert failed.failure_class == pb2.RUNTIME_FAILURE_CLASS_MODEL_LOAD
def test_inference_failure_is_classified(tmp_path):
engine = FailingEngine(ValueError("synthesis exploded"))
failed = _failure(
_run_direct(make_context(engine=engine), make_execute_request(tmp_path))
)
assert failed.stable_code == codes.INFERENCE_FAILED
assert failed.failure_class == pb2.RUNTIME_FAILURE_CLASS_INFERENCE
def test_gpu_oom_is_classified_as_gpu_resource(tmp_path):
engine = FailingEngine(RuntimeError("CUDA out of memory. Tried to allocate…"))
failed = _failure(
_run_direct(make_context(engine=engine), make_execute_request(tmp_path))
)
assert failed.stable_code == codes.GPU_OUT_OF_MEMORY
assert failed.failure_class == pb2.RUNTIME_FAILURE_CLASS_GPU_RESOURCE
def test_engine_input_rejection_is_invalid_input(tmp_path):
from services.tts_backend import TTSInputError
engine = FailingEngine(TTSInputError("text too long for this engine"))
failed = _failure(
_run_direct(make_context(engine=engine), make_execute_request(tmp_path))
)
assert failed.stable_code == codes.INPUT_REJECTED
assert failed.failure_class == pb2.RUNTIME_FAILURE_CLASS_INPUT
def test_url_handles_are_rejected_never_fetched(tmp_path):
request = make_execute_request(
tmp_path, input_handle="https://evil.example/input.txt", input_sha256=""
)
failed = _failure(_run_direct(make_context(), request))
assert failed.stable_code == codes.INPUT_HANDLE_INVALID
assert failed.failure_class == pb2.RUNTIME_FAILURE_CLASS_INPUT
def test_relative_output_handle_is_rejected(tmp_path):
request = make_execute_request(tmp_path, output_handle="relative/out.wav")
failed = _failure(_run_direct(make_context(), request))
assert failed.stable_code == codes.INPUT_HANDLE_INVALID
def test_model_digest_mismatch_is_rejected(tmp_path):
request = make_execute_request(tmp_path)
request.model.model_digest = "sha256:" + "f" * 64
failed = _failure(_run_direct(make_context(), request))
assert failed.stable_code == codes.INPUT_MODEL_DIGEST_MISMATCH
def test_non_ready_model_is_rejected(tmp_path):
installed = ModelInfo(
catalog_model_id=READY_MODEL.catalog_model_id,
model_version=READY_MODEL.model_version,
model_digest=READY_MODEL.model_digest,
precisions=READY_MODEL.precisions,
features=READY_MODEL.features,
state=STATE_INSTALLED,
)
context = make_context(inventory=FakeInventory(models=[installed]))
failed = _failure(_run_direct(context, make_execute_request(tmp_path)))
assert failed.stable_code == codes.INPUT_MODEL_NOT_READY
def test_unknown_model_is_rejected(tmp_path):
request = make_execute_request(tmp_path)
request.model.catalog_model_id = "who-dis"
failed = _failure(_run_direct(make_context(), request))
assert failed.stable_code == codes.INPUT_MODEL_UNKNOWN
def test_unknown_and_out_of_range_parameters_are_rejected(tmp_path):
unknown = make_execute_request(
tmp_path,
parameters={"exfiltrate": pb2.ParameterValue(string_value="x")},
)
assert _failure(_run_direct(make_context(), unknown)).stable_code == (
codes.INPUT_PARAMETER_UNKNOWN
)
out_of_range = make_execute_request(
tmp_path,
attempt_id="attempt-2",
parameters={"speed": pb2.ParameterValue(number_value=99.0)},
)
assert _failure(_run_direct(make_context(), out_of_range)).stable_code == (
codes.INPUT_PARAMETER_RANGE
)
def test_input_checksum_mismatch_is_rejected(tmp_path):
request = make_execute_request(tmp_path, input_sha256="0" * 64)
failed = _failure(_run_direct(make_context(), request))
assert failed.stable_code == codes.INPUT_CHECKSUM_MISMATCH
def test_empty_text_is_rejected(tmp_path):
request = make_execute_request(tmp_path, text=" ")
failed = _failure(_run_direct(make_context(), request))
assert failed.stable_code == codes.INPUT_TEXT_EMPTY
def test_unwritable_output_directory_is_local_storage(tmp_path):
locked = tmp_path / "locked"
locked.mkdir()
request = make_execute_request(tmp_path, output_handle=str(locked / "out.wav"))
locked.chmod(0o500)
try:
failed = _failure(_run_direct(make_context(), request))
finally:
locked.chmod(0o700)
assert failed.stable_code == codes.STORAGE_WRITE_FAILED
assert failed.failure_class == pb2.RUNTIME_FAILURE_CLASS_LOCAL_STORAGE
def test_duplicate_attempt_id_is_rejected(tmp_path):
context = make_context()
executor = context.executor()
first = make_execute_request(tmp_path)
assert terminal_of(list(executor.execute(first, None)))[0] == "completed"
duplicate = make_execute_request(tmp_path)
events = list(executor.execute(duplicate, None))
failed = _failure(events)
assert failed.stable_code == codes.INPUT_ATTEMPT_DUPLICATE
def test_slot_exhaustion_is_gpu_resource(tmp_path):
engine = SlowEngine(seconds=30)
context = make_context(engine=engine, slot_limit=1)
executor = context.executor()
hog = make_execute_request(tmp_path, attempt_id="hog")
hog_events = []
hog_thread = threading.Thread(
target=lambda: hog_events.extend(executor.execute(hog, None)), daemon=True
)
hog_thread.start()
assert engine.started.wait(5)
try:
crowded = make_execute_request(tmp_path, attempt_id="crowded")
failed = _failure(list(executor.execute(crowded, None)))
assert failed.stable_code == codes.GPU_SLOTS_EXHAUSTED
assert failed.failure_class == pb2.RUNTIME_FAILURE_CLASS_GPU_RESOURCE
finally:
context.registry.cancel("job-1", "hog")
hog_thread.join(timeout=10)
assert terminal_of(hog_events)[0] == "canceled"
def test_safe_detail_never_carries_local_paths(tmp_path):
engine = FailingEngine(RuntimeError(f"failed loading {tmp_path}/weights.bin"))
failed = _failure(
_run_direct(make_context(engine=engine), make_execute_request(tmp_path))
)
assert str(tmp_path) not in failed.safe_detail
assert "<path>" in failed.safe_detail
+40 -15
View File
@@ -460,30 +460,55 @@ class TaskExecutor:
@staticmethod
def _synthesize(backend, text: str, params: dict):
"""Call the engine through the same serial GPU gate local jobs use.
"""Render through the same seeded pipeline as local ``/generate``.
Held against the idle sweep for the duration: a long generation touches
the instance cache once, at the start, so on elapsed time alone it is
indistinguishable from a model nobody wants any more.
Do not reduce this to ``backend.generate()``. The control plane sends
a complete render contract (pinned gallery seed, synthetic reference,
quality controls, chunking, effects); calling the adapter directly
silently turns a selected gallery voice into a fresh random take.
"""
from services import tts_backend # noqa: PLC0415
from api.routers.generation import _run_backend_inference, _run_inference # noqa: PLC0415
kwargs = {
key: params[key]
for key in (
"ref_audio",
"ref_text",
"instruct",
"language",
"duration",
"description",
"speed",
)
if params.get(key) is not None
}
language = params.get("language")
ref_audio = params.get("ref_audio")
ref_text = params.get("ref_text")
instruct = params.get("instruct")
duration = params.get("duration")
num_step = params.get("num_step", 16)
guidance_scale = params.get("guidance_scale", 2.0)
speed = params.get("speed", 1.0)
denoise = params.get("denoise", True)
postprocess_output = params.get("postprocess_output", True)
used_seed = params.get("seed")
effect_preset = params.get("effect_preset", "broadcast")
max_chunk_chars = params.get("max_chunk_chars")
crossfade_ms = params.get("crossfade_ms")
try:
with tts_backend.engine_in_use(backend):
return backend.generate(text, **kwargs)
if isinstance(backend, tts_backend.OmniVoiceBackend):
# The OSS default engine has an extended native surface;
# preserving it is required for a gallery preview and a
# GPU-worker take to share the same voice identity.
return _run_inference(
backend._model, text, language, ref_audio, ref_text,
instruct, duration, num_step, guidance_scale, speed,
params.get("t_shift"), denoise, postprocess_output,
params.get("layer_penalty_factor"),
params.get("position_temperature"),
params.get("class_temperature"), used_seed,
effect_preset, max_chunk_chars, crossfade_ms,
)
return _run_backend_inference(
backend, text, language, ref_audio, ref_text, instruct,
duration, num_step, guidance_scale, speed, denoise,
postprocess_output, used_seed, effect_preset,
max_chunk_chars, crossfade_ms,
)
except Exception as exc:
from worker import errors as worker_errors # noqa: PLC0415
+18 -1
View File
@@ -657,8 +657,25 @@ class WorkerClient:
)
)
return
await self._send(pb.WorkerMessage(accepted=pb.TaskAccepted(ref=assignment.ref)))
# Reserve the slot BEFORE the accept-send await: awaiting yields to
# the event loop, and a concurrently delivered assignment would read
# the un-reserved counter and over-accept past capacity (#1536 — a
# capacity-1 worker accepted a second task on a slow runner). Message
# order on the stream survives the swap: _send enqueues synchronously
# (put_nowait before any suspension), so ACCEPTED is in the outbox
# before this handler ever yields to the just-created _run task.
self._running[key] = asyncio.create_task(self._run(assignment))
try:
await self._send(pb.WorkerMessage(accepted=pb.TaskAccepted(ref=assignment.ref)))
except BaseException:
# BaseException, not Exception: a handler CANCELLED mid-send must
# release the slot too, or the reserved task keeps running work
# the scheduler never saw accepted — and double-executes after
# reassignment. The stream-death case lands here as well.
task = self._running.pop(key, None)
if task is not None:
task.cancel()
raise
async def _run(self, assignment: pb.TaskAssignment) -> None:
key = self._key(assignment.ref)
+4
View File
@@ -57,6 +57,10 @@ REQUIRED_FEATURES = frozenset({
"task_progress_v1",
"task_inputs_v1",
"remote_model_download_v1",
# A generic backend.generate() call accepts the same wire shape but drops
# profile conditioning controls. Require the canonical worker render path
# so an older peer cannot successfully return a different voice.
"remote_tts_render_v1",
})
+1 -1
View File
@@ -15,7 +15,7 @@
},
"frontend": {
"name": "omnivoice-studio",
"version": "0.4.2",
"version": "0.5.0",
"dependencies": {
"@fontsource-variable/inter": "^5.2.8",
"@fontsource-variable/source-serif-4": "^5.2.9",
+106 -30
View File
@@ -12,7 +12,7 @@ env var that exempts trusted callers:
| Gate | Turn on with | Guards | Applies to |
|---|---|---|---|
| **Share PIN** | the in-app Network share toggle | casual LAN-share guests, one session | non-loopback **HTTP** |
| **API key** | `OMNIVOICE_API_KEY` env var on the backend | a durable remote credential | non-loopback **HTTP + WebSocket** |
| **API key** | `OMNIVOICE_API_KEY` env var on the backend | direct clients and first-party session bootstrap | non-loopback **HTTP + WebSocket** |
| **Trusted networks** | `OMNIVOICE_TRUSTED_NETWORKS` env var | *exempts* the two gates above | non-loopback **consumption** routes only |
Loopback traffic (`127.0.0.1`, `::1`, `localhost`) is **never** gated — local
@@ -25,7 +25,9 @@ tools keep working unchanged whichever gate is set.
> desktop-only even with a key (see [Admin routes](#admin-routes-and-server-mode)).
> Both gates can be active at once. The PIN and the API key are independent; when
> both are set, each is checked on the paths it covers.
> both are set, each is checked on the paths it covers. Session exchange validates
> the master key before the PIN gate so the UI can bootstrap safely; ordinary HTTP
> requests still require the PIN afterward, and the UI prompts for it next.
---
@@ -42,14 +44,14 @@ present it. Supply it any one of three ways:
| Where | How |
|---|---|
| Header | `X-VoiceStudio-Pin: <pin>` |
| Header | `X-OmniVoice-Pin: <pin>` |
| Query param | `?pin=<pin>` |
| Cookie | `ov_pin=<pin>` — the backend sets this automatically after the first valid PIN, so browser sessions only prove it once |
```bash
# From another device on the LAN — with the PIN
curl http://<host>:3900/v1/audio/voices \
-H "X-VoiceStudio-Pin: 123456"
-H "X-OmniVoice-Pin: 123456"
```
A missing or wrong PIN returns:
@@ -75,9 +77,11 @@ Notes on the PIN gate (`NetworkAccessMiddleware`, `backend/main.py`):
## API key
The API key is the durable credential for running the backend somewhere and
driving it remotely — a GPU box on your tailnet, a Docker container, a
reverse-proxied host. Set it on the **backend** process:
The API key is the backend's durable root credential for a GPU box, Docker
container, or reverse-proxied host. Direct API clients may send it on each
request. The first-party browser/Tauri UI instead exchanges it once for a
short-lived administrator session and never stores the master. Set it on the
**backend** process:
```bash
# Generate a strong key and start the backend with it
@@ -87,14 +91,16 @@ uv run uvicorn backend.main:app --host 0.0.0.0 --port 3900
```
While `OMNIVOICE_API_KEY` is set, every **non-loopback HTTP and WebSocket**
request must present it (the SPA shell paths below are the only HTTP exception).
Supply it any one of three ways:
request must present an accepted credential. SPA shell paths remain public;
`POST /api/auth/session` passes through the middleware only so its route can
validate the master and perform the one-time exchange. Direct-client
compatibility accepts:
| Where | How |
|---|---|
| Header | `Authorization: Bearer <key>`**preferred**; the one place a key isn't at risk of landing in a log |
| Cookie | `ov_key=<key>`set automatically after the first authenticated HTTP request; the safer fallback for browser WebSockets |
| Query param | `?api_key=<key>`last resort (browser WebSockets can't set headers). **A key in a URL leaks into proxy/access logs and browser history** — prefer the header or cookie |
| Header | `Authorization: Bearer <key>`**preferred** for scripts and SDKs |
| Legacy cookie | `ov_key=<key>`accepted only for compatibility and migrated by the first-party UI; the backend no longer creates it |
| Legacy query param | `?api_key=<key>`compatibility only. **A key in a URL leaks into proxy/access logs and browser history** |
```bash
# Prefer an encrypted transport (Tailscale Serve / TLS) for a real key; plain
@@ -133,6 +139,8 @@ code **1008** (policy violation) instead of a JSON body.
Notes on the API-key gate (`BearerKeyMiddleware`, `backend/main.py`):
- The key is compared in **constant time** and is **never logged**.
- The backend never copies the master into a response cookie. Browser clients
receive only `ov_session`, an opaque, HttpOnly, SameSite=Strict credential.
- The SPA shell paths bypass the gate on **HTTP** so a remote UI can load and
show what's wrong; WebSockets have no such exemption.
- **Plain HTTP is sniffable** — a Bearer key over `http://` on a hostile
@@ -140,6 +148,40 @@ Notes on the API-key gate (`BearerKeyMiddleware`, `backend/main.py`):
anything beyond a fully trusted LAN. See
[docs/remote-gpu.md](remote-gpu.md) for the full remote-backend setup.
### First-party administrator sessions
The bundled UI uses a narrower protocol:
1. `POST /api/auth/session` receives the master in an `Authorization` header
exactly once and selects `{"transport":"cookie"}` for exact same-origin
browsers or `{"transport":"bearer"}` for Tauri/cross-origin clients.
2. Cookie transport returns `204` and sets `ov_session` as HttpOnly,
SameSite=Strict, path `/`, with an eight-hour maximum lifetime. Bearer
transport returns an opaque `ovs_admin_session_…` value which the UI keeps
in **sessionStorage only**, bound to the exact backend base URL. Bearer JSON
responses include both `expires_at` and a bounded `expires_in`; the UI uses
the relative lifetime when available so clock skew between a remote GPU host
and the browser cannot reject a valid session. `expires_at` remains for
backward compatibility with older clients and servers.
3. `DELETE /api/auth/session` revokes the session. Removing or rotating
`OMNIVOICE_API_KEY`, backend restart, explicit logout, and the eight-hour
deadline also invalidate it.
The master is never written to localStorage/sessionStorage, never returned by
the backend, and never placed in a WebSocket URL. Legacy `ov_api_key` browser
storage is deleted before migration waits on the network. All auth responses,
including errors, carry `Cache-Control: no-store`.
Failed session exchanges are limited per client to ten attempts in a rolling
60-second window and then return `429` with `Retry-After`. A correct master key
is always evaluated and clears the failure window, so an attacker cannot lock
an operator out by deliberately exhausting the limit.
Cookie-authenticated mutations require both an exact allowed `Origin` and
`X-VoiceStudio-CSRF: 1`. Side-effectful GET actions additionally require the
browser's `Sec-Fetch-Site: same-origin`. Bearer/header clients are not subject
to the ambient-cookie CSRF check.
---
## Dictation WebSocket
@@ -149,13 +191,22 @@ own inline guard (`backend/api/routers/capture_ws.py`) *in addition to* the
API-key middleware. A non-loopback client reaches it only if it is **either**:
- on a [trusted network](#trusted-networks) (`is_local_host` passes), **or**
- presenting the **API key** — as `Authorization: Bearer <key>`, the `ov_key`
cookie, or `?api_key=<key>` (URL keys leak into logs — prefer the cookie).
- presenting a direct-client **API key** in `Authorization`, or through a
legacy `ov_key`/`?api_key=` transport.
```
ws://gpu-box:3900/ws/transcribe?api_key=<key>
```
That URL form is retained for non-browser compatibility only. The first-party
UI never constructs it. A bearer administrator session first calls
`POST /api/auth/ws-ticket` and puts only the returned `ws_ticket` in the URL.
Tickets are scoped to `/ws/transcribe` or `/ws/events`, expire after 30 seconds,
return the same bounded `expires_in`/`expires_at` pair, and are consumed
atomically at most once. Same-origin UI WebSockets use the
HttpOnly session cookie and must pass exact `Origin` validation; `null`, missing,
and lookalike origins are rejected.
The **share PIN does not authorize dictation** — the PIN gate is HTTP-only, and
the dictation guard checks only the API key (or trusted-network membership). A
LAN guest who has only entered a PIN can use the HTTP API but **not** live
@@ -196,10 +247,11 @@ time, so in production **restart the backend** to apply a change. Default empty
## Admin routes and server mode
Admin routes — `/system/*` (including `set-env`, **RCE-class**),
`/api/settings/*`, engine install/uninstall, media tools, MCP bindings — sit on
a stricter gate (`require_admin`, `backend/api/dependencies.py`) than
consumption. On the desktop build they are **true-loopback-only**: no PIN, key,
or trusted network reaches them from another machine.
`/api/settings/*`, engine selection/install/uninstall, media tools, MCP
bindings, pronunciation settings, and remote-worker management — sit on a
stricter gate (`require_admin`, `backend/api/dependencies.py`) than consumption.
On the desktop build they are **true-loopback-only**: no PIN, key, or trusted
network reaches them from another machine.
In **server mode** (`OMNIVOICE_SERVER_MODE=1`, the Docker image) the loopback
origin is unenforceable — NAT rewrites the source and even a
@@ -207,16 +259,23 @@ origin is unenforceable — NAT rewrites the source and even a
requirement is dropped (issue #261, else the operator is 403'd out of their own
`/system/*`). It is replaced by a **credential rule**, not removed:
- **No API key configured** → read-only admin discovery remains available for
the bare Docker bootstrap flow, but `POST`/`PUT`/`PATCH`/`DELETE` requests are
denied. Set `OMNIVOICE_API_KEY` before changing settings remotely.
- **A credential is configured** → admin requires the **API key** (`Authorization:
Bearer` / `?api_key` / `ov_key` cookie), or genuine loopback. The **6-digit
share PIN does not gate admin** (it is brute-forceable), and trusted-network
membership never does either. A **PIN-only** server-mode deployment therefore
allows remote read-only discovery but blocks remote mutations; remote writes
require the long API key. Discovery never returns the share PIN itself; only
loopback or a caller already authenticated with the API key can read it.
- **No credential configured** (neither API key nor share PIN) → read-only
admin discovery remains available for the bare Docker bootstrap flow, but
`POST`/`PUT`/`PATCH`/`DELETE` requests are denied. Side-effectful GET actions
are denied too: engine health may start a sidecar, deep diagnostics may load
a model, and LLM provider discovery makes a request with the saved provider
credential. Set `OMNIVOICE_API_KEY` before changing settings or triggering
those actions remotely.
- **An API key is configured** → admin requires that **API key** (direct-client
`Authorization` / legacy query or cookie), a valid short-lived administrator
session, or genuine loopback. The **6-digit share PIN does not gate admin**
(it is brute-forceable), and trusted-network membership never does either. A
**PIN-only** server-mode deployment therefore keeps admin routes loopback-only;
remote admin starts from the long API key.
Managed sidecar installation remains true-loopback-only even with an API key.
Its installer fetches mutable source and creates an editable environment, so it
must be run directly on that machine until the source supply chain is pinned.
Host paths are never selected through HTTP. The native Tauri process validates
model-cache and export destinations plus custom FFmpeg/FFprobe binaries, writes
@@ -261,13 +320,30 @@ the default list, so restate the loopback/Tauri origins alongside your own. (The
same origin.) If you only moved the Vite dev server's port, set
`OMNIVOICE_UI_PORT` instead and the default list follows it.
CORS wraps both authentication gates: credentialless browser preflights are
answered before PIN/API-key enforcement, and gate-generated `401` responses
retain CORS headers so the UI can read the actual failure and prompt for the
right credential.
TLS-terminating proxies must establish the effective scheme at the ASGI server
boundary. Uvicorn's proxy-header handling trusts loopback by default, which
covers Tailscale Serve; a custom proxy on another address must be listed with
`--forwarded-allow-ips=<proxy-ip>` (and proxy headers must remain enabled).
VoiceStudio deliberately does not trust a raw `X-Forwarded-Proto` header inside
the application: once Uvicorn accepts a trusted proxy, the resolved ASGI scheme
drives exact-Origin checks and the session cookie's `Secure` attribute.
For a public path prefix such as `/studio`, either strip that prefix before
forwarding or configure the ASGI `root_path` to the same value. WebSocket ticket
validation removes only that trusted, configured prefix; it never accepts an
arbitrary path merely because it ends in `/ws/events` or `/ws/transcribe`.
## Status codes
| Code | Meaning | What to do |
|---|---|---|
| **401** | Consumption auth failed — `{"detail": "PIN required"}` or `{"detail": "API key required"}`. | Supply the PIN / key (header, cookie, or query param above). A WebSocket surfaces this as close code **1008**. |
| **403** | Authorization failed: loopback/native access was required, a server-mode mutation lacked the API key, or a native path capability was invalid, expired, or for a different operation. | A PIN cannot grant admin or filesystem access. Run native operations from the desktop app; configure and present the API key for remote server-mode mutations; reopen the native picker if a one-shot capability expired. |
| **429** | **Not an auth failure.** The GPU pool is saturated (admission control) or a model download is rate-limited. Ships with `Retry-After` and `X-VoiceStudio-Retryable: true`. | Back off for `Retry-After` seconds and retry the identical request. |
| **403** | Authorization failed: loopback/native access was required, cookie Origin/CSRF validation failed, a server-mode mutation lacked an admin credential, or a native path capability was invalid/expired. | A PIN cannot grant admin or filesystem access. Re-authenticate the UI; scripts should use the API-key header; run native operations from the desktop app. |
| **429** | A failed administrator-session exchange exceeded its per-client limit, the GPU pool is saturated, or a model download is rate-limited. Ships with `Retry-After`; workload throttles also carry `X-VoiceStudio-Retryable: true`. | Back off for `Retry-After` seconds. For authentication, verify the master before retrying; a correct master is never locked out. |
---
+56
View File
@@ -0,0 +1,56 @@
# Benchmarks
Measured numbers per engine and device — how long a generation actually
takes on real hardware. Every number here is produced by the in-repo
harness, on named hardware, at a named version; nothing is estimated.
## How numbers are measured
```bash
# stop the app first — a running backend holds a model and skews numbers
uv run python scripts/bench_pipeline.py # everything
uv run python scripts/bench_pipeline.py tts # just the TTS stage
```
`scripts/bench_pipeline.py` profiles each pipeline stage one at a time,
memory-safely: it refuses to start a stage without enough free RAM and
unloads models between stages. See [performance.md](performance.md) for
what each stage spends its time on.
The `tts` stage emits the two values this table collects:
- **RTF** (real-time factor) — seconds of compute per second of generated
audio, printed next to each warm measurement. RTF < 1 means faster than
real time. Use the **short line (warm)** RTF for the table.
- **Peak VRAM** — printed on CUDA only. MPS is unified memory and CPU has
no VRAM; subprocess-isolated engines allocate outside the harness's view
(it prints `n/a` for them). Leave the column blank in all those cases.
## Results
No verified rows yet — this table fills from maintainer runs and community
submissions.
| Engine | Device | RTF (warm) | Peak VRAM (GB) | App version | Source |
|---|---|---|---|---|---|
| _none yet — contribute yours below_ | | | | | |
Column meanings: **Engine** — the TTS engine the harness resolved (printed
at stage start). **Device** — one string naming what ran the model, e.g.
`RTX 3060 12 GB`, `Apple M2 Pro`, `Ryzen 7 5800X (CPU)`. **RTF (warm)**
the short-line warm RTF from the harness. **Peak VRAM** — the harness's
CUDA peak, blank on MPS/CPU. **App version** — from `Settings → About`.
**Source** — a link to the PR that added the row.
## Contributing a row
1. Run the harness on an otherwise-idle machine (app stopped) and copy its
summary table.
2. Open a PR adding one row using the column meanings above, and paste the
raw harness output into the PR description — that PR link becomes the
row's **Source**.
3. One row per engine+device pair; a newer app version replaces the old row.
Numbers from different machines aren't directly comparable — that's fine.
The point is honest expectations ("this engine on this class of GPU ≈ this
fast"), not a leaderboard.
+61
View File
@@ -0,0 +1,61 @@
# Engine guides
One page per engine: what it's for, what it needs, how to enable it, and its
quirks. Select engines in **Model Catalogue → Engines** (or quick-switch with
<kbd>Ctrl</kbd>/<kbd>Cmd</kbd>+<kbd>E</kbd>), or pin one with
`OMNIVOICE_TTS_BACKEND` / `OMNIVOICE_ASR_BACKEND`.
The compute device (CUDA/ROCm/MPS/CPU) is auto-detected; pin it under
**Settings → Performance & Device** (or `OMNIVOICE_DEVICE`) if auto-detect
picks wrong — see [performance](../performance.md).
Measured speed/VRAM numbers live in [benchmarks](../benchmarks.md); what each
engine can do expressively in [expressive-speech](../expressive-speech.md);
sidecar disk footprints in [disk-usage](disk-usage.md); the bar a new engine
must clear in [engine-acceptance](../engine-acceptance.md).
New to VoiceStudio? Install the app first — [macOS](../install/macos.md)
(first launch needs the one-time right-click → **Open** Gatekeeper
approval), [Windows](../install/windows.md), [Linux](../install/linux.md),
[Docker](../install/docker.md).
## Text-to-speech
| Engine | Guide | Runs on | Cloning | Enabled by |
|---|---|---|---|---|
| VoiceStudio (OmniVoice) — **default** | [omnivoice](omnivoice.md) | CUDA · MPS · CPU | ✅ | installed by default |
| VoxCPM2 | [voxcpm2](voxcpm2.md) | CUDA · MPS · CPU | ✅ + voice design | `pip install "voxcpm>=2.0.3"` |
| MOSS-TTS-Nano | [moss-tts-nano](moss-tts-nano.md) | CUDA · CPU | ✅ (ref only) | clone + `uv pip install -e .` |
| KittenTTS | [kittentts](kittentts.md) | CPU | — (8 preset voices) | `pip install kittentts` |
| MLX-Audio (Kokoro, CSM, Dia, …) | [mlx-audio](mlx-audio.md) | Apple Silicon | model-dependent | `pip install mlx-audio` |
| CosyVoice 3 | [cosyvoice](cosyvoice.md) | CUDA · CPU | ✅ | clone + requirements |
| GPT-SoVITS | [gpt-sovits](gpt-sovits.md) | external server | ✅ | its own API server |
| Sherpa-ONNX | [sherpa-onnx](sherpa-onnx.md) | CUDA · CPU | — | `pip install sherpa-onnx` + model dir |
| IndexTTS 2.5 | [indextts](indextts.md) | CUDA · CPU | ✅ + emotion | one-click sidecar install |
| OmniVoice GGUF | [omnivoice-gguf](omnivoice-gguf.md) | CUDA · MPS · CPU | ✅ | bundled binary |
| Supertonic-3 | [supertonic3](supertonic3.md) | CPU | — (7 preset voices) | `uv sync --extra supertonic` + license |
| MOSS-TTS-v1.5 (8B) | [moss-tts-v15](moss-tts-v15.md) | CUDA · CPU | ✅ | clone + env var |
| dots.tts (2B) | [dots-tts](dots-tts.md) | CUDA · CPU (not Windows) | ✅ | clone + env var |
| OmniVoice (subprocess) | [omnivoice-subprocess](omnivoice-subprocess.md) | CUDA · MPS · CPU | ✅ | opt-in pick, no install |
| PocketTTS (Kyutai) | [pockettts](pockettts.md) | CPU (not Intel Mac) | ✅ | `uv sync --extra pockettts` + license |
| Confucius4-TTS | [confucius4-tts](confucius4-tts.md) | CUDA · CPU | ✅ | clone + env var |
## Speech-to-text
| Engine | Guide | Runs on | Best at | Enabled by |
|---|---|---|---|---|
| WhisperX | [whisperx](whisperx.md) | CUDA · CPU | dubbing (word timestamps + diarization) | installed by default |
| Faster-Whisper | [faster-whisper](faster-whisper.md) | CUDA · CPU | general transcription | installed by default |
| Faster-Whisper (isolated) | [faster-whisper-isolated](faster-whisper-isolated.md) | CUDA · CPU | unattended batches | opt-in pick |
| MLX Whisper | [mlx-whisper](mlx-whisper.md) | Apple Silicon | Mac default | `pip install mlx-whisper` |
| PyTorch Whisper | [pytorch-whisper](pytorch-whisper.md) | CUDA · MPS · CPU | ROCm hosts | installed by default |
| Parakeet TDT (NeMo) | [nemo-parakeet](nemo-parakeet.md) | CUDA · CPU | 25 languages, fast CPU | separate venv (never the app's) |
| Parakeet TDT (MLX) | [parakeet-mlx](parakeet-mlx.md) | Apple Silicon | dictation, 25 EU languages | default on mac-ARM source installs |
| Moonshine | [moonshine](moonshine.md) | CPU | edge/low-power, no timestamps | `pip install` (see guide) |
| FunASR (SenseVoice) | [funasr](funasr.md) | CUDA · CPU | 50+ languages, inline diarization | `pip install funasr` |
| Sherpa-ONNX dictation | [sherpa-onnx-asr](sherpa-onnx-asr.md) | CPU | live streaming dictation | curated model download |
| OpenAI-compatible (remote) | [openai-compatible-asr](openai-compatible-asr.md) | network | offloading to a server (audio leaves the machine) | Model Catalogue |
Speaker diarization is not an engine registry of its own — the dub pipeline
uses pyannote (HF-gated; see [diarization](../features/diarization.md)) and
FunASR can diarize inline with its `cam++` speaker model.
+61
View File
@@ -0,0 +1,61 @@
# VoiceStudio — Faster-Whisper (Crash-Isolated) Engine
The same CTranslate2 Whisper engine as [faster-whisper](faster-whisper.md),
run in a **separate child process** ("sidecar"). CTranslate2's GPU teardown
can segfault — the endemic faster-whisper crash — and a hung or crashed
transcribe in-process takes the whole backend down with it. Isolated, the
child can crash or be force-killed to reclaim a hung transcribe and its VRAM
while the backend stays up
([#730](https://github.com/debpalash/VoiceStudio/issues/730)).
There is nothing extra to install: the sidecar reuses the app's own venv —
only the process boundary is new.
## Selecting it
- **Model Catalogue → Engines**, ASR tab → **Use** on the crash-isolated row, or
- pin it with `OMNIVOICE_ASR_BACKEND=faster-whisper-isolated`.
It is never picked by auto-detect — it's an explicit opt-in escape hatch.
## Best at
- **Long batch runs** where one bad file must not kill the backend.
- Machines where in-process faster-whisper has crashed or hung before:
a sidecar crash fails only that job, and the next transcribe respawns a
fresh sidecar automatically.
## Platform support
Same as faster-whisper: CUDA float16 or CPU int8 on macOS, Windows, and
Linux. The sidecar picks cuda/cpu itself and walks the same
float16 → int8_float16 → int8 degrade chain on GPUs without efficient fp16
([#551](https://github.com/debpalash/VoiceStudio/issues/551)).
## Model selection
- `ASR_MODEL_FASTER` — the shared model selection, same as the in-process
engine: set it once and both variants load the same weights.
- `ASR_MODEL_FW` — optional sidecar-only override; when set it wins over
`ASR_MODEL_FASTER` for this engine. Default `large-v3`.
- `ASR_COMPUTE_TYPE` — optional: pin the sidecar to one CTranslate2 compute
type instead of the automatic degrade chain.
Weights download on first load — see
[downloading-models](../downloading-models.md).
## Trade-offs and quirks
- **Slightly slower per call** than in-process faster-whisper (IPC overhead);
the model stays warm inside the sidecar between calls, so the cost is per
request, not per chunk of audio.
- Word timestamps are Whisper-native (±100300 ms) — no forced alignment.
For dubbing lip-sync, use [whisperx](whisperx.md) or
[mlx-whisper](mlx-whisper.md).
- If the sidecar dies mid-transcription the job fails with a clear
"sidecar crashed" error and the backend stays up — retry to respawn.
- **cuDNN 8 is still required on CUDA** — same CTranslate2 requirement as the
in-process engine. It's checked up front so a missing cuDNN 8 shows as
"unavailable" in Model Catalogue → Engines instead of a sidecar that
silently fails every transcribe
([#1371](https://github.com/debpalash/VoiceStudio/issues/1371)).
+70
View File
@@ -0,0 +1,70 @@
# VoiceStudio — Faster-Whisper Engine
Faster-Whisper runs Whisper on CTranslate2 — the same transcription core
WhisperX uses, **without** the wav2vec2 forced-alignment pass. It's the safe
cross-platform fallback when whisperx isn't installed, and the capture/dictation
fallback on non-Apple machines.
## Selecting it
- **Model Catalogue → Engines**, ASR tab → **Use** on the Faster-Whisper row, or
- pin it with `OMNIVOICE_ASR_BACKEND=faster-whisper`.
Auto-detect only picks it when [whisperx](whisperx.md) is unavailable.
## Best at
- **Subtitles, dictation buffers, and batch transcription** where Whisper's
native word timing (±100300 ms) is good enough.
- For dubbing lip-sync, prefer [whisperx](whisperx.md) (or
[mlx-whisper](mlx-whisper.md) on Apple Silicon) — their forced alignment is
an order of magnitude tighter on word boundaries.
## Platform support
- **CUDA** — float16, with automatic degradation (below).
- **CPU** — int8 on macOS, Windows, and Linux.
- **Apple Silicon GPU / ROCm** — not supported: CTranslate2 has no Metal or
HIP build, so those hosts run on CPU
([#1529](https://github.com/debpalash/VoiceStudio/issues/1529)); auto-detect
routes them to mlx-whisper / pytorch-whisper instead.
## Model selection
`ASR_MODEL_FASTER` — default `Systran/faster-whisper-large-v3`. Accepts the
size aliases (`tiny``large-v3`, `distil-large-v3`) or any CTranslate2
Whisper repo on HF. Weights download on first load — see
[downloading-models](../downloading-models.md).
Segments are cleaned up by faster-whisper's built-in Silero VAD before
transcription.
## Degradation chains
- GPUs without efficient fp16 (older Maxwell/Pascal, GTX 16xx, or a
CTranslate2/cuDNN mismatch) fail at model construction with a compute-type
error; the engine walks float16 → int8_float16 → int8 instead of failing
every chunk ([#551](https://github.com/debpalash/VoiceStudio/issues/551)).
- A CUDA out-of-memory falls back to CPU (slower, same model and accuracy) —
flushing the resident TTS model frees VRAM for GPU-speed ASR
([#255](https://github.com/debpalash/VoiceStudio/issues/255)).
## Quirks
- **cuDNN 8 required on CUDA** — a missing cuDNN 8 would fast-fail the whole
process, so the engine checks up front and reports itself unavailable
instead ([#1371](https://github.com/debpalash/VoiceStudio/issues/1371)).
pytorch-whisper covers that case on torch's bundled cuDNN 9.
- On some hardened Linux kernels the CTranslate2 native library is rejected
with "cannot enable executable stack" (an OSError, not an ImportError) —
reported as unavailable rather than crashing engine selection
([#692](https://github.com/debpalash/VoiceStudio/issues/692)).
- CTranslate2's GPU teardown can rarely segfault the process at unload. If
you hit that, switch to the crash-isolated variant —
[faster-whisper-isolated](faster-whisper-isolated.md)
([#730](https://github.com/debpalash/VoiceStudio/issues/730)).
- Transcribes are time-bounded: `OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S`
(default 120 s per dub chunk) and `OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S`
(default 300 s whole-file).
Speed comparisons across engines live in [performance](../performance.md).
+62
View File
@@ -0,0 +1,62 @@
# VoiceStudio — FunASR (SenseVoice) Engine
FunASR drives Alibaba's SenseVoiceSmall with FSMN-VAD: an all-in-one
multilingual pipeline — transcription with punctuation and inverse text
normalization across **50+ languages**, plus optional **inline speaker
diarization** via the cam++ speaker model. It's the opt-in alternative to
WhisperX ([#182](https://github.com/debpalash/VoiceStudio/issues/182));
WhisperX remains the cross-platform default.
## Selecting it
- Install it into the app venv: `uv pip install funasr`.
- Then **Model Catalogue → Engines**, ASR tab → **Use** on the FunASR row, or
`OMNIVOICE_ASR_BACKEND=funasr`.
Auto-detect never picks it; it's an explicit opt-in.
## Best at
- **Multi-speaker transcription without any HuggingFace token.** This is the
only ASR engine with diarization built in: cam++ labels each sentence
(`Speaker 1`, `Speaker 2`, ...) in the same pass — no gated pyannote
model, no license click-through. Compare
[diarization](../features/diarization.md) for the pyannote/WhisperX route
and what each buys you.
- **Broad language coverage** beyond Whisper's strongest languages, with
punctuation included.
## Not suited for
- **Lip-sync dubbing** — FunASR returns sentence-level timestamps, not
word-level ones. Use [whisperx](whisperx.md) /
[mlx-whisper](mlx-whisper.md) when word timing matters.
## Platform support
CUDA or CPU, on macOS, Windows, and Linux.
## Model selection
| Variable | Default | Role |
| --- | --- | --- |
| `ASR_MODEL_FUNASR` | `iic/SenseVoiceSmall` | main ASR model |
| `ASR_FUNASR_VAD` | `fsmn-vad` | VAD segmentation model |
| `ASR_FUNASR_SPK` | `cam++` | speaker model; set to empty (`ASR_FUNASR_SPK=`) to disable diarization and use the dub pipeline's pyannote/heuristic path instead |
Weights download on first load (through FunASR's own model hub) — see
[downloading-models](../downloading-models.md).
## Quirks
- With the speaker model enabled, long recordings are transcribed in **one
call** and split by FunASR's internal VAD — cam++ assigns speaker cluster
IDs per call, so this is what keeps "Speaker 1" meaning the same person
across the whole file.
- The engine runs with `spk_mode="vad_segment"`: FunASR 1.3.1's default
(`punc_segment`) requires a separate punctuation model and crashes when
SenseVoice is loaded without one.
- SenseVoice's rich-token markup (language/emotion/event tags around the
text) is stripped from the output automatically.
- Language detection is automatic (`language: auto`); the detected language
is reported per file.
+78
View File
@@ -0,0 +1,78 @@
# VoiceStudio — GPT-SoVITS Engine
GPT-SoVITS (RVC-Boss) is one of the most popular open-source voice-cloning
systems (57k+ GitHub stars, MIT-licensed). It does zero-shot and few-shot
cloning with excellent naturalness in Chinese, English, Japanese, Cantonese,
and Korean, and it is very fast (RTF ~0.014 on suitable hardware).
Unlike VoiceStudio's other engines, GPT-SoVITS does not run inside the app.
It ships as a standalone API server, and VoiceStudio connects to it over
HTTP.
## When to pick it
- You already run (or want to run) a GPT-SoVITS server, e.g. with few-shot
fine-tuned voices.
- You need fast, natural cloning in zh/en/ja/yue/ko.
## Setup
1. Install and start the GPT-SoVITS API server (upstream project):
```bash
cd GPT-SoVITS
python api_v2.py -a 127.0.0.1 -p 9880 -c GPT_SoVITS/configs/tts_infer.yaml
```
2. Select the engine via **Model Catalogue → Engines** or
`OMNIVOICE_TTS_BACKEND=gpt-sovits`.
VoiceStudio marks the engine available only when the server responds
(2-second reachability probe).
## Configuration
| Variable | Default | Meaning |
| --- | --- | --- |
| `OMNIVOICE_GPTSOVITS_URL` | `http://127.0.0.1:9880` | API server URL |
| `OMNIVOICE_TRUSTED_NETWORKS` | (unset) | Required to allow a non-loopback server |
**Remote servers:** by default VoiceStudio only talks to loopback addresses
— part of the local-first guarantee. To point at a server on another
machine (e.g. a GPU box on your LAN), add its network to
`OMNIVOICE_TRUSTED_NETWORKS`; otherwise the connection is refused as an
untrusted endpoint.
Prefer `https://` (or a private tunnel such as Tailscale/WireGuard) for any
non-loopback server: with plain `http://` the text you synthesize and the
audio that comes back cross the network unencrypted. VoiceStudio does not
disable certificate verification, so a TLS endpoint needs a certificate the
system trusts.
## Behaviour notes
- Output is 32 kHz mono (server output is resampled if needed).
- Cloning passes your reference clip path and optional transcript to the
server; the reference path must be readable **by the server process**, so
remote servers need the clip on their own filesystem.
- Speed control is forwarded as the server's `speed_factor`.
- The GPU is whatever the GPT-SoVITS server itself uses (CUDA preferred);
VoiceStudio's side is just an HTTP client.
## Known limits
- Five languages only; for broader coverage use
[OmniVoice](omnivoice.md) ([languages.md](../languages.md)).
- No voice design; server availability is your responsibility — if the
server stops, generations fail with a "server not reachable" error.
## Troubleshooting
- "GPT-SoVITS server not reachable": start the server with the command
above, or fix `OMNIVOICE_GPTSOVITS_URL`.
- "endpoint is outside loopback or OMNIVOICE_TRUSTED_NETWORKS": see
Configuration above.
- Other issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [benchmarks.md](../benchmarks.md),
[expressive-speech.md](../expressive-speech.md).
+75
View File
@@ -0,0 +1,75 @@
# VoiceStudio — KittenTTS Engine
KittenTTS (KittenML) is the lightweight English "flash" tier: a 2580 MB
ONNX model with 8 preset voices that runs realtime on any CPU — no torch, no
CUDA, no GPU of any kind. Use it when you just need quick English narration
(voiceovers, demo reads, short phrases) with no reference sample.
## When to pick it
- English-only content where speed and a tiny install matter more than
cloning.
- Machines with no usable GPU.
The trade-off against [OmniVoice](omnivoice.md): no voice cloning, English
only — but a much faster and much smaller install.
## Setup
```bash
pip install kittentts
```
Then select the engine via **Model Catalogue → Engines** or
`OMNIVOICE_TTS_BACKEND=kittentts`.
## Voices
Eight preset voices, four male/female pairs:
```text
expr-voice-2-m expr-voice-2-f (default: expr-voice-2-f)
expr-voice-3-m expr-voice-3-f
expr-voice-4-m expr-voice-4-f
expr-voice-5-m expr-voice-5-f
```
An unknown voice id logs an info message and falls back to the default.
## Model selection
| Variable | Default | Meaning |
| --- | --- | --- |
| `OMNIVOICE_KITTENTTS_MODEL` | `KittenML/kitten-tts-mini-0.8` | HuggingFace checkpoint to load |
The ~80 MB model downloads from HuggingFace on first use (retried once on a
flaky connection). See [downloading-models.md](../downloading-models.md).
## Behaviour notes
- Output is 24 kHz mono.
- CPU-only by design — the ONNX graph has no CUDA/MPS path.
- Non-English `language` values are ignored with a log line pointing at
OmniVoice; reference audio is likewise ignored (no cloning).
- **Long-input hardening
([#1173](https://github.com/debpalash/VoiceStudio/issues/1173)):** the
shipped ONNX graph has a hard 512-token cap, and phonemization can expand
text massively (digits especially). VoiceStudio pre-measures every chunk
with the model's own tokenizer and splits oversized chunks at word
boundaries, so long or digit-heavy inputs no longer abort inside
onnxruntime with an opaque "invalid expand shape" error.
## Known limits
- English only; no cloning, no voice design, no emotion controls
(see [expressive-speech.md](../expressive-speech.md)).
- Preset voices only — speed is the one knob.
## Troubleshooting
- Engine unavailable: `pip install kittentts` into VoiceStudio's Python
environment and restart.
- Other issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [benchmarks.md](../benchmarks.md),
[disk usage](disk-usage.md).
+76
View File
@@ -0,0 +1,76 @@
# VoiceStudio — MLX-Audio Engine (Apple Silicon)
MLX-Audio (Blaizzy/mlx-audio) wraps 14+ TTS engines — Kokoro, CSM, Dia,
Qwen3-TTS, Chatterbox, MeloTTS, OuteTTS, and more — behind a single adapter
that runs on Apple's MLX framework. It is **Apple Silicon only**: the engine
is not shipped on Linux, Windows, or Intel Macs, and a stray wheel on those
platforms never reports as available
([#390](https://github.com/debpalash/VoiceStudio/issues/390)).
## When to pick it
- You're on an M-series Mac and want small, fast models tuned for it.
- You want one of the specific hosted models (Kokoro for small multilingual,
CSM for cloning, Qwen3-TTS for voice design, Dia for dialogue, …).
## Setup
```bash
pip install mlx-audio
```
Then select the engine via **Model Catalogue → Engines** or
`OMNIVOICE_TTS_BACKEND=mlx-audio`.
## Model selection
One backend hosts many models. The curated set:
| Key | Model | Niche |
| --- | --- | --- |
| `kokoro` (default) | `mlx-community/Kokoro-82M-bf16` | small multilingual |
| `csm` | `mlx-community/csm-1b-8bit` | voice cloning |
| `qwen3-tts` | `mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit` | voice design |
| `dia` | `mlx-community/Dia-1.6B` | dialogue |
| `chatterbox` | `mlx-community/Chatterbox-TTS-4bit` | expressive |
| `melotts` | `mlx-community/MeloTTS-English-v3-MLX` | lightweight VITS |
| `outetts` | `mlx-community/Llama-OuteTTS-1.0-1B-4bit` | LM-based |
Pick a model in the **Model Catalogue → Engines** curated picker
([#981](https://github.com/debpalash/VoiceStudio/issues/981)) or set
`OMNIVOICE_MLX_AUDIO_MODEL` to either a curated key (`kokoro`) or any full
HF repo id. The env var overrides the persisted UI choice.
## Behaviour notes
- Output is 24 kHz mono for most hosted models.
- **Cloning works only with the `csm` model** — it is the only curated model
confirmed to accept a reference clip. Other models silently ignore
reference audio, so the engine reports cloning support only when CSM is
selected (dub/batch jobs gate on this).
- Voice design (text description → voice) is available through the
Qwen3-TTS VoiceDesign model.
- Language support is per-model (Kokoro ~8 languages, others vary). An
unsupported language for Kokoro produces a clear error naming what it
does support ([#977](https://github.com/debpalash/VoiceStudio/issues/977))
— leave language on Auto or switch to a multilingual engine.
## Platform notes
This engine is exempt from cross-platform parity as a platform-only
capability behind explicit opt-in: it exists only where Apple's MLX runtime
exists. On any other platform the engine picker shows it unavailable with
the reason.
## Troubleshooting
- Unavailable on an M-series Mac: `pip install mlx-audio` into
VoiceStudio's Python environment; in a packaged app build, MLX's native
libraries may fail to load — the engine reports unavailable rather than
crashing.
- Other issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [benchmarks.md](../benchmarks.md),
[languages.md](../languages.md),
[downloading-models.md](../downloading-models.md),
[disk usage](disk-usage.md).
+57
View File
@@ -0,0 +1,57 @@
# VoiceStudio — MLX Whisper Engine
MLX Whisper runs Whisper on the Apple Silicon GPU via MLX. It exists because
CTranslate2 (whisperx / faster-whisper) has **no Metal build** — on a Mac
those engines transcribe on the CPU no matter what GPU is present. Measured
on an M2 with whisper-large-v3, one 30 s dub chunk: **90.4 s on WhisperX
(CPU) vs 20.5 s on MLX (GPU)** — which is why auto-detect picks MLX Whisper
on every Apple Silicon machine
([#1127](https://github.com/debpalash/VoiceStudio/issues/1127)).
## Selecting it
- Nothing to do on Apple Silicon — auto-detect prefers it there.
- Or explicitly: **Model Catalogue → Engines**, ASR tab → **Use**, or
`OMNIVOICE_ASR_BACKEND=mlx-whisper`.
## Best at
- **Dubbing on a Mac** — it layers the same wav2vec2 forced alignment
WhisperX uses on top of the GPU transcription, so word timing (±1030 ms)
and therefore lip-sync accuracy are unchanged. Same model, same alignment,
~4x the speed.
- **Dictation/capture** — the capture path automatically swaps in
`mlx-community/whisper-large-v3-turbo` (~5x faster than large-v3) unless a
sherpa dictation model or [parakeet-mlx](parakeet-mlx.md) is preferred.
## Platform support
**Apple Silicon only.** A shared platform gate refuses Linux, Windows, and
Intel Macs before any package import, so a stray `mlx-whisper` wheel on the
wrong platform never reports itself available
([#390](https://github.com/debpalash/VoiceStudio/issues/390)). All other
platforms use the CUDA/CPU engines instead.
## Model selection
- `ASR_MODEL` — default `mlx-community/whisper-large-v3-mlx`. Any MLX-format
Whisper repo works. Weights download on first load — see
[downloading-models](../downloading-models.md).
- `OMNIVOICE_ALIGN_DEVICE` — force the wav2vec2 aligner's device. The aligner
runs on MPS when it can and falls back to CPU; languages without a bundled
aligner (~20 major languages have one) keep Whisper's native word
timestamps.
## Quirks
- Audio is decoded through VoiceStudio's validated ffmpeg rather than the
bare `ffmpeg` PATH lookup mlx-whisper would do on its own — a clean
from-source install with no system ffmpeg works fine
([#479](https://github.com/debpalash/VoiceStudio/issues/479)).
- The model is warmed into unified memory in the background, so the first
transcribe after startup doesn't pay the load cost.
- In a packaged app, a native MLX library that fails to load is reported as
"unavailable" (with fallback to another engine) rather than crashing the
engine list.
Speed comparisons across engines live in [performance](../performance.md).
+48
View File
@@ -0,0 +1,48 @@
# VoiceStudio — Moonshine Engine
Moonshine is an edge-optimized ASR family built for CPU-only machines.
Unlike Whisper it processes variable-length audio (no padding everything to
30 s), which keeps latency low on short clips — sub-200 ms class on capture
buffers. It's the lightest local option for quick transcription on hardware
where even int8 whisper-large is too slow.
## Selecting it
- Install one of the runtimes into the app venv:
`uv pip install moonshine-onnx` (lighter, tried first) or
`moonshine-voice`.
- Then **Model Catalogue → Engines**, ASR tab → **Use** on the Moonshine row,
or `OMNIVOICE_ASR_BACKEND=moonshine`.
Auto-detect never picks it; it's an explicit opt-in.
## Best at
- **Quick notes and short-clip transcription on low-power CPU machines.**
- Environments where a sub-1 GB footprint matters more than word timing or
language coverage.
## Not suited for
- **Dubbing.** Output is plain text as a **single segment spanning the whole
file — no word or segment timestamps** — so there's nothing for lip-sync
or subtitle timing to work with. Use a Whisper-family engine or
[sherpa-onnx-asr](sherpa-onnx-asr.md) for those jobs.
- Multilingual work: results report English; for broad language coverage use
[whisperx](whisperx.md) or [funasr](funasr.md).
## Platform support
CPU only, by design — macOS, Windows, and Linux. It claims no GPU.
## Model selection
`ASR_MODEL_MOONSHINE` — default `moonshine/base`. Weights download on first
load — see [downloading-models](../downloading-models.md).
## Quirks
- The engine tries `moonshine_onnx` first and falls back to
`moonshine_voice` — installing either one is enough.
- Segment bounds are synthesized from the audio duration (start 0, end =
file length), since the model reports none.
+78
View File
@@ -0,0 +1,78 @@
# VoiceStudio — MOSS-TTS-Nano Engine
MOSS-TTS-Nano (OpenMOSS) is the low-resource, broad-language pick: a
100M-parameter autoregressive codec LM that runs realtime on a 4-core CPU —
no GPU required — with native 48 kHz output and 20 languages under an
Apache-2.0 license. It fills the "runs on a fanless laptop" tier while still
covering languages like Arabic, Hebrew, Persian, Korean, and Turkish.
## When to pick it
- CPU-only or low-power hardware, but you still need cloning and non-English
coverage.
- Your language is among: Chinese, English, German, Spanish, French,
Japanese, Italian, Hebrew, Korean, Russian, Persian, Arabic, Polish,
Portuguese, Czech, Danish, Swedish, Hungarian, Greek, Turkish.
## Setup
The package is **not on PyPI** — install it from the upstream repo into
VoiceStudio's Python environment:
```bash
git clone https://github.com/OpenMOSS/MOSS-TTS-Nano.git
cd MOSS-TTS-Nano
uv pip install -e .
```
Then select the engine via **Model Catalogue → Engines** or
`OMNIVOICE_TTS_BACKEND=moss-tts-nano`.
## Model selection
| Variable | Default | Meaning |
| --- | --- | --- |
| `OMNIVOICE_MOSS_TTS_MODEL` | `OpenMOSS-Team/MOSS-TTS-Nano` | HuggingFace checkpoint to load |
The first use downloads the weights (retried once on a truncated download).
See [downloading-models.md](../downloading-models.md).
## Behaviour notes
- **Cloning is reference-only**: pass a reference clip. Style instructions,
preset speakers, and speed control are not supported and are silently
ignored, so mixed-engine call sites keep working.
- The model emits 48 kHz stereo; VoiceStudio downmixes to mono, matching the
rest of the pipeline (the dub mixer treats TTS output as mono per
segment).
- Runs on CPU or CUDA.
## Upstream is unpinned
The upstream repo is installed straight from git with no pinned release, and
the model class it exports has changed before
([#1287](https://github.com/debpalash/VoiceStudio/issues/1287)). VoiceStudio
therefore verifies that a usable model class actually exists — not just that
the package imports — before reporting the engine as ready. If the engine
shows unavailable with a "does not expose a usable model class" message,
pull the latest upstream and re-run `uv pip install -e .`, or open an issue
with the version you have.
## Known limits
- No voice design, no instruct, no speed control — cloning from a reference
clip only.
- Quality sits below the large engines; see
[benchmarks.md](../benchmarks.md).
## Troubleshooting
- "moss_tts_nano package not installed": run the clone + `uv pip install -e .`
steps above.
- Entry-point errors after an upstream update: see "Upstream is unpinned"
above.
- General issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [languages.md](../languages.md),
[expressive-speech.md](../expressive-speech.md),
[disk usage](disk-usage.md).
+58
View File
@@ -0,0 +1,58 @@
# VoiceStudio — Parakeet TDT (NVIDIA NeMo) Engine
NVIDIA's Parakeet TDT via the NeMo toolkit: a FastConformer encoder with a
Token-and-Duration Transducer decoder. It beats Whisper large-v3 on English
benchmarks (~6% WER) and supports **25 (mostly European) languages** with
automatic language detection. The 0.6B model is fast even on CPU — measured
RTF 0.080.23 on an Apple Silicon M2 CPU (2026-07-02), ~20x faster than
faster-whisper large-v3 int8 on the same host.
## Do not install NeMo into the app venv
`nemo_toolkit`'s ASR extras pin `transformers>=4.57,<4.58`, which conflicts
with VoiceStudio's own `transformers>=5.3` requirement and **will break the
backend** (ImportError on startup) if installed into the shared venv. There
is currently no safe in-app install path for this engine; in-app isolation
is tracked separately.
If you want the Parakeet models without a separate environment, use these
instead — same model family, no NeMo dependency:
- **Apple Silicon:** [parakeet-mlx](parakeet-mlx.md) (installed by default on
mac-ARM source installs).
- **Any platform, CPU:** [sherpa-onnx-asr](sherpa-onnx-asr.md) — its default
dictation model is an int8 ONNX export of Parakeet TDT v3.
## Selecting it
Only meaningful if you've set up `nemo_toolkit[asr]` in a **separate,
dedicated Python environment** that runs the backend:
- **Model Catalogue → Engines**, ASR tab → **Use** on the Parakeet TDT row, or
- `OMNIVOICE_ASR_BACKEND=nemo-parakeet`.
Auto-detect never picks it; it's an explicit opt-in.
## Best at
- **English and European-language transcription** where WER matters more
than word-level subtitle timing.
- **CPU-only hosts** — faster than realtime without any GPU.
## Platform support
CUDA or CPU (the old hard CUDA gate was removed — see the RTF numbers
above). Availability is a pure dependency check on `nemo.collections.asr`.
## Model selection
`ASR_MODEL_NEMO` — default `nvidia/parakeet-tdt-0.6b-v3`. Weights download
on first load — see [downloading-models](../downloading-models.md).
## Quirks
- Output is a **single segment** for the whole file (NeMo doesn't VAD-split
like Whisper), with word timestamps when the model exposes them — fine for
dictation and plain transcripts, not ideal for long-form subtitles.
- The detected language isn't exposed cleanly by NeMo, so results report
`en` regardless of the actual (auto-detected) language.
+83
View File
@@ -0,0 +1,83 @@
# VoiceStudio — OmniVoice GGUF Engine
OmniVoice GGUF runs the same OmniVoice model as the [default
engine](omnivoice.md), but through a bundled native binary
(`bin/omnivoice-tts-<platform>`) loading quantized GGUF weights. It is
hardware-adaptive: a probe picks the quantization that fits your machine, so
small GPUs and CPU-only hosts get a working OmniVoice instead of a paging,
timing-out one.
## When to pick it
- Your GPU is below the default engine's 6 GB VRAM floor.
- CPU-only machines that still want OmniVoice's voice and language coverage.
- You want generation isolated in a separate process (a crash or leak never
takes the app down — each generation spawns the binary fresh).
## Quantization selection
Weights come from the `Serveurperso/OmniVoice-GGUF` HuggingFace repo, pinned
to an exact revision. The hardware probe selects:
| Hardware | Quant | Approx. VRAM use |
| --- | --- | --- |
| 12 GB+ VRAM | BF16 | ~1.6 GB (quality-first) |
| 412 GB VRAM | Q8_0 | ~945 MB (recommended balance) |
| 14 GB VRAM | Q4_K_M | ~659 MB (minimal footprint) |
| CPU-only | Q4_K_M | RAM-bound, latency-tolerable |
You can override the selection from Settings; overrides are allow-listed
against the same table (an F32 reference quant, ~3.2 GB, is override-only).
## Setup
Nothing to install: installer and CI builds bundle the binary for your
platform. Select the engine via **Model Catalogue → Engines** or
`OMNIVOICE_TTS_BACKEND=omnivoice-gguf`. The quant weights download on first
use (see [downloading-models.md](../downloading-models.md)) — install them
ahead of time from **Model Catalogue → Models** if you want the first
generation to be quick; a long first render is the download, not a hang.
**Source checkouts:** the repo ships zero-byte placeholders in `bin/` — real
binaries come from CI or the installer. The engine detects a placeholder and
reports unavailable with instructions
([#1172](https://github.com/debpalash/VoiceStudio/issues/1172)) instead of
failing at spawn time; build one with
`scripts/build-omnivoice-tts.sh --platform <slug>` or use the default
in-process engine.
## Integrity and self-healing
Before reporting ready, the engine:
- verifies the binary against the SHA-256 manifest (`bin/checksums.sha256`);
- detects macOS Gatekeeper quarantine and prints the exact
`xattr -cr '/Applications/VoiceStudio.app'` fix;
- restores a missing execute bit (a git clone or zip extract on POSIX can
drop `+x`, which used to surface as a permission error mislabeled as
out-of-memory — [#437](https://github.com/debpalash/VoiceStudio/issues/437)).
The chmod runs only after the SHA check confirms it's the right file.
## Behaviour notes
- Output is 24 kHz mono — same model, same rate as in-process OmniVoice.
- Cloning from a reference clip (with optional transcript) and style
instructions are supported; no voice design.
- Same multilingual surface as OmniVoice ([languages.md](../languages.md)).
- Because generation runs in another process, the app's own GPU counters
don't see its allocations — diagnostics label it accordingly.
| Variable | Default | Meaning |
| --- | --- | --- |
| `OMNIVOICE_GGUF_GENERATE_TIMEOUT_S` | (generous built-in) | Per-generation timeout for the spawned binary |
## Troubleshooting
- "GGUF binary missing": this build doesn't bundle the runtime for your
platform — use the default engine.
- Checksum mismatch or quarantine messages: follow the printed fix, or
reinstall.
- Other issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [benchmarks.md](../benchmarks.md),
[performance.md](../performance.md), [disk usage](disk-usage.md).
+74
View File
@@ -0,0 +1,74 @@
# VoiceStudio — OmniVoice Engine (default)
OmniVoice (k2-fsa/OmniVoice) is VoiceStudio's default TTS engine — the one a
fresh install uses without any configuration. It does zero-shot voice cloning
across 600+ languages and outputs 24 kHz mono audio. Voice cloning, dubbing,
and dictation all run on it out of the box.
## When to pick it
- You want cloning plus the broadest language coverage (see
[languages.md](../languages.md)).
- You have a GPU (CUDA or Apple Silicon MPS) with ~6 GB VRAM or more.
- You just installed VoiceStudio — it's already selected.
For low-VRAM or CPU-only machines, the
[OmniVoice GGUF](omnivoice-gguf.md) variant runs the same model through a
quantized native binary with a much smaller memory footprint.
## Requirements
- Runs on CUDA, MPS (Apple Silicon), or CPU — auto-detected.
- Recommended VRAM floor: **6 GB** on a dedicated GPU. This is the only
engine with a measured floor: on 4 GB cards (GTX 1650 Ti, Quadro P2000 —
issues [#1226](https://github.com/debpalash/VoiceStudio/issues/1226) /
[#1222](https://github.com/debpalash/VoiceStudio/issues/1222)) the driver
pages to system RAM and a render that should take seconds runs for minutes
until the compute budget kills it. The UI warns before you wait; nothing
hard-blocks, since short inputs can still fit.
- No extra install — the model ships with the app and downloads its weights
on first use (see [downloading-models.md](../downloading-models.md)).
## Selecting the engine
OmniVoice is the default, so normally there is nothing to do. If you switched
away and want it back:
- **Model Catalogue → Engines**, or
- set `OMNIVOICE_TTS_BACKEND=omnivoice`.
The env var overrides the persisted UI choice.
## Behaviour notes
- Weights load lazily on first use and are shared with the rest of the app
(dubbing, dictation) — the model is never double-loaded.
- On CUDA the model runs fp16 with `torch.compile`; a speech recognizer is
co-loaded for the cloning path.
- Output is 24 kHz mono; the shared mastering chain (highpass + compressor)
is tuned for this rate and applied automatically.
- Cloning takes a short reference clip (`ref_audio`); an optional transcript
of the clip improves conditioning.
## Known limits
- No voice design from a text description — use [VoxCPM2](voxcpm2.md) for
that.
- Below the 6 GB VRAM floor, expect very slow renders or budget timeouts;
prefer [OmniVoice GGUF](omnivoice-gguf.md) or a CPU engine such as
[PocketTTS](pockettts.md).
## Troubleshooting
- "Too heavy for the available compute" on a small GPU: see the VRAM floor
above — switch to OmniVoice GGUF or close other GPU apps.
- First generation is slow: the first call downloads multi-GB weights. To
keep the first render quick, install the model ahead of time from
**Model Catalogue → Models** — a long first generate is almost always the
download, not a hang.
- General install issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [benchmarks.md](../benchmarks.md),
[performance.md](../performance.md),
[expressive-speech.md](../expressive-speech.md),
[disk usage](disk-usage.md).
+58
View File
@@ -0,0 +1,58 @@
# VoiceStudio — Parakeet TDT v3 (MLX) Engine
NVIDIA's Parakeet TDT v3 on the Apple Silicon GPU, via the small pure-Python
`parakeet-mlx` package. It gives Macs the Parakeet tier CUDA/CPU users get
through NeMo or sherpa-onnx: **25 European languages**, word timestamps from
the TDT decoder itself (no wav2vec2 alignment pass needed), ~1.2 GB download,
~2 GB unified memory, dictation-grade speed on the GPU.
Unlike [nemo-parakeet](nemo-parakeet.md) it needs no `nemo_toolkit` (whose
transformers pin conflicts with the app's) — it is **installed by default on
Apple Silicon source installs since 0.3.22**.
## Selecting it
- **Model Catalogue → Engines**, ASR tab → **Use** on the Parakeet TDT v3
(MLX) row, or `OMNIVOICE_ASR_BACKEND=parakeet-mlx`.
- **Dictation prefers it automatically**: once the model weights are
installed (Model Catalogue → Models — the auto-pick never triggers a
download), live dictation/capture uses it whenever your system language is
one of the 25 covered European languages. Other languages keep the
multilingual Whisper engine, so dictation coverage never regresses.
## Best at
- **Live dictation on a Mac** — TDT decoding is fast enough for the capture
path, at Parakeet's better-than-Whisper English WER.
- **European-language transcription** with word timestamps at a fraction of
whisper-large-v3's memory and compute.
For languages outside the 25 (CJK, Arabic, ...), use
[mlx-whisper](mlx-whisper.md) instead.
## Platform support
**Apple Silicon only** — the same shared MLX platform gate as mlx-whisper
refuses Linux, Windows, and Intel Macs before any import
([#390](https://github.com/debpalash/VoiceStudio/issues/390)). It runs on the
unified-memory GPU; there is no CPU tier.
## Model selection
`ASR_MODEL_PARAKEET_MLX` — default `mlx-community/parakeet-tdt-0.6b-v3`.
Weights download on first load — see
[downloading-models](../downloading-models.md).
## Quirks
- Long files are processed in 120 s chunks internally to bound unified-memory
use; short dictation buffers and dub chunks are unaffected.
- Parakeet v3 auto-detects among its 25 languages but doesn't expose the
pick, so the reported language is the one you requested (or none) — it is
never hardcoded to English.
- Word timestamps are merged from the decoder's subword tokens — good for
subtitles and dictation; for lip-sync-critical dubbing the wav2vec2-aligned
engines ([mlx-whisper](mlx-whisper.md), [whisperx](whisperx.md)) remain the
accuracy tier.
Speed comparisons across engines live in [performance](../performance.md).
+82
View File
@@ -0,0 +1,82 @@
# VoiceStudio — PocketTTS Engine
PocketTTS (kyutai-labs/pocket-tts, 100M parameters) is the fastest-CPU-render
pick: small, low-latency, CPU-only, with zero-shot voice cloning from a
reference clip. It covers six languages — English, French, German,
Portuguese, Italian, Spanish — with one model per language, and measures
roughly 89x real-time on an Apple M3 Pro.
It complements the quality engines: where they fall back to CPU, PocketTTS
is built for it. CPU-only is deliberate — upstream observes no GPU speedup
for this model.
## When to pick it
- CPU-only machines that need fast rendering *and* voice cloning.
- Latency-sensitive use (dictation-style, short utterances) in one of the
six languages.
## Setup
1. Install the optional dependency:
```bash
uv sync --extra pockettts
```
(Or enable it from **Model Catalogue → Engines**.)
2. **Accept the license in-app**
([#1306](https://github.com/debpalash/VoiceStudio/issues/1306)). The code
is MIT and the weights are CC-BY-4.0, but the weights are **gated on
HuggingFace** behind an access agreement with an acceptable-use clause.
VoiceStudio surfaces this before first use: the engine stays unavailable
until you review and accept in **Model Catalogue → Engines → PocketTTS**.
You also need HuggingFace access to the gated repo (see
[downloading-models.md](../downloading-models.md) for token setup).
3. Select the engine via **Model Catalogue → Engines** or
`OMNIVOICE_TTS_BACKEND=pockettts`.
## Platform notes
- Works on Linux, Windows, macOS Apple Silicon — CPU only everywhere.
- **Not available on Intel Macs**: the required PyTorch version has no
macOS x86_64 wheel. The engine reports this plainly instead of failing
mid-install.
## Behaviour notes
- Output is 24 kHz mono.
- Six languages, one model per language, chosen by the `language` you
request; cloning takes a short reference clip.
- Runs in a crash-isolated sidecar process (parent Python environment): a
wedged generation is hard-killed by a watchdog and its memory reclaimed —
something an in-process engine cannot do.
- The first use downloads the gated weights; the sidecar heartbeats
progress during the download so the watchdog doesn't fire.
| Variable | Default | Meaning |
| --- | --- | --- |
| `OMNIVOICE_POCKETTTS_RECV_TIMEOUT_S` | `600` | Sidecar response deadline in seconds (min 30; cold loads download weights) |
## Known limits
- No voice design, no emotion controls
(see [expressive-speech.md](../expressive-speech.md)).
- Six languages only — for broader coverage use
[OmniVoice](omnivoice.md) ([languages.md](../languages.md)).
- Revoking the license acceptance takes effect immediately, without a
restart — subsequent generations refuse.
## Troubleshooting
- "pocket_tts package not installed": run the `uv sync` above.
- "license not accepted": open **Model Catalogue → Engines → PocketTTS**
and review/accept.
- Timeouts on a slow connection: raise
`OMNIVOICE_POCKETTTS_RECV_TIMEOUT_S` for the first (download-heavy) run.
- Other issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [benchmarks.md](../benchmarks.md),
[performance.md](../performance.md), [disk usage](disk-usage.md).
+61
View File
@@ -0,0 +1,61 @@
# VoiceStudio — PyTorch Whisper Engine
Whisper through the plain `transformers` pipeline, riding torch itself. No
extra install — transformers ships with the app — and because it runs on
torch's own stack (including torch's bundled cuDNN 9), it works on machines
where the CTranslate2 engines can't load. It is also the engine that
genuinely uses **AMD ROCm** GPUs, so auto-detect picks it on ROCm hosts
([#1529](https://github.com/debpalash/VoiceStudio/issues/1529)).
## Selecting it
- **Model Catalogue → Engines**, ASR tab → **Use** on the PyTorch Whisper
row, or `OMNIVOICE_ASR_BACKEND=pytorch-whisper`.
- Auto-detect picks it on ROCm, and as the last resort everywhere else.
## Best at
- **ROCm dubbing/transcription** — the only Whisper engine that uses the HIP
GPU (CTranslate2 has no HIP build, MLX is Apple-only).
- **Rescue engine** when whisperx/faster-whisper can't load — e.g. the
missing-cuDNN-8 case
([#255](https://github.com/debpalash/VoiceStudio/issues/255)) — since it
needs neither CTranslate2 nor cuDNN 8.
For lip-sync-grade word timing prefer [whisperx](whisperx.md) or
[mlx-whisper](mlx-whisper.md); this engine returns the pipeline's own word
timestamps.
## Platform support
CUDA, Apple Silicon (MPS), ROCm (HIP), and CPU — wherever torch runs, on
macOS, Windows, and Linux.
## Model selection
`OMNIVOICE_PYTORCH_ASR_MODEL` — default `openai/whisper-large-v3-turbo`. Any
transformers-format Whisper repo works. Weights download on first load — see
[downloading-models](../downloading-models.md).
## VRAM preflight
whisper-large-v3-turbo needs roughly 3.2 GiB before generation adds its
workspace; loading it onto a nearly-full card "succeeds" and then the first
transcribe OOMs with zero segments. So on CUDA the engine checks free VRAM
against a 5 GB budget before loading and uses the CPU instead when the card
is too full (flush the TTS model to restore GPU-speed ASR). Disable with
`OMNIVOICE_ASR_VRAM_PREFLIGHT=0`.
## Quirks
- If the pipeline fails to import (`AutoFeatureExtractor` errors), the cause
is either an incomplete transformers install or a torch/torchvision
version mismatch — the error message names the exact reinstall command;
the trio has to move together at the pinned versions
([#549](https://github.com/debpalash/VoiceStudio/issues/549),
[#1376](https://github.com/debpalash/VoiceStudio/issues/1376)).
- Transcribes are time-bounded like every local engine:
`OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S` (default 120 s per dub chunk),
`OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S` (default 300 s whole-file).
Speed comparisons across engines live in [performance](../performance.md).
+71
View File
@@ -0,0 +1,71 @@
# VoiceStudio — Sherpa-ONNX Dictation Engine
The k2-fsa/sherpa-onnx ONNX runtime as a **live dictation** engine: small
int8 models that transcribe faster than realtime on CPU, with identical
behavior on macOS (arm64 + x86_64), Windows, and Linux — no CUDA dependency.
Streaming models emit partial text frame-by-frame as you speak; offline
models re-transcribe a growing buffer on a short cadence, so you see live
partials either way.
## Selecting it
- Ensure `sherpa-onnx` is installed (`uv add sherpa-onnx` on source installs).
- Pick a dictation model in the app (Model Catalogue → Models lists the
curated set below), or **Model Catalogue → Engines**, ASR tab → **Use**, or
pin `OMNIVOICE_ASR_BACKEND=sherpa-onnx-asr`.
- `OMNIVOICE_SHERPA_ASR_MODEL` selects the model — default
`sherpa-parakeet-tdt-v3`.
## Best at
- **Live dictation on CPU** — the whole point of this engine. Fast partials,
automatic endpointing on silence, no GPU required.
- It also honors the regular offline `transcribe` contract, so any of its
models can transcribe a file — plain text, single segment, no word
timestamps, which makes it a dictation/notes tool rather than a dubbing
engine.
## The 7 curated models
| Id | Type | Languages | Download |
| --- | --- | --- | --- |
| `sherpa-parakeet-tdt-v3` (default) | offline | 25 European languages | 0.67 GB |
| `sherpa-parakeet-tdt-v2` | offline | English | 0.66 GB |
| `sherpa-zipformer-bilingual-zh-en` | streaming | Chinese + English | 0.20 GB |
| `sherpa-paraformer-bilingual-zh-en` | streaming | Chinese + English | 0.24 GB |
| `sherpa-zipformer-en-20m` | streaming | English | 0.044 GB |
| `sherpa-zipformer-zh-14m` | streaming | Chinese | 0.025 GB |
| `sherpa-whisper-tiny` | offline | 90+ languages (auto-detect) | 0.104 GB |
Sizes are measured on-disk download sizes. Weights are int8 ONNX checkpoints
that download on first use through the same HF cache as everything else —
see [downloading-models](../downloading-models.md). Peak RAM for the 0.6B
Parakeets is noticeably higher than their download size (onnxruntime's arena
allocator holds onto freed blocks).
## Platform support
CPU on every platform, by the strict cross-platform default-parity rule.
`OMNIVOICE_SHERPA_ASR_PROVIDER` can override the ONNX provider on a verified
GPU build, but the default never diverges.
## Tuning
- `OMNIVOICE_SHERPA_ASR_THREADS` — decode threads (default 2; the 0.6B
Parakeets automatically use up to 4 when the host has the cores, so decode
keeps ahead of the speaker).
- `OMNIVOICE_DICTATION_ENDPOINT_R1` / `OMNIVOICE_DICTATION_ENDPOINT_R2`
streaming endpoint rules in seconds (defaults 1.0 / 0.6: text commits
~0.6 s after you stop speaking). Applied without a restart.
## Quirks
- The recognizer is **pre-warmed in the background** so the first dictation
session doesn't pay the 1.32.5 s ONNX session load
([#888](https://github.com/debpalash/VoiceStudio/issues/888)); it's then
shared warm across sessions.
- On Apple Silicon, installing the [parakeet-mlx](parakeet-mlx.md) model
makes dictation prefer the GPU Parakeet automatically for the 25 covered
languages; an explicitly selected sherpa model still wins.
- The offline `transcribe` path reports `language: auto` — per-file language
detection is only meaningful for the Whisper Tiny model.
+75
View File
@@ -0,0 +1,75 @@
# VoiceStudio — Sherpa-ONNX Engine
Sherpa-ONNX (k2-fsa/sherpa-onnx) is a unified C++ ONNX runtime that wraps
20+ TTS model families (VITS, MeloTTS, Piper, Kokoro, Matcha, and more)
behind one API, with pre-built wheels for Linux, Windows, and macOS (x86 and
ARM). You bring the model: point VoiceStudio at any downloaded sherpa-onnx
TTS model directory.
## When to pick it
- You want a specific community model (e.g. a Piper or VITS voice for your
language) that no other engine hosts.
- You need a dependable CPU engine with optional CUDA acceleration.
## Setup
1. Install the runtime:
```bash
pip install sherpa-onnx
```
2. Download a TTS model from the
[sherpa-onnx releases](https://github.com/k2-fsa/sherpa-onnx/releases)
and unpack it somewhere permanent.
3. Point VoiceStudio at the model directory and restart:
```bash
export OMNIVOICE_SHERPA_MODEL=/path/to/model-dir
```
4. Select the engine via **Model Catalogue → Engines** or
`OMNIVOICE_TTS_BACKEND=sherpa-onnx`.
The directory must contain `model.onnx` and `tokens.txt`. Sherpa-ONNX ships
no bundled default model, so the engine reports unavailable — with the
reason — until `OMNIVOICE_SHERPA_MODEL` points at a valid directory. (Before
this gate, selecting the engine unconfigured produced a failure mislabeled
as out-of-memory —
[#919](https://github.com/debpalash/VoiceStudio/issues/919).)
## Configuration
| Variable | Default | Meaning |
| --- | --- | --- |
| `OMNIVOICE_SHERPA_MODEL` | (unset) | Directory containing `model.onnx` + `tokens.txt` |
## Behaviour notes
- Output defaults to 22.05 kHz (the VITS default); once a model is loaded,
its own sample rate is used.
- CPU is the universal baseline; the CUDA onnxruntime provider is available
on Linux/Windows installs.
- **No cloning**: voices come from the model itself. Multi-speaker VITS
models select a voice by numeric speaker id; speed is supported.
- Languages depend entirely on the model you download.
## Known limits
- One model at a time — switching models means changing
`OMNIVOICE_SHERPA_MODEL` and restarting.
- No voice design, no reference-audio cloning, no emotion controls
(see [expressive-speech.md](../expressive-speech.md)).
## Troubleshooting
- "OMNIVOICE_SHERPA_MODEL not set" / "No model.onnx in …": follow Setup
above — the variable must point at the *unpacked* model directory, not
the archive.
- Other issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [benchmarks.md](../benchmarks.md),
[languages.md](../languages.md),
[disk usage](disk-usage.md).
+76
View File
@@ -0,0 +1,76 @@
# VoiceStudio — Supertonic-3 Engine
Supertonic-3 (Supertone Inc.) is a ~99M-parameter ONNX TTS engine covering
31 languages with 7 preset voices at native 44.1 kHz. It is CPU-only by
design — pure ONNX Runtime on the CPU execution provider, with no CUDA or
MPS path in the upstream SDK — and runs in its own sidecar process so
crashes and cold init never block the rest of VoiceStudio.
## When to pick it
- Broad language coverage on machines with no usable GPU.
- Preset-voice narration at a higher sample rate than the default engine.
## Setup
1. Install the optional dependency into VoiceStudio's environment:
```bash
uv sync --extra supertonic
```
(Or enable it from **Model Catalogue → Engines**, which installs the
pinned `supertonic` wheel for you.)
2. **Accept the license in-app.** First use is gated behind an explicit
acceptance dialog: the inference SDK is MIT, but the model weights are
**OpenRAIL-M**, which carries use restrictions. The engine stays
unavailable until you review and accept in **Model Catalogue → Engines →
Supertonic-3**.
3. Select the engine via **Model Catalogue → Engines** or
`OMNIVOICE_TTS_BACKEND=supertonic3`.
The first synthesis cold-downloads ~400 MB of model weights, pinned to an
exact HuggingFace revision SHA so the bytes match what the SDK was validated
against. See [downloading-models.md](../downloading-models.md).
## Voices
Seven preset voices are surfaced: `M1` (default), `M3`, `M4`, `M5`, `F3`,
`F4`, `F5`. The SDK itself accepts the full `M1``M5` / `F1``F5` set if a
caller passes one explicitly; unknown ids fall back to the default with a
log line.
## Behaviour notes
- Output is 44.1 kHz mono.
- Runs as a long-lived sidecar in the parent Python environment (its
dependencies — onnxruntime, numpy, soundfile — already match
VoiceStudio's pins); subsequent calls reuse the warm ONNX session.
- `speed` is clamped to 0.72.0; quality steps clamp to 512.
- Language is an ISO 639-1 code; Auto engages the SDK's multilingual
fallback.
## Known limits
- **No cloning and no voice design** — preset voices only. Dub/batch jobs
that need cloning won't select it.
- CPU-only: hardware acceleration is a property of the upstream SDK, not a
VoiceStudio limitation.
- OpenRAIL-M weights are not covered by VoiceStudio's blanket
commercial-use statement — review the model license terms in the
acceptance dialog.
## Troubleshooting
- "supertonic package not installed": run the `uv sync` above or enable
from the Model Catalogue.
- "license not accepted": open **Model Catalogue → Engines → Supertonic-3**
and accept.
- Other issues: [install/troubleshooting.md](../install/troubleshooting.md).
See also: [benchmarks.md](../benchmarks.md),
[languages.md](../languages.md),
[expressive-speech.md](../expressive-speech.md),
[disk usage](disk-usage.md).
+76
View File
@@ -0,0 +1,76 @@
# VoiceStudio — VoxCPM2 Engine
VoxCPM2 (OpenBMB) is the studio-quality option: native 48 kHz output,
zero-shot voice cloning, and — uniquely among VoiceStudio's engines —
**voice design**: creating a synthetic voice from a text description
("young female, warm tone, British accent") with no reference audio at all.
## When to pick it
- You want voice design without a reference clip.
- You want the highest output sample rate (48 kHz vs OmniVoice's 24 kHz).
- Your language is among its 30 supported languages: Arabic, Burmese,
Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew,
Hindi, Indonesian, Italian, Japanese, Khmer, Korean, Lao, Malay,
Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish,
Tagalog, Thai, Turkish, Vietnamese.
## Requirements
- Python ≥ 3.10, PyTorch ≥ 2.5.
- CUDA ≥ 12 recommended for full speed; MPS (Apple Silicon) and CPU also
work.
## Setup
Install the package into VoiceStudio's Python environment:
```bash
pip install "voxcpm>=2.0.3"
```
That is a version **floor**, not a pin — an older install still works, but
the engine logs an upgrade hint at load time. Then select the engine via
**Model Catalogue → Engines** or `OMNIVOICE_TTS_BACKEND=voxcpm2`.
## Model selection
| Variable | Default | Meaning |
| --- | --- | --- |
| `OMNIVOICE_VOXCPM_MODEL` | `openbmb/VoxCPM2` | HuggingFace checkpoint to load |
The first use downloads a multi-GB checkpoint from HuggingFace. A download
interrupted near the end used to abort the load outright
([#1224](https://github.com/debpalash/VoiceStudio/issues/1224)); the load is
now retried once with a fresh client. See
[downloading-models.md](../downloading-models.md).
## Behaviour notes
- **Voice design:** provide a description and no reference audio.
- **Cloning:** the reference clip is prepared before use (edge-silence trim
and length cap) so dead air in a raw clip doesn't condition the output; on
any prep problem the raw clip is used as-is.
- **Style instructions** are passed as an inline prefix to the text.
- VoxCPM2 emits mastered, studio-grade audio, so VoiceStudio **skips its
shared mastering chain** (which is tuned for 24 kHz engines) — only benign
loudness normalization applies.
- A trailing-silence guard trims long near-silent tails from generations,
keeping a short natural tail.
## Known limits
- Slower than the lightweight CPU engines — see
[benchmarks.md](../benchmarks.md) and [performance.md](../performance.md).
- Language coverage is 30 languages; for anything else use the default
[OmniVoice](omnivoice.md) engine ([languages.md](../languages.md)).
## Troubleshooting
- Engine shows unavailable: the `voxcpm` package isn't installed — run the
`pip install` above and restart VoiceStudio.
- Repeated first-download failures: check connectivity/HF access, then see
[install/troubleshooting.md](../install/troubleshooting.md).
See also: [expressive-speech.md](../expressive-speech.md),
[disk usage](disk-usage.md).
+80
View File
@@ -0,0 +1,80 @@
# VoiceStudio — WhisperX Engine
WhisperX is the default ASR engine on CUDA and plain-CPU hosts: faster-whisper
(CTranslate2) transcription plus a **wav2vec2 forced-alignment** pass that
snaps word boundaries to ±1030 ms (Whisper's own timestamps are ±100300 ms).
That word timing is what dubbing lip-sync depends on, which is why auto-detect
prefers it wherever CTranslate2 can use the GPU.
## Selecting it
- **Model Catalogue → Engines**, ASR tab → **Use** on the WhisperX row, or
- pin it with `OMNIVOICE_ASR_BACKEND=whisperx` (the env var always wins over
the Settings pick; with neither set, auto-detect chooses per-hardware).
## Best at
- **Dubbing** — the forced alignment is the accuracy tier lip-sync needs.
- **Batch transcription** with word-level subtitles.
- Multi-speaker work: it pairs with pyannote speaker diarization — see
[diarization](../features/diarization.md).
## Platform support
| Host | What happens |
| --- | --- |
| NVIDIA CUDA | GPU, float16 (degrades automatically, see below) |
| CPU (any OS) | int8 — works, but slow for large-v3 |
| Apple Silicon | CPU only — CTranslate2 has no Metal build, so auto-detect prefers [mlx-whisper](mlx-whisper.md) there ([#1127](https://github.com/debpalash/VoiceStudio/issues/1127)) |
| AMD ROCm | CPU only — CTranslate2 has no HIP build, so auto-detect prefers [pytorch-whisper](pytorch-whisper.md) there ([#1529](https://github.com/debpalash/VoiceStudio/issues/1529)) |
## Model selection
- `ASR_MODEL_WHISPERX` — default `large-v3`. Accepts the usual size aliases
(`tiny``large-v3`, `distil-large-v3`) or a full HF repo id. Weights
download on first load — see [downloading-models](../downloading-models.md).
- `OMNIVOICE_ALIGN_DEVICE` — force the wav2vec2 aligner's device. Aligners
exist for ~20 major languages; other languages keep Whisper's native word
timestamps instead of failing.
## VRAM preflight and degradation
Loading fp16 large-v3 onto a nearly-full 8 GB card dies as a *native* CUDA
abort — no Python exception, the whole backend goes down
([#723](https://github.com/debpalash/VoiceStudio/issues/723)). So before every
load the engine checks free VRAM against per-compute-type budgets
(float16 5.0 GB, int8_float16 3.5 GB, int8 3.0 GB, scaled down for smaller
models) and degrades the compute type — or falls to CPU int8 — instead of
starting a load that would kill the process. Disable with
`OMNIVOICE_ASR_VRAM_PREFLIGHT=0`.
Two more fallback chains run at load time:
- GPUs without efficient fp16 (older Maxwell/Pascal, GTX 16xx) raise a
compute-type error — the engine retries int8_float16, then int8
([#551](https://github.com/debpalash/VoiceStudio/issues/551)).
- A genuine CUDA OOM retries on CPU int8, so dubbing still completes
(slower, same model and accuracy).
## Quirks
- **cuDNN 8 required on CUDA.** CTranslate2 links cuDNN 8; if it's missing the
process fast-fails with no traceback, so the engine is reported unavailable
up front and selection falls through to pytorch-whisper, which uses torch's
own cuDNN 9 ([#1371](https://github.com/debpalash/VoiceStudio/issues/1371)).
- On some hardened Linux kernels CTranslate2's native library is rejected with
"cannot enable executable stack" — reported as unavailable, not a crash
([#692](https://github.com/debpalash/VoiceStudio/issues/692)).
- A partially-installed environment (interrupted sync, antivirus quarantine)
can break WhisperX's deep import chain (whisperx → pyannote →
lightning_fabric). The engine is then reported unavailable with a repair
hint — reinstall, or `uv sync --reinstall` on a source checkout
([#1185](https://github.com/debpalash/VoiceStudio/issues/1185)).
- Audio is decoded through VoiceStudio's validated ffmpeg, not a bare `ffmpeg`
PATH lookup ([#479](https://github.com/debpalash/VoiceStudio/issues/479)).
- Transcribes are time-bounded: each dub chunk by
`OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S` (default 120 s), whole files by
`OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S` (default 300 s). Raise them for very
long files on slow hardware.
Speed comparisons across engines live in [performance](../performance.md).
+29
View File
@@ -0,0 +1,29 @@
# Hosted Voice integration
VoiceStudio remains local-first. `/profiles` and `/generate` keep their local
SQLite and on-device synthesis behaviour unless a caller explicitly asks for a
hosted operation. No profile or generation is uploaded merely because hosted
configuration exists.
To enable the optional adapter, configure the backend environment:
```text
VSS_HOSTED_API_BASE=http://127.0.0.1:8080
VSS_HOSTED_API_TOKEN=<scoped API credential>
VSS_HOSTED_PROJECT_ID=<hosted project id>
VSS_HOSTED_MODEL_ID=<approved TTS model id>
VSS_HOSTED_MODEL_VERSION=<approved model version>
VSS_HOSTED_BASE_VOICE_ID=<model-approved base voice>
VSS_HOSTED_CONSENT_TEXT_VERSION=oss-spoken-consent-v1
```
First record ownership consent in the local profile UI, then explicitly call
`POST /profiles/{profile_id}/hosted-sync`. The adapter uploads the reference
recording through hosted Artifact grants and creates a consent-backed
`/v1/voices` record; it never sends a local path or a consent recording. The
returned hosted ID is stored only as local synchronization metadata.
Call `POST /generate` with `hosted=true` and that synchronized `profile_id` to
use the hosted durable `/v1/jobs` path. The adapter stages text as an Artifact,
polls the durable Job, and downloads the result only through a temporary grant.
Without `hosted=true`, `/generate` stays entirely on-device.
+7
View File
@@ -214,6 +214,13 @@ Two paths are worth persisting across container restarts:
The running version is now shown in **Settings → About → Version** (read live
from the backend), so the web UI no longer displays a dash in Docker.
- **Checking which version is running:** `docker exec <container> python3 -c "import importlib.metadata; print(importlib.metadata.version('omnivoice'))"`, or hit the `/health` endpoint — it returns `{"status": "ok", "device": ..., "version": "0.3.x"}`. Use the container name listed by `docker compose ps` (or `omnivoice` for the `docker run` examples).
- **Watching startup:** the port answers within about a second of container
start, but heavy initialization (PyTorch, API routes, database migration)
continues in the background. During that window `/health` returns **503**
with the current step, and `GET /startup/progress` returns the full
step-by-step ledger (`status`, current `step`/`label`, per-step states) —
useful when a start seems slow and you want to see where it actually is.
The Docker `HEALTHCHECK` flips healthy only once `/health` is 200.
- **"Loopback origin required" errors (and a blank version):** the desktop
build restricts the `/system/*` and `/api/settings/*` routes to a loopback
origin, but Docker's NAT makes every request look non-loopback, so the gate
Binary file not shown.

After

Width:  |  Height:  |  Size: 218 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 137 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.2 MiB

+4
View File
@@ -80,6 +80,7 @@ None of them are required — the defaults are chosen for the common case.
| Variable | Default | What it does |
|---|---|---|
| `OMNIVOICE_DEVICE` | `auto` | Pin the compute device (`cuda` / `rocm` / `xpu` / `mps` / `cpu`) instead of auto-detect. Same control lives in **Settings → Performance & Device** (the env var wins over the UI pick). Honored only for devices the host actually has — a family that isn't detected is noted and ignored, never obeyed blindly. Applies at the next backend start. |
| `OMNIVOICE_IDLE_TIMEOUT_S` | `900` | Seconds of idle before the TTS model unloads to free memory. Raise it (e.g. `3600`) if you generate in bursts and dislike the ~8 s reload; lower it on tight-memory machines. |
| `OMNIVOICE_SIDECAR_IDLE_TIMEOUT_S` | `300` | Same idea for sidecar engines (IndexTTS 2.5 etc.). |
| `OMNIVOICE_LLM_CONCURRENCY` | `6` | Parallel LLM translation calls during a dub. Raise for a fast API endpoint, lower if your provider rate-limits. |
@@ -229,6 +230,9 @@ uv run python scripts/bench_pipeline.py tts clone # just these stages
If you report a performance issue, pasting its table (plus your platform and
RAM/VRAM) turns a guessing game into a bisect.
Measured results per engine/device — and how to contribute yours — live in
[benchmarks.md](benchmarks.md).
## Things that look like knobs but aren't
- **Deleting and re-adding a voice** doesn't speed anything up; the reference
+28 -17
View File
@@ -24,13 +24,14 @@ loopback-only exactly as before.
┌──────────────┐ tailnet (WireGuard) ┌─────────────────────┐
│ laptop │ ws/https to MagicDNS URL │ gpu-box │
│ VoiceStudio UI │ ──────────────────────────▶ │ VoiceStudio backend │
│ (thin client) │ Authorization: Bearer … │ OMNIVOICE_API_KEY set │
│ (thin client) │ short-lived session/ticket │ OMNIVOICE_API_KEY set │
└──────────────┘ └─────────────────────┘
```
The desktop app *is* the thin client — there is no separate binary. You set a
**Backend URL** and an **API key** in Settings, and every request (including
the dictation and TTS WebSockets) is sent to the remote with the key attached.
The desktop app *is* the thin client — there is no separate binary. You enter a
**Backend URL** and an **API key** in Settings. The key is exchanged once for a
short-lived session; ordinary HTTP requests use that session and WebSockets use
path-bound, single-use tickets. The master is never stored or put in a URL.
## 1. On the GPU box: run the backend with a key
@@ -50,11 +51,11 @@ the backend's CORS allow-list to include that origin — see
[Browsers from another origin (CORS)](api-auth.md#browsers-from-another-origin-cors);
neither server mode nor trusted networks covers CORS.
When `OMNIVOICE_API_KEY` is set, **every non-loopback HTTP and WebSocket
request must present it**, as `Authorization: Bearer <key>`, `?api_key=<key>`
(browser WebSockets can't set headers), or the `ov_key` cookie the backend
sets after the first authenticated request. Loopback traffic on the box
itself is never gated, so local tools keep working.
When `OMNIVOICE_API_KEY` is set, every non-loopback request needs an accepted
credential. Scripts should use `Authorization: Bearer <key>`. Legacy
`?api_key=` and `ov_key` transports remain accepted for compatibility, but the
backend no longer creates a master-key cookie and the bundled UI uses only
short-lived sessions. Loopback traffic on the box itself remains ungated.
## 2. Reach it over Tailscale
@@ -79,6 +80,11 @@ Serve terminates on the node and forwards from `127.0.0.1`, so to the backend
the request looks like loopback — which is why the **API key is still
required** in that path (the bearer gate doesn't rely on the source address
for non-local exposure; set the key and it always applies to keyed clients).
Uvicorn trusts proxy headers from loopback by default, so Serve's forwarded
HTTPS scheme becomes the authoritative ASGI scheme and browser session cookies
receive `Secure`. For a non-loopback reverse proxy, explicitly configure
Uvicorn's `--forwarded-allow-ips=<proxy-ip>`; the application never trusts an
arbitrary `X-Forwarded-Proto` header itself.
> **Do not use `tailscale funnel`** (public-internet exposure) for this. Even
> with a key, a voice-cloning backend should not be on the open internet.
@@ -90,10 +96,11 @@ Settings → Sharing → **Remote backend**:
- **Backend URL**: the MagicDNS URL from step 2 (with `:3900` if you didn't
use Serve, or no port if you did).
- **API key**: the value of `OMNIVOICE_API_KEY` from step 1.
- **Test connection** hits `{url}/health` and shows the remote's version and
device.
- **Save & reload** stores both in this browser/app and restarts the UI
against the remote. The URL must be a full `http://` or `https://` URL
- **Test connection** hits the auth-exempt `{url}/health` with no credential,
then exchanges the entered key for a session if health succeeds.
- **Save & reload** stores only the URL and restarts the UI against the remote.
The key input is cleared after its single exchange. The URL must be a full
`http://` or `https://` URL
(`gpu-box:3900` alone is rejected), and saving a URL that hasn't passed
**Test connection** asks for confirmation first — a wrong base would leave
the app unable to reach any backend until you change it back here.
@@ -111,12 +118,13 @@ https://gpu-box.your-tailnet.ts.net/#api_key=<key>
Use the fragment (`#`, not `?`) deliberately: fragments are never sent to the
server, so the key stays out of the GPU box's and any reverse proxy's request
logs. The key is stored for that browser and the fragment is scrubbed from the
address bar (so it doesn't linger in history or get re-applied on a reload). If
logs. The fragment is scrubbed synchronously, then the key is exchanged once
for an eight-hour maximum session; the master is not stored. If
your key contains `+`, `&`, `#`, or `=`, URL-encode it (e.g. `#api_key=a%2Bb`);
keys from `secrets.token_urlsafe` (above) need no encoding.
Thereafter the UI loads normally with the key attached to every request. If a
request ever 401s again (wrong/rotated key), you're prompted to re-enter it. The
Thereafter the UI loads normally with the short-lived session. Cross-origin
bearer sessions are tab-scoped; closing the tab requires re-entry. If a request
401s again (expired/wrong/rotated key), you're prompted to re-enter it. The
same gate shows a LAN-share **PIN** prompt instead when network sharing — not a
remote key — is what's gating access.
@@ -125,6 +133,9 @@ remote key — is what's gating access.
- **Plain HTTP is sniffable.** A bearer key over `http://` on a hostile
network can be read off the wire. Use Tailscale (WireGuard-encrypted) or
Tailscale Serve (TLS) for anything beyond a fully trusted LAN.
- The first-party UI never persists `OMNIVOICE_API_KEY`, never creates a URL
containing it, and never puts its administrator session in a WebSocket URL.
WebSocket tickets expire after 30 seconds and work once for one path.
- The API key and the LAN-share **PIN** are independent: the PIN guards a
casual share session, the key is the durable remote credential. Either can
be active; both are checked when set.
+26 -8
View File
@@ -1,12 +1,18 @@
# Remote GPU workers
Run OmniVoice on this machine, but hand individual jobs to GPUs on your other
Run VoiceStudio on this machine, but hand individual jobs to GPUs on your other
machines. Results come back here.
This is **opt-in and off by default**. Until you turn it on and approve a
worker, nothing leaves your computer, no port is opened, and the app behaves
exactly as it did before.
Worker management is an admin surface. In Docker/server mode, viewing status
works during bare bootstrap, but joining, enabling, approving, issuing keys,
disconnecting, or removing machines remotely requires `OMNIVOICE_API_KEY`.
The share PIN and trusted-network exemptions authorize playback, not worker
administration.
> **Not the same as [Remote backend](remote-gpu.md).** That points this app at
> a backend running somewhere else, so the whole app — your projects, your
> voices, your history — lives on that machine. This keeps everything here and
@@ -17,7 +23,7 @@ exactly as it did before.
## What you need
* OmniVoice on both machines, on versions no more than two releases apart.
* VoiceStudio on both machines, on versions no more than two releases apart.
* The worker machine must be able to **reach** this one over the network. Same
LAN is enough at home; across networks, a VPN such as
[Tailscale](https://tailscale.com/) is the reliable answer. The worker dials
@@ -147,6 +153,15 @@ fallback is reported once. ASR, diarization and translation also remain local. D
runs here, deliberately and permanently, because there latency *is* the
feature. The remaining operations are being ported one at a time.
### Voice identity parity
For TTS, the worker receives the complete local rendering contract: the voice
profile's reference audio and transcript, its pinned seed, model quality
controls, text chunking/crossfade settings, and output effect preset. The
worker runs the same native or generic rendering pipeline as local
`/generate`; selecting a gallery voice therefore does not turn it into a new
random voice merely because it was rendered on another GPU.
The picker knows this. It resolves against the surface you are on, so a chosen
worker reads **Local** on a tab whose work has no remote path yet and names the
reason, instead of showing a green dot next to a GPU that receives nothing. The
@@ -157,10 +172,11 @@ The Dictation surface states that it always uses this machine without showing
the generic "not ported yet" notice.
For protocol development, a task can also be placed by hand with
`POST /workers/tasks` — a **development-only** endpoint. It is loopback-only,
`POST /workers/tasks` — a **development-only** endpoint. It is admin-gated,
sits behind the same opt-in as everything else here, takes a mandatory
deadline, submits one task and waits for it. It is not a stable API and goes
away once generation routes itself.
deadline, submits one task and waits for it. On desktop that means loopback;
in server mode a remote caller needs `OMNIVOICE_API_KEY`. It is not a stable
API and goes away once generation routes itself.
## How work is placed
@@ -197,16 +213,18 @@ The row tells you what happened in words — "Paused after 3 failures … retryi
in 45s" — and **Resume** clears it immediately when you've fixed the machine.
**You quit the app mid-task.** Remote work keeps running on the worker. On next
launch OmniVoice recovers those tasks and reconciles with each worker about
launch VoiceStudio recovers those tasks and reconciles with each worker about
what is genuinely still in flight.
**Version or feature mismatch.** The protocol keeps a two-release compatibility
window, but release numbers alone do not prove that a worker understands every
additive command. Registration therefore also declares named features for task
inputs, progress leases, and remote model downloads. A worker outside the
inputs, progress leases, remote model downloads, and the voice-identity render
pipeline. A worker outside the
version window, or one missing a required feature, is refused with
`UPGRADE_REQUIRED` and an update instruction before any task runs. It can never
silently render without reference audio or leave a download stuck at 0%.
silently render without reference audio, substitute a different voice, or leave
a download stuck at 0%.
Every remote failure includes a concrete next step. Capacity, missing models,
expired leases or sessions, authentication, rejected inputs, and result upload
Binary file not shown.

Before

Width:  |  Height:  |  Size: 268 KiB

After

Width:  |  Height:  |  Size: 185 KiB

Some files were not shown because too many files have changed in this diff Show More