Opt-in watch folder on the batch queue: pick a directory once and new videos are auto-enqueued with the last Add-to-queue settings, with pause/stop controls and copy-in-progress protection. Files stream to the loopback backend as bytes; paths never leave the app. Also gives the batch queue a reachable UI entry point and streams multipart uploads to disk. Maintainer fix: the watched directory handle is opened with full share mode on Windows so users can rename or delete the folder while it is watched, matching macOS/Linux behaviour, with a cross-platform regression test. Thanks @mvanhorn!
Fixes #1711.\n\nGives only the OmniVoice subprocess a 120-second readiness budget while retaining the shared 30-second default for all other sidecars, with regression coverage.
* fix(setup): tolerate reserved memory in the RAM preflight, add OMNIVOICE_RAM_PREFLIGHT=0 escape hatch (#1618)
An "8 GB" machine reports ~7.8 GB usable (firmware/iGPU/kernel
reservations), so comparing OS-reported RAM against the marketing-size
8 GB threshold hard-blocked exactly the boundary hardware the minimum is
meant to admit — with no way past the wizard. Both thresholds are now
compared with a 7% reserved-memory allowance, and
OMNIVOICE_RAM_PREFLIGHT=0 downgrades a genuine fail to a warning for
users who accept the OOM risk (same opt-out shape as
OMNIVOICE_ASR_VRAM_PREFLIGHT).
Regression tests: backend/tests/test_ram_preflight_1618.py.
Docs: troubleshooting §1c.
* review: hermetic preflight stubs in tests; correct the escape-hatch doc
Greptile P1: the Settings panel can't set OMNIVOICE_RAM_PREFLIGHT (and the
blocker appears before setup completes anyway) — the doc now points at
PowerShell / shell env only.
CodeRabbit: stub _network_check and media_tools.summary so each RAM
assertion stays fast and offline (26s -> 6s locally).
* feat(settings): compute-device override — auto | CUDA | ROCm | XPU | MPS | CPU
Auto-detect stays the default; the override kills the 'auto-detect picked
wrong' issue class. Applied at the single choke point (_probe()'s family
selection) so routing, get_best_device(), and every badge inherit it.
Resolution: OMNIVOICE_DEVICE env > Settings pick (prefs.json) > auto (#981
pattern). An override can steer, never invent hardware: a family the host
lacks is noted and ignored; cpu is always honorable. Applies at next
backend start (host caps are immutable per process — same restart contract
as the rest of the Performance tab, RestartBadge shown).
GET/PUT /api/settings/compute-device (admin-gated) reports resolved vs
applied so the panel shows restart-required truthfully and disables itself
under an env pin instead of pretending.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): entry for the compute-device override (#1557)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(device-override): harvest — the override reaches CT2 ASR, full i18n, honest edge states
- _ctranslate2_cuda_ok() and the ASR sidecar now gate on the probe's family,
so a cpu pin (or ROCm host) can never hand CTranslate2 a CUDA device —
the override reaches every CT2 loader through one shared gate
- override_ignored exposed by the API and shown by the panel (env pin naming
a device this machine lacks: auto is in effect, restart won't change it)
- all 8 panel strings + 5 device-family labels translated into all 21
locales; failed saves keep their error visible through the re-sync
- test isolation: cleanup drops OMNIVOICE_DEVICE before re-probing so no
overridden caps leak into later tests; panel tests wait for loaded state
- xpu/intel search keywords; oxfmt formatting
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(device-override): round 2 — fail-safe probe fallbacks, complete i18n, combined pin state
- a broken capability probe now means CPU everywhere (CT2 gate + ASR
sidecar) — never a torch-derived guess that would bypass a cpu pin or
re-open #1529 on ROCm; regression test added
- env-pinned AND not-detected shows both facts in one subtitle
- device_load_failed/perf_save_failed translated into all 21 locales;
CJK/th/vi/ar strings no longer say literal 'Auto'
- test_ctranslate2_never_gets_cuda_on_a_rocm_build pins the probe family
(it was order-dependent on the lru_cache before)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): pin the probe family in the faster-whisper OOM-fallback test
Same class as the rocm-build test: it mocked torch but not the probe the
new override gate consults first, so on a cpu-family CI host the CUDA
fallback chain under test was unreachable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(gallery): save gallery voices as profiles, with validated audio references
Work-in-progress lifted from the concurrent gallery session at the
owner's request (its uncommitted working tree, preserved verbatim from
base 92b1ee5d; safety snapshot remains at rescue/gallery-wip):
- gallery voices can be saved as local profiles: audio is copied into
the profile store with content-addressed filenames, existing profiles
are detected and refreshed only when the source clip changed
- backend/core/audio_validation.py: symlink-rejecting, root-contained
resolution for persisted profile WAV references, with tests
- archetype/community routers and the Voice Gallery UI updated for the
save-as-profile handoff (spec: docs/specs/longform/26-gallery-use-handoff.md)
- locale updates for the new gallery strings across all 21 files
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: drop a stray local screenshot script that rode in with the tree copy
* fix(community): explain the tolerated Content-Length parse failure; drop an unused import
CodeQL on #1542: the empty except now says why it is safe (the streamed
byte counter enforces the same cap regardless), and the test file loses
an unused Path import.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(gallery): review findings — copy outside the write lock, no stale completions
CodeRabbit on #1542, all findings addressed:
- the profile-audio copy stages to a .part temp BEFORE BEGIN IMMEDIATE
and publishes via atomic os.replace inside it — other backend writers
no longer block for the duration of an audio copy; a mid-copy failure
leaves no temp droppings and no profile row (both pinned by tests)
- VoiceGallery async ops carry per-operation generation tokens: a
preview or save-as-profile that resolves after unmount (or after a
newer operation) can no longer play audio, redirect into a workspace,
or touch state — three fail-before regression tests
- VoiceGalleryActions imports the page at test runtime; the e2e locator
uses a stable data-testid instead of a translated string; symlink
tests skip cleanly where the OS can't create symlinks; the changelog
line carries its PR ref
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: static ffmpeg fallback when the chocolatey feed is down
Third feed outage to break a PR run (2026-07-20, 2026-07-28, today —
three attempts, three 'installed 0/1'). Chocolatey is a distribution
channel, not the dependency: after the retry loop exhausts, fetch the
static gyan.dev build from its GitHub release mirror and put it on
PATH — same binary, no feed in the path. URL verified live (HTTP 200).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* test(asr): parse the guard, don't grep it
Two Majors CodeRabbit raised on #1523 — which I merged before reading
them, so this is the follow-up rather than a fix on the branch.
- Approved files were matched by BASENAME, so any future
`<anything>/asr_backend.py` was exempt from the guard it exists to
enforce. Matching is by relative path now; a decoy
`backend/engines/asr_backend.py` calling the selector is caught.
- Detection was a line regex, wrong in both directions: it missed
`import get_active_asr_backend as pick` and fired on the name inside
docstrings and comments. It walks the AST now, alias-aware, so only
real calls count.
Both verified by planting the exact bypasses: an aliased call in
services/tts_backend.py and the decoy module above. Neither was caught
before this change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(asr): resolve the selector's bindings before flagging a call
CodeRabbit, #1524: matching any call named get_active_asr_backend also
reported a local helper or an unrelated object's method that happens to
share the name. False positives are how a guard stops being believed —
people add allowlist entries for code that was never the bug.
Bindings are resolved first now: a bare call counts only if the name was
imported FROM services.asr_backend, an attribute call only if it hangs
off a module alias for it. Six shapes are pinned in the suite — direct,
aliased and module-attribute calls flagged; a same-named local function,
an unrelated method, and the name inside a docstring not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The recurrence guard from #1512 scanned api/routers only. A service or
engine module that transcribes on a request's behalf skips ensure_loaded()
just as thoroughly, so the guard could be sidestepped by moving the call
one module down the stack — verified: adding a get_active_asr_backend()
call to services/tts_backend.py passes the router scan and fails this one.
The broader scan is the one thing #1519 did better than the fix that
landed in #1515; absorbing it here rather than leaving it in a PR that
now conflicts. Thanks @ahov520.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(launchpad): quieter, borderless design refresh
The launchpad carried decoration from an earlier direction: icon chips,
corner-hung count badges, a permanently visible filled arrow, uppercase
mono card titles, and a dotted stipple divider — plus a frame that had
been invisible since the app-wide border tokens were zeroed.
Rework it around what the borderless direction actually implies:
- Feature tiles get a whisper-faint surface instead of a dead frame, and
read as three bands (bare glyph + count / title + arrow / description).
`--card-hue` is spent sparingly — the glyph at rest, the surface, count
and arrow only once raised. Titles move to sans sentence case; counts
are plain tabular numerals. Lift softened 4px -> 2px, coloured glow ->
neutral shadow, plus an explicit focus ring and a staggered entrance.
- Hero drops the boxed "646" pill and the filled A/B-Compare button for
quiet type, with a hairline standing in for the separation.
- Section labels trade the dotted stipple for a single fading hairline;
rows are transparent until hover and reveal "Open" on hover/focus (it
stays in the DOM, so AT and keyboard always reach it).
- Hero, tiles, recent files, callout and project lists now share one
1180px column — previously only the top half was capped, so lists ran
edge-to-edge on a wide display while the deck stayed centred.
Two bugs found and fixed while doing it:
- Buttons that had `border border-solid border-transparent` removed fell
back to the UA default border and rendered a visible 1px outline. They
now carry `border-0` explicitly.
- `.lp-animate` used `animation-fill-mode: both`, so after the entrance
it kept owning `transform` — and animation-origin declarations outrank
normal ones, which silently killed the card hover lift. Now `backwards`,
which still holds the from-state through the stagger delay.
Also drops CSS the page has not rendered since #904: the cursor-spotlight
layer, the breath ring, and the per-card waveform strip.
Verified with headless renders at 1600/1280/940 and the empty state.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(dictation): decode Wayland portal signals and show the capture pill
The GlobalShortcuts portal declares Activated/Deactivated as
(o session, s shortcut_id, t timestamp, a{sv} options). We decoded the
timestamp as u32, so zbus rejected every signal with
Signature mismatch: got `(osta{sv})`, expected `(osua{sv})`
and the press was dropped as an invalid signal. Registration succeeded
and the desktop even reported the bound chord back, so the hotkey looked
wired up while doing nothing at all — on every Wayland compositor, for
the whole life of the feature (#1490). Decode the 64-bit timestamp, and
keep the 32-bit spelling as a fallback so a non-conforming portal
degrades to working rather than to silence.
With presses arriving, the second half of the failure showed: nothing
had shown the widget window since it became a hidden recorder host, so a
capture ran with no pill on screen — and a mic or Accessibility failure
rendered into a window nobody could see. Add show_dictation_pill, which
bottom-centres the capsule on the monitor under the pointer and shows it
without taking focus (Windows keeps SW_SHOWNOACTIVATE so paste still
lands in the user's document), and call it from the widget for every
state but idle. Wayland denies clients their own placement, so the
compositor picks the spot there; the pill still appears.
dispatch_dictation_capture now logs whether a press was emitted or
queued — a press that reaches Rust and produces nothing was otherwise
indistinguishable from one the compositor never delivered.
Tests: portal signals decode at both timestamp widths (the 64-bit case
fails before this change with the exact production error); pill
placement centres, respects a second monitor's origin, and clamps rather
than going off-screen; the widget shows for a state needing the user,
stays hidden while idle, and never shows for a press that arrives while
dictation is disabled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: sync in-progress workspace changes
Uncommitted work already in the tree, checkpointed so the branch matches
the local machine:
- Remote GPU workers: join-from-the-app flow, one-time secrets, QR join
codes, a Compute control in the status bar, and the device-list
Workers panel (#1516)
- Model Catalogue workspace, with Settings pointing at it
- Settings sidebar search and keyboard navigation
- Demo assets for dubbing, dictation and voice design, plus the scripts
that render them
- Backend: validation-error handling, ASR request-path degradation, and
the accompanying tests
- CHANGELOG entries for the above and for the Wayland dictation fix
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(tests): follow Engines to the Model Catalogue, and green the sweep
- test_supertonic3 asserted the license gate points at "Settings" while
the engine now names Model Catalogue → Engines, which is where the
accept button actually lives. The assertion follows the move; what it
pins is unchanged — the hint must name a place the user can reach it.
- Carries the CJK allowlist entries for the rendered dub bundle (#1517)
and the regenerated route snapshot for /workers/agent (#1516), both of
which this branch inherits from the workspace sync.
- docs/install/linux.md: the dictation capsule is bottom-anchored
everywhere except Wayland, where the protocol gives applications no
say in their placement. Documented rather than left as a surprise
(CodeRabbit).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: stop a flaky dependency fetch from failing green runs
en-core-web-sm resolves to a direct GitHub release URL, and github.com
intermittently answers `http2 error: refused stream before processing
any application logic`. uv's own three retries all land within the same
few seconds and fail together, so the whole job dies on a dependency
that has nothing to do with the change under test — it cost #1518 and
#1517 an otherwise-green run tonight.
Two changes: back off between whole `uv sync` attempts, which is what
actually clears it, and pass --no-sync to the pytest steps. `uv run`
re-resolves the environment before running, so every test step was a
fresh chance to hit the same fetch even though the install step had
already synced — that is exactly how #1518 failed, in the isolated
backend/tests step, with all 5467 tests already passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: one retry seam for every uv sync, not just the job that failed last
en-core-web-sm resolves to a direct GitHub *release* URL rather than a
package index, and github.com intermittently answers `http2 error:
refused stream before processing any application logic`. uv's own
retries all land inside the same ~10 seconds and fail together, so a job
dies on a dependency unrelated to the change under test. Tonight that
cost four otherwise-green runs across #1515, #1517 and #1518 — and the
first fix only covered the Tests job, so the next failure simply moved
to Smoke (Linux), which syncs separately.
The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh
backs off between whole attempts (15s, 45s, 90s) and every workflow that
syncs now goes through it — ci.yml (tests + the platform matrix),
release.yml, security.yml, evals.yml. It still fails loudly after four
attempts, so a genuinely broken lockfile is not disguised as a flake.
The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke
matrix has always set, which is part of why it was the one that kept
dying; it has them now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(ci): pin the Intel-Mac contract by intent, not by command spelling
test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted
the literal line `run: uv sync --extra pockettts`, so routing every sync
through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The
contract it exists to protect is that the pockettts extra installs ONLY
on backend_supported legs — which the regex now pins, while leaving how
the sync is invoked free to change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: keep every uv run out of the resolver, and bound the retry budget
CodeRabbit, #1517:
- `uv run` re-resolves before running, so the smoke suite, the
worker-artifact tests, the release test run and the eval run were each
a fresh chance to hit the flaky direct-URL fetch outside the retry
loop. All of them pass --no-sync now; the environment is already
synced by the step that owns the retries. security.yml's
`uv run --with pip-audit` is deliberately left alone — it layers an
ephemeral package rather than running the project's own tests.
- The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a
120 s timeout on the smoke matrix). Three attempts and 60 s of total
backoff outlast the refusals actually observed while staying well
inside the jobs' timeout-minutes.
- The Intel-Mac contract test pinned the smoke command literally too, so
--no-sync tripped it exactly like the sync line did. Same fix: assert
the contract (smoke runs only on backend_supported legs), not its
spelling.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Validate HTTPS probe destinations against the shipped origin allowlist and clean up the related CodeQL test findings. Reviewed by Greptile; CodeRabbit was harvested but rate-limited. CI, Security, and cross-platform smoke checks are green.
Decode MLX Whisper audio once through VoiceStudio's validated bundled ffmpeg and reuse the waveform for forced alignment. This restores source-install transcription on Apple Silicon without requiring system ffmpeg.\n\nThanks @gambletan!
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
The report arrived as a raw traceback pasted out of a log file — because that is the only place the reason existed. A failed chapter was a red row and the word "failed"; the SSE event carried the literal string "chapter failed to render", the symptom the user could already see.
- Both the per-chapter and the terminal all-failed events are now built with the shared `core.failure` builder: sanitized text, a guaranteed non-empty reason, error class, docs deeplink and hint. `error` keeps mirroring `reason`, so older frontends and the Stories exporter are unaffected. The chapter list shows the reason inline, full text in the row tooltip.
- **A silent infinite hang on the same path.** asyncio refuses to put `StopIteration` into a Future — `_copy_future_state` raises TypeError inside the event loop's own callback, so the `run_in_executor` future is never completed and the caller waits forever, with no error, event or timeout. Reachable from ordinary input: VoxCPM's `next_and_close` is a bare `next(gen)`, so a generator ending without yielding raises it straight into a GPU-pool worker. Fixed at the pool boundary (`_ResilientGpuPool.submit` plus a matching `_cpu_pool` subclass) as `WorkerStopIteration(RuntimeError)`, keeping the original as `__cause__`.
- An all-failed render now marks the job failed. It previously returned without touching job history, so the row stayed `running` and the next startup's orphan sweep read a hopeless render as interrupted and offered it as resumable (Greptile P1).
Tests: 6 pool-guard tests (two hang without the fix, bounded so a hang fails rather than stalls CI), longform e2e extended for both event shapes, the empty-`str(exc)` floor, and both sides of the job-status branch, plus the chapter-list component.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`tests/backend/conftest.py` purges `main` / `core*` / `api*` / `services*` from `sys.modules` after every test it owns, and `backend/tests/` modules bind `import services.model_manager as mm` at COLLECTION time — so in a combined `pytest tests/ backend/tests/` session the alias and the live module are two different objects. Two tests read the wrong one; CI's split invocation hid both.
- `test_lifespan_shutdown_mid_load_is_clean_and_clears_sentinel` patched the stale alias then drove main's lifespan, which loads through the live module. It now runs inside a purge/restore context and resolves from `sys.modules` after importing main, so the failure is deterministic standalone.
- The autouse `_clean_model_manager_shutdown_state` fixture cleaned only `sys.modules` while `test_shutdown_state_isolation.py` dirtied the alias. It now cleans the live module and any module-typed alias in the requesting test module — the idiom already used there for `asr_backend` — taking both the `import x.y as z` binding (package attribute) and the `sys.modules` entry.
- New fail-before/pass-after pair guards the alias half of the fixture in isolation.
Test-infrastructure only; no production code touched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: reset model-manager shutdown state between backend tests
Two leaks, one of them mine.
1. `model_manager._shutting_down` is a module-global Event and the GPU pool is a
module-global executor. Any test that runs the app lifespan flips both on the
way out (begin_shutdown + _reset_gpu_pool) and nothing puts them back —
correct in production, where the process is ending; wrong across a combined
session. A test arriving with the flag set finds a shut-down executor, so its
first run_in_executor raises "cannot schedule new futures after shutdown",
which the preload path classifies as benign and swallows. The symptom is a
load that silently never starts. Reset before AND after: before so an
inherited flag cannot decide the test, after so a test that legitimately
shuts down does not hand it on.
2. tests/test_torch_compile_path_gate.py assigned services.settings_store into
sys.modules directly instead of via monkeypatch.setitem. That leaks
process-wide out of collection and breaks every later import of the real
module. backend/tests/test_no_module_stubs.py exists to catch exactly that,
and caught it — I introduced it two commits ago while isolating the Settings
gate for a review finding.
#1269 stays open for its last failure, which is a different root cause:
test_lifespan_shutdown_mid_load fails because a reload fixture in tests/
replaces services.model_manager, so the test patches one module object while
main's lifespan uses another (verified: `same=False`). That is the duplicate-
module class, not a state leak, and needs its own fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test: let the shutdown-state reset fail loudly
CodeRabbit Major: the broad try/except meant a reset that raised left the next
test with stale shutdown or executor state — precisely the order-dependent
failure the fixture exists to remove, while looking like it had worked. That is
the same silent-fail-open shape as the watermark and ffmpeg bugs fixed earlier
in this cycle.
If reset_shutdown_flag() or _reset_gpu_pool() can raise, that is a real problem
in model_manager and it should be loud.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test: assert the shutdown-state reset fixture actually resets (CodeRabbit)
Ordered pair: one test leaves the module globals exactly as the lifespan
leaves them, the next asserts it arrived clean — delete the fixture and the
second fails. Plus a mechanical guard that the reset stays un-swallowed, so
a future try/except cannot make the fixture look like it worked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(engines): path-aware GPU-pool slot in SubprocessASRBackend.transcribe()
transcribe() had the same on-pool self-deadlock that generate() had (fixed in
#1298): a bare no-op submitted to the GPU pool, but run_transcribe_guarded
dispatches it via run_in_executor(_gpu_pool), already on a pool worker, so on
a 1-worker (MPS) pool the no-op queued behind the job running it and timed
out before the sidecar spawned. IsolatedFasterWhisperBackend on MPS hit this
on every transcription.
Mirror generate()'s path-aware slot block (on-pool skip via
running_on_gpu_pool; off-pool _occupy hold) in transcribe(). The pattern is
duplicated rather than extracted into a shared helper to avoid reworking
generate(), which just shipped (#1298) with a CodeQL fix; extracting a shared
contextmanager is a clean follow-up. Regression test added (transcribe
dispatched on a pool worker).
* fix(engines): import threading in subprocess_asr
transcribe()'s off-pool slot hold uses threading.Event(), but the module never
imported threading — every subprocess-ASR transcribe raised NameError, and the
three round-trip tests failed in CI.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(engines): cover the off-pool transcribe branch
Only the on-pool path had a test, so `threading.Event()` in the off-pool
branch shipped with `threading` never imported — every direct caller hit
NameError before the sidecar started. Both bots caught it on review; nothing
in the suite did. A branch with no test is how a one-word bug reaches CI.
Also asserts the slot is genuinely released afterwards. Fails without the
import fix; the pre-existing on-pool test still passes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(engines): make SubprocessBackend.generate() path-aware on the GPU pool
generate()'s slot handling had two bugs:
1. On-pool self-deadlock: /v1/audio/speech and /generate (and audiobook,
dub, batch) dispatch generate() via run_on_gpu_pool_guarded, already on
a pool worker, so the inner slot submit queued behind the very job
running it on a 1-worker (MPS) pool and timed out before the sidecar
spawned. Every subprocess engine surfaced the in-process 300s-abandon
instead of synthesizing.
2. Off-pool no hold: the off-pool slot was a bare no-op that released the
worker before _spawn(), so off-pool callers (engine self-test,
diagnostics) could synthesize concurrently with a pool job and
over-subscribe the GPU.
Make the slot block path-aware: on-pool callers skip (the outer
run_on_gpu_pool_guarded already holds _running for the whole sidecar
exchange); off-pool callers hold a real slot for the whole synthesis via
an _occupy task that blocks the worker until _held is set in the finally.
Single release point in the finally.
Regression tests: generate dispatched on a pool worker (on-pool skip) and
a concurrent pool job blocked during an off-pool generate (off-pool hold).
Both verified fail-before / pass-after.
Supersedes #1296 (on-pool-skip-only). Closes#1295, #1297.
* Address review: couple on-pool skip to the pool prefix; fix comment
/simplify + /code-review flagged that the on-pool skip keyed on the literal
"gpu-pool" string, decoupled from _build_gpu_pool's thread_name_prefix. A
rename would silently re-introduce the exact self-deadlock this PR fixes (and
the tests can't catch it, since they hardcode the prefix). Centralise the
prefix in _GPU_POOL_THREAD_PREFIX + a running_on_gpu_pool() helper, used by
_build_gpu_pool, the skip in generate(), and _heal_tts_placement.
Also fix the comment: the Settings engine self-test rejects subprocess-isolated
engines with a 400, so the only real off-pool caller is the diagnose.py
deep-synth probe.
* fix(engines): bind slot_future before the off-pool branch
CodeQL py/uninitialized-local-variable (error, blocking CI). `_held is not
None` does imply slot_future was assigned, so the current code is correct —
but the two are only coupled by convention, which the analyser cannot see and
a third exit path would quietly break. Binds it to None up front and guards
the cancel.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(engines): make the slot-hold regression deterministic and leak-free
CodeRabbit, valid on both counts. The test used sleep(0.8)/sleep(0.5) as
synchronization — the tests/** contract forbids it, and on a slow runner the
marker could be enqueued before the generator had reserved anything, so the
assertion passed for the wrong reason. It now waits on an event signalled when
the slot task actually starts, and asserts "did not run" via a result()
timeout rather than a bare sleep.
Cleanup moved into finally: an assertion failure used to leak the sidecar
process and the pool thread into the rest of the session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(engines): add omnivoice-subprocess, a crash-isolated TTS engine
The default in-process OmniVoice engine runs on the GPU ThreadPoolExecutor.
When a generate or load exceeds its execution budget the pool is "reset", but
the abandoned worker thread cannot be killed (Python cannot interrupt a native
torch/MPS call), so it keeps holding the device until it finishes on its own
and later synths queue behind it and hang. The reset restores pool capacity
but not the device. This is the residual root cause behind the closed#730
and #1190: the messaging/reset mitigations address the symptom, not the
device-holding zombie.
Add an opt-in `omnivoice-subprocess` engine that runs the same model in a
child process via SubprocessBackend. A child process can be hard-killed: on a
recv-timeout the watchdog calls proc.kill(), reclaiming VRAM/device, and the
next request transparently respawns a fresh sidecar. The in-process engine
remains the default, so existing users see no change; this is an opt-in for
unattended / scheduled / reaction-triggered synthesis where a stuck job must
self-recover instead of hanging until a manual restart.
Base-class and mitigation changes that ship with it:
- SubprocessBackend.generate() now consumes non-terminal {"op":"progress"}
frames a sidecar emits during a cold load (previously the first cold
generate after spawn failed, then worked on retry). Additive: engines that
reply with audio directly are unaffected.
- recv_timeout_s is overridable per engine (default 60s unchanged); the new
engine sets it to the generate budget so a long-but-valid synth is not
falsely killed while a wedged one still is.
- make_room_before_generate(): free idle GPU memory before a warm, heavy
generate. The cold-load path already evicted; the warm path skipped it, so a
long synth on a VRAM-tight MPS box could contend its way into the budget.
Verified end-to-end against the live model (cold / warm / recovery-after-kill)
and under a sustained + concurrent-pressure soak: killed-worker recovery 5/5,
chunked long text 9/9, no memory leak.
* Address review: install_hint + move make_room into get_model
- Add `omnivoice-subprocess` to `_INSTALL_HINTS`; the
test_install_hints_cover_all_registered_backends gate requires every
registered backend to carry one (this was the CI failure).
- Move the warm-generate VRAM eviction out of the /generate and
/v1/audio/speech routes and into get_model()'s warm-return path, so EVERY
native TTS generate is covered (REST, WS TTS, dub, batch, audiobook), not
just the two REST routes. Drops the now-redundant per-route wiring.
(Greptile P1: the per-route placement missed the other generation surfaces.)
* Address review: drop dead long-text eviction path; log probe failure
- _should_make_room_for_generate: the long-text headroom boost became dead
code once the eviction moved into get_model() (which has no text), so the
long-text branch never fired. Removed the text param, the long-text
threshold/multiplier branch, and the now-unused _env_float helper. The core
RAM-tight gate (the part that matters on a starved box) is unchanged.
- Log the available_memory probe failure at debug instead of silently
swallowing it (CodeRabbit: silent swallow breaks the debug trail).
- Tests updated for the text-agnostic policy.
* fix(engines): stop subprocess generate() self-deadlock on 1-worker pools
SubprocessBackend.generate() acquires a GPU-pool slot for accounting, but
/v1/audio/speech and /generate dispatch backend.generate() via
run_on_gpu_pool_guarded, i.e. already ON a pool worker. On a 1-worker pool
(MPS) the inner pool.submit queued behind the very job running it and
slot_future.result(timeout=10) raised before the sidecar ever spawned, so
omnivoice-subprocess (and every other subprocess engine on MPS) surfaced the
in-process 300s-abandon instead of synthesizing.
Skip the slot acquisition when current_thread() is already a gpu-pool worker;
the outer guard already accounts for the slot. Direct callers (off the pool)
still acquire one. Regression test added (generate on a pool worker).
* Address review: reword slot-skip comment (fixes watermark-coverage CI) + simplify
- The slot-skip comment said "dispatch backend.generate() via", and
test_watermark_route_coverage's _SYNTH_CALL regex matches the literal
backend.generate( anywhere in a module, so it counted subprocess_backend.py
as a synthesis producer that must reference mark_synthetic (it doesn't — the
routes apply mark_synthetic; the engine sits below the chokepoint, like
tts_backend.py). Reworded to "dispatch generate() via".
- Fold in the simplify refinement: single negated predicate, import+pool
moved into the acquire branch.
* fix(errors): key the no-report rule on the shutdown marker, not on 503
Self-caught regression from #1272. Suppressing the "Report" action for every
503 was too broad: 503 is also how a real engine-load timeout and an
unavailable engine are reported (#1246, #1260, and #1277, filed hours ago
with a 503 in its very title). That would have removed the report button from
exactly the class of failure users need to be able to file — silencing real
bugs to hide a benign one.
The backend now tags only the shutdown case with a `[shutting_down]` marker,
following the existing `[clone_ref_unusable]` convention, and the UI matches
the marker instead of the status. A 503 that is a genuine failure keeps its
report action.
Localized message added to all 21 locales; the backend test pins the marker
so the cross-layer contract can't be renamed away.
Fail-before verified: the two "still offers Report for a 503" tests fail
against the shipped blanket-503 version.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* i18n(de): formal register for the two new German strings
Review: the surrounding errors.* strings and update.retry use "Sie"; both new
strings used the informal form. Turkish is consistently informal (install,
retry) so it stays as-is.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(dub): purge order was hash-dependent, and the cap could evict a live marker
main went red on my own test. Two distinct defects, both mine.
1. `targets = set(job_ids)` made iteration order depend on PYTHONHASHSEED, so
which markers a cap-forced trim discarded was luck. That is why the test
passed locally and failed in CI — verified: the old code passes at seeds
0/7/42 and fails at 12345. Now a de-duplicated list in caller order.
2. The size cap could evict markers the CURRENT purge had just recorded. Those
are the newest and the likeliest to still be held by a running job, so
dropping one is precisely the resurrection this mechanism exists to prevent.
A 'clear history' larger than the cap forced exactly that. The cap now never
touches the current purge, making the real bound cap + one purge — stated
plainly rather than implied.
Tests now pin both: identical survivors across runs, and an oversized purge
keeping all of its own markers. Verified across six hash seeds; full suite green
under 12345, the seed that reddened main.
* feat(update): announce updates as a toast with actions, not a wall of text
A version's release notes ARE the whole changelog section — v0.4.1's was 42
bullets. Any surface that renders them inline becomes unusable: older builds put
them in a blocking OS dialog that filled the screen and had to be dismissed
before the app could be touched.
Removing that dialog left the opposite failure. The only remaining signal was a
6-pixel dot beside the version number in the footer, which is easy to never
notice — so users either got shouted at or told nothing.
A toast is the middle: it names the version, offers Install and restart /
What's new / Later, and leaves. The notes stay one click away in Settings →
Updates, where there is room. Keyed by version so the 6-hourly re-check
replaces rather than stacks, and it never auto-dismisses — an update the user
hasn't answered is still true. Install declines while a generation is running,
since the relaunch would lose it.
Strings added to all 21 locales, translated rather than English-filled.
Tests pin the shape that matters: the toast takes no notes prop at all, renders
under 200 characters, and cannot stack duplicates.
* fix(update): don't relaunch while work is in flight; share the busy check
Review of #1272 found the restart guard was `dubStep === 'generating'` and
nothing else. Installing an update relaunches the process, so that permitted
throwing away a dub upload, a transcription, a translation, an export or a
standalone TTS synth. The same narrow check was written twice — in the new
toast and in UpdatesPanel — so the two could also drift apart.
Replaced with a single `isAppBusy(state)` in utils/appBusy, unioning the
signals the store actually has: the dub state machine, the floating status
pill (which every long background operation already pushes to), and
`ttsGenerating` — a new transient store field mirroring useTTS's local
`isGenerating`, which lived in a hook where no global check could see it.
`dubStep === 'editing'` is deliberately not busy: it waits on the user, and
counting it would block updates for as long as a transcript stays open.
Also from review:
- A failed lazy import of the toast fell into the outer catch, whose
setUpdateIdle() erased the update that had just been found. The
announcement is optional; the available state is not.
- "Dismiss" was machine-translated into the employment sense — terminate an
employee — in de/ja/ru/zh-CN/zh-TW, on close buttons and one aria-label.
Swept every locale rather than the four lines that were flagged.
- Hoisted the test's mock state with vi.hoisted.
CodeRabbit's changelog finding is declined: Highlights bullets carry no issue
refs by design, which tests/test_changelog_style.py enforces.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(shutdown): a quit mid-generate is not a 500, and not a bug report (#1276)
#1174 made a model load interrupted by shutdown benign for the background
preload, but a *request* that triggered a load took the generic unhandled-
exception path: crash log, ERROR traceback, and an error-journal entry that
feeds the bug-report pipeline. Quitting the app with a generate queued
surfaced "500 Internal Server Error: model load skipped: backend shutting
down" and offered to file a GitHub issue for a normal teardown.
Nothing failed — the process is exiting. The handler now answers 503 with
Retry-After and an actionable detail, ahead of the crash-log/journal writes.
Frontend half of the same bug: toastErrorWithReport offered "Report" for any
error. A 503 means "not now, try again" by definition, so it now shows the
backend's message without the report action — covering a still-warming
backend too, not just this shutdown path.
Fail-before/pass-after tests on both sides, including that the shutdown case
leaves no crash-log entry and no journal record.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(models): repair a half-downloaded model whichever way it reports (#1273)
transformers has two unrelated wordings for "this snapshot has no weight
shard", sharing no words:
hub load "<repo> does not appear to have a file named …"
local dir "Error no file named model.safetensors, … found in directory …"
The self-heal and the failure classifier both matched only the first. The
second is what a load of a *subfolder* inside a cached snapshot raises —
exactly where an interrupted download leaves a half-written repo — so the
reporter got neither the automatic repair (delete broken entries →
re-download → retry) nor an actionable hint, just a raw 500. Their disk had
10.6 GB free, i.e. a download that ran out of room.
The phrase list now lives once in core.failure, so the healer and the error
text cannot disagree about what an interrupted download looks like. Both
fragments of the second wording must match — "no file named" alone is
ordinary English and must not claim the class.
Also hardens the #1276 handler found by running both suites in one session:
services.model_manager can be imported under two module names, making two
distinct ModelLoadInterruptedByShutdown classes and breaking a bare
isinstance — which silently restored the 500 this fix exists to remove. Now
matched by isinstance OR class name, with a test that raises a same-named
class from a different module.
Verified against the full tests/ + backend/tests/ session: the only remaining
failures are the four pre-existing #1269 isolation leaks, unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(update): count synths in flight at the chokepoint, not in one caller
Review round 2. Both Greptile P1s were the same weakness in my first pass:
tracking the synth in useTTS meant only the Generate tab counted (voice
previews, the compare modal, the stories editor and profile previews call
generateSpeech directly and were invisible), and a boolean meant two
overlapping syntheses cleared each other — whichever settled first reported
"idle" while the other was still running, so Install and restart discarded it.
Moved to `api/generate.ts`, around the one `/generate` call all seven paths
share, as a count with a `finally` release (so an abort or a network error
frees it too). A future synth caller is covered without opting in.
Fail-before verified: the overlap test fails against the boolean version.
Also from review:
- UpdatesPanel re-reads the busy state at click time; the render-time snapshot
only exists to disable the button, and work can start after the last render.
- The 503 now carries the allowed-origin CORS headers. Without them the
browser reports a bare CORS failure and the actionable detail — the whole
point of the fix — never reaches the user. Both error responses build them
through one helper now.
- Log `request.url.path`, not the full URL: a query string can carry tokens
and newlines.
- `update.busy` said "finish your dub first" in all 21 locales, but the guard
now covers uploads, transcription, translation, export and synth. Rewritten
as work-in-progress wording, translated per locale.
- ru `common.dismissStatus` was "reject status"; missed in the earlier sweep.
Declined: CodeRabbit's request for a `(#NNN)` ref on the Highlights bullet —
Highlights carry no refs by design (CLAUDE.md), and tests/test_changelog_style.py
enforces it. Also declined gating the cache re-download behind a confirmation:
the self-heal already existed and already ran for the sibling wording, this
only stops it missing half the class, and it stays behind the existing
_hf_offline() check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(update): grey out Install while work is running
CI lint caught `busy` as unused after the click-time check moved into the
handler — and that exposed a real gap: `busy` had never been wired to
anything. The comment claimed it disabled the button; it didn't. Clicking
Install during a synth just bounced a toast back.
Now it does what it said: the Install and Restart buttons are disabled while
work is in flight, with `update.busy` as the tooltip. The click-time read
stays the authority, since work can start between the last render and the
click — the disabled state is the explanation, not the safety.
Adds the first UpdatesPanel test. The case it pins hardest is the inverse of
the bug: a busy predicate stuck at true would make the app permanently
un-updatable, which is worse than what this fixes. So it asserts enabled when
idle and during dub 'editing' (which waits on the user), disabled during an
upload, a translation, and one or more synths.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
- fallback preflight pins the candidate backend id (Greptile) — also the
CI empty-cache failure: deep-import fall-through tests get the
asr_model_installed fixture
- locale ratchet: improvement warns instead of failing CI (CodeRabbit)
- stream-exit ASR unload runs on the GPU pool fire-and-forget instead of
blocking the event loop (CodeRabbit)
- E741 rename; M4A ftyp sniff scope documented (CodeRabbit)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Harvested and verified every CodeRabbit/Greptile finding from PRs #1175,
#1189, #1192, #1195: 16 real ones fixed (fallback ASR preflight bypass,
VRAM release on stream exit, typed 409 parity, uv env independence,
path-privacy in errors, MCP clone_voice hardening, CaptureWidget WS
guard, test hygiene), 4 refuted with evidence, rest documented as
deliberate design or deferred.
Deterministic CI replaces hand-enforcement: tests/test_changelog_style.py
(quiet one-liner format) and tests/test_locale_parity.py (21-locale
key/placeholder lockstep with a ratchet baseline) — the latter surfaced
and fixes 151 already-broken locale strings. CodeRabbit/Greptile carry
the house rules via .coderabbit.yaml + greptile.json; CLAUDE.md gains
the harvest-before-merge and never-accept-as-is rules.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- validate managed binaries before exec; 0-byte GGUF placeholders fail
actionably instead of "Exec format error" (#1172)
- KittenTTS: tokenizer-measured chunking to the ONNX 512-token cap;
clear 400 for unspeakable input (#1173)
- clean SIGTERM during weight load: shutdown-aware loader, benign
cancelled-load classification, lifespan hardening, scoped log
silencers (transformers load + alembic fileConfig) (#1174)
- broken ASR deep-imports (lightning_fabric) mark the engine
unavailable with a repair hint and fall through (#1185)
- uv cache + managed Python follow the chosen install drive on
Windows (cherry-picked cross-drive class fix + spaces/D: tests) (#1186)
- adaptive silence-removal ladder for quiet clone references; localized
actionable error for truly silent clips, all 21 locales (#1188)
- CHANGELOG: consolidated Unreleased into the quiet one-liner style
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The curated default `sherpa-parakeet-tdt-v3` installs cleanly, loads without
error, and returns an empty token list for clear speech — while whisper and
zipformer transcribe the same bytes. Ruled out: quantisation (fp32 fails too),
sherpa-onnx version (1.13.3 and 1.13.4), decoding method (greedy and
modified_beam), and model_type (explicit and auto-detect). The defect is inside
sherpa-onnx's NeMo-TDT decoder and cannot be fixed here by configuration.
The previous commit made that survivable — a silent model falls back to the
capture ASR engine for the session. This makes it stop recurring.
Hard-coding a different default per OS would be a guess: there is evidence for
Windows only, and blanket-changing the curated pick could regress a platform
that works. So the app observes rather than predicts. When a session hears
speech-level audio and the model returns nothing, that model is demoted ON THIS
MACHINE and is no longer auto-selected; dictation routes to the capture ASR
engine instead. Self-correcting wherever the breakage actually is, a no-op
everywhere it isn't, and no round trip repeated on every later session.
Demotion is never a dead end: explicitly choosing a model in Settings clears it,
so a sherpa upgrade that fixes the decoder is reachable and the user stays in
charge.
Verified end-to-end on a clean backend: with parakeet pinned and no demotions,
one speech session demoted it ("demoted on this machine" logged), after which
dictation_model_id() returns None (capture ASR engine) while the pref is
untouched — and re-picking the model restores it. 12 tests cover the round
trip, per-model scoping, and the user-regains-control path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two bugs that together made dictation look completely dead on Windows.
1. Silent model. The curated default `sherpa-parakeet-tdt-v3` downloads, loads
with zero errors, and is correctly detected as a TDT model — then returns an
empty token list for clear speech. Measured on one 18.9s WAV, same machine,
same sherpa-onnx:
sherpa-whisper-tiny -> "Alright, here we are. I hope that's all..."
sherpa-zipformer-en-20m -> "ANTS BOTH IN WHAT DISGUISED THIS THAT..."
parakeet-tdt-v3 int8 -> '' <-- the default
parakeet-tdt-v3 fp32 -> ''
parakeet-tdt-v2 int8 -> ''
Ruled out: quantisation (fp32 fails too), sherpa-onnx version (1.13.3 and
1.13.4), decoding method (greedy and modified_beam), and `model_type`
(explicit and auto-detect). The fault is inside sherpa-onnx's NeMo-TDT
decoder, so the app cannot fix it by config — but it must not present it to
the user as "dictation is broken". The offline handler now tracks whether it
ever heard speech-level audio; if the model returns nothing anyway, it falls
back to the capture ASR engine for that session and returns a `warning` +
`model_silent` naming the model that failed. The pref is left alone, so the
user stays in control. A quiet user still yields a quiet result.
2. Stranded pill. Error states deliberately never auto-dismiss so a failed
paste keeps the transcript on screen. But the mic / model-missing / server /
connection paths have no transcript to rescue, so they parked the widget on
the user's screen forever ("the dictation bubble is permanently sticking
when it's not used"). One effect now dismisses an error pill that carries no
text, covering every existing path and any added later; delivery failures
keep their text and their stay-up behaviour.
`is_model_silent` is extracted so the decision is unit-tested rather than
buried in the socket handler.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Dubbing's vocal-separation step died with `TypeError: An asyncio.Future,
a coroutine or an awaitable is required` whenever the backend ran on an
event loop without native async-subprocess support — notably the Windows
`SelectorEventLoop` that uvicorn forces under `--reload` (`bun desktop`).
On that loop `spawn_subprocess` returns a thread-backed `_AsyncCompatProc`
whose `.stderr` is a plain SYNC pipe, but `run_proc_streaming_stderr`
unconditionally did `await asyncio.wait_for(p.stderr.read(256), ...)`. A
sync `.read()` returns bytes, not a coroutine, so `wait_for` raised. The
extract step survived only because it uses `communicate()` (async-wrapped).
The wrapper now advertises `uses_sync_pipes`; the streaming reader keys off
it and, on that loop, runs the process to completion via the wrapper's async
`communicate()` and replays stderr as the same `('stderr', line)` events —
no live progress bar on the degraded loop, but demucs actually runs. The
native async path (Proactor/posix — every release build, macOS, Linux) is
untouched. Regression test covers the fallback (fails before, passes after)
and pins the native path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When the backend is torn down while the TTS model is loading — the user closes
the window during first-load, uvicorn is told to stop, or a port bind fails —
transformers' weight materializer (which runs in its OWN thread pool) raises
`RuntimeError: cannot schedule new futures after interpreter shutdown`. That was
caught by the generic model-load handler and logged as "Model loading failed"
with a full traceback + a phantom /model/status error, so a normal shutdown
looked like a crash in the backend-crash report.
Recognise the interpreter-shutdown teardown (walking __cause__/__context__,
since transformers' lazy-import machinery buries it several layers deep, and
matching "interpreter shutdown" specifically so an ordinary single-pool reset
is NOT silenced) and log it calmly instead. Unit-tested.
Note: this is about how a shutdown-during-load is *reported*; the port-bind
collision that triggered it in testing is a separate, multi-instance race the
Rust bootstrap's attach-or-reclaim logic already prevents in normal use.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The bootstrap fix hid the console windows the Rust shell spawns, but the running
backend is a second, larger source: it makes 70+ subprocess spawns (ffmpeg for
dubbing, engine sidecars, yt-dlp, demucs, WhisperX), only one of which set any
creation flag — and third-party libraries (imageio-ffmpeg, yt-dlp) shell out to
ffmpeg themselves. Since the backend runs console-less, on Windows each console
child got a fresh cmd window flashed on screen during generation/dubbing.
Patch subprocess.Popen once at backend startup (before anything spawns) to OR
CREATE_NO_WINDOW into every child's creationflags — covering run/call/
check_output (all built on Popen) and every stdlib-using library, from a single
auditable choke point. Composes with the existing CREATE_NEW_PROCESS_GROUP and
honours an explicit CREATE_NEW_CONSOLE. No-op on macOS/Linux. Unit-tested.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Self-hosting behind a reverse proxy or on a LAN used to force a blunt choice:
OMNIVOICE_SERVER_MODE (trust every non-loopback source) or the API-key/PIN gates
— which a proxy that strips the Authorization header breaks for browser clients
entirely.
Add OMNIVOICE_TRUSTED_NETWORKS (comma-separated CIDRs) whose addresses are
treated as trusted by the CONSUMPTION gates (PIN/API-key middleware, dictation
WebSocket) via is_local_host — a LAN/proxy client is exempted from consumption
auth. Admin routes (require_loopback → /system/set-env, /api/settings/*) stay
true-loopback-only (is_loopback, not is_local_host) to preserve the two-tier
privilege model: consumption trust ≠ admin trust (RCE-class surface). Opt-in,
default empty → zero behavior change. The granular companion to server-mode (#261).
Tests: is_loopback / is_local_host / require_loopback contract for trusted CIDRs,
adjacent subnets, malformed entries, the two-tier split (trusted-network rejected
by the admin gate), and the default (no-trust) case. Docs + CHANGELOG.
Only the TTS model (~2.4 GB) is required on first run; ASR models are
per-platform curated picks (curated_on in models.yaml) installed on demand.
Every transcription surface returns a typed asr_model_missing error with a
one-click download CTA instead of silently pulling multi-GB Whisper weights.
Settings -> Models is a grouped, platform-aware catalog. New guided
permissions UX (wizard System Check + Settings -> Permissions + mic
pre-flight) with native mic-state checks and OS settings deep-links. New
parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated
capture preference so multilingual dictation never regresses). Docs:
expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ.
Hardening: preflight fails open for custom model pins, ROCm curation no
longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged
consent key, CaptureWidget setup race fixed, offline-cache CI simulation
fixes so empty-cache runners stay green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>