d91beef0fd314250d8d9b94de86dfea019a8bd96
20
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
549a56fc2d | fix: remove Windows sidecar supervisor hop | ||
|
|
4e5e8d1f89 |
Allow OmniVoice slow sidecar startup (#1743)
Fixes #1711.\n\nGives only the OmniVoice subprocess a 120-second readiness budget while retaining the shared 30-second default for all other sidecars, with regression coverage. |
||
|
|
a8371baaa8 |
fix(desktop): serialize backend lifecycle (#1635)
Closes #1635. |
||
|
|
b79ba9bd3b |
docs(readme): lead with download + first clone; seed benchmarks page (#1555)
* docs(readme): lead with download + first clone; seed benchmarks page Quickstart (installers, install guides, a three-step first-clone walkthrough) moves above What's-new/Features in both READMEs — visitors get the action before the pitch. New docs/benchmarks.md anchors measured per-engine/device numbers on the bench_pipeline.py harness, community-contributed, no estimates. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): entry for the README conversion restructure (#1555) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): emit RTF + CUDA peak VRAM; guard NaN RAM; define the benchmarks schema Bot harvest on #1555: the tts stage now prints RTF per warm measurement and CUDA peak VRAM (None elsewhere — no made-up zeros), the stage floor refuses unmeasurable RAM instead of sailing past a NaN comparison (FLOOR_GB=0 overrides), docs/benchmarks.md columns map 1:1 to what the harness prints, and the download badges say they open the release page. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): link palash.dev from the maker section Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): name the resolved engine, track VRAM from resolution, comment the guards Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): the quick-switch gif is the hero image The hero shows motion now; the Launchpad screenshot moves into the 0.5.0 What's-new slot so nothing appears twice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): peak VRAM is reserved memory; adapter engines name their model Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): subprocess-isolated engines report VRAM n/a, not a parent-side zero Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): out-of-process detection is declarative; sherpa rows name their model 'runs_out_of_process' is now a TTSBackend attribute set by SubprocessBackend AND omnivoice-gguf (which inherits TTSBackend directly but spawns a binary per generate — the isinstance check missed it). Duck-typed for the same module-purge reason as _is_subprocess_isolated. Sherpa-onnx identity comes from _model_dir's basename when _model_id is absent. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bench): backends self-report model identity via TTSBackend.model_identity() Greptile enumerated the adapter engines one at a time (mlx _model_id, sherpa _model_dir, cosyvoice env-only) — the attribute sniffing rots per engine. The hook fixes the class: each multi-model backend reports its own identity, the profiler just asks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
41722afe3b |
refactor(launchpad): quieter, borderless design refresh (#1515)
* refactor(launchpad): quieter, borderless design refresh The launchpad carried decoration from an earlier direction: icon chips, corner-hung count badges, a permanently visible filled arrow, uppercase mono card titles, and a dotted stipple divider — plus a frame that had been invisible since the app-wide border tokens were zeroed. Rework it around what the borderless direction actually implies: - Feature tiles get a whisper-faint surface instead of a dead frame, and read as three bands (bare glyph + count / title + arrow / description). `--card-hue` is spent sparingly — the glyph at rest, the surface, count and arrow only once raised. Titles move to sans sentence case; counts are plain tabular numerals. Lift softened 4px -> 2px, coloured glow -> neutral shadow, plus an explicit focus ring and a staggered entrance. - Hero drops the boxed "646" pill and the filled A/B-Compare button for quiet type, with a hairline standing in for the separation. - Section labels trade the dotted stipple for a single fading hairline; rows are transparent until hover and reveal "Open" on hover/focus (it stays in the DOM, so AT and keyboard always reach it). - Hero, tiles, recent files, callout and project lists now share one 1180px column — previously only the top half was capped, so lists ran edge-to-edge on a wide display while the deck stayed centred. Two bugs found and fixed while doing it: - Buttons that had `border border-solid border-transparent` removed fell back to the UA default border and rendered a visible 1px outline. They now carry `border-0` explicitly. - `.lp-animate` used `animation-fill-mode: both`, so after the entrance it kept owning `transform` — and animation-origin declarations outrank normal ones, which silently killed the card hover lift. Now `backwards`, which still holds the from-state through the stagger delay. Also drops CSS the page has not rendered since #904: the cursor-spotlight layer, the breath ring, and the per-card waveform strip. Verified with headless renders at 1600/1280/940 and the empty state. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(dictation): decode Wayland portal signals and show the capture pill The GlobalShortcuts portal declares Activated/Deactivated as (o session, s shortcut_id, t timestamp, a{sv} options). We decoded the timestamp as u32, so zbus rejected every signal with Signature mismatch: got `(osta{sv})`, expected `(osua{sv})` and the press was dropped as an invalid signal. Registration succeeded and the desktop even reported the bound chord back, so the hotkey looked wired up while doing nothing at all — on every Wayland compositor, for the whole life of the feature (#1490). Decode the 64-bit timestamp, and keep the 32-bit spelling as a fallback so a non-conforming portal degrades to working rather than to silence. With presses arriving, the second half of the failure showed: nothing had shown the widget window since it became a hidden recorder host, so a capture ran with no pill on screen — and a mic or Accessibility failure rendered into a window nobody could see. Add show_dictation_pill, which bottom-centres the capsule on the monitor under the pointer and shows it without taking focus (Windows keeps SW_SHOWNOACTIVATE so paste still lands in the user's document), and call it from the widget for every state but idle. Wayland denies clients their own placement, so the compositor picks the spot there; the pill still appears. dispatch_dictation_capture now logs whether a press was emitted or queued — a press that reaches Rust and produces nothing was otherwise indistinguishable from one the compositor never delivered. Tests: portal signals decode at both timestamp widths (the 64-bit case fails before this change with the exact production error); pill placement centres, respects a second monitor's origin, and clamps rather than going off-screen; the widget shows for a state needing the user, stays hidden while idle, and never shows for a press that arrives while dictation is disabled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: sync in-progress workspace changes Uncommitted work already in the tree, checkpointed so the branch matches the local machine: - Remote GPU workers: join-from-the-app flow, one-time secrets, QR join codes, a Compute control in the status bar, and the device-list Workers panel (#1516) - Model Catalogue workspace, with Settings pointing at it - Settings sidebar search and keyboard navigation - Demo assets for dubbing, dictation and voice design, plus the scripts that render them - Backend: validation-error handling, ASR request-path degradation, and the accompanying tests - CHANGELOG entries for the above and for the Wayland dictation fix Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tests): follow Engines to the Model Catalogue, and green the sweep - test_supertonic3 asserted the license gate points at "Settings" while the engine now names Model Catalogue → Engines, which is where the accept button actually lives. The assertion follows the move; what it pins is unchanged — the hint must name a place the user can reach it. - Carries the CJK allowlist entries for the rendered dub bundle (#1517) and the regenerated route snapshot for /workers/agent (#1516), both of which this branch inherits from the workspace sync. - docs/install/linux.md: the dictation capsule is bottom-anchored everywhere except Wayland, where the protocol gives applications no say in their placement. Documented rather than left as a surprise (CodeRabbit). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: stop a flaky dependency fetch from failing green runs en-core-web-sm resolves to a direct GitHub release URL, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own three retries all land within the same few seconds and fail together, so the whole job dies on a dependency that has nothing to do with the change under test — it cost #1518 and #1517 an otherwise-green run tonight. Two changes: back off between whole `uv sync` attempts, which is what actually clears it, and pass --no-sync to the pytest steps. `uv run` re-resolves the environment before running, so every test step was a fresh chance to hit the same fetch even though the install step had already synced — that is exactly how #1518 failed, in the isolated backend/tests step, with all 5467 tests already passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: one retry seam for every uv sync, not just the job that failed last en-core-web-sm resolves to a direct GitHub *release* URL rather than a package index, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own retries all land inside the same ~10 seconds and fail together, so a job dies on a dependency unrelated to the change under test. Tonight that cost four otherwise-green runs across #1515, #1517 and #1518 — and the first fix only covered the Tests job, so the next failure simply moved to Smoke (Linux), which syncs separately. The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh backs off between whole attempts (15s, 45s, 90s) and every workflow that syncs now goes through it — ci.yml (tests + the platform matrix), release.yml, security.yml, evals.yml. It still fails loudly after four attempts, so a genuinely broken lockfile is not disguised as a flake. The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke matrix has always set, which is part of why it was the one that kept dying; it has them now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ci): pin the Intel-Mac contract by intent, not by command spelling test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted the literal line `run: uv sync --extra pockettts`, so routing every sync through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The contract it exists to protect is that the pockettts extra installs ONLY on backend_supported legs — which the regex now pins, while leaving how the sync is invoked free to change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: keep every uv run out of the resolver, and bound the retry budget CodeRabbit, #1517: - `uv run` re-resolves before running, so the smoke suite, the worker-artifact tests, the release test run and the eval run were each a fresh chance to hit the flaky direct-URL fetch outside the retry loop. All of them pass --no-sync now; the environment is already synced by the step that owns the retries. security.yml's `uv run --with pip-audit` is deliberately left alone — it layers an ephemeral package rather than running the project's own tests. - The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a 120 s timeout on the smoke matrix). Three attempts and 60 s of total backoff outlast the refusals actually observed while staying well inside the jobs' timeout-minutes. - The Intel-Mac contract test pinned the smoke command literally too, so --no-sync tripped it exactly like the sync line did. Same fix: assert the contract (smoke runs only on backend_supported legs), not its spelling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
77d6507318 | fix(pockettts): recheck consent after synthesis queue | ||
|
|
7c64a7270c | fix: make resolve heartbeat shutdown race-free | ||
|
|
e80fcab8fa |
fix(engines): join the resolve heartbeat instead of only signalling it
_beat() can be past its stop.wait() and already committed to a write when the context exits. Signalling alone lets that write land after _run_on_gpu_pool's _job pops the ident — the pop exists so a stale beat cannot vouch for a later job on the same reused worker ident, and a post-pop write resurrects exactly what it was there to prevent. The wedge detector then reads a heartbeat the next job never emitted and keeps extending a stuck one. Joining orders the last write before the pop. Bounded, so a wedged writer degrades to the previous behaviour rather than blocking. The regression forces the interleaving with a parking map rather than waiting on the scheduler, so it fails deterministically without the join. |
||
|
|
c46ad1bed8 |
fix(engines): a venv resolution that takes minutes is not a stalled generation (#1414)
_spawn() calls venv_python(), and on a cold first run that is not cheap: the probe spawns each candidate interpreter to import the engine, and if none is installed it runs the whole uv venv + uv pip install bootstrap — bounded at 900s by design, because installing torch takes minutes. All of it happens on a GPU-pool worker inside a generate request whose execution budget is 300s. Nothing along the way reported progress, so the budget expired part-way through the install and the job was abandoned. The first generation that triggers a bootstrap could never succeed, on any hardware, and the message blamed the hardware anyway. The sidecar's own cold load already heartbeats for exactly this reason (#1367); resolution is the step before it that never did. Pool jobs only, and crediting the resolving thread rather than the beater — an off-pool ident is not tracked by the clock, and a pool worker reusing it would inherit unearned extension (#1379). |
||
|
|
5cab8e0149 |
feat: rename the product to VoiceStudio (previously OmniVoice-Studio)
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
|
||
|
|
08a1b8f92c |
fix(generate): a healthy model download extends the budget it was blowing (#1367) (#1379)
Every subprocess engine cold-loads its model inside the synthesize handler, so a first-use generate on a slow connection spent its whole 300s execution budget downloading — then failed blaming the hardware, while the sidecar's watchdog was being fed progress frames the entire time. The outer clock now listens to that evidence: SubprocessBackend forwards each sidecar progress frame to a per-worker-thread heartbeat (pool jobs only, cleared when the job ends), and the guarded waiter extends the deadline past the soft budget only while heartbeats stay fresher than MODEL_LOAD_HEARTBEAT_GRACE_S, bounded by MODEL_LOAD_EXTRA_TIMEOUT_S. The extension is logged once. A silent job still dies at the original deadline; stopped heartbeats kill within the grace; another thread's heartbeat is no alibi; caller cancellation cancels and consumes the abandoned future. 11 regression tests, headline case verified failing before. |
||
|
|
0e9733bc16 |
fix(engines): path-aware GPU-pool slot in SubprocessBackend.generate() (on-pool skip + off-pool hold) (#1298)
* fix(engines): make SubprocessBackend.generate() path-aware on the GPU pool generate()'s slot handling had two bugs: 1. On-pool self-deadlock: /v1/audio/speech and /generate (and audiobook, dub, batch) dispatch generate() via run_on_gpu_pool_guarded, already on a pool worker, so the inner slot submit queued behind the very job running it on a 1-worker (MPS) pool and timed out before the sidecar spawned. Every subprocess engine surfaced the in-process 300s-abandon instead of synthesizing. 2. Off-pool no hold: the off-pool slot was a bare no-op that released the worker before _spawn(), so off-pool callers (engine self-test, diagnostics) could synthesize concurrently with a pool job and over-subscribe the GPU. Make the slot block path-aware: on-pool callers skip (the outer run_on_gpu_pool_guarded already holds _running for the whole sidecar exchange); off-pool callers hold a real slot for the whole synthesis via an _occupy task that blocks the worker until _held is set in the finally. Single release point in the finally. Regression tests: generate dispatched on a pool worker (on-pool skip) and a concurrent pool job blocked during an off-pool generate (off-pool hold). Both verified fail-before / pass-after. Supersedes #1296 (on-pool-skip-only). Closes #1295, #1297. * Address review: couple on-pool skip to the pool prefix; fix comment /simplify + /code-review flagged that the on-pool skip keyed on the literal "gpu-pool" string, decoupled from _build_gpu_pool's thread_name_prefix. A rename would silently re-introduce the exact self-deadlock this PR fixes (and the tests can't catch it, since they hardcode the prefix). Centralise the prefix in _GPU_POOL_THREAD_PREFIX + a running_on_gpu_pool() helper, used by _build_gpu_pool, the skip in generate(), and _heal_tts_placement. Also fix the comment: the Settings engine self-test rejects subprocess-isolated engines with a 400, so the only real off-pool caller is the diagnose.py deep-synth probe. * fix(engines): bind slot_future before the off-pool branch CodeQL py/uninitialized-local-variable (error, blocking CI). `_held is not None` does imply slot_future was assigned, so the current code is correct — but the two are only coupled by convention, which the analyser cannot see and a third exit path would quietly break. Binds it to None up front and guards the cancel. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(engines): make the slot-hold regression deterministic and leak-free CodeRabbit, valid on both counts. The test used sleep(0.8)/sleep(0.5) as synchronization — the tests/** contract forbids it, and on a slow runner the marker could be enqueued before the generator had reserved anything, so the assertion passed for the wrong reason. It now waits on an event signalled when the slot task actually starts, and asserts "did not run" via a result() timeout rather than a bare sleep. Cleanup moved into finally: an assertion failure used to leak the sidecar process and the pool thread into the rest of the session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
47c51698c9 |
feat(engines): add omnivoice-subprocess, a crash-isolated (killable) TTS engine (#1292)
* feat(engines): add omnivoice-subprocess, a crash-isolated TTS engine The default in-process OmniVoice engine runs on the GPU ThreadPoolExecutor. When a generate or load exceeds its execution budget the pool is "reset", but the abandoned worker thread cannot be killed (Python cannot interrupt a native torch/MPS call), so it keeps holding the device until it finishes on its own and later synths queue behind it and hang. The reset restores pool capacity but not the device. This is the residual root cause behind the closed #730 and #1190: the messaging/reset mitigations address the symptom, not the device-holding zombie. Add an opt-in `omnivoice-subprocess` engine that runs the same model in a child process via SubprocessBackend. A child process can be hard-killed: on a recv-timeout the watchdog calls proc.kill(), reclaiming VRAM/device, and the next request transparently respawns a fresh sidecar. The in-process engine remains the default, so existing users see no change; this is an opt-in for unattended / scheduled / reaction-triggered synthesis where a stuck job must self-recover instead of hanging until a manual restart. Base-class and mitigation changes that ship with it: - SubprocessBackend.generate() now consumes non-terminal {"op":"progress"} frames a sidecar emits during a cold load (previously the first cold generate after spawn failed, then worked on retry). Additive: engines that reply with audio directly are unaffected. - recv_timeout_s is overridable per engine (default 60s unchanged); the new engine sets it to the generate budget so a long-but-valid synth is not falsely killed while a wedged one still is. - make_room_before_generate(): free idle GPU memory before a warm, heavy generate. The cold-load path already evicted; the warm path skipped it, so a long synth on a VRAM-tight MPS box could contend its way into the budget. Verified end-to-end against the live model (cold / warm / recovery-after-kill) and under a sustained + concurrent-pressure soak: killed-worker recovery 5/5, chunked long text 9/9, no memory leak. * Address review: install_hint + move make_room into get_model - Add `omnivoice-subprocess` to `_INSTALL_HINTS`; the test_install_hints_cover_all_registered_backends gate requires every registered backend to carry one (this was the CI failure). - Move the warm-generate VRAM eviction out of the /generate and /v1/audio/speech routes and into get_model()'s warm-return path, so EVERY native TTS generate is covered (REST, WS TTS, dub, batch, audiobook), not just the two REST routes. Drops the now-redundant per-route wiring. (Greptile P1: the per-route placement missed the other generation surfaces.) * Address review: drop dead long-text eviction path; log probe failure - _should_make_room_for_generate: the long-text headroom boost became dead code once the eviction moved into get_model() (which has no text), so the long-text branch never fired. Removed the text param, the long-text threshold/multiplier branch, and the now-unused _env_float helper. The core RAM-tight gate (the part that matters on a starved box) is unchanged. - Log the available_memory probe failure at debug instead of silently swallowing it (CodeRabbit: silent swallow breaks the debug trail). - Tests updated for the text-agnostic policy. * fix(engines): stop subprocess generate() self-deadlock on 1-worker pools SubprocessBackend.generate() acquires a GPU-pool slot for accounting, but /v1/audio/speech and /generate dispatch backend.generate() via run_on_gpu_pool_guarded, i.e. already ON a pool worker. On a 1-worker pool (MPS) the inner pool.submit queued behind the very job running it and slot_future.result(timeout=10) raised before the sidecar ever spawned, so omnivoice-subprocess (and every other subprocess engine on MPS) surfaced the in-process 300s-abandon instead of synthesizing. Skip the slot acquisition when current_thread() is already a gpu-pool worker; the outer guard already accounts for the slot. Direct callers (off the pool) still acquire one. Regression test added (generate on a pool worker). * Address review: reword slot-skip comment (fixes watermark-coverage CI) + simplify - The slot-skip comment said "dispatch backend.generate() via", and test_watermark_route_coverage's _SYNTH_CALL regex matches the literal backend.generate( anywhere in a module, so it counted subprocess_backend.py as a synthesis producer that must reference mark_synthetic (it doesn't — the routes apply mark_synthetic; the engine sits below the chokepoint, like tts_backend.py). Reworded to "dispatch generate() via". - Fold in the simplify refinement: single negated predicate, import+pool moved into the acquire branch. |
||
|
|
83f943bead |
fix: bot-review harvest (16 findings) + deterministic style/locale CI + reviewer configs
Harvested and verified every CodeRabbit/Greptile finding from PRs #1175, #1189, #1192, #1195: 16 real ones fixed (fallback ASR preflight bypass, VRAM release on stream exit, typed 409 parity, uv env independence, path-privacy in errors, MCP clone_voice hardening, CaptureWidget WS guard, test hygiene), 4 refuted with evidence, rest documented as deliberate design or deferred. Deterministic CI replaces hand-enforcement: tests/test_changelog_style.py (quiet one-liner format) and tests/test_locale_parity.py (21-locale key/placeholder lockstep with a ratchet baseline) — the latter surfaced and fixes 151 already-broken locale strings. CodeRabbit/Greptile carry the house rules via .coderabbit.yaml + greptile.json; CLAUDE.md gains the harvest-before-merge and never-accept-as-is rules. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
933743e336 |
fix: resolve open issue batch — #1172 #1173 #1174 #1185 #1186 #1188
- validate managed binaries before exec; 0-byte GGUF placeholders fail actionably instead of "Exec format error" (#1172) - KittenTTS: tokenizer-measured chunking to the ONNX 512-token cap; clear 400 for unspeakable input (#1173) - clean SIGTERM during weight load: shutdown-aware loader, benign cancelled-load classification, lifespan hardening, scoped log silencers (transformers load + alembic fileConfig) (#1174) - broken ASR deep-imports (lightning_fabric) mark the engine unavailable with a repair hint and fall through (#1185) - uv cache + managed Python follow the chosen install drive on Windows (cherry-picked cross-drive class fix + spaces/D: tests) (#1186) - adaptive silence-removal ladder for quiet clone references; localized actionable error for truly silent clips, all 21 locales (#1188) - CHANGELOG: consolidated Unreleased into the quiet one-liner style Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
dc1d36fe5f |
refactor(models): model-management v2 cleanup (mm2, all tiers) (#428)
One coherent lifecycle surface over the in-process model, diarization, and subprocess sidecars; fixes the engine-switch VRAM leak; tightens download robustness. Backend-only, response shapes preserved, no new deps. Tier 1 — correctness: - MM2-01: get_active_tts_backend() caches one instance per backend id and unload()s the outgoing engine on switch (fixes the VRAM leak behind #278); adds reset_active_backend(). - MM2-02: OmniVoiceBackend.unload() releases the shared model_manager singleton + free_vram(); SubprocessBackend.unload() -> unload_sidecar(self.id), inherited by all sidecar engines. Idempotent + preload-safe. - MM2-03: /model/loaded ASR row reports the real device + a note explaining the disabled unload button. Tier 2 — single surface: - MM2-04: new services/model_lifecycle.py owns list_loaded/unload/unload_all/ free_vram; system.py routers are thin delegations (shapes unchanged). - MM2-05: idle timeouts (in-process + sidecar) resolve via prefs.resolve (env wins, no restart); removed the duplicated _IDLE_TIMEOUT_SECONDS. Tier 3 — robustness/observability: - MM2-06: _install_cooldowns swept (1h TTL) + cleared on success — bounded. - MM2-07: per-extension weight floors (onnx 64KB, tensors 5MB) OR the original >=5MB catch — small ONNX no longer false-flagged, #352 still caught. - MM2-08: indextts GPU sidecar self-reports vram_mb in pong; parent surfaces it in list_live_sidecars (0 = CPU/unmeasured). - MM2-09: is_cached scan_cache_dir->disk fallback logs WARNING w/ exc type (#117/#118), was invisible at DEBUG. Tests: tests/test_mm2_lifecycle.py (15). Full suite: 1379 passed. Plan/summary: .planning/quick/260613-mm2-clean-model-management-v2/. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
599f3bcc5c |
feat(engines): on-demand unload of subprocess-engine sidecars (Action 13) (#406)
Completes the dynamic engine load/unload slice. The idle reaper (#401) frees sidecar VRAM after 5 min; this adds a user-initiated "free VRAM now" path so multi-engine users don't have to wait: - subprocess_backend: `list_live_sidecars()`, `unload_sidecar(id)`, `unload_all_sidecars()` via a shared `_force_reap(predicate)` — busy-guarded exactly like the idle reaper (non-blocking lock; a sidecar mid-synth is skipped, never interrupted; next request respawns it). - system.py: `/model/loaded` now surfaces live sidecars as unloadable rows; `/model/unload/{sidecar:<id>|sidecars}` frees one or all. The existing generic flush panel picks these up with zero frontend change. Also refresh CLAUDE.md stale version notes: main is 0.3.6 (latest release v0.3.5 + 1 patch); the v0.3.0-as-unreleased framing in the project/cadence notes is corrected to the v0.3.x continuous-to-main reality. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
34c8ab2409 |
feat(engines): idle-reap subprocess-engine sidecars to free VRAM (Wave 13) (#401)
Parity Action 13 (dynamic load/unload), subprocess-engine half. A subprocess engine's sidecar holds a process — and, for GPU engines, VRAM — for the life of the backend, even after the user switches engines. The default in-process OmniVoice model already idle-unloads (model_manager.idle_worker); this gives the subprocess engine class the same treatment. subprocess_backend gains a background reaper (lazy daemon thread, started on first spawn) that shuts down sidecars idle past OMNIVOICE_SIDECAR_IDLE_TIMEOUT_S (default 300 s; <= 0 disables). The next request transparently respawns one via the existing dead-process relaunch. Safety: the reaper only acts while holding the per-backend lock acquired NON-blockingly, so it can never run mid-op — if an op holds the lock it skips that backend this round. Reuses the idempotent shutdown() (which doesn't take the lock, so no re-entrancy). Each backend tracks last-use and registers in a weak live-set. Scope: subprocess engines only (the heavy, VRAM-holding, process-isolated class). In-process non-default engines and cross-engine VRAM preemption remain TODO — get_active_tts_backend returns a fresh instance per call, so those need an instance-tracking refactor. 6 reaper tests via the stdlib echo sidecar (no torch): kills idle, respawns, skips busy (lock held), recent-use kept, disabled at <=0, ignores dead. The 3 subprocess suites pass together (24). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e862f0faf0 |
feat(asr): crash-isolated faster-whisper subprocess backend (Wave 4.2) (#393)
* feat(asr): crash-isolated faster-whisper subprocess backend (Wave 4.2) Native ASR engines (faster-whisper / CTranslate2) can segfault on GPU teardown — a process-level crash that kills the whole backend. Running the engine in a child process turns that into a failed job: the sidecar dies, the parent raises a decorated error (engine id + device), and the next request respawns a fresh sidecar. - services/subprocess_asr.py: SubprocessASRBackend reuses SubprocessBackend's wire protocol + lifecycle — including respawn-on-dead-process (_spawn relaunches when the child isn't alive) and GPU-slot acquire/release — adding a 'transcribe' op (the TTS 'generate' surface is stubbed). IsolatedFasterWhisperBackend wraps faster-whisper using the PARENT venv (already a dep — only the process boundary is new); opt-in via OMNIVOICE_ASR_BACKEND=faster-whisper-isolated. - engines/_asr_sidecar/main.py: the faster-whisper runner (stdlib wire protocol; torch/CT2 import lazily so the ready handshake fits the timeout). - engines/_echo/main.py: a 'transcribe' echo op so the round-trip + crash recovery are testable without a real engine. - asr_backend._REGISTRY is now a lazy dict (mirrors the TTS registry) so the isolated backend lists/resolves without importing the subprocess stack unless selected. Tests (echo sidecar, stdlib-only): round-trip, single long-lived sidecar across calls, crash-mid-transcribe → decorated error + backend healthy + next call respawns, registry exposure, generate-not-supported. Spec 7 / parity program Wave 4.2. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(asr): deterministic crash test + drift marker for lazy ASR registry (Wave 4.2 CI) CI surfaced two issues: - The echo crash test relied on the crash-AFTER-reply hook, whose reply may still reach the parent (timing-dependent) — and a leaked OMNIVOICE_ECHO_CRASH from a sibling subprocess test poisoned the non-crash tests. Fix: a deterministic OMNIVOICE_ECHO_CRASH_NO_REPLY hook that exits BEFORE replying (guaranteed dead pipe → decorated error), and the asr fixture clears both crash envs so the round-trip/two-call tests can't inherit a leak. - check-docs-drift's _ASR_MARKER didn't match the new lazy registry line (_LazyASRRegistry({); updated the marker + the self-test fixture. Verified the no-reply crash hook by driving the sidecar directly (reply=None, exit 1); drift self-test + real-repo check green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(asr): allowlist the 'segments' op so transcribe replies aren't dropped (Wave 4.2 CI) The parent's PARENT_INBOUND_OPS frozenset gated inbound sidecar frames but never included 'segments' — the ASR transcribe reply op. _recv() dropped the frame as disallowed, tail-recursed, hit EOF, and returned None, so every transcribe surfaced as a bogus 'sidecar crashed mid-transcription'. TTS ('audio') was allowlisted; ASR ('segments') was missed. Add it (and list 'transcribe' in the informational SIDECAR_INBOUND_OPS), update the exact-shape allowlist test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0fc5ea6cf3 |
Phase 2 Plan 02-01: SubprocessBackend primitive (Wave 1 of Phase 2) (#97)
* Phase 2 Plan 02-01: SubprocessBackend primitive + echo sidecar + ENGINE-05 wrap
Lands the durable SubprocessBackend primitive — the architectural keystone
that Plans 02-03 (IndexTTS migration), Phase 3 (Supertonic-3), and
Phase 4 (GGUF / Singing) plug into.
Files added:
- backend/services/subprocess_backend.py — base class owning spawn,
shutdown, _send/_recv (length-prefixed JSON), GPU-slot acquire-release,
atexit teardown, stderr drain, op allowlist (T-02-04), and 64 MB
frame cap (T-02-01). No multiprocessing — subprocess.Popen
exclusively so subclasses can target a *different* venv's interpreter
(Locked Decision D4 / Pitfall 1).
- backend/engines/_echo/main.py — permanent CI regression sidecar.
Stdlib-only, runs under the parent's sys.executable. Implements
ready/ping-pong/synthesize/shutdown plus test-only probe_env and
emit_unknown ops for env-forwarding and op-allowlist tests. DO NOT
DELETE — the round-trip test depends on this file.
- tests/backend/services/test_subprocess_backend.py — 13 tests:
round-trip, health_check, no-zombie, shutdown idempotency, env
forwarding (HF_TOKEN/HF_HOME/HF_ENDPOINT/HF_HUB_CACHE), oversize
frame, short read, op-allowlist drop, op-allowlist constant shape,
sidecar-crash recovery, no-multiprocessing grep gate, MAX_FRAME_BYTES.
- tests/backend/services/test_tts_backend_registry.py — 6 tests for
list_backends() resilience + shape + isolation_mode + last_error
caching + existing-engines preservation + install_hint passthrough.
Files modified:
- backend/services/tts_backend.py:
* Adds module-level _LAST_ERRORS dict for ENGINE-06.
* Rewrites list_backends() to wrap each is_available() in try/except
so one broken engine cannot blank the picker (ENGINE-05).
* Adds last_error + isolation_mode keys to each response entry
(ENGINE-06 UI in Plan 02-04 consumes via the same /engines route).
* Uses a duck-typed _is_subprocess_isolated marker rather than
issubclass(cls, SubprocessBackend) because test fixtures (token
resolver suite) purge sys.modules["services"] between tests and the
re-imported SubprocessBackend would be a different class object.
Threat-model mitigations (Plan 02-01 frontmatter):
T-02-01 DoS via length-prefix → MAX_FRAME_BYTES = 64 * 1024 * 1024
T-02-02 GPU slot leak on sidecar death → try/finally in generate
T-02-03 token bytes in stderr → drained via parent logger
(HFTokenRedactor from Phase 1 already on root)
T-02-04 unknown ops from compromised sidecar → PARENT_INBOUND_OPS
allowlist, unknown frames logged and dropped
T-02-05 Tauri group-kill scope → start_new_session=True on Unix /
CREATE_NEW_PROCESS_GROUP on Windows
Verification:
- 337 passed, 6 skipped, 12 xfailed, 1 xpassed (full suite,
`uv run pytest tests/ --ignore=tests/manual`)
- All 19 new tests pass on macOS Apple Silicon
- Smoke tests still pass: `uv run pytest tests/smoke/ -q` → 4 passed
- SoniTranslate untouched (D1 locked decision)
- Zero new Python dependencies
Closes part of ENGINE-01 + ENGINE-05.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(02-01): plan summary — public API, invariants, deviations
Documents the SubprocessBackend public API so Plan 02-03 (IndexTTS) and
Phase 3 (Supertonic-3) authors don't need to re-read the source.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|