Commit Graph
443 Commits
Author SHA1 Message Date
Palash Debnath f2302e8c95 fix(desktop): don't adopt a backend running stale code (#1796)
Exports failed with a 422 naming a field the current app never sends — twice, from different users. The cause was the attach handshake: if something already answers on the backend port and reports a matching version, the app adopts it and skips the source sync a normal launch performs. A version string holds steady for a whole release cycle, so a same-version process can still be running weeks-old code, and that code then serves a current UI.

The handshake now compares a fingerprint of the shipped Python sources, read from the same response as the version so a dropped probe can't masquerade as a missing field. A backend predating the mechanism is treated as stale; one that is current but started outside the app is still accepted. Refusals are logged with a greppable marker, since this class previously took two reports and a code audit to identify.

Fixes #1770. Closes the duplicate report tracked in #1792.
2026-09-04 03:22:41 +05:30
Palash Debnath 7f7a4c5f83 fix(settings): make the generation budget reachable and honest (#1797)
The compute-time error told users to raise a generation timeout that had no control anywhere in the app — the only knob was an environment variable, and on Windows the docs explicitly warn against the usual way of setting one. Both budgets are now editable in Settings under Performance & Device, persisted and applied on the next start.

Two defects found in review and fixed here rather than shipped: an explicit universal budget silently overrode a separately saved CPU budget, so the CPU row would have looked like it worked and done nothing; and a value already set in the environment shadowed the saved preference while the panel still reported success. A shadowed row now says so instead. Long-input warnings also fire on Apple Silicon, which gets the accelerated budget and was the device in one of the duplicate reports.

Fixes #1787. Closes the reports tracked in #1774 and #1778.
2026-09-04 01:55:06 +05:30
Matt Van Horn 1515b46adb feat(batch): watch-folder auto-ingest (#1768)
Opt-in watch folder on the batch queue: pick a directory once and new videos are auto-enqueued with the last Add-to-queue settings, with pause/stop controls and copy-in-progress protection. Files stream to the loopback backend as bytes; paths never leave the app. Also gives the batch queue a reachable UI entry point and streams multipart uploads to disk. Maintainer fix: the watched directory handle is opened with full share mode on Windows so users can rename or delete the folder while it is watched, matching macOS/Linux behaviour, with a cross-platform regression test. Thanks @mvanhorn!
2026-09-03 18:47:32 +05:30
Matt Van Horn a95041f1e6 feat(studio): speech-to-speech voice changer (#1765)
Add a bounded, local-first speech-to-speech Convert workflow with shared ASR/TTS admission, duration matching, stale-request cancellation, profile conditioning, watermarking, persistence, localization, and regression coverage.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
2026-09-02 09:45:40 +05:30
Matt Van Horn 4053397921 feat(dub): karaoke word-highlight caption burn-in (#1764)
Adds opt-in word-timed ASS karaoke captions while preserving the existing line-caption default.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
2026-09-02 08:32:43 +05:30
Palash Debnath 08569397d3 fix(asr): secure configured endpoints and refresh guidance (#1751)
Refreshes README and linked docs with accurate installation, platform, privacy, API, and model-license guidance; documents local gigastt; and pins OpenAI-compatible ASR traffic to the configured secure origin. Closes #1736.
2026-09-02 04:41:36 +05:30
Palash Debnath 1d06c6f079 fix: accept ASR-detected dub source codes (#1755)
Accepts every Whisper language code persisted by ASR, including three-letter Cantonese yue, so subsequent dubbing uploads no longer fail validation. Closes #1737.
2026-09-02 02:26:56 +05:30
Palash Debnath 497d57ee62 Show complete engine disk costs before install (#1728)
Closes #1718.

Adds structured pre-install and measured post-install disk costs, complete sidecar preflight accounting, strict authorization for recursive scans, and localized catalogue details with confidence and deduplication context.
2026-09-01 23:49:05 +05:30
Palash Debnath cbc89e2d15 fix(diagnostics): track live execution evidence lifecycle 2026-08-30 22:09:05 +05:30
Palash Debnath 2be73b43e3 feat: expose engine execution evidence 2026-08-30 20:40:35 +05:30
Palash Debnath 3b6e15dad5 fix(macos): isolate OmniVoice MPS generation 2026-08-28 19:21:06 +05:30
Palash Debnath de51120d6a fix(models): handle load OOMs safely (#1696)
* fix(models): handle load OOMs safely (#1695)

* fix(dub): sanitize streamed generation failures

* style(ui): format readiness checklist
2026-08-28 16:00:24 +05:30
Palash Debnath 8cc7c88692 fix(dub): restore audio when the media preview falls back (#1693)
Preserve audible, language-matched playback when WebView video decoding falls back, keep waveform timing aligned with the selected dub, prevent recursive recovery failures, and disable stale audio caching.

Closes #1692.
2026-08-28 13:37:34 +05:30
Palash Debnath 0268f43e7f fix(dub): keep language and media tools ready (#1679)
Fixes #1677 and #1678.

Publishes first-run media tools to the live backend, provides precise cross-platform missing-process guidance, and keeps localized source-language selection available before transcription. Includes regression coverage and deterministic model-store test isolation.
2026-08-28 06:32:49 +05:30
Palash Debnath 1103898d7b Merge remote-tracking branch 'origin/main' into fix/runtime-stability-1652
# Conflicts:
#	CHANGELOG.md
2026-08-27 22:30:58 +05:30
Palash Debnath 0e47053823 fix(dub): scope imported speaker clones to matched cues 2026-08-27 21:57:37 +05:30
Palash Debnath bea72cadbe Merge remote-tracking branch 'origin/main' into fix/runtime-stability-1652
# Conflicts:
#	CHANGELOG.md
2026-08-27 21:50:51 +05:30
Palash Debnath 6624b3fb19 Merge remote-tracking branch 'origin/main' into fix/dub-srt-voices-1660
# Conflicts:
#	CHANGELOG.md
2026-08-27 21:50:46 +05:30
Palash Debnath c7890d2a5a fix(dub): close final media and retry review gaps 2026-08-27 21:25:56 +05:30
Palash Debnath 633e991edc Merge remote-tracking branch 'origin/main' into fix/runtime-stability-1652
# Conflicts:
#	CHANGELOG.md
2026-08-27 21:09:34 +05:30
Palash Debnath 64b3785b0b Merge remote-tracking branch 'origin/main' into fix/docker-admin-auth-1651
# Conflicts:
#	CHANGELOG.md
2026-08-27 21:09:30 +05:30
Palash Debnath 4b3c62d783 Merge remote-tracking branch 'origin/main' into fix/dub-srt-voices-1660
# Conflicts:
#	CHANGELOG.md
2026-08-27 21:09:27 +05:30
Palash Debnath 86f6c9ec8e fix(auth): reject blank server admin keys 2026-08-27 20:46:24 +05:30
Palash Debnath e1a76f9aca fix(dub): close shared workflow review gaps 2026-08-27 20:43:21 +05:30
Palash Debnath 422dbd1313 fix(dub): address cast, language, and cancellation review 2026-08-27 20:25:16 +05:30
Palash Debnath 0ee9bc35d0 fix(runtime): prevent overlapping native work and false crashes 2026-08-27 20:20:58 +05:30
Palash Debnath a1cd15964c Merge remote-tracking branch 'origin/main' into fix/voice-clone-reference-lease-1668 2026-08-27 20:10:51 +05:30
Palash Debnath 49c175301a Merge remote-tracking branch 'origin/main' into fix/dub-srt-voices-1660 2026-08-27 20:10:07 +05:30
Palash Debnath b92d35ac5d feat: add local speech platform (#1671)
Fixes #1646
2026-08-27 19:51:07 +05:30
Palash Debnath b6bd125f23 fix(dub): keep preview extraction off event loop 2026-08-27 19:45:31 +05:30
Palash Debnath 1d445855d2 fix(dub): make language workflow reliable 2026-08-27 19:43:49 +05:30
Palash Debnath e93b6366e6 fix(dub): preserve voices across SRT imports 2026-08-27 19:14:56 +05:30
Palash Debnath 89cee3f824 fix(generate): lease ad-hoc references across abandoned jobs 2026-08-27 19:06:51 +05:30
Palash Debnath 5a615d2c66 feat(workers): package headless GPU nodes (#1638) (#1648)
Closes #1638.\n\nPackages headless GPU workers with durable enrollment, bounded artifact handling, cross-platform lifecycle cleanup, and regression coverage. Incorporates CodeRabbit, Greptile, CodeQL, and platform-CI findings before merge.
2026-08-24 16:32:56 +05:30
Palash Debnath afa361913c fix(generate): surface the real cause on streaming failures (#1633)
Classify and journal local and remote streaming generation failures, return actionable scrubbed guidance, and keep exception details, tokens, and user paths out of logs and NDJSON responses.
2026-08-23 15:55:05 +05:30
Palash Debnath 7718a7a10b fix: cross-platform dictation delivery (#1610)
Makes dictation delivery, capture, recovery, model fallback, AEC, and localized status behavior reliable across macOS, Windows, and Linux.
2026-08-21 03:33:50 +00:00
Palash Debnath 89d585a36e perf(dub,stream): reuse cached segments, batch the default engine, report real TTFA (#1620)
Reuses verified cached segments, safely batches default-engine dubbing, and reports synthesis-only TTFA/RTF.
2026-08-20 21:12:04 +00:00
Paolo Antinorianddebpalash 3223a20f88 fix(openai-compat): reuse cached engine instances in _resolve_engine (#1614)
* fix(openai-compat): reuse cached engine instances in _resolve_engine

The direct engine-ID path in /v1/audio/speech constructed a fresh
backend per request (return cls()). For SubprocessBackend engines that
meant: a new sidecar process, a full torch import and an engine model
reload on EVERY request (measured ~28s floor per pockettts request on
an M3 Pro), plus another atexit hook registration each time — exactly
what get_engine_instance_for()'s docstring warns against.

Route the explicit-ID path through the same cached-singleton seam the
active-engine path already uses. Unknown/unavailable IDs keep their
400s; tts-1/tts-1-hd and the OmniVoiceBackend special case are
unchanged.

* fix(openai-compat): unload the outgoing engine on explicit-ID switches

Review follow-up (Greptile/CodeRabbit on #1614): caching instances without
a switch rule would let each distinct explicit engine ID stay resident,
accumulating sidecars / multi-GB in-process models. Mirror
get_active_tts_backend's MM2-01 switch rule: a different explicit ID
(omnivoice included, which resolves to the active engine) unloads the
outgoing instance first, best-effort.

* fix(openai-compat): evict via the shared single-engine-resident seam, not a router-local cache

The explicit-ID unload cache (13c14e2c) kept its own instance ref keyed by
model id. The shared engine cache is deliberately keyed by CLASS (registry
rebinds, idle sweeps and engine_memory eviction all mutate it), so the
router's id-keyed ref could go stale and keep serving an instance the
lifecycle system no longer tracked — caught by
test_openai_speech_toggle_off_sends_raw_text in full-suite order, and it
also introduced a novel unload path that ignored the
OMNIVOICE_SINGLE_ENGINE_RESIDENT opt-out.

Drop the router-local cache entirely: _resolve_engine returns the shared
cached singleton (get_engine_instance_for), and create_speech calls
evict_other_tts_engines(backend.id) before warming the engine — the exact
seam /generate uses. That covers every transition (explicit id → explicit
id, explicit id → tts-1/omnivoice aliases), honors the policy opt-out, and
leaves no per-router state to drift. Regression pinned at the route level in
test_speech_request_evicts_other_resident_engines.

* chore(changelog): trim the #1614 entry to the one-liner limit

415 chars against the 400 the style test allows — CI would have failed on it.

---------

Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com>
2026-08-20 19:14:10 +00:00
Palash Debnath 2d37627ab2 fix(setup): tolerate reserved memory in the RAM preflight, add OMNIVOICE_RAM_PREFLIGHT=0 escape hatch (#1621)
* fix(setup): tolerate reserved memory in the RAM preflight, add OMNIVOICE_RAM_PREFLIGHT=0 escape hatch (#1618)

An "8 GB" machine reports ~7.8 GB usable (firmware/iGPU/kernel
reservations), so comparing OS-reported RAM against the marketing-size
8 GB threshold hard-blocked exactly the boundary hardware the minimum is
meant to admit — with no way past the wizard. Both thresholds are now
compared with a 7% reserved-memory allowance, and
OMNIVOICE_RAM_PREFLIGHT=0 downgrades a genuine fail to a warning for
users who accept the OOM risk (same opt-out shape as
OMNIVOICE_ASR_VRAM_PREFLIGHT).

Regression tests: backend/tests/test_ram_preflight_1618.py.
Docs: troubleshooting §1c.

* review: hermetic preflight stubs in tests; correct the escape-hatch doc

Greptile P1: the Settings panel can't set OMNIVOICE_RAM_PREFLIGHT (and the
blocker appears before setup completes anyway) — the doc now points at
PowerShell / shell env only.
CodeRabbit: stub _network_check and media_tools.summary so each RAM
assertion stays fast and offline (26s -> 6s locally).
2026-08-20 19:03:50 +00:00
debpalash 605236566c fix: align generation routing and CPU deadlines 2026-08-20 09:59:47 +05:30
debpalash 2048d2793a Merge PR #1604: keep healthy CPU synthesis past five minutes
# Conflicts:
#	CHANGELOG.md
2026-08-20 09:17:30 +05:30
debpalash 7efae54cf8 fix: budget the selected engine device 2026-08-20 09:13:18 +05:30
debpalash 3a1013527f Merge PR #1601 follow-up: normalize original-only exports 2026-08-20 09:04:00 +05:30
debpalash c198a8349a fix: allow bounded CPU synthesis time 2026-08-20 09:01:43 +05:30
debpalash c4f3ca457d Merge latest PR #1599 original-only normalization 2026-08-20 09:00:43 +05:30
debpalash e1e3a477a7 fix: normalize original-only dub exports 2026-08-20 09:00:34 +05:30
debpalash e20add344c Merge PR #1601: close integration review findings 2026-08-20 09:00:01 +05:30
debpalash 1044483edf Merge latest PR #1599 trusted output paths 2026-08-20 08:56:29 +05:30
debpalash a04c972b71 fix: keep dub export paths trusted 2026-08-20 08:56:20 +05:30
debpalash 2dcfd0bb55 Merge remote-tracking branch 'contributor/fix/watermark-prefetch-cold-start' into fix/integration-eventbus-loop-clean 2026-08-20 08:53:51 +05:30