Commit Graph
460 Commits
Author SHA1 Message Date
Palash Debnath 7f7a4c5f83 fix(settings): make the generation budget reachable and honest (#1797)
The compute-time error told users to raise a generation timeout that had no control anywhere in the app — the only knob was an environment variable, and on Windows the docs explicitly warn against the usual way of setting one. Both budgets are now editable in Settings under Performance & Device, persisted and applied on the next start.

Two defects found in review and fixed here rather than shipped: an explicit universal budget silently overrode a separately saved CPU budget, so the CPU row would have looked like it worked and done nothing; and a value already set in the environment shadowed the saved preference while the panel still reported success. A shadowed row now says so instead. Long-input warnings also fire on Apple Silicon, which gets the accelerated budget and was the device in one of the duplicate reports.

Fixes #1787. Closes the reports tracked in #1774 and #1778.
2026-09-04 01:55:06 +05:30
Matt Van Horn 999345de41 feat(dub): realtime dub preview (#1769)
Opt-in live preview for dub segments: edits debounce into a streamed /ws/tts synthesis played through the chunk player, with cancellation preserved through buffered playback. Maintainer fixes: /ws/tts added to the backend ticket allowlist (feature was dead off-loopback), handshake failures surface a toast, loopback-only plaintext refusal reverted to keep the documented remote-GPU setup working, PCM16 decode hardened. Thanks @mvanhorn!
2026-09-03 18:27:07 +05:30
Matt Van Horn a95041f1e6 feat(studio): speech-to-speech voice changer (#1765)
Add a bounded, local-first speech-to-speech Convert workflow with shared ASR/TTS admission, duration matching, stale-request cancellation, profile conditioning, watermarking, persistence, localization, and regression coverage.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
2026-09-02 09:45:40 +05:30
Matt Van Horn 4053397921 feat(dub): karaoke word-highlight caption burn-in (#1764)
Adds opt-in word-timed ASS karaoke captions while preserving the existing line-caption default.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
2026-09-02 08:32:43 +05:30
Palash Debnath 08569397d3 fix(asr): secure configured endpoints and refresh guidance (#1751)
Refreshes README and linked docs with accurate installation, platform, privacy, API, and model-license guidance; documents local gigastt; and pins OpenAI-compatible ASR traffic to the configured secure origin. Closes #1736.
2026-09-02 04:41:36 +05:30
Palash Debnath 6f80110a42 fix(windows): await timed-out Job teardown 2026-09-02 03:15:16 +05:30
Palash Debnath c5ac548100 Merge remote-tracking branch 'origin/main' into fix/windows-direct-job-owner-1734
# Conflicts:
#	CHANGELOG.md
2026-09-02 03:14:09 +05:30
Palash Debnath 00d923b4fa fix: repair incomplete Sherpa model caches (#1753)
Repairs missing and zero-byte Sherpa ONNX cache assets before model loading, with offline regression coverage. Closes #1733.
2026-09-02 01:42:03 +05:30
Palash Debnath 549a56fc2d fix: remove Windows sidecar supervisor hop 2026-09-02 01:38:05 +05:30
Palash Debnath 4e5e8d1f89 Allow OmniVoice slow sidecar startup (#1743)
Fixes #1711.\n\nGives only the OmniVoice subprocess a 120-second readiness budget while retaining the shared 30-second default for all other sidecars, with regression coverage.
2026-09-02 00:22:05 +05:30
Palash Debnath 497d57ee62 Show complete engine disk costs before install (#1728)
Closes #1718.

Adds structured pre-install and measured post-install disk costs, complete sidecar preflight accounting, strict authorization for recursive scans, and localized catalogue details with confidence and deduplication context.
2026-09-01 23:49:05 +05:30
Palash Debnath 9447ace2c3 fix: invalidate unloaded ASR evidence 2026-08-30 22:21:18 +05:30
Palash Debnath cbc89e2d15 fix(diagnostics): track live execution evidence lifecycle 2026-08-30 22:09:05 +05:30
Palash Debnath bab794e13b fix: report only observed engine execution 2026-08-30 21:12:54 +05:30
Palash Debnath 503e4ffde8 Merge remote-tracking branch 'origin/main' into fix/engine-execution-evidence-1717 2026-08-30 21:10:29 +05:30
Palash Debnath 80aafa3a53 fix: preserve startup and media failure evidence (#1722)
* fix: preserve startup and media failure evidence

* fix(dictation): bind readiness to installed revision

* test(dictation): align cache resolution contract

* fix: keep retry failures actionable
2026-08-30 20:49:53 +05:30
Palash Debnath 0ec9c074e8 fix: distinguish opaque loaded sidecars 2026-08-30 20:49:26 +05:30
Palash Debnath 2be73b43e3 feat: expose engine execution evidence 2026-08-30 20:40:35 +05:30
Palash Debnath 3b6e15dad5 fix(macos): isolate OmniVoice MPS generation 2026-08-28 19:21:06 +05:30
Palash Debnath de51120d6a fix(models): handle load OOMs safely (#1696)
* fix(models): handle load OOMs safely (#1695)

* fix(dub): sanitize streamed generation failures

* style(ui): format readiness checklist
2026-08-28 16:00:24 +05:30
Palash Debnath 0268f43e7f fix(dub): keep language and media tools ready (#1679)
Fixes #1677 and #1678.

Publishes first-run media tools to the live backend, provides precise cross-platform missing-process guidance, and keeps localized source-language selection available before transcription. Includes regression coverage and deterministic model-store test isolation.
2026-08-28 06:32:49 +05:30
Palash Debnath 1103898d7b Merge remote-tracking branch 'origin/main' into fix/runtime-stability-1652
# Conflicts:
#	CHANGELOG.md
2026-08-27 22:30:58 +05:30
Palash Debnath c7890d2a5a fix(dub): close final media and retry review gaps 2026-08-27 21:25:56 +05:30
Palash Debnath 633e991edc Merge remote-tracking branch 'origin/main' into fix/runtime-stability-1652
# Conflicts:
#	CHANGELOG.md
2026-08-27 21:09:34 +05:30
Palash Debnath 4b3c62d783 Merge remote-tracking branch 'origin/main' into fix/dub-srt-voices-1660
# Conflicts:
#	CHANGELOG.md
2026-08-27 21:09:27 +05:30
Palash Debnath e1a76f9aca fix(dub): close shared workflow review gaps 2026-08-27 20:43:21 +05:30
Palash Debnath 422dbd1313 fix(dub): address cast, language, and cancellation review 2026-08-27 20:25:16 +05:30
Palash Debnath 0ee9bc35d0 fix(runtime): prevent overlapping native work and false crashes 2026-08-27 20:20:58 +05:30
Palash Debnath a1cd15964c Merge remote-tracking branch 'origin/main' into fix/voice-clone-reference-lease-1668 2026-08-27 20:10:51 +05:30
Palash Debnath 49c175301a Merge remote-tracking branch 'origin/main' into fix/dub-srt-voices-1660 2026-08-27 20:10:07 +05:30
Palash Debnath b92d35ac5d feat: add local speech platform (#1671)
Fixes #1646
2026-08-27 19:51:07 +05:30
Palash Debnath 53331a6f5a fix(dub): normalize uploaded video for preview 2026-08-27 19:45:10 +05:30
Palash Debnath 1d445855d2 fix(dub): make language workflow reliable 2026-08-27 19:43:49 +05:30
Palash Debnath 89cee3f824 fix(generate): lease ad-hoc references across abandoned jobs 2026-08-27 19:06:51 +05:30
Palash Debnath a8371baaa8 fix(desktop): serialize backend lifecycle (#1635)
Closes #1635.
2026-08-24 18:11:37 +05:30
Palash Debnath 5a615d2c66 feat(workers): package headless GPU nodes (#1638) (#1648)
Closes #1638.\n\nPackages headless GPU workers with durable enrollment, bounded artifact handling, cross-platform lifecycle cleanup, and regression coverage. Incorporates CodeRabbit, Greptile, CodeQL, and platform-CI findings before merge.
2026-08-24 16:32:56 +05:30
Palash Debnath ef1cb57944 fix: route OmniVoice to ROCm GPUs (#1647) 2026-08-24 03:25:56 +05:30
Palash Debnath 7718a7a10b fix: cross-platform dictation delivery (#1610)
Makes dictation delivery, capture, recovery, model fallback, AEC, and localized status behavior reliable across macOS, Windows, and Linux.
2026-08-21 03:33:50 +00:00
Palash Debnath 89d585a36e perf(dub,stream): reuse cached segments, batch the default engine, report real TTFA (#1620)
Reuses verified cached segments, safely batches default-engine dubbing, and reports synthesis-only TTFA/RTF.
2026-08-20 21:12:04 +00:00
Palash Debnath 43f1d46fe6 fix(indextts): accept the config name upstream ships, and keep long text alive (#1619)
* fix(indextts): accept the config name upstream ships, and keep long text alive

Two independent defects, both reported on a working IndexTTS 2.5 install.

Install always failed. IndexTeam/IndexTTS-2.5 ships the model config as
config.yaml — at the pinned revision d0aa86e7 and at HEAD; config_v2_5.yaml
exists in no upstream revision. VoiceStudio demanded that name, so
_weights_floor_ok never found it and the install died claiming 'the download
was likely interrupted' when the download had been perfect. The only way
through was to hand-rename the file. Both names are accepted now, in the
installer and on the load path, so installs created with the workaround keep
working without a reinstall.

Long text was killed at 60s. infer() is one blocking upstream call that puts
nothing on the wire, and IndexTTS was the only sidecar still on the 60s
recv_timeout_s class default while pockettts and omnivoice-subprocess had both
raised theirs. Raising the default alone does not fix it — which is why the
reporter's RECV_TIMEOUT_S=3600 edit didn't help: progress frames are also what
report activity to the GPU pool's execution clock (#1367), so a silent sidecar
still trips the outer generate budget. The sidecar now heartbeats every 5s
while infer() runs (and during the cold model construction), _send takes a
lock so the beat thread can't interleave framing, and the deadline rises to
900s via OMNIVOICE_INDEXTTS_RECV_TIMEOUT_S.

test_indextts25_health_requires_25_config_name asserted the bug — that a
checkout holding only config.yaml is unhealthy — so it is rewritten to the
corrected contract, including that a genuinely truncated download is still
caught.

Fixes #1611

* test(indextts): follow the installed config name in the sidecar loader tests

Two more tests encoded the config_v2_5.yaml assumption, both asserting
cfg_path against a directory where no config existed at all — so they were
pinning the literal name rather than the resolution. They now lay down a real
checkpoints/ tree and assert the resolved path, including that a checkout
carrying the pre-fix hand-renamed config still resolves.

Caught by the full suite; the targeted runs during development did not reach
tests/backend/services/.

* test(indextts): event-driven heartbeat tests, real interleaving proof, precedence pin

Review round on #1619 — all four findings taken.

- The docs line naming 0.5.1 is version-neutral now ('Earlier installs') —
  version labels are the owner's call.
- The heartbeat tests waited on wall-clock sleeps; they now block on a
  per-write Event with a bounded deadline, so scheduler load can't flake them.
- The _send test asserted the lock EXISTS — a tautology. It now drives four
  concurrent writers through a stream that yields between every byte and
  asserts every frame decodes; verified fail-before by removing the lock
  (torn frame) and pass-after.
- The precedence test deleted config.yaml before creating the renamed one, so
  reversed precedence still passed. Both files now coexist for the assertion;
  verified fail-before by reversing _CFG_NAMES.
2026-08-20 19:43:52 +00:00
Palash Debnath 54a88f694b fix(watermark): run AudioSeal eagerly instead of through torch.compile (#1617)
* fix(watermark): run AudioSeal eagerly instead of through torch.compile

AudioSeal vendors moshi's @torch_compile_lazy on SEANetEncoder.forward, so
the first embed of a session — not the model load, which #1576's prefetch
already warms — called torch.compile and dropped into Inductor's C++ codegen.
On a macOS arm64 deployment that compile raised CppCompileError on 10/10
takes: the embed fail-opened and the audio shipped UNMARKED, an EU AI Act
Art. 50(2) provenance gap, after burning 30-40s on the first take and 5-8s on
each later one.

The compile is pure cost even where it succeeds. Measured on an M3 (5s of
24kHz audio, three consecutive embeds): compiled 9.70/0.26/0.23s vs eager
0.30/0.28/0.27s — a ~10s first-embed tax to save ~0.03s afterwards, on CPU
work already bounded by the 30s chunk loop. Both embed and detect now run
inside audioseal's own no_compile() switch, restored on the way out (it is a
process global, and other models are entitled to compile).

Verified end-to-end: first embed 9.70s -> 0.26s, watermark still round-trips
at confidence 1.0 with the OmniVoice message intact.

Fixes #1615

* fix(watermark): collapse the eager-guard globals into one lock-guarded state

CodeQL flagged _eager_saved's module-level initializer as dead, and it was
right: depth 0->1 always writes the field before depth 1->0 reads it, so the
None at import was never observed. Depth and saved-value are only meaningful
together and only under _eager_lock, so they become one dict rather than two
module scalars — which also drops the global statement.

Also splits three semicolon-joined statements in the regression test (Ruff
E702, CodeRabbit).

Mutation re-checked after the refactor: a naive no_compile() body still fails
with 'compile was handed back mid-embed'.
2026-08-20 18:16:00 +00:00
debpalash aa7c2f5801 fix: fail open on watermark dispatch deadlines 2026-08-20 10:32:02 +05:30
debpalash 918c400f29 fix: preserve audio on watermark teardown cancellation 2026-08-20 09:48:43 +05:30
debpalash 76d16ac1bd fix: fail open on watermark shutdown submission race 2026-08-20 09:43:58 +05:30
debpalash 7efae54cf8 fix: budget the selected engine device 2026-08-20 09:13:18 +05:30
debpalash ed6d7a9652 fix: preserve explicit generation watchdog 2026-08-20 09:06:50 +05:30
debpalash c198a8349a fix: allow bounded CPU synthesis time 2026-08-20 09:01:43 +05:30
debpalash 2dcfd0bb55 Merge remote-tracking branch 'contributor/fix/watermark-prefetch-cold-start' into fix/integration-eventbus-loop-clean 2026-08-20 08:53:51 +05:30
debpalash 93616a9c2a fix(watermark): fail open while pool drains 2026-08-20 08:53:42 +05:30
debpalash daefad8769 Merge remote-tracking branch 'contributor/fix/watermark-prefetch-cold-start' into fix/integration-eventbus-loop-clean
# Conflicts:
#	CHANGELOG.md
2026-08-20 08:49:38 +05:30