d91beef0fd314250d8d9b94de86dfea019a8bd96
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a95041f1e6 |
feat(studio): speech-to-speech voice changer (#1765)
Add a bounded, local-first speech-to-speech Convert workflow with shared ASR/TTS admission, duration matching, stale-request cancellation, profile conditioning, watermarking, persistence, localization, and regression coverage.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com> |
||
|
|
4db02d0c97 |
fix(test): watermark producer scan must match code, not prose (#1564)
* fix(test): watermark producer scan must match code, not prose
|
||
|
|
c643706d07 |
feat(workers): make a remote GPU actually run a task, end to end
Selecting a remote worker repainted a badge and nothing else. The cause was not subtle: `scheduler.submit` had no production caller, and `routing.decide()` was read only by the status endpoint that paints the header. Remote execution was a complete, tested pipeline with no producer at its head. This adds the producer and fixes the defects that made the pipeline unable to carry a real job: - Nothing routed to the scheduler. Adds `POST /workers/tasks` (loopback-gated, **development-only** until the gateway lands) and `Scheduler.wait`, backed by per-task futures rather than the unregisterable `on_change` listener list. - Every task over two minutes died. No worker ever sent `TaskProgress`, so the 120s progress lease expired mid-render — including during the cold model load, which happens after `TaskStarted`. Workers now report progress and emit a keepalive, bounded by the phase's absolute budget so it renews the lease without deleting the only enforced bound in the system. - The executor rebuilt its engine per task (`return cls()`), so every job paid a cold load. Engines now share one instance cache with the router, resolved by the assignment's engine — never `get_active_tts_backend()`, which returns the worker machine's own Settings preference and would silently run the wrong engine. - One lease expiry took a worker offline permanently: parked slots were never reclaimed. Parks now expire on a TTL, and are deliberately NOT reconciled against the worker's own load report — at a ceiling of one the only task such a worker can report is the wedged one, so "busy" would drop the park and the next idle heartbeat would hand out a slot with a live GPU thread (#730/#1190). - A worker that dropped and reconnected mid-render had every liveness frame discarded: task frames were fenced on the live session epoch, which bumps on every reconnect, while the worker echoes the ref stamped at dispatch. The control plane then expired a task whose GPU was still rendering, and swallowed the failure report when it went wrong. Fenced per attempt instead. - A result from one worker could commit another's task, after which the owner's real delivery arrived as a duplicate and its audio was discarded. "Unknown attempt" and "another worker's attempt" are no longer the same answer. - An oversized result was a poison pill, re-sent identically on every reconnect and permanently disconnecting the worker. It is now a terminal `RESULT_TOO_LARGE`, which is also classified — it was falling through to TRANSIENT and retrying a re-render that could never fit. - `_store_inline` joined the artifact directory with worker-supplied ids, and `os.path.join` discards its prefix on an absolute component. Paths are now minted control-plane-side and resolved through `core.path_security`. - Remote synthesis bypassed `mark_synthetic`, and the guard that exists to catch exactly that walked only `backend/api` and `backend/services` — so it stayed green while a fourth unmarked producer shipped. Marking moved to the worker's tensor stage; the guard now walks `backend/worker` too. Also adds pre-rendered voice previews (`services/gallery.py`), so browsing the gallery no longer needs a GPU or a downloaded model. The manifest is verified against the updater's release key already baked into the binary; a fresh install hears voices without downloading 2.4GB first, and everything falls back to local rendering when the gallery is unreachable. Verified on hardware, not just in CI: 1728 characters submitted to an RTX 4090 returned 105.94s of 24kHz audio in 23.9s, committed and served from the artifact store. Not yet done, and deliberately not claimed: the keepalive fix cannot be exercised end-to-end on fast hardware, because any job long enough to reach the 120s lease produces audio past the 8MiB inline cap. Chunked `UploadResult` has to land first. Pinning to the worker the user chose is also still absent, so "Remote" reaches a remote GPU but not necessarily the one on the badge. |
||
|
|
3a7368cb26 |
fix(watermark): route every synthetic-audio producer through one mark_synthetic chokepoint (#1169)
POST /v1/audio/speech returned synthetic audio without the AudioSeal provenance watermark while /generate marked the same text — and the audit behind #1169 found the same class of gap in five more producers. EU AI Act Art. 50(2) (applicable 2026-08-02, expressly carved out of the open-source exemption in Art. 2(12)) makes machine-readable marking of synthetic audio a provider obligation, so per-door coverage gaps are compliance bugs. Root cause: coverage grew call-site-by-call-site (three separate embed_watermark calls) with nothing forcing a new producer to opt in. The fix, per the whole-class rule: * services/watermark.py grows `mark_synthetic(wav, sr, *, context, force=False)` — THE named chokepoint (delegates to embed_watermark: pref-gated, no-op without AudioSeal, never raises / degrades to unmarked) — plus `will_mark()` for cache-key derivation. * Gaps closed at the tensor stage, before any encoding: - openai_compat `_run_tts` (the reported gap; all response_formats) - tts_stream /ws/tts (per-sentence, before PCM16 framing) - generation stream=true preview chunks (marked streamed copy; the saved take keeps its single whole-take mark in finalize) - batch dub pipeline (assembled track, before WAV write / aac mux) - longform chapter render — covers /audiobook, /longform/render (Stories), /audiobook/preview and /audiobook/resume; the chapter cache key now carries a watermark tag so stale unmarked cache entries can never be served for a marked-on render - dub preview-segment (docstring had declared it exempt) - archetype render (served preview + materialized profile reference) * Already-covered paths (generate finalize, dub segments, persona bundles) migrated onto the same chokepoint. * Documented non-producers/gaps instead of fake coverage: /stories/encode is a pure transcoder of user uploads (must NOT mark); the opt-in SoniTranslate sidecar synthesizes+muxes externally and is a documented provenance gap at its route. * Regression tests: tests/test_synthetic_audio_watermark_1169.py runs detect_watermark() on the actual response audio of every producing route (real watermark service, fake AudioSeal nets, fake engine; verified fail-before on the pre-fix tree — 11 of 13 fail). * Recurrence-proofing: tests/test_watermark_route_coverage.py structurally asserts every synthesis call site references mark_synthetic (justified allowlist), producers keep their call, and embed_watermark is never called outside the chokepoint. Settings toggle semantics are unchanged (the issue's two legal questions stay flagged for a lawyer, deliberately unanswered here). |