Commit Graph
5 Commits
Author SHA1 Message Date
Palash Debnath afa361913c fix(generate): surface the real cause on streaming failures (#1633)
Classify and journal local and remote streaming generation failures, return actionable scrubbed guidance, and keep exception details, tokens, and user paths out of logs and NDJSON responses.
2026-08-23 15:55:05 +05:30
debpalash c198a8349a fix: allow bounded CPU synthesis time 2026-08-20 09:01:43 +05:30
debpalash 544d345069 fix(security): keep stream failures out of logs 2026-08-09 23:18:16 +00:00
debpalash 96a6c7574a fix(security): stabilize streamed failure responses 2026-08-09 21:50:42 +00:00
14d2cae1e8 feat(tts): streaming playback preview — audio starts on the first chunk while the rest renders (#1088)
* feat(tts): streaming playback preview — audio starts on the first chunk while the rest renders

Long scripts meant staring at a spinner until the ENTIRE render finished.
POST /generate now takes stream=true (NDJSON: start → N × chunk → done),
synthesizing the existing Wave 1.2 sentence-boundary text chunks
sequentially and yielding each chunk's audio (base64 PCM16, with the same
effect chain the final take gets) the moment it's rendered — engine-
agnostic by construction, no per-engine token streaming.

The saved take is untouched: the raw chunks go through the SAME concat →
effect-chain → watermark → save → history → retention pipeline as the
classic path (extracted verbatim into _finalize_generation, now shared by
both), so the on-disk file is byte-identical to a non-streamed render
(regression-tested). [pause]-marker inputs and short single-chunk texts
keep their unchanged single-shot pipeline and stream as one chunk. Each
chunk gets its own generate budget on the GPU pool (model load already
warmed under the #1039 load budget before streaming starts), so a long
script can't time out merely for being long.

Frontend: when the preview would auto-play anyway and Web Audio exists
(all three Tauri webviews), useTTS drives the stream and the global
mini-player (#1042) starts playback from the first received chunk —
scheduled AudioBufferSource nodes with the backend's linear crossfade,
live seek within the buffered region, progressive peaks, growing
duration, and a "Streaming preview…" label that flips to "Generated
audio" on completion. Underruns (synthesis slower than playback) wait at
the buffered edge and resume when the next chunk lands. Any MID-stream
failure (in-band error event, transport drop, Web Audio failure) tears
down silently and falls back to the classic whole-file flow — the user
sees nothing beyond the old wait. Pre-stream HTTP errors keep their
ApiError identity so real 400/503s surface exactly as before, without a
wasted second render.

Tests: backend — incremental delivery proven at the ASGI boundary (fake
engine with per-chunk delay; TestClient/httpx buffer whole bodies and
can't observe it), final-file identity vs the classic path, mid-stream
error → error event with no done/history/file, single-chunk short text,
classic default unaffected. Frontend — NDJSON client, chunk player
(seek/stop/label flip), PCM decode, peaks, fallback error taxonomy.
Verified end-to-end over real HTTP (uvicorn): 1.0s chunk-arrival spread,
saved file byte-identical to the classic take.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style: oxfmt line-length fix in streamingTts.js

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(tts): cover the native OmniVoice model path in streaming preview tests

Per-chunk generate calls with duration=None + saved-take identity vs the
classic render, mirroring the pluggable-engine coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(tts): make streaming preview tests hermetic — pin outputs dir + history DB

The three streaming tests read the saved take back through
core.config.OUTPUTS_DIR while the router writes it through
api.routers.generation.OUTPUTS_DIR. Those are separate module bindings, and a
full-suite run can split them apart — an earlier test that reloads
core.config/main under a tmp data dir (test_dub_transcribe's app_client) moves
one binding and not the other, so the save lands in one dir and the read-back
looks in another: the CI FileNotFoundError under outputs/<uuid>.wav.

Add an autouse fixture that pins BOTH OUTPUTS_DIR bindings and the history DB
(via the ensure_schema.__globals__ get_db seam the takes suite already uses
against the #909/#932 module-purge leak) at per-test throwaway paths, so each
test is hermetic and order-independent. Deterministic fail-before (divergent
bindings pre-seeded) → 3 failures identical to CI; pass-after → green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: mergetest <nizam4103@gmail.com>
2026-07-12 12:10:47 +05:30