dba8934aa2fc8700a33b7adce6a6737f958ae22f
148
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b7cecde57e |
feat(setup): faster downloads by default + prominent, encouraged HF-token entry (#669)
Two changes that make first-run downloads faster and easier to speed up further.
1. Segmented (multi-connection) downloader is now ON by default. The app forces
the legacy-LFS path (HF_HUB_DISABLE_XET=1) for clear progress, but that path
is single-stream and slow — which is why downloads felt sluggish. The built-in
IDM/uGet-style segmented accelerator (parallel byte-ranges, live speed/ETA)
was already implemented but defaulted OFF. Flip it ON: it only engages when
Xet is inactive (the default), and ANY failure falls back to snapshot_download
("can never compromise a correct install"). Pure-httpx, cross-platform,
auth-safe (token never forwarded to a CDN). Override with
OMNIVOICE_SEGMENTED_DOWNLOAD=0.
2. The Hugging Face token field is now a prominent, always-visible card right
above Continue — was a collapsed "advanced" fold almost nobody opened. A free
token gives authenticated downloads (higher rate limits, fewer stalls), so it
pairs with change #1 to keep the parallel fetch from getting throttled. The
card leads with the speed benefit, shows a saved-state, and adds a one-click
"Get one free →" link to huggingface.co/settings/tokens.
Docs: downloading-models.md updated — the legacy-LFS section now documents the
default-on segmented accelerator + the HF-token speed tip, and the tuning table
reflects OMNIVOICE_SEGMENTED_DOWNLOAD=0 as the disable knob (docs-sync).
Test: test_segmented_download_default.py pins the new default ON and that the
env override still disables it; existing FDL-08 behavior tests stay green.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
81a007552b |
fix(generate): classify a bad-instruct error as a 400, not a 500 "ran out of memory" (#664) (#665)
A user typed free-form prose ("Speak with high energy … like a podcast host")
into the voice-design instruct field and got a **500** whose message read "TTS
engine stopped mid-generation. This usually means it ran out of memory. Try the
Flush button …" — with the real cause ("Unsupported instruct items found …")
buried as the underlying error. The user is told to Flush for an OOM that never
happened; the actual problem is a rejected instruct.
Root cause: `_resolve_instruct` raises on unknown/conflicting instruct items, but
by the time the error reaches `_oom_friendly_reraise` it's no longer a bare
`ValueError` (a lower layer wraps it), so the route's `except ValueError -> 400`
guard misses it and it falls through to the generic OOM `RuntimeError`. v0.3.7
has had that guard since v0.3.6 yet still produced the OOM message — proving the
error arrives wrapped, so type-based detection is insufficient.
Fix: in `_oom_friendly_reraise`, detect the instruct-validation **message
signature** ("unsupported instruct items" / "conflicting instruct items" / "in a
single instruct") regardless of exception type and re-raise a clean `ValueError`,
so the route returns a **400 with the instruct guidance** instead of a 500 OOM.
This is version-independent and complements the client-side guard (#658/#612):
it also covers API/MCP callers and stored profiles whose instruct slips through.
Test: two cases in test_generation_audio_guard.py — a bare instruct `ValueError`
and one wrapped in a `RuntimeError` both reclassify to a ValueError without the
"ran out of memory" text; the generic OOM path is unchanged.
Closes #664
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
252f0d4fac |
fix(asr): bound whole-file transcription so a stall isn't reported as "can't reach backend" (#656)
A Windows/CUDA user (Vietnam) hit "Can't reach the local backend" only when dubbing/transcribing. Their log proves the backend started fine — model loaded, preload complete, 25 models — and the log ends right after `whisperx transcribing …tmp.wav`. The backend was alive; the *transcription* stalled (large-v3 ASR contending with the resident TTS model for VRAM on an 8 GB-class GPU), which the UI surfaces as an unreachable backend. Root cause (class, not instance): the chunked dub pipeline already bounds each chunk (OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S), but the *whole-file* transcribe paths ran unbounded: - dub QC re-transcribe (dub_export) - dictation (capture) - OpenAI-compat /audio/transcriptions A slow/stuck transcribe on any of these hung the request AND held a GPU-pool worker — indistinguishable from a dead backend. Fix: add run_transcribe_guarded() in services/asr_backend.py — a shared asyncio.wait_for wrapper (ASRTimeoutError, a TimeoutError subclass) with a generous env-tunable bound (OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S, default 300 s). On timeout the request returns 504 with actionable guidance (backend is alive; free VRAM / pick a smaller ASR model / use CPU; restart to clear the stuck worker) instead of hanging forever. Wired into all three whole-file paths. Docs: new troubleshooting §14 — "Can't reach the local backend during transcription/dubbing" — explains it's ASR weight/VRAM pressure, not a network/ mirror problem, and corrects the misconception that a "Network → Restricted/Global mirror" Settings toggle exists (the Network control is LAN sharing). Serves the #602/#585/#567 "can't reach backend" cluster. Test: backend/tests/test_asr_transcribe_timeout.py — slow fn raises ASRTimeoutError with the actionable message, fast fn passes through, subclass-of-TimeoutError so the openai_compat broad catch still maps to 504. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b15acbaae9 |
fix(dub+generate): yt-dlp 403 player-client fallback (#625) + non-finite audio guard (#629) (#635)
Two independent fixes from issue triage; no version bump. #625 — yt-dlp 403 on the media download (some videos serve formats signature-protected to the default player client) is not transient, so the existing broken-pipe retry (#579) kept 403ing. The URL download now escalates the YouTube player client (tv → android → web_safari) on a 403 before giving up; a 403 no longer counts against the transient-retry budget. #629 — a numerical glitch in the model (seen on MPS) could leave NaN/inf samples that write an unreadable WAV; a downstream decode then failed with an opaque "ffmpeg returned error code: 183 / Invalid data", surfaced to the user as a misleading "ran out of memory". Sanitize non-finite samples to silence in _apply_effect_chain (single chokepoint, covers the raw path too) so the WAV is always decodable, and classify a decode/ffmpeg failure as unreadable-audio rather than OOM in _oom_friendly_reraise. Tests: 403 escalation order + success-on-alternate-client; NaN/inf sanitize + finite-passthrough + decode-error classification. Full suite 1851 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a63c8e851b |
fix(dub): speaker-aware re-split so merged speaker turns separate (#486) (#616)
Segmentation groups words into sentences BEFORE diarization, so a two-speaker exchange can land in one segment; assign_speakers_* then only relabels it with the majority speaker, losing the turn boundary (the second half of #486 — the per-speaker voice auto-assign was fixed in #490). Add a post-diarization pass that re-splits any segment whose words span >1 speaker at the word-level boundary, assigning each piece its speaker: - backend/services/segmentation.py: resplit_segments_by_diarization / resplit_segments_by_turns + a pure _resplit_core. Single-speaker segments are returned BYTE-FOR-BYTE UNCHANGED (same dict/id/text/start/end) — the no-single-speaker-regression guarantee. Pieces keep the segment's outer start/end (preserving onset-snap) and use word times for interior splits, so they exactly cover the original span. A lone mis-attributed word is smoothed, not split (diarization noise). - backend/api/routers/dub_core.py: accumulate global-timeline words alongside segments; apply the re-split after both the pyannote and FunASR-turns assign. Heuristic fallback (no word-speaker data) is untouched. 8 regression tests pin the invariant + the split/3-way/noise-smoothing/label behaviour. Full suite: 1836 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1fe68ba11e |
fix(setup): weight-aware install-state so truncated model cache isn't read as installed (#622) (#626)
A first-run user whose model download was interrupted after the config/ tokenizer files landed but before the weight shard got stranded on the Models & Engines page: GET /models computed "installed" purely from cache size on disk, so a size-positive-but-weight-less cache reported installed=true, the wizard hid the re-download button, and the model manager (Settings → Models) that could repair it was unreachable behind the wizard gate. Make install-state weight-aware. The boolean weight-floor scan now lives in models.py (the lowest module in the setup import graph) as snapshot_has_weights() + cache_is_complete(); list_models() and recommendations() downgrade a truncated cache to installed=false (+ an explicit incomplete=true on /models), so the existing "install" action re-appears and the user can re-download in-wizard. Fixes the whole class, not just /models: download.py's install-time validator now delegates to the same shared scan (one source of the floors, can't drift), matching the load-time repair in model_manager.py (#581/#606). config_only repos (pyannote/speaker-diarization-3.1 — a pipeline whose real weights live in referenced sub-repos and whose own cache is legitimately tiny) carry a new config_only:true hint in models.yaml and are exempt, so they're not false-flagged as incomplete. Tests: tests/test_mm2_lifecycle.py — snapshot_has_weights truncated-vs-complete, cache_is_complete on a truncated weight repo + config-only exemption, and list_models downgrading a size-positive truncated cache to installed=false / incomplete=true. Full backend suite green (1832 passed). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1575baca36 |
fix(design): heal validator-rejecting instruct on design voices (#594/#571/#596) (#600)
* fix(design): heal validator-rejecting instruct on design voices (#594/#571/#596) A designed voice could persist an `instruct` the engine validator rejects — either the literal "[object Object]" from a pre-fix build (#550) or freeform prose typed into the style field — so every Generate/Dub that used the voice failed with `Unsupported instruct items found in …` (400/500, and "Can't reach the local backend" when it tore down mid-render). Migration 0006 only *blanked* "[object Object]", which silently discarded the design — an Indonesian female voice then rendered male (#594). Fix the whole class by healing at every seam and rebuilding from the authoritative source (the design's saved `vd_states` category picks): - omnivoice/utils/voice_design.py: add sanitize_instruct / instruct_from_vd_states / heal_design_instruct — forgiving (never raise), drop poison/prose to valid tags, and rebuild tags from vd_states when the stored value is unusable. - profiles.py: sanitize + rebuild at save (POST) and sanitize at edit (PUT), so no poisoned instruct can ever be persisted again. - generation.py + dub_generate.py: heal whenever a profile drives synthesis, so legacy poisoned rows resolve to valid tags instead of 400-ing. - migration 0007: heal existing profiles in place (recovers gender/age/pitch from vd_states), self-contained (frozen vocab snapshot) so it never drags torch into startup; supersedes 0006's blanking. Backward-compatible. Tests: unit coverage for the healer, a migration test driving 0006->0007 on the real schema, a parity guard so the frozen snapshot can't drift, and two API guards. Corrected one existing test that had encoded the #594 behaviour. Resolves #571, #594, #596; removes a major driver of the "Can't reach backend" reports. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(cjk): allowlist migration 0007's frozen dialect-tag snapshot (#564) The 0007 instruct-heal migration carries a frozen copy of the design-tag whitelist (incl. Chinese dialect tags) so it stays self-contained; add it to the hardcoded-CJK allowlist like omnivoice/utils/voice_design.py. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
31ba6d3d27 |
fix(transcribe): surface the real ASR-load failure instead of a generic "stream dropped" (#578) (#608)
When WhisperX (or any ASR backend) failed to load its model, the transcribe SSE stream dead-ended on a generic "Transcribe stream dropped … Likely ASR backend failed to load" message with no actionable cause. Two root causes, both fixed: 1. WhisperX loads lazily inside transcribe(), so a load failure (faster-whisper weights, CTranslate2/cuDNN mismatch, torch-2.6 weights-only VAD regression) was buried in per-chunk errors and retried on every chunk. Added ASRBackend.ensure_loaded() (no-op default; WhisperX triggers its lazy loader) and call it in the transcribe pre-flight so the genuine cause surfaces once, up front, as a structured error event. 2. The pre-flight and audio-load error paths closed the SSE stream with a bare `error` and no terminal `done`, so the browser's native EventSource connection-drop could race and win against the structured error — discarding the real cause. Every terminal error now emits `done`, and the frontend latches the structured cause so a connection drop can't overwrite it with the generic message. Adds a fail-before/pass-after regression test driving the stream's async generator through the ASR-load-failure path; updates the existing #516 fake backend to the new ensure_loaded() contract; CHANGELOG ### Fixed entry. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8f2c4bbc5c |
fix(tts): NFC-normalize text + dense-script-aware chunking for long-form quality (#502/#505) (#587)
Two defensive fixes for non-Latin / long-form synthesis quality: #502 (Vietnamese clone distorted/unintelligible): the /generate text path never NFC-normalized its input, so pasted decomposed (NFD) Vietnamese — base letter + combining diacritic instead of the single composed codepoint — reached the tokenizer/model as two characters and rendered as garbled speech. Normalize the input text to NFC at the endpoint (no-op for already-composed text), mirroring what the duration estimator already does so the estimate and synthesis agree. #505 (long-form 5+ min degrades — repeated/skipped/mispronounced): the chunker split purely by character count (800), but CJK/kana/Hangul pack ~1 char = 1 syllable, so an 800-char chunk is ~4-5 minutes of audio in a single shot — past the model's reliable range, where it starts repeating/skipping. When a chunk is predominantly dense-script, cap it to max_chars/2.5 so each chunk's spoken length stays bounded; Latin/spaced text is unchanged. Dense-script detection is by code point (no literal CJK in source — no-literal-CJK gate stays clean). Tests: _dense_char_count, _effective_max_chars (shrink-when-dense, unchanged-for- Latin, disabled-passthrough, floor), and that a 400-CJK-char string now splits (was one chunk) while a Latin paragraph still doesn't. Note: #502's exact distortion still wants a user sample to fully confirm; this is the defensive NFC fix that's correct regardless. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7393ae80f5 |
feat(design): seed pin / re-roll for designed voices (#526) (#577)
Voice design rolled a brand-new random seed on every synth, so tweaking an attribute also re-rolled the whole base timbre — you could never iterate on the "same voice, slightly different". #526 asks for the seed to be shown with a "keep this seed" control. - Backend: `/generate` already accepted `seed` and echoed `X-Seed`, but left `used_seed=None` when nothing supplied one (non-deterministic, unreproducible, empty X-Seed). Now it materializes a concrete random seed when none resolves, so every take is reproducible and the real seed is always returned and stored — this also helps the clone/profile paths, not just design. - Frontend: new store slice (`designSeed`, `keepSeed`); the design synth reuses the pinned seed when "keep this seed" is on (via `pickDesignSeed`) and reads the authoritative seed back from `X-Seed`. Design tab gains a Seed field + "keep this seed" checkbox + "New seed" (re-roll) button. Test: `pickDesignSeed` (pin when kept+valid, re-roll otherwise, range guard). i18n keys added to en.json (other locales fall back; parity probe is advisory). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8ed76b40a6 |
fix(lang): propagate profile/request language into generation + longform (#533/#505/#502) (#565)
Non-English voices drifted to English/wrong-language because the request's or profile's language wasn't reaching the model: - #533: generate_speech() read instruct/ref_text/seed from a resolved profile row but never row['language'] (and collapsed Auto→None), so a German archetype previewed in German yet generated in English on the user's own call (and via Docker/API). Fall back to the profile's stored language when the request didn't pin one; an explicit non-Auto request language still wins. (Frontend already sets the dropdown on profile-select; this is the authoritative backend fix.) - #505 (B2): the audiobook/longform synth hardcoded language=None, so the engine re-autodetected per chunk and a non-English clone flipped language mid-render. Add _resolve_default_language (request → profile → autodetect) and thread the resolved language through _build_synth/_prepare_synth/_render_longform_sse, the three longform request models, the preview path, and the resume manifest. Genuine Auto/unset behavior is unchanged. - #502 (partial): the duration estimator weights combining marks (U+0300–036F) at 0.0, so NFD/decomposed text under-allocated frames → rushed audio. NFC- normalize text at the estimator entry — fixes the whole diacritic-script class (no-op for precomposed text). (The residual "distorted" core still needs the reporter's sample; tracked separately.) Tests (fail-before/pass-after): profile language reaches the engine (German→de; explicit/Auto override semantics); longform synth gets the resolved language (→ja), not None; NFD vs NFC duration parity (Korean Hangul diverges ~3x pre-fix). Full suite: tests/ 1740 passed, backend/tests/ 114 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a7ab148483 |
fix(asr): float16-unsupported GPUs fall back to int8 instead of "no segments" (#561)
#551: both CTranslate2 ASR backends request compute_type="float16" on CUDA with NO fallback. On GPUs without efficient fp16 (older Maxwell/Pascal, GTX 16xx) or a CTranslate2/cuDNN binary mismatch, WhisperModel/whisperx.load_model raise a ValueError at construction — which escaped the existing OOM-only `except RuntimeError`, so every chunk failed and the user got "Transcription produced no segments". Add a per-device compute_type fallback chain (cuda: float16 → int8_float16 → int8; cpu: int8 → float32) to both backends + the ASR sidecar, alongside (not replacing) the existing OOM→CPU path, with an ASR_COMPUTE_TYPE override for exotic hardware (documented in README). Also in the same ASR-robustness pass: - #549: PyTorchWhisperBackend._ensure_pipe wraps the transformers pipeline load and re-raises an actionable error (reinstall transformers / use faster-whisper) instead of a bare "Could not import module 'AutoFeatureExtractor'". - #516: the /dub/transcribe SSE generator is wrapped so it can NEVER close without a terminal event — any unanticipated exception now yields a structured `error` (with build_failure's hint) + `done`, turning "stream dropped, likely ASR failed" into the real cause + Retry. - failure.py: COMPUTE_TYPE_UNSUPPORTED + TRANSFORMERS_IMPORT classes so the no-segments toast is actionable. Tests (fail-before/pass-after): float16-unsupported → int8 for both WhisperX + FasterWhisper; a generic non-OOM RuntimeError still raises; classify() maps the two new classes; the SSE stream always terminates with error→done. 7 + 1 passed, 17 in the failure suite (no regression). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3656f0a4ef |
fix(profiles): decouple design-profile save from TTS render (#476) (#488)
* fix(profiles): decouple design-profile save from TTS render (#476) Saving a design voice profile forced a full TTS model load + inference to render a deterministic identity sample. On a fresh model-less image (Docker first-run) that 503'd, so the save failed. A secondary guard also rejected an all-Auto design (empty instruct) with a 422. Saving a design profile is now a pure persistence operation: - The seed-42 identity sample render is attempted opportunistically but is non-fatal — if the engine isn't ready the row is persisted with ref_audio_path=NULL (sample pending). The row's vd_states + instruct already make the voice fully usable (generation.py falls back to instruct-only conditioning for design profiles with no ref audio). - The sample is rendered lazily + cached on the first GET /profiles/{id}/audio request; if the engine is still unavailable that path returns a precise "model not ready — finish setup / download a model" 503. - The all-Auto (empty-instruct) design is now saveable (vd_states still required). Adds tests/test_profile_design_save_decouple.py (top-level tests/, asyncio.run per test) covering: design save with model unavailable creates the row instead of 503-ing; all-Auto design is saveable; the pending sample materializes on first /audio request. Updates the unification spec (docs-sync). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): contain profile-audio paths under VOICES_DIR (CodeQL CWE-22) The lazy design-sample path was built as `os.path.join(VOICES_DIR, f"{profile_id}.wav")` / `os.path.join(VOICES_DIR, audio_file)` where profile_id is the request path param — CodeQL flagged 5 high-severity path-injection alerts (profiles.py + the taint flowing into archetypes.py's torchaudio save). Add `_safe_voice_path()` (basename + safe-char sanitise + realpath containment, mirroring core.config.dub_seg_path) and route both the read and lazy-render sites through it; a traversal id now 404s instead of escaping VOICES_DIR. Regression test covers the containment guard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): use CodeQL-recognized path-injection guards (CWE-22) The previous `_safe_voice_path()` helper was correct (basename + realpath containment) but CodeQL's taint tracking didn't propagate the barrier through the function return, so the 5 path-injection alerts persisted. Switch to guards CodeQL recognizes, inline at each file-op site: - validate `profile_id` against the generated-id charset (`[A-Za-z0-9_-]{1,64}`) with `re.fullmatch` and 404 on mismatch (covers the `f"{profile_id}.wav"` render path); - read only `os.path.join(VOICES_DIR, os.path.basename(name))` so a stored/derived filename is always a direct child of VOICES_DIR (covers the read + the taint flowing into archetypes.py's torchaudio save). Drop the helper. Test now asserts a traversal/separator/NUL profile_id 404s at the guard. Same security property, recognized by CodeQL. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): inline realpath+commonpath containment for CodeQL (CWE-22) CodeQL didn't recognize the earlier sanitizers — neither the helper (barrier hidden behind a function return) nor os.path.basename / a cross-function regex guard cleared the 5 path-injection alerts. Use the canonical, CodeQL-recognized form INLINE at each file-op site: resolve the path with os.path.realpath (which collapses any `..`) and confirm os.path.commonpath((base, path)) == base before the read / the render, returning 404 / raising on escape. Same property the helper had, now in a shape CodeQL's taint tracking follows. Keeps the profile_id charset guard as defense-in-depth. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): route design-sample path through shared _voices_path guard (#476) The inline realpath+commonpath containment in get_profile_audio and _materialize_design_sample wasn't recognized by CodeQL as a path-injection sanitizer (5 new high-severity py/path-injection alerts at the file-op sites, incl. archetypes.py mkdir via the rendered Path). Both now reuse the existing _voices_path() helper, which applies the os.path.basename() barrier plus symlink-resolved containment — the same guard the consent endpoint uses and that CodeQL already accepts. Behavior is unchanged: the DB columns only ever hold bare {profile_id}.wav filenames, so basename() is a no-op here. Tests: tests/test_profile_design_save_decouple, test_profile_unification, test_profile_consent, test_archetype_blank_guard — 25 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1650a121db |
fix(dub): auto-assign per-speaker voices in multi-speaker dubbing (#486) (#490)
Multi-speaker dubs detected speakers and built per-speaker clones (Voice
dropdown showed "From Video → Speaker N"), but most segments stayed on
"Default" voice and had to be set by hand — inconsistently across runs.
Root cause: after diarization, dub_core stamped each long line (the
default-on per-segment-ref path) with `auto-seg:{id}` as its profile_id.
The dub editor's Voice <select> (and the Cast panel) only render `auto:`
options, so an `auto-seg:` value matched no <option> and silently showed
"Default". Short lines (<3s) fell through to `auto:{speaker}`, which DID
render — hence "sometimes the cloned voice is picked".
Fix: bind every segment to the UI-visible `auto:{speaker}` whenever its
detected speaker has a clone; only fall back to `auto-seg:{id}` when the
speaker has no per-speaker clone at all. The per-segment-ref quality win
is preserved: dub_generate's `auto:` branch now transparently prefers
THIS segment's own per-segment ref (segment_clones[seg_id]) when present,
else the per-speaker clone. Manual overrides and the no-clone path are
untouched; existing jobs that persisted `auto-seg:` ids still resolve.
Tests: tests/test_dub_multispeaker_voice_486.py — assignment binds to
auto:{speaker} (not auto-seg:), never clobbers manual overrides, falls
back to auto-seg: only when the speaker has no clone; generate-time
resolution prefers per-segment ref then per-speaker clone. Green
alongside the existing dub generate/incremental/segmentation suites.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
4e3136c1e0 |
fix(translate): guess source language from text instead of defaulting to "en" (#478)
When neither the request nor the job carries a detected source language, _resolve_source_lang() silently fell back to "en". For non-English audio (e.g. Korean) this produced en -> en, which has no Argos package and failed every segment — even though WhisperX had detected the language correctly (e.g. "Detected language: ko (0.98)"). Add a last-resort script-based guess (ko/ja/zh/ru/ar) from the segment text so the bare "en" fallback no longer breaks non-English dubbing. Co-authored-by: stronghamjji <289942360+stronghamjji@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
18793a99a7 |
feat(audiobook): durable crash-resume for interrupted longform renders (#470)
* feat(audiobook): durable crash-resume for interrupted longform renders
Chapter WAVs were already content-addressed (a re-run reused finished chapters),
but resume only worked if the user could re-submit the EXACT script — impossible
for Stories, whose plan is compiled from cast+lines. This persists the plan
itself so an interrupted render is resumable without the original input.
- New services/longform_resume.py (pure file/JSON): on render start, write a
resume.json manifest (compiled plan + render params + title) into the job work
dir, atomically; clear it on successful completion. read/has/clear/build
helpers, schema-versioned (a foreign/corrupt manifest is ignored, never
resumed).
- _render_longform_sse: accepts an optional job_id + resume flag (resume reuses
the original job row + cached chapters instead of creating a new one); writes
the manifest at start, clears it on done. Both front doors (/audiobook,
/longform/render) unchanged for callers.
- GET /audiobook/jobs — lists interrupted renders (running/failed longform jobs
that still have a manifest; a job left "running" across an app restart is
interrupted by definition), with title + total/done chapter counts for the UI.
- POST /audiobook/resume/{job_id} — rebuilds the plan from the manifest and
replays _render_longform_sse under the original job_id; the content-addressed
cache makes finished chapters instant, so only the unrendered ones synthesize.
404 on unknown id / missing manifest.
Resume durability is best-effort — a manifest failure never blocks the render.
The resume UI affordance is a follow-up (the endpoints are ready for it).
Tests: tests/test_longform_resume.py (7, pure manifest round-trip / version &
corrupt rejection / atomic write — monkeypatches OUTPUTS_DIR, no global
core.config stub so the shared tests/ session isn't polluted) +
backend/tests/test_audiobook_resume_api.py (6, config-stub: jobs-list with
progress, failed-included, done/manifestless/non-longform excluded, resume
404s). 13 passed. CJK green. Stale module docstring updated.
* fix(audiobook): confine resume paths — py/path-injection (CodeQL) + quality
The default-setup CodeQL (security-and-quality suite) flagged the crash-resume
work: longform_resume built filesystem paths from job_id, which on the
POST /audiobook/resume/{job_id} endpoint is a request-supplied path param →
py/path-injection (10 high-severity sinks: open/replace/remove/makedirs/isfile).
- longform_resume.work_dir now confines like profiles._voices_path: reject an
unknown job_type or an id that isn't a bare safe token (^[A-Za-z0-9_-]{1,64}$),
then realpath + startswith(OUTPUTS_DIR + os.sep) — a crafted id (`../`, NUL,
separators) can never escape OUTPUTS_DIR. Returns None on violation; all
callers (manifest_path/read/write/clear/has) degrade gracefully.
- The resume endpoint also gates the path-param id up front (404 on a bad
token) — barrier at the source as well as the sink.
Also cleared the quality alerts the same diff introduced:
- py/repeated-import: the 4 inline `from services import longform_resume` calls
collapse to one module-top import (it's pure, no torch).
- py/empty-except: the best-effort manifest blocks now logger.debug instead of
a bare `pass`.
13 resume tests still pass; all job ids in tests are safe tokens.
* fix(audiobook): sanitize resume job_id at the source (path + log injection)
The first CodeQL pass wasn't enough: resume made job_id request-controlled, so
it tainted not just the manifest paths but the EXISTING work-dir join and the
progress log lines too (py/path-injection + py/log-injection, ~14 alerts).
Fix at the source so the whole dataflow is clean:
- _render_longform_sse strips job_id to a safe token (`re.sub` removing anything
but [A-Za-z0-9_-], capped 64) right after it's resolved — no path separator,
no CR/LF can survive, whether the id came from the resume path param or a
fresh uuid.
- The work dir now routes through longform_resume.work_dir, which adds the
proven os.path.basename(seg)==seg barrier (the shape CodeQL accepts in
_voices_path) on top of the realpath+startswith confinement — so the join and
every path derived from it (meta/concat/out) is sanitized.
- The best-effort manifest-write log no longer interpolates the raw exception
(uses exc_info); clear_manifest's OSError handler returns instead of bare pass
(py/empty-except).
13 resume tests still pass.
* fix(audiobook): launder resume job_id via trusted FS scan (CodeQL path/log-injection)
The custom realpath/regex barriers weren't in CodeQL's recognized sanitizer set,
so the request-supplied resume job_id kept tainting the work-dir/manifest paths
and the progress logs. Switch to the pattern CodeQL does accept — launder the id
through a trusted filesystem enumeration:
- longform_resume.scan_resumable() lists resumable jobs by scanning OUTPUTS_DIR
for <type>_<id>/resume.json; every id it returns is sourced from os.listdir
(never request input).
- POST /audiobook/resume/{job_id} now only resumes an id that scan_resumable()
reports (membership match), and uses the (job_type, job_id) pair FROM that
trusted list for everything downstream — so nothing request-controlled reaches
a filesystem path or a log line.
- GET /audiobook/jobs lists from scan_resumable() too (filesystem-sourced ids).
work_dir keeps the realpath+startswith+basename confinement as genuine defense;
the render path's job_id is now always either a fresh uuid or a laundered id.
13 resume tests still pass.
* fix(audiobook): exact-match allowlist on the work-dir name (CodeQL path-injection)
The remaining 4 path-injection alerts were inside work_dir: I validated job_id
with an anchored regex but then joined a DIFFERENT f-string (`{job_type}_{job_id}`),
so CodeQL didn't carry the sanitization to the joined value. Mirror the pattern
the repo's _safe_cover_path uses (which CodeQL accepts): validate the WHOLE
joined component against an exact-match allowlist regex (_SAFE_SEG_RE), then
confine with os.path.commonpath containment (the recognized barrier) instead of
startswith. 13 resume tests still pass.
* fix(audiobook): basename-sanitize the work-dir name for CodeQL path-injection
The exact-match regex alone wasn't credited; route the joined value through os.path.basename() first — the sanitizer CodeQL recognizes (mirrors _safe_cover_path) — then the regex + commonpath. Functionally identical (no separator in the name) but clears the 4 remaining alerts. 13 tests pass.
* fix(audiobook): allow-list membership guard launders resume job_id (CodeQL)
The next(... if pair[1]==job_id) comparison-select didn't sanitize for CodeQL. Build a dict of resumable ids from the trusted scan and gate with 'if job_id not in resumable' — the membership barrier CodeQL recognizes — then use job_id directly downstream. 13 tests pass.
* fix(audiobook): eliminate request→path flow in resume (definitive CodeQL fix)
Five rounds of recognized path-injection barriers (regex, basename, exact-match,
commonpath, membership-guard) still left CodeQL flagging the resume job_id →
work-dir/manifest/log flow. Remove the flow entirely instead of guarding it:
- scan_resumable() now returns {job_type, job_id, manifest_path} where
manifest_path is built from the os.listdir dir name (trusted), plus
load_manifest_file(path) / discard_manifest_file(path) that operate on those
trusted paths. The request job_id is used ONLY to *select* a scan entry, never
to build a path.
- POST /audiobook/resume/{job_id} reads the manifest via the trusted scan path
and renders under a FRESH server uuid (job_id=None). The chapter cache is
content-addressed (keyed by chapter content, not the job id), so finished
chapters still hit instantly — resume works, but the request's id never names
a work dir, output file, or log line.
- The interrupted job's manifest is discarded (trusted path) once the fresh-id
resume kicks off, so it stops showing as resumable.
Net: no request-controlled value reaches any file operation or log on the
render path (job_id there is always a server uuid). work_dir keeps its
confinement barriers as defence-in-depth. 13 resume tests pass.
|
||
|
|
2b8c8aec7c |
fix: actionable errors for non-executable engine binary (#437) + unreachable backend (#438/#454/#466) (#471)
Two reliability bugs from open issues, both first-run papercuts where the error told the user the wrong thing. #437 — `[Errno 13] Permission denied: bin/omnivoice-tts-linux-x86_64`: a git clone / zip extract on POSIX can drop the bundled binary's execute bit. It only surfaced at spawn time, and the generic synth handler then mislabeled it as "ran out of memory" and told the user to flush the model. - omnivoice_gguf.is_available() now self-heals: after the SHA check confirms the binary is the right file, it adds +x (best-effort) on POSIX; if it can't, it returns a clear "isn't executable — run chmod +x <path>" message instead of a spawn-time crash. No-op on Windows. - generation.py classifies PermissionError / EACCES / "Permission denied" as its own case ("a bundled binary lost its execute bit — reinstall or chmod +x"), so it never again masquerades as OOM. #438/#454/#466 — bare "Failed to fetch" / "NetworkError": when the local backend is still starting, crashed, or the dev server dropped, fetch() throws a TypeError that propagated raw to the user. - client.ts apiFetch now catches the thrown fetch and raises an ApiError with an actionable message ("Can't reach the local OmniVoice backend — it may still be starting up… restart the app or check Settings → Logs"), status:0 to mark a transport failure vs an HTTP error. Tests: client.test.ts +1 (thrown fetch → ApiError status 0 + actionable text); 3 pass. CJK guard green. |
||
|
|
35c063ae52 |
feat(persona): /personas export·import·inspect router + wiring (#29 slice B) (#461)
* feat(persona): .ovsvoice build/parse core + embed_watermark(force=) (#29 slice A) Extends the merged persona-bundle nucleus (constants, normalize_spdx, build_manifest, build_consent_json) with the model-coupled core that the export/import router (next slice) will sit on: - `build_persona_bundle(profile, *, license_spdx, tags, include_reference, embed_fn, …)` → assembles the .ovsvoice ZIP in memory: a watermarked preview.wav (24 kHz mono 16-bit, downmixed + resampled + trimmed ≤8 s), manifest.json, a legacy-shaped metadata.json (so an older OmniVoice can still import the ref audio), optional consent.json, and the raw ref/locked/consent members unless include_reference=False (privacy / preview-only, A12). Raises NoPreviewSource (router → 503) when no source clip is readable (A2-A5). - `parse_persona_bundle(bytes)` → validates the ZIP, prefers manifest.json and falls back to legacy metadata.json, resolves audio members by prefix (last-wins, B9; member names never build paths — zip-slip safe), normalizes the SPDX id, flags preview-only / future-schema_version. Raises BundleError(400|413) for B1-B11. No DB, no file writes. - `ParsedPersona` dataclass with `extract_member(prefix, dest_path)` — the router derives dest_path from the server-generated id, never the member name. - `embed_watermark(..., *, force=False)`: keyword-only flag that bypasses the user's invisible-watermark preference for the mandatory persona preview, but still no-ops without AudioSeal. All existing positional call sites are unchanged (default force=False) — default cross-platform behaviour identical. All heavy imports (torch/torchaudio/watermark/audio_io) are lazy so the module stays model-free at collection (avoids the local torch/Triton segfault). tests/test_persona_bundle.py: +31 cases — parse validation (manifest/legacy selection, preview-only, future-schema, missing/malformed/no-audio → 400, oversize → 413, bad-SPDX normalize, last-wins dup, advisory consent), build round-trip (identity fields, metadata sibling, no-source → NoPreviewSource, include_reference=False, stereo/off-rate downmix+resample), and the force= unit (D1/D3). 25 pure cases pass locally; the 6 torchaudio-coupled cases run on CI (local torch+pytest segfault is pre-existing). CJK guard green. * feat(persona): /personas export·import·inspect router + wiring (#29 slice B) Thin HTTP layer over the persona_bundle service (slice A), registered in main.py next to the legacy marketplace router: - POST /personas/export/{id} → builds the .ovsvoice off the event loop (run_in_executor) and streams it (application/zip, .ovsvoice filename; empty name → persona_<id>). 404 when the profile is missing; NoPreviewSource → 503 (no readable source audio); any other build error → 503 with a generic message (no raw exception text in the body). - POST /personas/import → parse (BundleError → its HTTP status), extract audio members to server-named files ({id}{ext}/{id}_locked{ext}/{id}_consent{ext} — never the member name, zip-slip safe via profiles._voices_path), 17-column INSERT (legacy 13 + the 4 consent columns), event_bus emit after commit. Verified-own-voice is granted ONLY with a real recording ≥ floor AND non-empty consent_text AND consent.json present (forgery guard, B12-B16). Rollback: every written file is deleted on any extraction/INSERT failure; id-collision retries once (renaming the on-disk files to the new id). Accepts legacy .omnivoice too (case-insensitive extension guard). - POST /personas/inspect → manifest + consent summary with NO DB row and NO file extracted (import-preview UI). backend/tests/test_personas_api.py: 13 cases (config-stub pattern → mounts only the router, no main/torch import) — export 404; import bad-ext/non-zip/missing- manifest 400; round-trip row+file under server name; case-insensitive ext; forgery-unverified; verified-with-recording; short-recording-unverified; preview-only-as-ref; legacy .omnivoice; inspect no-write + consent summary. 13 passed locally. CJK guard green. |
||
|
|
ca8a2e8eb8 |
feat(audiobook): PDF ingest for /audiobook/import (ebook-in core value) (#459)
The audiobook importer accepted .txt/.md/.epub but not PDF — the single most common "ebook in" format. Add a pure `pdf_to_chapter_script(data)` that extracts the text layer page-by-page and runs it through the existing chapterizer, so PDFs land in the same `# Heading` + body grammar EPUB and plaintext already produce (one front door onto the unchanged render pipeline). - Dep: `pypdf>=4.0` — pure-Python, MIT, zero native deps, so PDF import behaves identically on macOS/Windows/Linux (default-feature cross-platform rule). EPUB + plaintext stay stdlib-only; only PDF needs a real parser. - Robustness, surfaced as actionable 400s rather than silent empty imports: corrupt file, password-protected (empty-password decrypt attempted first), scanned/image-only (no text layer → clear "scanned PDF" message), and a page-count ceiling. A single unparseable page is skipped, not fatal. - Route: `.pdf` branch in audiobook_import; frontend accept filter + api-client doc updated to `.txt,.md,.epub,.pdf`. tests/test_longform_import.py: 5 PDF cases (extract+chapterize, no-marker single chapter, corrupt, image-only, page-cap) using a hand-built in-memory PDF — no PDF-authoring test dep, mirroring the in-memory-EPUB approach. 16 passed; frontend suite 401; CJK guard green. |
||
|
|
4531e999b1 |
feat(capture): opt-in LLM refinement on REST /transcribe (parity with live dictation) (#457)
The live-dictation socket (capture_ws) already runs the final transcript through the configured local LLM (disfluency/self-correction/punctuation cleanup, Wave 2.1). The REST /transcribe endpoint — the MCP / CLI / file-upload surface — only did the always-on hallucination-loop collapse, so agentic and batch callers couldn't get the same cleaned output. Add an opt-in `refine` form flag that runs the identical `maybe_refine` pipeline off-thread: - OFF by default → existing MCP/CLI callers keep raw-only output and pay no LLM latency (backward-compatible). - Honours the user's Settings → Dictation-refinement config and silently passes through when no LLM backend is configured (cross-platform default parity — identical no-op everywhere with no LLM). - Raw `text` is always returned; `refined_text` is added only when the LLM actually changed the text — same contract the socket emits. tests/test_capture_refine.py: 13 cases — flag-off no-call, refined_text on change, no-op/identical omission, and flag parsing. maybe_refine is patched at its source module since the handler imports it lazily. |
||
|
|
d20c24e1e1 |
feat(longform): two-pass loudnorm measure orchestrator + wiring (#28 slice 2) (#455)
* feat(longform): two-pass loudnorm measure orchestrator + wiring (#28 slice 2) Completes accurate ACX/podcast mastering end-to-end (builds on the pure builders from #28 slice 1). - `services/loudness.py` — `measure_loudness(ffmpeg, concat, preset, *, job_id)`: runs ffmpeg's measure pass, parses the loudnorm JSON → MeasuredLoudness. **Never raises** — skip / non-zero rc / rc None / asyncio.TimeoutError / spawn OSError / empty or unparseable stderr / silent program all WARN + return None → single-pass fallback (a slow/broken measure degrades the master, never aborts the render). Logs rc + a static message only, never the raw stderr (path-safe / local-first). UTF-8 decode with replacement (Windows-cp safe). - `_render_longform_sse` (audiobook.py): between the concat write and the mux, when `loudness` is a known preset (acx/podcast; same `.lower()`/no-strip gate as the builders) → emit a `mastering` event, measure, and pass `measured` into `build_render_cmd` (two-pass apply; `None` → single-pass). `done` gains a `loudness` block {preset, target_i, target_tp, two_pass, measured_i} ONLY for a requested preset — off/None paths keep the byte-identical legacy `done` shape. Both front doors (/audiobook + /longform/render) get it via the shared generator. Chapter cache key is deliberately untouched (loudness-agnostic → acx/off reuse the same cached WAVs; no re-render, no cache-layout break). Tests: `test_loudness.py` (14 — happy fixture, skip-without-spawn for off/ unknown/whitespace/None, non-zero/None rc, timeout-not-propagated, OSError, empty/unparseable stderr, non-UTF-8 stderr, job_id+argv forwarding) + 2 e2e cases (mastering event + done.loudness present for acx; absent for off). Orch tests run locally (stubbed run_ffmpeg, no torch); e2e on CI. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(loudness): lazy-import run_ffmpeg so the measure stub survives sys.modules purges test_loudness monkeypatched services.loudness.run_ffmpeg, but the route-shape fresh_app fixture purges services.* from sys.modules, so under the full-suite ordering the patch missed the re-imported module → real ffmpeg ran → 3 failures. Lazy-import run_ffmpeg inside measure_loudness and patch it at its source (services.ffmpeg_utils.run_ffmpeg) so the stub is always picked up at call time. Verified by running the purging suite + test_loudness together (31 pass). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2a1c3eee3d |
feat(routing): synth-time no-silent-fallback gating at all TTS entry points (#21 follow-up) (#440)
Closes the last #21 gap: a per-request engine=/model= override bypasses the /engines/select host-gate, so an engine that can't use this host's GPU could still be triggered at synth time and silently fall back to CPU (or die mid- synth). Now enforced at every TTS synth entry point, reusing the SAME probe + resolver — never re-deriving routing. Shared helpers (services/engine_routing.py): - `routing_notice(result)` → (status, reason) to surface, or None. Fires for cpu_fallback (always) and accelerated-with-caveat (driver/arch); silent for cpu_only / clean-accelerated / n/a. - `header_safe_reason(reason)` → scrubbed + ASCII-sanitized (headers are latin-1; a non-ASCII device name would 500 otherwise) + ≤256 chars. No regex. Entry points: - REST `POST /generate` (generation.py): after engine resolution, resolve routing once; `unavailable` → 400; cpu_fallback / accelerated-caveat → 200 + `X-OmniVoice-Routing` + `X-OmniVoice-Routing-Reason` headers on the WAV StreamingResponse; benign → no headers. Covers OmniVoice + adapter branches. - OpenAI-compat `POST /v1/audio/speech` (openai_compat.py): same gate + same headers; the tts-1/tts-1-hd alias inherits the active engine's routing. - WebSocket `/ws/tts` (tts_stream.py): no headers → frames. `unavailable` → `{"type":"error",...}` + skip stream; cpu_fallback / caveat → one `{"type":"routing","status","reason"}` frame before any audio. - `select_engine` response now echoes routing_status / effective_device / routing_reason (PR #432 added the gate; this adds the fields so the UI can warn on a cpu_fallback pick). New fields on SelectEngineResponse. Frontend: `useTTS` reads the X-OmniVoice-Routing header and shows a one-time, non-blocking toast (in-memory de-dup by status — a 50-clip batch fires once, no localStorage). i18n keys `tts.routingFallback`/`tts.routingCaveat`. Tests: routing_notice + header_safe_reason (ASCII/length/scrub) unit tests; REST synth gate (unavailable→400, cpu_fallback→headers, cpu_only→none) via the fake-engine harness with a mocked host; select response routing fields. Deferred (small follow-up): dub-pipeline ASR routing note on the preflight_error SSE channel — separate path, not a TTS synth entry point. No frontend /ws/tts client exists today (the routing frame serves external API consumers). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c6a55794da |
feat(routing): active-engine GPU verdict in preflight + diagnose (#21 PR 4/5) (#433)
Surfaces a routing verdict for the CURRENTLY-SELECTED TTS engine in the two system-health surfaces, so a CPU fallback / unavailable-GPU is heard about before a slow or failed synth — the no-silent-fallback contract, read-only. - `tts_backend.active_routing()` + `gpu_routing_verdict()`: the active engine's routing derived from list_backends() (byte-identical to the matrix) plus the host compute summary (family + VRAM from the canonical probe). Never raise. - `/system/diagnose` gains a `gpu_routing` check: accelerated→ok, accelerated-with-caveat / cpu_fallback→warn (+ actionable hint), cpu_only→ok (no-GPU host is the expected normal state — never noise-warns), unavailable→ fail, no-engine→warn. ASCII-safe detail strings (the text dump enforces ASCII). - `/setup/preflight` gains an "Active engine routing" check + an explicit `gpu_routing` object on PreflightResponse (a real field — the response has no extra="allow", so it would otherwise be dropped). `device` gains `gpu_family` (ROCm-vs-CUDA aware) + `vram_gb`. New `GpuRouting` schema. Tests: gpu_routing_verdict (host + active-engine + degraded), diagnose status mapping across all 6 states + never-raises, preflight gpu_routing object + check + device.gpu_family. Existing diagnose/preflight tests stay green (checks are additive; the report's top-level key set is unchanged). Deferred (documented): synth-time routing headers/WS-frames at the 3 synth entry points. Selection is already hard-gated (PR 3 select_engine), and the matrix (PR 5) + this preflight/diagnose verdict surface the situation — the synth-time signal is incremental belt-and-suspenders for the env-var-pinned edge and is best validated interactively. Tracked as a #21 follow-up. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8c8d525397 |
feat(routing): wire effective-device into /engines + select gate (#21 PR 3/5) (#432)
* feat(routing): wire effective-device + routing_status into /engines (#21 PR 3/5) Surfaces the PR-1 probe + resolver through the engine registries so the matrix UI (PR 5) and the no-silent-fallback gates can consume it. - `engine_routing.routing_fields()`: shared helper returning the three serialization-ready keys, centralizing the scrub rule — routing_reason is scrubbed via `core.scrub.scrub_text` only when truthy, so a None reason stays JSON `null` (never coerced to ""). - TTS/ASR `list_backends()` each gain `effective_device` / `routing_status` / `routing_reason`, computed from a SINGLE `detect_host_caps()` call per request (host caps are constant per process). ASR is brought to full TTS parity: it now also carries `install_hint` / `last_error` / `isolation_mode` and a SCRUBBED `reason` (closing a pre-existing ASR token-leak gap) — an identical 11-key shape across families. ASR also gains the same is_available()-raises resilience TTS has (degrade to available:false, never 500). - LLM `list_backends()` reaches 11-key parity too but emits literal `effective_device:"network"` / `routing_status:"n/a"` / `routing_reason:null` (NOT via resolve_routing — LLM runs no local GPU model). `LLMBackend.gpu_compat = ()`. "network" is a label, not a probe — nothing here touches the network. - `select_engine` host-routing gate: refuses a pick whose `routing_status` is `unavailable` on this host (400 with an actionable detail), while ALLOWING `cpu_fallback` (it runs, just slower). LLM is never gated. Defensive `.get` so legacy payloads still select. New typed `SelectEngineResponse`. Tests: 11-key shape across all 3 families, well-formed tts/asr routing keys (+ None-not-"" contract), LLM network/n/a labels, select gate (block unavailable / allow cpu_fallback / never-gate LLM). Updated the registry exact-shape test for the 3 new keys. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(cjk): allowlist docs/specs/ in the hardcoded-CJK guard PR #429 merged the longform design specs, which legitimately quote functional CJK (test-fixture descriptions, CosyVoice speaker IDs, multilingual sample text). The CJK guard scans every tracked file, so those docs turned main red. Specs are documentation, not shipped UI strings — allowlist the docs/specs/ prefix, matching the individually-allowlisted docs already in the set. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
000010ebb8 |
feat(dub): dedicated Dub home (projects/history) + project rename (#435)
The dub Projects + History rail (WorkspaceProjects/WorkspaceHistory) used to
sit beside the editor at all times. Now it's a landing: shown only when no
project is being edited (dubStep === 'idle'); opening/creating one switches to
a full-width editor. (The global Sidebar is already hidden in dub mode, so the
studio-right rail is the only surface — no Sidebar change needed.)
Adds project rename:
- backend: PATCH /projects/{id} updates just the name (400 on empty, 404 on
missing) — lighter than PUT which rewrites the whole state blob.
- api: renameProject(id, name); App.jsx renameProject handler (updates the
active-project label + refreshes the list).
- UI: inline rename on each project card (pencil → edit → Enter/Save / Esc).
Verified: PATCH create→rename→list / 400 / 404; frontend typecheck:ci clean.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
dc1d36fe5f |
refactor(models): model-management v2 cleanup (mm2, all tiers) (#428)
One coherent lifecycle surface over the in-process model, diarization, and subprocess sidecars; fixes the engine-switch VRAM leak; tightens download robustness. Backend-only, response shapes preserved, no new deps. Tier 1 — correctness: - MM2-01: get_active_tts_backend() caches one instance per backend id and unload()s the outgoing engine on switch (fixes the VRAM leak behind #278); adds reset_active_backend(). - MM2-02: OmniVoiceBackend.unload() releases the shared model_manager singleton + free_vram(); SubprocessBackend.unload() -> unload_sidecar(self.id), inherited by all sidecar engines. Idempotent + preload-safe. - MM2-03: /model/loaded ASR row reports the real device + a note explaining the disabled unload button. Tier 2 — single surface: - MM2-04: new services/model_lifecycle.py owns list_loaded/unload/unload_all/ free_vram; system.py routers are thin delegations (shapes unchanged). - MM2-05: idle timeouts (in-process + sidecar) resolve via prefs.resolve (env wins, no restart); removed the duplicated _IDLE_TIMEOUT_SECONDS. Tier 3 — robustness/observability: - MM2-06: _install_cooldowns swept (1h TTL) + cleared on success — bounded. - MM2-07: per-extension weight floors (onnx 64KB, tensors 5MB) OR the original >=5MB catch — small ONNX no longer false-flagged, #352 still caught. - MM2-08: indextts GPU sidecar self-reports vram_mb in pong; parent surfaces it in list_live_sidecars (0 = CPU/unmeasured). - MM2-09: is_cached scan_cache_dir->disk fallback logs WARNING w/ exc type (#117/#118), was invisible at DEBUG. Tests: tests/test_mm2_lifecycle.py (15). Full suite: 1379 passed. Plan/summary: .planning/quick/260613-mm2-clean-model-management-v2/. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
4cc55ab852 |
Fast model downloads: Xet fast path + accurate progress (FDL W0–W2 + W4) (#424)
* feat(downloads): Xet fast path + accurate progress (FDL W0–W2)
Make model downloads fast and show accurate downloaded/remaining/speed.
Research confirmed hf-xet already implements the IDM/uGet technique
(content-defined chunking, parallel byte-range gets, dedup, resume), and
the spike found all 25 catalog repos are Xet-backed — so the win is
driving Xet well + accurate progress, not a custom downloader.
W1 — maximize + guarantee Xet:
- pin huggingface_hub>=1.7 + hf-xet>=1.1 (was transitive); no hf_transfer
- drive snapshot_download with explicit tqdm_class + max_workers + endpoint
- opt-in HF_XET_HIGH_PERFORMANCE / HDD sequential-write knobs (default off)
- /system/info reports fast_download {xet_enabled, xet_version, high_perf}
W2 — accurate progress:
- dry_run preflight -> install_plan event (exact total/cached/remaining)
- utils/download_aggregator.py: one overall bar; byte bars (by id) vs the
"Fetching N files" count bar; windowed rate; emits one 'aggregate' event
- frontend overall bar (speed/remaining/ETA), cached-skip, ⚡ fast badge
Known limit (verified live): under Xet+hf_hub 1.7.2 per-file byte bars
never advance/close via tqdm, so mid-download the bar is file-granular and
bytes flush to the exact total on completion. Classic-LFS/mirror repos get
true byte progress (W4).
Drive-by: download.py used os.walk without importing os (latent NameError
in _validate_snapshot_has_weights on every install) — fixed.
Tests: tests/backend/setup/test_download_preflight.py (10). Spike + plan
under .planning/quick/260613-fdl-fast-model-downloads/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in mirror + cancel + docs (FDL W4)
- mirror (FDL-10): snapshot_download(endpoint=) honours prefs hf_endpoint /
env HF_ENDPOINT on preflight + download (per-call, no process-wide env).
Documented as the classic-LFS path (no Xet) for restricted networks.
- cancel (FDL-11): POST /models/install/cancel {repo_id} stops further
retries at the next boundary, emits install_cancelled, clears the cooldown
(cancel is intent, not failure). Frontend treats it as a terminator.
- docs (FDL-12): docs/downloading-models.md (Xet fast path, progress
semantics + byte-speed limitation, opt-in tuning, mirror, cancel,
troubleshooting) + README pointer. Docs-sync rule satisfied.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(planning): model-management v2 cleanup plan (mm2)
GSD plan for cleaning the model-management subsystem: registry unload-on-
switch + per-engine unload() (fixes VRAM leak), model_lifecycle facade,
unified idle/timeout config, bounded cooldowns, sidecar VRAM self-report,
cache-fallback logging. Planning artifact only — no code.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(downloads): reconcile with main's HF_HUB_DISABLE_XET; honest status
Rebasing onto main surfaced that main forces HF_HUB_DISABLE_XET=1 (classic
LFS) because Xet progress bypasses the tqdm hook — the same limitation found
here. Reconcile instead of fight:
- /system/info fast_download now reports runtime truth: xet_installed +
xet_active (installed AND not HF_HUB_DISABLE_XET) + xet_enabled alias. The
⚡ badge only shows when Xet actually runs; startup log says
"downloads: Xet disabled → legacy LFS".
- complete(): clear the rate window before the final flush so crediting the
full size in one step can't emit an absurd instantaneous rate.
- docs/downloading-models.md rewritten: default is legacy LFS for accurate
progress; Xet is opt-in via HF_HUB_DISABLE_XET=0. hf-xet pin stays (ready
for a future Xet progress hook).
W2 (preflight total/remaining + aggregate bar + exact completion) is the
value on either path; W1's "maximize Xet" is dormant by main's design.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in segmented multi-connection accelerator (FDL W3)
Since main forces Xet off (HF_HUB_DISABLE_XET=1), the default path is
single-stream legacy LFS — so a segmented downloader is the way to get BOTH
parallel speed and live byte progress.
- services/segmented_download.py: async multi-connection Range downloader for
one file — parallel byte-ranges, resume (.part + manifest), per-segment
short-read truncation guard, optional sha256/etag verify, cancel, and a
single-stream fallback when the server won't range. Auth-safe: the HF
Authorization header is sent only to huggingface.co/hf.co and never
forwarded to a CDN host on redirect (unit-tested).
- dispatch (download.py): opt-in via prefs segmented_downloader / env
OMNIVOICE_SEGMENTED_DOWNLOAD (default off). When on and Xet inactive,
fetches each file into the HF cache mirroring hf_hub_download (blobs +
snapshot symlinks + refs/main), feeding real bytes to the aggregator. Any
failure falls back to snapshot_download — never breaks a correct install.
- fix: complete() was adding a full total on top of accumulated segmented
bytes (2x); now replaces byte bars so the sum is exactly total.
Verified live (accelerator on): real byte progress to ~16.6 MB/s, final
bytes==total, /models installed=True, delete frees correctly.
Tests: test_segmented_download.py (7) + aggregator double-count regression.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(downloads): relocate FDL tests to top-level; loop-isolate segmented test
CI runs the full suite, which exposed a pre-existing test-isolation leak:
several tests/backend/** fixtures purge core.*/services.* from sys.modules
under a temp OMNIVOICE_DATA_DIR and never restore, leaving core.config/core.db
bound to a dead temp dir. It only bites when collection order puts a purging
test ahead of a real-DB reader (test_longform_jobs). Adding tests under
tests/backend/setup/ reordered collection and tripped it.
Fix without touching the shared (fragile) fixtures or risking class-identity
breakage from a blanket sys.modules restore:
- move the two FDL test files to top-level tests/ (tests/test_fdl_*.py) so
tests/backend/** collection order is identical to main — longform passes.
- rewrite the segmented test to run each case under asyncio.run() (fresh loop)
instead of asyncio.get_event_loop(), which an earlier async test can leave
closed in the full suite.
Full suite green locally: 1364 passed, 0 failed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
e196d790cc |
fix(longform): evict oldest chapters from the render cache (review fast-follow) (#423)
The content-addressed longform_cache/ accumulated uncompressed chapter WAVs across every render with no bound (a review finding). Add prune_cache_dir() — LRU-by-mtime eviction down to a 2 GB ceiling (OMNIVOICE_LONGFORM_CACHE_MAX_GB); best-effort, never raises. Called at the start of each render job, before its chapters are written, so the fresh ones are never the eviction target. Tests: under-cap no-op, evicts-oldest-keeps-newest, missing-dir safe. 38 green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dde43de5a4 |
feat(longform): pronunciation lexicon — per-render word respelling (PR 8a) (#419)
Lets a render correct hard-to-say words (e.g. {"GIF":"jiff","Dr":"Doctor"}).
Backend wiring; the editor UI folds into the full-width Audiobook redesign.
- services/pronunciation.py (parallel-built, 19 tests): apply_lexicon —
whole-word, case-insensitive, longest-first, word-boundary, single ReDoS-safe
re.sub pass; + normalize/load/save_lexicon (JSON).
- synthesize_chapter gains a `lexicon` kwarg, applied to each span's text before
chunk splitting (None/empty = no-op → backward compatible).
- _render_chapter_cached folds the normalized lexicon into the chapter cache key
(a lexicon edit re-renders); threaded through _render_longform_sse + the
/audiobook, /audiobook/preview, /longform/render request models.
Tests: synthesize_chapter respells via lexicon; pronunciation module (19);
75 related backend tests green.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
18e4c2347a |
fix(longform): correctness + robustness fixes from adversarial review (#418)
* fix(longform): correctness + robustness fixes from adversarial review Fixes the confirmed findings from a multi-agent review of the convergence: HIGH (correctness/output): - MP3 + cover produced a corrupt file (-map 2:v -c:v copy is invalid for mp3). Cover art is now embedded for M4B only; mp3 skips it (m4b is the cover format). - Chapter cache key omitted ref_text — editing only a profile's ref_text served stale audio. ref_text is now part of the voice signature. - Preview wrote audiobook_cache/ but the render reads longform_cache/ (rename missed in PR 5) → cache-warming silently broke. Unified to longform_cache/. Robustness (DoS/OOM guards): - /audiobook/import caps upload at 64 MB; epub_to_chapter_script bounds per-entry (25 MB) and cumulative (300 MB) uncompressed reads (zip-bomb guard). - /longform/render rejects > 10,000 chapters (422). Frontend leaks: - StoriesEditor.removeTrack revokes the line's preview blob URL. - AudiobookTab revokes the cover blob URL on replace/unmount. Deferred fast-follows (also from review): render-cache disk eviction; restoring the standalone chapter cue-sheet export (needs chapter times in the done event). Tests: mp3-drops-cover, epub entry/total caps, import + chapter-count limits; updated the cache-hit test for the 4-field voice sig. 70 backend + 334 frontend green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(longform): pass EPUB caps as params, not monkeypatch (CI import-path fix) The cap tests monkeypatched module constants, but in the full-suite CI context the module loads under a different import path so the patch missed the function (it used the real 300 MB cap → tests failed). epub_to_chapter_script now takes max_entry_bytes/max_total_bytes kwargs (default to the constants); tests pass small values directly — deterministic regardless of import path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
36e7fb12fc |
feat(longform): job library — finished books/stories in Projects (PR 7/8) (#417)
Surfaces finished Audiobook + Story renders so they're re-downloadable from the Projects view — closing the resume/history loop of the convergence. Backend (new, no migration — reads existing job_store rows): - routers/longform_jobs.py: GET /longform/jobs lists finished audiobook/story jobs newest-first, recovering output/chapters/duration from each job's persisted 'done' SSE event. Pure build_longform_library() over the job_store callables; defensive (skips unparseable jobs, never 500s). Registered in main.py. Frontend: - Projects.jsx: new "Audiobooks" category fed by /longform/jobs; each row opens the rendered file (/audio/<output>) with type/chapters/duration. Offline-safe (empty on fetch failure). en.json keys added. Built via parallel worktree agent; backend tests/test_longform_jobs.py (9) green; 334 frontend tests + build clean. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
00f400e4c7 |
fix(stories): thread per-line speed through the shared renderer (PR 6/8) (#416)
PR 5 moved Stories' full export to /longform/render but dropped per-line **speed** — the old client export sent each line's speed to /generate; the converged path silently ignored it. This restores it end-to-end. - Span gains an optional `speed`; synthesize_chapter passes it to the injected synth (signature now `synth(text, voice_id, speed)`); both engine paths (OmniVoice model + generic TTSBackend) forward it to generate(speed=…). - chapter_cache_key now includes speed (a speed change re-renders; tuples accept an optional 4th element so existing 3-tuple callers/tests still work). - LongformSpan + /longform/render carry speed; storyToSpans emits each line's speed onto its spans. Emotion note: per-line tone is already model-native via inline tags ([laughter] etc.) inserted into the text, so no separate emotion→instruct plumbing is needed — the dead `emotion` store field stays unused/superseded. Tests: storyToSpans speed passthrough (8); cache-key speed sensitivity; synth stubs updated for the 3-arg signature. 65 backend + 334 frontend green; build clean. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0f67895585 |
feat(stories): full export → shared server-side renderer (PR 5/8) (#413)
* feat(stories): full export → shared server-side renderer (PR 5/8)
The convergence core. Stories' full export no longer stitches audio in the
browser (Web Audio, capped by RAM, no resume/loudness/markers) — it compiles
cast + lines into a chapter/span plan and streams through the same chapterized
renderer the Audiobook tab uses.
Backend:
- Extracted the audiobook SSE job into a shared `_render_longform_sse(plan, …)`
generator (resume cache, per-chapter fault isolation, mux). /audiobook is now
a thin caller.
- New POST /longform/render — accepts a pre-built {chapters:[{title,spans:
[{voice_id,text,pause_ms_after}]}]} plan (+ format/loudness/cover/metadata) and
renders it. Pause-only spans (empty text) are kept as silence. job_type=story.
- Shared content-addressed cache renamed longform_cache (one render per unique
chapter across both front doors).
Frontend:
- storyToSpans(tracks, cast) — pure compiler: `# ` lines → chapters; each line
resolves its cast/override voice; inline [voice:]/[pause] split into spans;
pauses fold into the previous span.
- StoriesEditor.generateAll now posts via longformRender and downloads the
server file (chaptered M4B / MP3). Single-line preview stays client-side;
stems export unchanged. Format select WAV→M4B.
Deferred to PR 6 (with the component split): per-line regenerate, emotion→instruct.
Tests: storyToSpans (7) — cast resolution, chapters, per-line + inline voice,
pause folding, empty-drop. 64 backend + 333 frontend green; build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): confine cover_path to OUTPUTS_DIR + don't leak exception text (CodeQL)
- _safe_cover_path() restricts the user-supplied cover to OUTPUTS_DIR before it
reaches ffmpeg (py/path-injection).
- SSE error events now emit a generic message and log the detail server-side
(py/stack-trace-exposure); empty best-effort excepts annotated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): cover path via basename+fixed dir (clears CodeQL py/path-injection)
CodeQL didn't recognize realpath+startswith as a barrier; os.path.basename is a
recognized sanitizer. Covers only come from /audiobook/cover (OUTPUTS_DIR/
audiobook_covers), so rebuilding from the basename onto that fixed dir is both
CodeQL-clean and strictly tighter — no caller path can escape it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): regex-allowlist cover filename (clears CodeQL py/path-injection)
basename alone wasn't a barrier CodeQL credits. Restrict the cover name to the
exact pattern /audiobook/cover emits (12 hex + jpg/jpeg/png) before joining onto
the fixed covers dir — an anchored-regex guard CodeQL recognizes as sanitizing,
and strictly tighter than before.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): commonpath-confine resolved cover path (CodeQL py/path-injection)
Add an os.path.realpath + os.path.commonpath containment check on the resolved
cover path (the barrier static analysis recognizes), on top of the regex
allowlist + basename. Defense in depth; the path provably cannot escape
OUTPUTS_DIR/audiobook_covers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ea6833138b |
feat(audiobook): text + EPUB import → auto-chapter (PR 4/8) (#412)
Spec PR 4. A front door onto the existing chapter parser: import a file, get a
chapter-delimited script in the editor.
Backend (new services/longform_import.py — pure, stdlib only, no new dep):
- chapterize_plaintext(text): inserts `# ` headings ahead of short standalone
chapter-title lines (Chapter/Part/Prologue/…); no-op if the text already has
H1s; long "Chapter …" sentences stay prose. ReDoS-safe (anchored, per-line).
- epub_to_chapter_script(bytes): parses EPUB (zipfile + ElementTree +
html.parser) in spine order → `# Title` + stripped body per document; skips
empty/nav pages; the heading becomes the chapter title (not narrated). Raises
ValueError on a malformed EPUB. ET.fromstring annotated `# nosec B314` (local
user file, no external-entity expansion).
- POST /audiobook/import (UploadFile) → {text, chapters}.
Frontend: an Import button (.txt/.md/.epub) that fills the script editor.
Tests: tests/test_longform_import.py (9) incl. an in-memory synthetic EPUB
(spine order, empty-doc skip, tag stripping, bad-zip). 64 backend + 326 frontend
green; build clean; en.json valid.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
7af5143fac |
feat(audiobook): per-chapter preview + resume + chapter fault-isolation (PR 3/8) (#411)
* feat(audiobook): per-chapter preview + resume + chapter fault-isolation (PR 3/8) Builds on the shared core (#408) and metadata UI (#409). Chapter-level control, the spec's PR 3. Shared core: - chapter_cache_key(spans, sr, engine_id, voice_sig) — deterministic content hash of a chapter's audio inputs. Same inputs → reuse; any change (text, voice, order, pauses, sr, engine, resolved-voice signature) → re-render. Backend (audiobook router): - Chapter WAVs are now content-addressed in OUTPUTS_DIR/audiobook_cache. A re-run after a failure/interruption reuses already-rendered chapters and only synthesizes the missing/changed ones (resume). Job emits `cached` per chapter and `cached_chapters`/`failed_chapters` on done. - Per-chapter fault isolation: a chapter that throws emits `chapter_error` and the job continues; the m4b assembles from the successful chapters. Re-running retries only the failed (un-cached) chapters. - POST /audiobook/preview — render a single chapter to audition it; shares the same cache so a preview warms the full run and a re-preview is instant. - _build_synth now exposes resolve + engine_id; _prepare_synth unifies the omnivoice/generic paths for both the job and preview. Frontend: - Plan view: a ▶ preview button per chapter with inline playback. - Done panel: "reused N chapters" + "N failed — click Create to retry" notes. Tests: chapter_cache_key determinism + sensitivity (8); preview validation + cache-hit-skips-synth (3). 55 backend + 326 frontend green; build clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(audiobook): mark cache-key SHA1 usedforsecurity=False (bandit B324) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
086ac08592 |
feat(audiobook): metadata, cover art, format + loudness UI (PR 2/8) (#409)
Surfaces the shared-render-core capabilities (PR 1, #408) in the Audiobook tab. Backend: - POST /audiobook/cover — multipart cover upload (jpg/png, 8 MB cap), returns a server-side path passed back as cover_path. Unit-tested via the handler directly (no main+torch import). Frontend: - api/audiobook.ts: AudiobookGenerateBody (format/loudness/cover_path/metadata) + audiobookUploadCover(file). - AudiobookTab: format select (M4B/MP3), loudness select (off/ACX/podcast, default off), and a "Cover & details" panel — cover picker with preview + title/author/narrator/year/genre/description. On create, the cover uploads first, then the job runs with metadata + format + loudness. - en.json: audiobook.* keys for the new controls. Tests: tests/test_audiobook_cover.py (4) green; frontend vitest 326 green; prod build clean; CJK + i18n-parity gates pass. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e9481ef307 |
feat(longform): shared render core — loudness, metadata, cover art (PR 1/8) (#408)
First slice of the Stories+Audiobook convergence (spec: docs/specs/2026-06-13-stories-audiobook-maturity.md). Both features will compile to one server-side chapterized renderer; this lands the shared pure builders and wires them behind Audiobook. New `backend/services/longform_render.py` (all pure, unit-tested without ffmpeg/torch): - build_ffmetadata(chapters, global_meta) — FFMETADATA1 with an optional global tag block (title/author→artist/narrator→composer/year→date/genre/description→ comment) + chapter table. - build_loudnorm_filter(preset) — `-af loudnorm` for ACX (~-19 LUFS, -3 dBTP) or podcast (-16 LUFS); off/unknown → None. Opt-in, so default behavior stays platform-identical. - validate_cover_image — jpg/png + 8 MB cap guard. - build_render_cmd — generalizes the m4b mux: m4b|mp3, optional cover (attached_pic) + loudness, bitrate validated. - build_concat_list — moved here. `services/audiobook.py`: build_chapter_ffmetadata / build_m4b_cmd / build_concat_list are now backward-compatible wrappers over the core (existing imports + tests unchanged). `POST /audiobook`: now accepts optional `format` (m4b|mp3), `loudness`, `cover_path`, and `metadata` and passes them through — backend-complete; the UI for these lands in PR 2. Tests: tests/test_longform_render.py (28) + existing test_audiobook.py (11) green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
599f3bcc5c |
feat(engines): on-demand unload of subprocess-engine sidecars (Action 13) (#406)
Completes the dynamic engine load/unload slice. The idle reaper (#401) frees sidecar VRAM after 5 min; this adds a user-initiated "free VRAM now" path so multi-engine users don't have to wait: - subprocess_backend: `list_live_sidecars()`, `unload_sidecar(id)`, `unload_all_sidecars()` via a shared `_force_reap(predicate)` — busy-guarded exactly like the idle reaper (non-blocking lock; a sidecar mid-synth is skipped, never interrupted; next request respawns it). - system.py: `/model/loaded` now surfaces live sidecars as unloadable rows; `/model/unload/{sidecar:<id>|sidecars}` frees one or all. The existing generic flush panel picks these up with zero frontend change. Also refresh CLAUDE.md stale version notes: main is 0.3.6 (latest release v0.3.5 + 1 patch); the v0.3.0-as-unreleased framing in the project/cadence notes is corrected to the v0.3.x continuous-to-main reality. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6704d062fc |
fix(persona): preserve design kind + vd_states across share/import (Wave 5 §R3) (#405)
The persona-gallery surface already exists (VoiceGallery Community zone + community.py manifest + marketplace .omnivoice bundles). The blocker for §R3's 'synthetic-only' gate was data integrity: a *designed* persona lost its kind='design' (and vd_states) when imported from the community gallery or round-tripped through a bundle — silently demoting it to a clone. - community.py /use: a 'preset' (rendered from instruct) imports as kind='design'; a 'voice' (real reference clip) as 'clone'. - marketplace.py: extract a pure _bundle_metadata() (dedupes export+publish) that captures kind + vd_states; import restores them. Old bundles without the keys import as 'clone' (backward-compatible). This makes 'accept only designed/synthetic voices' enforceable instead of everything defaulting to clone. No new persona-gallery feature was built — that would duplicate the existing community/marketplace surface. 4 torch-free tests (isolated DB): _bundle_metadata captures design + defaults to clone; import round-trip preserves design kind+vd_states; legacy bundle → clone. docs §R3 status updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9441274ab6 |
feat(audiobook): synth job → chapterized m4b, SSE progress (Wave 5) (#403)
Completes the audiobook backend: POST /audiobook renders each chapter through the active TTS engine (synthesize_chapter + chunked_tts), writes per-chapter WAVs, then muxes a chapterized m4b (FFMETADATA1 chapters via build_m4b_cmd + concat demuxer). Progress streams as SSE (started/chapter/assembling/done/ error), recorded to job_store. ffmpeg-gated — emits an error event and stops when ffmpeg is absent (m4b is the only output). - services/audiobook.build_concat_list: pure ffmpeg concat-list builder with proper single-quote escaping (no arg injection). Unit-tested. - router: voice resolution (compact form of generation.py's locked/design/ clone cases) cached per id; OmniVoice native model path + generic TTSBackend path; chapter synthesis runs on the GPU pool, ffmpeg via run_ffmpeg. Reuses the tested building blocks from #402 (parser, synthesize_chapter, FFMETADATA + m4b argv builders) — the new router glue is thin and import-checked by CI. Deferred: epub/pdf ingest, ACX loudnorm mastering, crash-resume, UI. 15 audiobook tests (added concat-list); docs §R3 updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
34b47282af |
feat(audiobook): chapterized audiobook core + plan preview (Wave 5) (#402)
* feat(audiobook): chapterized audiobook core + plan preview (Wave 5) First cut of the long-form vertical (parity §R3). Engine-agnostic core in services/audiobook.py: - parse_audiobook_script: pure parser. Markdown '# H1' headings → chapters; inline [voice:NAME] switches the narrator; [pause …] is delegated to the shared omnivoice.utils.text.parse_pause_markers so audiobooks and single-shot synthesis keep one pause dialect. Returns a chapter/span plan. - synthesize_chapter: orchestration via an injected synth(text, voice) callable (reuses chunked_tts split + crossfade, stitches inter-span silence) — so it's unit-testable with a stub backend, no model/GPU. - build_chapter_ffmetadata + build_m4b_cmd: pure FFMETADATA1 [CHAPTER] builder and faststart-m4b concat-demux argv (bitrate-validated, no injection). POST /audiobook/plan returns the parsed plan (no TTS/ffmpeg, no side effects). Deferred (follow-ups): the streaming synth job + chapterized-m4b run, epub/pdf ingest (new dep), ACX loudnorm mastering, crash-resume, UI. 14 tests: parser (chapters/voice/pause/intro/empties/to_dict), FFMETADATA offsets+escaping, m4b argv + bitrate guard, and stub-synth orchestration (span+silence stitching, voice threading). docs §R3 status updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): linear-time regexes (CodeQL ReDoS) CodeQL flagged polynomial backtracking on user-provided input in three regexes reachable from the new POST /audiobook/plan endpoint: - _VOICE_RE: \s*(...)\s* → single [^\]]* class, stripped in code. - _HEADING_RE: trailing [ \t]* removed; title captured greedily + stripped. - _PAUSE_RE (omnivoice/utils/text.py): the numeric spec is now an atomic group (?>…) so its leading \s+ can't backtrack against the trailing \s*. Behavior-preserving (Python >=3.11 already required); 14 pause tests + 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): require non-space heading title start (CodeQL ReDoS) The previous _HEADING_RE '[ \t]+(.+)' still let the leading whitespace class and the title '.+' both match the same tab run (overlap → polynomial). Anchor the title capture with \S so the two can't overlap. 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): exclude '[' from voice-tag content (CodeQL ReDoS) [^\]]* still matched '[', so a run of nested [voice: prefixes produced overlapping finditer match attempts → O(n^2). Excluding both brackets ([^\]\[]) makes matches non-overlapping and linear. A voice name never contains a bracket. 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e8705a106d |
feat(dictation): opt-in NLMS AEC for dictate-over-playback (Wave 8b) (#399)
* feat(dictation): opt-in NLMS AEC for dictate-over-playback (Wave 8b) Dictating while OmniVoice plays audio (TTS preview, dub, video) leaks the loudspeaker signal into the mic, and the streaming ASR transcribes that bleed. Browser echoCancellation varies per platform/webview — it can't be a cross-platform default — so this adds a server-side canceller that behaves identically everywhere. services/aec.py ports Patter's NlmsEchoCanceller (MIT): a time-domain NLMS adaptive filter with a Geigel double-talk detector, warm-up step ramp, and far-end staleness pass-through. /ws/transcribe gains an opt-in '?aec=1[&sr=]' mode: frames are raw int16 mono PCM tagged with a 1-byte prefix (0x00 mic, 0x01 playback reference); the mic is cleaned against the reference before buffering, and the cleaned PCM is muxed via stdlib wave (not ffmpeg). Without the param the protocol and behaviour are byte-for-byte unchanged. Backend ships dark (no new deps — numpy already pinned); frontend far-end streaming is a follow-up. Tests cover echo attenuation, double-talk preservation, cold/stale pass-through, param validation, and the framing helpers — all pure-numpy/stdlib so they skip the torch ASR stack. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(capture_ws): stubs accept the new pcm_sr kwarg _transcribe_buffer/_transcribe_buffer_full gained an optional pcm_sr kwarg for the AEC PCM path; the protocol-test stubs had fixed signatures and raised TypeError on it, so the handler sent 'error' instead of 'final'. Accept **kw in the stubs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4aa4d983aa |
feat(settings): Hugging Face mirror (HF_ENDPOINT) for restricted networks (Wave 4.3) (#391)
The model manager already lists/deletes cached models; this adds the remaining high-value slice — an in-app HF mirror setting so users behind restricted networks (e.g. the Great Firewall) can route downloads through hf-mirror.com or any HF_ENDPOINT. Persisted to the durable per-user env (survives Tauri/Finder launches); HF reads HF_ENDPOINT at import, so the override applies on restart (surfaced in the UI). - GET/PUT /api/settings/hf-mirror (loopback-gated): presets (official + hf-mirror.com), http(s) validation, empty clears to official. - Models-tab panel with quick-picks + free-text field + restart note. (Skipped 'hf cache verify' — version-fragile across huggingface_hub releases and low value vs the mirror, which the China/Russia network research flagged as the real gap.) 3 endpoint tests (default, set+trim+clear, non-http rejection). Spec §R4(c) / parity program Wave 4.3. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3161166328 |
fix: stale-chunk preload recovery (#380) + surface unsupported-GPU-arch in notifications (#284) (#385)
- #380: vite:preloadError (old hashed assets after an update) triggers a one-time reload to pick up the fresh manifest; session flag prevents loops - #284: check_device_compatibility's warning (e.g. Blackwell sm_120 on a pre-cu128 torch) now appears in the notification panel as an error with the pip fix — a log line never reached affected users while synthesis silently produced noise. Cached once per process. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f5579b40aa |
fix: issue-triage batch — timeline box flicker, truncated-model detection, stale history pruning (#381)
- #373: drop will-change:transform on the segment lane (persistent compositor layer made the semi-transparent boxes vanish during playback/drag on some Windows GPUs) + raise region alpha 0.30→0.45 - #352: validate a finished snapshot actually contains weights (>5 MB file) so interrupted downloads fail at install time with a re-download hint; loader translates the opaque transformers error into the same guidance - GET /history prunes rows whose audio file is gone instead of serving dead 404 players forever Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
66f2ea7e50 |
feat(profiles): unified profile model — kind discriminator + stored design params (spec P3) (#376)
Migration 0005_unified_profiles (0004 taken by mcp bindings):
- voice_profiles.kind TEXT DEFAULT 'clone' ('clone' | 'design'), backfilled
- voice_profiles.vd_states TEXT NULL — JSON of design category picks
- mirrored in _BASE_SCHEMA; idempotent _has_column guards; downgrade drops
POST /profiles:
- ref_audio now optional; kind + vd_states form fields with validation
(clone requires audio; design requires vd_states JSON object + instruct)
- design profiles render a deterministic identity sample (seed 42) through
the shared archetype renderer — one TTS code path
POST /generate:
- profile resolution branches on profile.kind (authoritative) instead of
the brittle is_locked/instruct inference; legacy pre-0005 rows keep the
old inference as fallback; history.mode records profile.kind
Frontend:
- 'Save design as profile' in the Design tab (vd_states + buildDesignInstruct)
- selecting a design profile restores its sliders (vd_states) for re-editing
Also unforks the alembic chain (0004_mcp + my 0004 both revised 0003 →
multiple heads broke alembic upgrade head and the 0003 migration tests).
Tests: tests/test_profile_unification.py — validation, design-create with
mocked renderer, migration up/backfill/downgrade. 18/18 profile tests,
312/312 frontend, related backend suite green.
Note: docs/specs/voice-studio-unification.md (on feat/studio-ux-overhaul)
still says 0004 — renumber to 0005 when branches meet.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d6562d6f30 |
feat(dub): Smart Fit phase B — per-segment video retime export, drift absorption, fitted subtitles (#350)
* feat(dub): Smart Fit phase B — per-segment video retime export, drift absorption, fitted subtitles Executes the video side of the Smart Fit plans persisted by Phase A (job["fit_plans"], #347) at export and preview time. Backend: - services/video_retime.py (new, clean-room): two-tier retime executor. ≤48 chunks → the proven single-pass split/trim/setpts/concat filter_complex; above → batches of 40 chunks rendered to intermediate slices (identical libx264 medium/crf20 params, keyframe at t=0) joined losslessly with the concat demuxer. Slices are CFR-resampled (fps=) because setpts leaves VFR-ish timestamps that broke tpad and drifted a frame per retimed chunk on ffmpeg 7.x. Temp slices cleaned on success AND failure/abort. - Drift absorption: fitted track longer than retimed video → freeze-frame tail (tpad=stop_mode=clone) predicted into the last slice / single-pass graph, with residual mux-side tpad; video longer → silence-pad the dub audio chain (apad=whole_dur). ±50 ms tolerance. - VFR guard: probe r_frame_rate vs avg_frame_rate; normalise with fps= before trim/setpts; probe failure degrades gracefully. - Plan resolution: _video_retime_plan_for spans legacy video_stretch_plans (byte-identical resolution + command construction) and fit_plans, gated on the track's own timing_strategy so stale plans never retime a track re-generated under another strategy. - Fitted subtitles: /dub/srt + /dub/vtt accept ?lang= and serve cue times from fitted_segments for Smart Fit tracks; _write_burn_srt does the same for burn-in. burn_subs+retime is now allowed for smart_fit (burn runs AFTER the retime graph); still rejected for legacy stretch_video. - /dub/preview-video resolves the same plan so in-app preview matches export. - Fallback ladder: batch encode failure/timeouts → un-retimed export with a structured core.failure warning (X-Dub-Export-Warning header + job["last_export_warning"]); concat join rejection → one single-pass retry while ≤96 chunks; abort → 409 + proc kill via run_ffmpeg job_id registration (/dub/abort reaches export encodes now) + temp cleanup. Frontend: - Export drawer passes ?lang= on subtitle exports and shows an i18n'd re-encode cost note (~0.5–2× video length on CPU) when a retiming strategy is active — translated in all 21 locales. Tests: tests/test_smart_fit_export.py — plan resolution, batch math, graph parity + new stages, fitted-cue SRT/VTT/burn selection, burn policy, VFR detection; ffmpeg-gated integration renders both executor tiers (batch size forced to 2) and the real /dub/download endpoint, ffprobing durations within ±50 ms across both pad branches. All existing dub export/subtitle/preview/timing tests pass unchanged. Refs docs/competitive-analysis.md Action 1 (dub-length fitting v2); completes Smart Fit (Phase A = #347). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): sanitize Smart Fit retime work paths at every sink (CodeQL py/path-injection) The job_id-derived retime work path (retimed_*.mp4 / preview_retimed_*.tmp.mp4) flowed unguarded from dub_export into prepare_smart_fit_video / render_retimed_video and their derived slice/concat paths and ffmpeg argv. Apply the repo's proven inline realpath+startswith containment pattern (helpers/commonpath are not recognized — see #309/#328/#329/#348): - dub_export.py: validate work_path against DUB_DIR at both construction sites (export + preview) and pass the validated realpath onward. - video_retime.py: make both entry points self-defending — realpath + DUB_DIR containment on out_path/work_path before any derivation, raising RetimeError(stage="plan") on escape; slices_dir/slice_path/list_path and RetimeDecision.file_path now all derive from the sanitized value. DUB_DIR is read via module attribute so test fixtures reloading core.config work. - ffmpeg_utils.py: document that all caller-assembled argv paths are realpath-validated upstream. - tests: sandbox DUB_DIR in the executor integration tests (tmp_path) so the new guard sees the test workspace. No behavior change for valid (server-built) paths — the guard only fires on traversal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(smart-fit): patch DUB_DIR on video_retime's own config ref — survives suite-wide reload The retime guard reads video_retime._config.DUB_DIR at call time; the sandbox fixture patched a fresh 'import core.config' instead. Another test reloads core.config in the full suite, so the two module refs diverged — the patch missed and the guard rejected the test's tmp paths (green in isolation, red in CI's full run). Patch the exact ref the guard dereferences. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): resolve DUB_DIR live at call time in retime guards — survive full-suite reload The path-containment guards bound DUB_DIR via a module-level 'from core import config as _config'. Other tests importlib.reload() core.config (sandboxing OMNIVOICE_DATA_DIR), after which the guard checked containment against a stale DUB_DIR while dub_export built the path under the reloaded one — every retime path then 'escaped the dub workspace' (green file-alone, red full-suite: the 5 integration failures CI hit). Re-import DUB_DIR locally in each guard so it always reads the current sys.modules value; simplify the sandbox fixture to patch the canonical module. Verified: full backend suite green on the Smart Fit tests (the 2 remaining settings_store failures are pre-existing on main, unrelated — local data-dir artifact). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): clear CodeQL alerts on Smart Fit export — job_id allowlist, proc-registry decouple - py/path-injection (8, video_retime.py): validate job_id with a strict inline regex allowlist (re.fullmatch [A-Za-z0-9_-]{1,64}) at the entry of dub_download and dub_preview_video, before it reaches any filesystem path or ffmpeg argv. The existing realpath containment guards stay as defense-in-depth; the regex barrier is the sanitizer CodeQL recognizes through the service-module call chain. - py/log-injection (4): newline-strip job_id inline at the logger calls in ffmpeg_utils.run_ffmpeg and the two retime-fallback logger.error sites in dub_export. - py/empty-except (3): best-effort cleanup os.remove handlers now log the OSError at debug instead of bare pass (video_retime + both dub_export mux finally blocks; _discard_tmp too for consistency). - py/cyclic-import (2): break the dub_pipeline <-> ffmpeg_utils cycle for real — the subprocess registry (register_proc/unregister_proc/ kill_job_procs/has_active_procs + state) moves to a new stdlib-only leaf module services/proc_registry.py. ffmpeg_utils now imports it at module top (no lazy import); dub_pipeline re-exports every name so dub_core aliases and tests keep working unchanged. No behavior change for valid inputs; invalid job ids now get a clean 400 instead of a 404/containment error. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): address #350 review — cancelled-vs-failed retime, logged best-effort excepts, redacted probe logs, narrowed test assert - rc<0 (killed by user cancel) now raises RetimeError(stage='aborted') instead of reporting an ordinary render failure (CodeRabbit) - best-effort cleanup/QC-event excepts log at debug instead of bare pass (CodeQL empty-except x3) - probe failure logs use basename, not full user paths (CodeRabbit/CodeQL) - test_render_cleans_slices_on_failure asserts RetimeError, not Exception Rebuttals (no change needed, see PR comment): fitted-cue subtitles track the fitted AUDIO timeline which is correct even on retime fallback; the planner only emits stretch ratios >1 so the early-exit guard is a true no-op check; '\'' is ffmpeg's own utility quoting for concat lists; has_active_procs is an intentional re-export (noqa'd). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
825f4f7ac6 |
feat(dub): regenerate subtitle timeline on the fitted timeline (Wave 3.1) (#371)
Smart Fit Phase A (planner) + the export-side video retime + audio stretch already shipped (#347 + dub_export stretch filter). The last piece of Spec 1 was the subtitle timeline: under stretch_video the dubbed audio plays at FITTED positions, but the standalone SRT/VTT export still used the original segment times — so external subtitles drifted against the dubbed video. - services/fitted_subtitles.py (pure, tested): map_time_to_fitted() + fitted_cues() remap original cue times onto the same per-chunk {orig→new, stretch_ratio} plan the video stretch uses, with a monotonicity guard. - dub_export SRT + VTT endpoints: when a job used stretch_video, cues are regenerated from the plan (subtitles track actual dub placement); no plan → original times, unchanged. New optional ?lang= selects the track. 7 pure tests (chunk-bound mapping, linear interpolation, unit-rate tail, fitted cues, monotonicity, empty-plan identity). Spec 1 (remaining) / parity program Wave 3.1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a12492af07 |
feat(dub): second-pass ASR QC — flag lines whose dub drifts from target (Wave 3.3) (#370)
After a dub is generated, re-recognize the synthetic audio and compare what
the ASR heard against what we asked the TTS to say. Lines that drift are
flagged for the user to re-listen / re-dub — turning subtitle timing and
pronunciation from trusted math into measured truth, and doubling as an
automatic dub-quality check.
Design delta from pyvideotrans (which lets recognized text REPLACE the
subtitles wholesale): we keep the generated text authoritative and use the
second pass only for MEASUREMENT — a per-line drift score + measured
start/end that feed the incremental re-dub loop, never silently overwriting
the translation.
- services/dub_qc.py (pure, tested): word_error_rate (normalized token edit
distance, case/punct-insensitive, script-agnostic) + score_dub (matches
recognized segments to dub segments by time overlap, concatenates the
hypothesis, scores drift, derives measured bounds).
- POST /dub/qc/{job_id}: runs the active ASR backend on the dubbed track in
the GPU pool, annotates each segment with qc_drift/qc_flagged/
qc_recognized/qc_measured_start-end (non-destructive — content untouched),
persists, emits a qc_done job event. Opt-in, never fatal.
- Frontend: dubQc() API fn + a red 'Verify' badge on flagged segment rows
(en.json keys; other locales fall back).
12 pure scoring tests (identical/substitution/empty/no-overlap/multi-segment
matching/measured-timing); endpoint validated in CI.
Spec 5 / parity program Wave 3.3.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
8cce99298a |
feat(dub): per-segment clone references (Wave 3.2) (#369)
Cut each long-enough dub segment's clone reference from the isolated vocals
at that segment's own timestamps, so the dub of each line carries the
prosody/emotion of its source line — finer than one reference per speaker.
Reimplemented from the clean-room spec (pyvideotrans per-line ref idea); our
design delta is a quality floor with fallback.
- services/speaker_clone.py: extract_segment_refs() keyed by segment id;
reference transcript is the SOURCE text (text_original), since the vocals
slice is source-language audio. Floor at MIN_SEGMENT_REF_DURATION_S=3.0
(not the per-speaker 5.0, which most dialogue lines fall under) — shorter
lines are omitted and fall back to the per-speaker clone, so it's a strict
improvement, never a regression.
- dub_core: run extraction at transcribe (per_segment_refs query param,
default on), store job['segment_clones'], default each unassigned
segment's profile_id to 'auto-seg:{id}' when it has its own ref, else the
existing 'auto:{speaker}'. Forcing per-speaker (per_segment_refs=false)
is supported for long-form consistency.
- dub_generate _gen: resolve 'auto-seg:' from segment_clones, ahead of the
per-speaker 'auto:' path. profile_id is already a fingerprint field, so
flipping the mode re-dubs automatically (no _GEN_INPUT_FIELDS change).
7 pure tests over a synthetic vocals wav (own-ref for long lines,
short-line omission/fallback, source-text transcript, bounds clamping,
floor boundary). Pipeline wiring validated in CI.
Spec 4 / parity program Wave 3.2.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|