d91beef0fd314250d8d9b94de86dfea019a8bd96
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
283ef36b13 |
feat(dub): Voice match toggle — per-line prosody vs one consistent reference per speaker (#1147)
* feat(dub): Voice match toggle — per-line prosody vs one consistent reference per speaker Owner report: "still 4 segments different in voice as they are 4 times done from each segment?" — Wave 3.2 clones each dub line from a reference cut from its OWN source audio (great prosody match), but the voice IDENTITY drifts line to line, and heuristic-diarized jobs have no pooled speaker clones to anchor it. The precedence was hardcoded; now it's a per-dub-job setting. DubRequest.voice_match: - "per_line" (DEFAULT, unchanged): segment clip preferred, speaker clone fallback — byte-identical to the previous behaviour. - "consistent": ONE reference per speaker for the whole dub. `auto:` bindings use the pooled speaker clone; when none exists (heuristic diarization skips extraction entirely — the key case) a deterministic pick among that speaker's segment clips (longest ≥3 s, tie-break lowest segment id) is reused for every line. Server-default self `auto-seg:` bindings join the pick (they're what prepare stamps on heuristic jobs — the Voice dropdown can't even render them, so no user choice is overridden); explicit CROSS auto-seg bindings still honour their clip. The shared pick is multi-use, so it stays warm in the clone-prompt cache (#1132 cache_ref semantics) at both the main generate and the OOM-retry call site. voice_match is part of the segment fingerprint when non-default (mixed in like track_lang, so all stored hashes keep their values): flipping the toggle marks segments stale instead of letting "Regen changed" splice mixed-identity voices (#281 class). The client sends the mode on both /tools/incremental recompute paths. UI: a compact Voice-match Segmented control next to the Timing picker in the dub panel, persisted in the prefs slice; labels + tooltips in all 21 locales. Tests: resolution through the real dub_generate path for both modes (incl. the 4-segment heuristic job unifying on one ref — fail-before/pass-after), pick determinism + tie-breaks, schema validation, fingerprint semantics, and frontend store→request wiring. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(changelog): Voice match toggle entry under Unreleased (#1147) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d90cfde1bb |
feat(dub): per-language translations + per-track caches — switching languages stops destroying work (P1) (#958)
* feat(dub): per-language translations + per-track caches — switching languages stops destroying work (P1) Multi-language dubbing translated per language (#957) but stored everything in single-slot state, so tracks silently destroyed each other's work: P1.2 — per-language translation storage (additive): - Frontend keeps every translation in s.translations[langCode] alongside the legacy s.text slot (still = the shown language). Translate All writes both; the new store action switchDubLangCode swaps text through the map on a user-driven language switch (non-destructive; restore paths keep the plain setter); manual edits / restore-original update the current language's entry; merge joins per-language texts, split drops them. Rides project save/load inside dubSegments — legacy projects behave exactly as before. - Backend mirrors it as job["segments_i18n"] = {lang: {segKey: text}} (segKey = stable id, index for id-less legacy rows), written by _sync_job_segments; job["segments"] stays byte-identical for every existing consumer. /dub/srt|vtt?lang= and subtitle burn-in now emit THAT language's text when present — ExportModal's "all dubs" batch stops producing N identical files. Legacy jobs without the field fall back to today's output. P1.3 — per-track WAV cache + fingerprints: - Per-segment WAVs are language-keyed (seg_{lang}_{id}.wav). The partial-regen read path falls back to legacy seg_{id}.wav ONLY while the job has no other-language track — single-language jobs keep their whole on-disk cache; multi-track jobs stop splicing the last-generated language into the current track. Read-only endpoints (segment preview, clips zip) gained ?lang= with the permissive legacy fallback they always had. - Fingerprints include the track language (segment_fingerprint(track_lang=…), /tools/incremental lang=…) and live in job["seg_hashes_by_lang"]; the flat job["seg_hashes"] stays as the current track's mirror so the done event, history restore and older frontends read it unchanged. A legacy flat map is attributed to the job's last-generated language (dropped when unknown) and reads stale once — the safe direction. seg_wav_kind is per-track too. - The frontend stores fingerprints per language and judges "Regen N changed" against the ACTIVE track; project save/load and dub-history restore carry all tracks' hashes (segHashesByLang / seg_hashes_by_lang, additive). Tests: fail-before regression coverage — two-track regen never splices the other language's audio (sample-level assert on the mixed track), legacy single-track cache reuse + multi-track gate, per-lang seg_hashes with flat mirror + migration semantics, /dub/srt|vtt?lang= emitting different text per track with legacy fallbacks, per-lang burn-in, /tools/incremental lang scoping, and 14 frontend tests for translations round-trips, per-track fingerprints and legacy-project behaviour. Full backend + frontend suites, typecheck, lint and format:check green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add per-language storage + per-track caches under [Unreleased] (#958) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4b21f82619 |
feat(dub): Smart Fit timing strategy — planner, fingerprints, generate path (phase A) (#347)
* feat(dub): Smart Fit planner, fit fingerprints, shared ffmpeg stretch helpers - services/fit_planner.py: pure, I/O-free planner for dub-length fitting v2 — slack absorption (gap guard), audio-only band (<=1.2x), geometric 50/50 audio/video split capped at 1.5x / 2.0x, residual overflow accounting, and a stretch_video-compatible video_plan + fitted timeline cursor. Clean-room reimplementation from a published description. - services/incremental.py: fit_fingerprint() over the fit params with the same _canon_value canonicalisation as segment hashes (#281 class). Fit params stay OUT of segment_fingerprint — a fit change re-mixes, never re-TTSes. - services/ffmpeg_utils.py: move _atempo_chain/_pitch_preserving_stretch out of the dub_generate router (lazy torch/numpy imports) so the Phase B export pipeline can reuse them; add probe_duration() ffprobe helper. - schemas/requests.py: timing_strategy gains "smart_fit"; optional fit_options knob overrides default server-side. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(dub): smart_fit branch in the generate path TTS loop unchanged (dur_s=None, natural-rate WAVs on disk). After the loop, plan_fit() decides per segment; the mix loop applies audio_rate via the pitch-preserving atempo pipe (linear-interp fallback), trims residual overflow with the existing fades, and places audio at the planned new_start on a fitted-length canvas. Truthful fit_status entries (audio_rate / video_ratio / overflow_s) feed the row badges. Persists job["fit_plans"][lang] = {plan (exact _build_video_stretch_filter_graph shape), fitted_segments (cue times from ACTUAL stretched sample positions), total/orig duration, params, fit_fp} and mirrors fit_fp on dubbed_tracks[lang]. video_stretch_plans untouched. Strategy-transition guard: job["seg_wav_kind"] records whether on-disk seg WAVs are natural or slot-squeezed; a smart_fit partial regen over slotted (or unknown) WAVs forces one full regen instead of double-compressing. Old strategies and old persisted jobs are byte-identical (all new reads via .get()). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ui): Smart Fit option in the dub timing picker (all 21 locales) - prefsSlice: TimingStrategy union gains 'smart_fit'; optional FitOptions overrides (null by default — backend defaults apply identically on every platform); persisted alongside timingStrategy. - DubTab: Segmented gains Smart Fit with i18n label + tooltip. - useDubWorkflow: sends fit_options only when set and strategy is smart_fit. Default strategy stays 'concise' — no default behaviour change on any platform. - locales: dub.timing_smart_fit{,_title} translated in all 21 languages. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(dub): fit planner unit + golden suites, smart_fit generate-path integration - test_fit_planner.py: threshold boundaries (0.9/1.0/1.2/1.21/4.0), cap saturation -> overflow, slack absorption incl. gap guard, last-segment tail, cursor monotonicity, allow_video_retime=False, video_plan fed straight into _build_video_stretch_filter_graph, fit_fingerprint canonicalisation (int vs float, omitted vs default — the #281 class) and a pinned stable digest. - tests/fixtures/fit_planner/*.json: 4 golden FitPlans; algorithm drift is a deliberate fixture diff, never a silent change. - test_smart_fit_generate.py: hermetic end-to-end runs (mock TTS, no ffmpeg) covering audio-only stretch, hybrid timeline growth + persisted plan shape, fit_options override, strict_slot->smart_fit forced regen then zero-TTS fit-only re-mix, and concise back-compat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(competitive): dub-length fitting row reflects Smart Fit Phase A Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(incremental): mark fingerprint hashes usedforsecurity=False — dedup keys, not security (Bandit) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
48ae4dae1d |
fix(dub): re-dub honors transcript edits — fingerprints canonicalised, preview cache-busted, atomic mux (#281) (#329)
* fix(dub): re-dub honors transcript edits — fingerprints canonicalised, preview cache-busted, mux made atomic (#281) Three symptoms, three causes: 1. Edited line, unchanged result: the dubbed preview-video URL was identical across re-dubs, so the WebView kept serving the previous dub. A generation nonce now cache-busts the preview after every completed generation. 2. Preview stuck loading forever: overlapping preview requests ran ffmpeg against the same output path and the mtime cache check saw the half-written file as valid. The mux now runs under a per-path lock, writes to a temp file, and os.replace()s into place. 3. One edit re-dubs all lines: server-side fingerprints were computed from pydantic-parsed segments (defaults filled in) but recomputed client-side from raw dicts (keys omitted), so every segment always looked stale and incremental degraded to a full re-dub. Values are now canonicalised on the backend and the frontend builds generation inputs through one shared helper (utils/segments.js) for both the generate request and the incremental plan. Fixes #281 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Uncontrolled data used in path expression' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(dub): realpath containment for job-derived preview paths (CodeQL) Request-supplied job_id/lang flowed into the preview mux output path. Both now pass a realpath containment guard against DUB_DIR (the file's existing per-segment pattern) and lang is allowlist-validated before it lands in a filename. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): inline the containment guard — CodeQL can't track it through a helper Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> |
||
|
|
1edd35cfd0 |
Per-segment audio effects DSP preset selector (closes #67, rebased from #68) (#109)
* Add per-segment audio effects DSP preset selector to dub pipeline * Add shape assertions to podcast, warm, and bright preset tests * Fix raw preset semantics, add preset validation, update docs, remove duplicate sys.path * Narrow OOM catch to model.generate only in dub_generate * Preserve original OOM exception context in dub_generate * Bind effect_preset to _gen via explicit parameter to avoid loop capture * Catch RuntimeError instead of torch.mps.MPSError for MPS OOM --------- Co-authored-by: 4shil <166588383+4shil@users.noreply.github.com> |
||
|
|
52d68d05dc | refactor: update backend architecture, expand frontend state management, and synchronize voice-pro research modules. |