d91beef0fd314250d8d9b94de86dfea019a8bd96
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a4d9d9f128 |
feat(dub): underrun fill — short dubbed lines are slowed toward their slot instead of leaving dead air (#1137)
* feat(dub): underrun fill — short dubbed lines are slowed toward their slot instead of leaving dead air The dub pipeline has always handled audio that is too LONG for its slot (atempo compression, Smart Fit's audio/video split, trims). Audio that is too SHORT was start-aligned and abandoned — and that is the common case, not the corner: translations routinely speak faster than the source delivery. Measured on a real 4-segment dub, 8.8 of 18.7 seconds of original speech time had no dubbed voice. What fills those holes is the separated bed's under-speech residue (37% of the original energy, measured), so the user hears them as BOTH "little silences" AND "the music is numbed" — and sees them as lip-sync failure, since the mouth keeps moving after the dub stopped. The fill: when a line's natural duration covers less than UNDERRUN_TOLERANCE (95%) of its slot, slow it toward the slot with the same pitch-preserving atempo pipe the compression path uses, bounded at min_audio_rate (default 0.85x — comfortably natural; atempo handles <1 natively). Wired into both fitting strategies: - fit_planner._fit_one: need < 1 now resolves to audio_rate=max(need, floor), status "audio_slowed" — planner stays a pure function; golden fixtures regenerated per their own instructions (10 substantive lines: five underrun segments across four scenarios flip to audio_slowed@0.85). - dub_generate smart_fit branch: applies the rate in both directions (the target formula was already direction-agnostic). - dub_generate strict_slot branch: mirror of its compression arm. - stretch_video and concise strategies deliberately untouched (natural-rate by design / never-intervene by design). OMNIVOICE_UNDERRUN_MIN_RATE overrides the floor (1.0 disables; clamped to atempo's sane range). The per-segment fit badge shows "slowed N.NNx" with a tooltip, translated in all 21 locales. Tests: planner contracts (fill bounded by floor, tolerance zone untouched, disable switch, empty-audio guard), the flipped unit/golden/integration expectations updated with the rationale, and the existing smart_fit integration test now exercises the fill through the real mix loop (its seg0 comes out audio_slowed@0.85 end to end). Full suite: 2987 backend + 1236 frontend. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(dub): strict-slot slow-downs report themselves honestly (review); ru pitch wording Review round on #1137: - Greptile P1 "slowdown reports fits" — REAL: the strict_slot underrun fill fell through to the unconditional {"status": "fits"} entry, so a slowed segment's badge hid the applied rate (and compression_applied mislabeled it). The branch now emits {"status": "audio_slowed", "audio_rate": …} like the smart_fit path — same honesty contract everywhere. - Greptile P1 "padded audio hides underruns" — REFUTED with evidence: nothing pads strict-slot audio before the check (_load_entry_wav returns the natural-length WAV; only error/silence slots are slot-sized, and those are synthetic silence by design). On-disk segment WAVs measure both shorter and longer than their slots, which pre-padding would make impossible. - CodeRabbit: Russian tooltip now says "высота тона сохранена" (pitch), not "высота сохранена" (height). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4b21f82619 |
feat(dub): Smart Fit timing strategy — planner, fingerprints, generate path (phase A) (#347)
* feat(dub): Smart Fit planner, fit fingerprints, shared ffmpeg stretch helpers - services/fit_planner.py: pure, I/O-free planner for dub-length fitting v2 — slack absorption (gap guard), audio-only band (<=1.2x), geometric 50/50 audio/video split capped at 1.5x / 2.0x, residual overflow accounting, and a stretch_video-compatible video_plan + fitted timeline cursor. Clean-room reimplementation from a published description. - services/incremental.py: fit_fingerprint() over the fit params with the same _canon_value canonicalisation as segment hashes (#281 class). Fit params stay OUT of segment_fingerprint — a fit change re-mixes, never re-TTSes. - services/ffmpeg_utils.py: move _atempo_chain/_pitch_preserving_stretch out of the dub_generate router (lazy torch/numpy imports) so the Phase B export pipeline can reuse them; add probe_duration() ffprobe helper. - schemas/requests.py: timing_strategy gains "smart_fit"; optional fit_options knob overrides default server-side. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(dub): smart_fit branch in the generate path TTS loop unchanged (dur_s=None, natural-rate WAVs on disk). After the loop, plan_fit() decides per segment; the mix loop applies audio_rate via the pitch-preserving atempo pipe (linear-interp fallback), trims residual overflow with the existing fades, and places audio at the planned new_start on a fitted-length canvas. Truthful fit_status entries (audio_rate / video_ratio / overflow_s) feed the row badges. Persists job["fit_plans"][lang] = {plan (exact _build_video_stretch_filter_graph shape), fitted_segments (cue times from ACTUAL stretched sample positions), total/orig duration, params, fit_fp} and mirrors fit_fp on dubbed_tracks[lang]. video_stretch_plans untouched. Strategy-transition guard: job["seg_wav_kind"] records whether on-disk seg WAVs are natural or slot-squeezed; a smart_fit partial regen over slotted (or unknown) WAVs forces one full regen instead of double-compressing. Old strategies and old persisted jobs are byte-identical (all new reads via .get()). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ui): Smart Fit option in the dub timing picker (all 21 locales) - prefsSlice: TimingStrategy union gains 'smart_fit'; optional FitOptions overrides (null by default — backend defaults apply identically on every platform); persisted alongside timingStrategy. - DubTab: Segmented gains Smart Fit with i18n label + tooltip. - useDubWorkflow: sends fit_options only when set and strategy is smart_fit. Default strategy stays 'concise' — no default behaviour change on any platform. - locales: dub.timing_smart_fit{,_title} translated in all 21 languages. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(dub): fit planner unit + golden suites, smart_fit generate-path integration - test_fit_planner.py: threshold boundaries (0.9/1.0/1.2/1.21/4.0), cap saturation -> overflow, slack absorption incl. gap guard, last-segment tail, cursor monotonicity, allow_video_retime=False, video_plan fed straight into _build_video_stretch_filter_graph, fit_fingerprint canonicalisation (int vs float, omitted vs default — the #281 class) and a pinned stable digest. - tests/fixtures/fit_planner/*.json: 4 golden FitPlans; algorithm drift is a deliberate fixture diff, never a silent change. - test_smart_fit_generate.py: hermetic end-to-end runs (mock TTS, no ffmpeg) covering audio-only stretch, hybrid timeline growth + persisted plan shape, fit_options override, strict_slot->smart_fit forced regen then zero-TTS fit-only re-mix, and concise back-compat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(competitive): dub-length fitting row reflects Smart Fit Phase A Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(incremental): mark fingerprint hashes usedforsecurity=False — dedup keys, not security (Bandit) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |