Files
VoiceStudio/tests/fixtures/fit_planner
a4d9d9f128 feat(dub): underrun fill — short dubbed lines are slowed toward their slot instead of leaving dead air (#1137)
* feat(dub): underrun fill — short dubbed lines are slowed toward their slot instead of leaving dead air

The dub pipeline has always handled audio that is too LONG for its slot
(atempo compression, Smart Fit's audio/video split, trims). Audio that is too
SHORT was start-aligned and abandoned — and that is the common case, not the
corner: translations routinely speak faster than the source delivery.
Measured on a real 4-segment dub, 8.8 of 18.7 seconds of original speech time
had no dubbed voice. What fills those holes is the separated bed's
under-speech residue (37% of the original energy, measured), so the user
hears them as BOTH "little silences" AND "the music is numbed" — and sees
them as lip-sync failure, since the mouth keeps moving after the dub stopped.

The fill: when a line's natural duration covers less than UNDERRUN_TOLERANCE
(95%) of its slot, slow it toward the slot with the same pitch-preserving
atempo pipe the compression path uses, bounded at min_audio_rate (default
0.85x — comfortably natural; atempo handles <1 natively). Wired into both
fitting strategies:

- fit_planner._fit_one: need < 1 now resolves to audio_rate=max(need, floor),
  status "audio_slowed" — planner stays a pure function; golden fixtures
  regenerated per their own instructions (10 substantive lines: five
  underrun segments across four scenarios flip to audio_slowed@0.85).
- dub_generate smart_fit branch: applies the rate in both directions (the
  target formula was already direction-agnostic).
- dub_generate strict_slot branch: mirror of its compression arm.
- stretch_video and concise strategies deliberately untouched (natural-rate
  by design / never-intervene by design).

OMNIVOICE_UNDERRUN_MIN_RATE overrides the floor (1.0 disables; clamped to
atempo's sane range). The per-segment fit badge shows "slowed N.NNx" with a
tooltip, translated in all 21 locales.

Tests: planner contracts (fill bounded by floor, tolerance zone untouched,
disable switch, empty-audio guard), the flipped unit/golden/integration
expectations updated with the rationale, and the existing smart_fit
integration test now exercises the fill through the real mix loop (its seg0
comes out audio_slowed@0.85 end to end). Full suite: 2987 backend + 1236
frontend.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(dub): strict-slot slow-downs report themselves honestly (review); ru pitch wording

Review round on #1137:

- Greptile P1 "slowdown reports fits" — REAL: the strict_slot underrun fill
  fell through to the unconditional {"status": "fits"} entry, so a slowed
  segment's badge hid the applied rate (and compression_applied mislabeled
  it). The branch now emits {"status": "audio_slowed", "audio_rate": …} like
  the smart_fit path — same honesty contract everywhere.
- Greptile P1 "padded audio hides underruns" — REFUTED with evidence: nothing
  pads strict-slot audio before the check (_load_entry_wav returns the
  natural-length WAV; only error/silence slots are slot-sized, and those are
  synthetic silence by design). On-disk segment WAVs measure both shorter and
  longer than their slots, which pre-padding would make impossible.
- CodeRabbit: Russian tooltip now says "высота тона сохранена" (pitch), not
  "высота сохранена" (height).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 19:09:02 +05:30
..