Commit Graph
1 Commits
Author SHA1 Message Date
Palash DebnathandClaude Fable 5 8cce99298a feat(dub): per-segment clone references (Wave 3.2) (#369)
Cut each long-enough dub segment's clone reference from the isolated vocals
at that segment's own timestamps, so the dub of each line carries the
prosody/emotion of its source line — finer than one reference per speaker.
Reimplemented from the clean-room spec (pyvideotrans per-line ref idea); our
design delta is a quality floor with fallback.

- services/speaker_clone.py: extract_segment_refs() keyed by segment id;
  reference transcript is the SOURCE text (text_original), since the vocals
  slice is source-language audio. Floor at MIN_SEGMENT_REF_DURATION_S=3.0
  (not the per-speaker 5.0, which most dialogue lines fall under) — shorter
  lines are omitted and fall back to the per-speaker clone, so it's a strict
  improvement, never a regression.
- dub_core: run extraction at transcribe (per_segment_refs query param,
  default on), store job['segment_clones'], default each unassigned
  segment's profile_id to 'auto-seg:{id}' when it has its own ref, else the
  existing 'auto:{speaker}'. Forcing per-speaker (per_segment_refs=false)
  is supported for long-form consistency.
- dub_generate _gen: resolve 'auto-seg:' from segment_clones, ahead of the
  per-speaker 'auto:' path. profile_id is already a fingerprint field, so
  flipping the mode re-dubs automatically (no _GEN_INPUT_FIELDS change).

7 pure tests over a synthetic vocals wav (own-ref for long lines,
short-line omission/fallback, source-text transcript, bounds clamping,
floor boundary). Pipeline wiring validated in CI.

Spec 4 / parity program Wave 3.2.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 12:50:35 +05:30