Cut each long-enough dub segment's clone reference from the isolated vocals
at that segment's own timestamps, so the dub of each line carries the
prosody/emotion of its source line — finer than one reference per speaker.
Reimplemented from the clean-room spec (pyvideotrans per-line ref idea); our
design delta is a quality floor with fallback.
- services/speaker_clone.py: extract_segment_refs() keyed by segment id;
reference transcript is the SOURCE text (text_original), since the vocals
slice is source-language audio. Floor at MIN_SEGMENT_REF_DURATION_S=3.0
(not the per-speaker 5.0, which most dialogue lines fall under) — shorter
lines are omitted and fall back to the per-speaker clone, so it's a strict
improvement, never a regression.
- dub_core: run extraction at transcribe (per_segment_refs query param,
default on), store job['segment_clones'], default each unassigned
segment's profile_id to 'auto-seg:{id}' when it has its own ref, else the
existing 'auto:{speaker}'. Forcing per-speaker (per_segment_refs=false)
is supported for long-form consistency.
- dub_generate _gen: resolve 'auto-seg:' from segment_clones, ahead of the
per-speaker 'auto:' path. profile_id is already a fingerprint field, so
flipping the mode re-dubs automatically (no _GEN_INPUT_FIELDS change).
7 pure tests over a synthetic vocals wav (own-ref for long lines,
short-line omission/fallback, source-text transcript, bounds clamping,
floor boundary). Pipeline wiring validated in CI.
Spec 4 / parity program Wave 3.2.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>