d91beef0fd314250d8d9b94de86dfea019a8bd96
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d959aae41b |
fix(dub): dialogue starts stop snapping to footsteps — sustained-energy onsets, bounded snap (#963) (#967)
* fix(dub): dialogue starts stop snapping to footsteps — sustained-energy onsets, bounded snap distance (#963) Field report #963 (point 3): dubbed speakers start seconds early or late. The reporter's own theory was right on the money — 'when a noise is heard (a sigh or footsteps), it's interpreted as the start of the conversation.' The #280 onset snapper took the FIRST 20 ms frame above an adaptive RMS threshold as the speech onset, so any transient qualified; it also had no snap-distance bound (a wrong onset could move a start by the whole segment minus 0.3 s) and ran even when Demucs had failed and the 'vocals' track was really the raw mix, where every ambient sound is a candidate. Three layered guards, all pure NumPy (no new deps): - Sustained energy: an onset only counts when >=160 ms of the following 300 ms stays above the threshold. Footsteps/door thuds light up one or two frames and die; syllables keep the energy up. - Bounded snap distance: shifts beyond 1.5 s are only trusted when the skipped span is (near-)silent — that is exactly the genuine #280 whisper start-stretch on the vocals track (Demucs removed the music, leaving real silence), so long trims over silence still work in full. Long jumps over audible content (e.g. quiet speech under the relative threshold) are refused instead of playing the dub seconds late; an isolated transient in the span (<10% audible frames) doesn't block it. - Source-aware: snapping now runs only on the separated vocals track. dub_core detects the Demucs fallback (vocals_path == audio_path, see dub_pipeline) at both call sites and passes separated_vocals=False on mixed audio, disabling snapping — whisper's own timestamps beat a confidently wrong snap when music/ambience is sustained energy too. Tests (tests/test_onset_align.py, fail-before/pass-after): transient burst rejected at detect- and snap-level, transient-only window yields no onset, long jump over audible content refused, bounded shift over audible lead still allowed, >1.5 s trim over true silence still snaps (#280 regression guard), mixed-audio mode is a no-op. 28 pass in the file; full dub-adjacent suites green. Credit: theory and repro description by the #963 reporter. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add onset-snap robustness under [Unreleased] (#967) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
65fc5245dc |
feat(dub): timeline segment editor — drag, snap-to-onset, keyboard a11y (#280) (#348)
* feat(dub): full-track speech-onset detection + GET /dub/onsets/{job_id} (#280)
detect_speech_onsets() lists every speech rise across the track (frame RMS,
adaptive threshold, 150ms hysteresis) — powers the timeline editor's
snap-to-onset ticks. Route prefers the Demucs vocals stem, falls back to the
mix, and caches onsets.json per job (mtime-invalidated).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): timeline editor math core — windowing, snap, clamp, fingerprint-safe commit (#280)
Pure helpers for the segment track: binary-search windowing, snapTime with
deterministic ties, neighbour/min-duration clamps with Alt-overlap (<=200ms),
commitMoveResize with fingerprint parity (move touches only start/end; resize
sets speed exactly like the old Regions handler and DELETES the key at 1.0 so
_canon_value's missing-vs-1.0 hashing can't mark untouched segments stale),
and overlap detection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): SegmentTrack editing lane replaces the Regions plugin (#280)
Custom DOM segment boxes (6px edge handles, body-drag move, speaker colors,
stale/fresh tint, hatched overlap warning) virtualized by time over a single
{pxPerSec, scrollLeft} alignment source read off WaveSurfer's wrapper.
Snap-to-onset ticks on a viewport-sized canvas light up in snap range;
Ctrl/Cmd-wheel zooms centered on the cursor; double-click plays the slot via
playRange (timeupdate watcher pauses at slot end). Roving-tabindex listbox
keyboard model (arrows / Enter / Shift / Alt / Delete / S) with polite
aria-live announcements. WebKit fallback keeps a self-scrolling lane at a
fixed px/sec. timeline.* strings translated in all 21 locales.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): wire timeline editor — per-gesture undo, id fix, table selection sync (#280)
segmentMoveResize() pushes undo ONCE per gesture (drag commits on pointerup;
keyboard nudges coalesce per focus session) and matches by String(id) — the
old parseInt('seg-3_a') path edited the wrong segment after a split. Commits
go through commitMoveResize for fingerprint parity, and the existing
recomputeIncremental effect picks up every commit. Clicking a timeline box
scrolls + highlights its row in DubSegmentTable; 'preview dub here' parks
the player at the slot start, then synthesizes the line.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(dub): inline the onsets-cache containment guard — CodeQL can't track helpers
Same lesson as #328/#329: the realpath+startswith sanitizer must sit at
the sink, not behind a function return. Unused helper removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
3c780dced9 |
feat(dub): speech-onset alignment + regional dialect targeting (#280) (#330)
Items 1 and 2 from the improvement list: 1. Synchronization — Whisper-family ASR stretches segment starts back over leading non-speech (intro music, silence), so the dub starts at 0:00 while the speaker starts at 0:02-0:03. New onset_align service snaps each segment start forward to the first audible vocal onset (adaptive RMS threshold over the Demucs-isolated vocals when available). Forward-only and conservative: never moves a start earlier, ignores sub-100ms shifts, preserves minimum duration, leaves silent-window segments untouched. Pure NumPy — identical across platforms. 2. Accent/vocabulary by country — a Dialect picker in the Dub panel (BCP-47 codes per target language) injects a regional instruction into LLM translation prompts (OpenAI/Ollama engines and the Cinematic refine pass): Argentina yields 'Vos sos muy listo', not 'Tú eres muy listo'. Non-LLM engines show a clear hint that the dialect needs an LLM. New i18n keys translated in all 21 locales. Item 3 (segment rectangles: move/crop/stretch on the timeline) is a larger editor feature and stays open on #280. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: mergetest <test@local> |