d91beef0fd314250d8d9b94de86dfea019a8bd96
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
018cdcb47f |
fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store (#1149)
* fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store Review round on #1148, both findings real: - Greptile P1 "missing translations leave mixed text": the in-browser translations map can be PARTIAL (tracks generated before per-language persistence, partial regens); the non-destructive switch then left those rows in the previous language under a single-language preview. New GET /dub/segments-text/{job}?lang= exposes segments_i18n (the authoritative per-language map every generate rebuilds); the tab click hydrates only the gap rows, failure-silent, and skips stale responses if the user switched again mid-fetch. - CodeRabbit "clear stale dialect": the dropdown paths each cleared a non-matching dubDialect by hand; the guard now lives inside switchDubLangCode so every caller (dropdown, multi-language loop, preview tabs, future ones) inherits it. Matching dialects survive. Tests: endpoint (i18n map served, never-generated track -> empty map, legacy job -> empty map), hydration (stored rows swap instantly, missing row hydrates from the mock backend and is cached into translations), dialect guard (cleared on mismatch, kept on match). Suites: dub sweep 262, frontend 1253, both green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(api): register /dub/segments-text in the route-inventory snapshot The inventory guard caught the new endpoint exactly as designed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dc3527ab05 |
fix(dub): stereo, full-band music bed — separate the HQ extraction, pin the mix to stereo (#1138)
* fix(dub): stereo, full-band music bed — separate the HQ extraction, pin the mix to stereo Owner asked for a channels/Hz/samples comparison of a dub against its original to tune generation toward the source. The measurements found a class, not a knob: L/R correlation: original 0.754, dub 1.000 (mono in a stereo container) stereo width (S/M): 0.375 vs 0.003 LUFS: -17.8 vs -17.2 (already fine) Two stacked causes: 1. INGEST: Demucs separated audio.wav — the 16 kHz MONO extraction made for ASR. The music bed therefore inherited mono AND an 8 kHz bandwidth ceiling at its source (Demucs upsamples to 44.1 kHz internally, so the stems LOOKED like 44.1k stereo files while carrying neither). Ingest now extracts a second full-quality file (44.1 kHz stereo, pcm_s16le) just for separation; ASR keeps its 16 kHz mono file; Demucs cost is ~unchanged (it resampled to 44.1 kHz internally either way). Best-effort: if the HQ extraction fails, separation falls back to the ASR file — exactly the old behavior. The stem-move path follows the input's basename. 2. MIX: amix negotiates ONE channel layout across inputs, and the synthesized voice is mono — so even a true-stereo bed was collapsed at the mix. bed_mix_filter now pins BOTH legs to stereo (aformat=channel_layouts=stereo); upmixing the mono voice duplicates it dead-center, which is where dubbed dialogue belongs anyway. Verified with real ffmpeg: the new graph preserves a stereo bed's width through the mix (and the ingest test pins that demucs receives audio_hq.wav with -ac 2 -ar 44100 while ASR keeps -ac 1 -ar 16000). Both tests fail with their half of the fix reverted. Full suite: 2989 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(dub): pre-HQ stem caches are not reused (review) Greptile P1, real: the content-hash cache restores a previous job's stems for the same video and skips Demucs — so every video processed BEFORE the HQ-extraction change would keep its 16 kHz-mono-derived bed forever, and the fix would never apply to exactly the videos users re-upload to hear the difference. find_cached_job now requires the audio_hq.wav marker in the cached job dir; older candidates are skipped with a log line and separation reruns once at full quality. Regression test covers both directions of the gate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a7efaa24f7 |
fix(dub): background bed no longer plays quiet and muffled — cancel amix normalization, mix at 48 kHz (#1136)
* fix(dub): background bed no longer plays quiet and muffled — cancel amix normalization, mix at 48 kHz Reported live: "background music is so much not like the original." Two stacked fidelity bugs in every bed-mix site, measured with real ffmpeg on a real dub job: 1. LEVEL — ffmpeg's amix NORMALIZES its inputs, so the per-site weight strings meant "favor dialogue slightly" but actually played the music bed at ~57% of its original level (batch.py stacked an explicit volume=0.15 under the same normalization, leaving its bed near 8%). 2. BANDWIDTH — the voice track is synthesized at 24 kHz and amix negotiates one common rate, so the 44.1 kHz bed was silently downsampled to 24 kHz: everything above 12 kHz (cymbals, air, brightness) vanished. Six call sites carried six hand-rolled variants of the same filter string (dub_export x5, batch x1) with inconsistent input ordering — the same copy-divergence pattern that orphaned the clone-prompt cache (#1130). They now share one builder, services.ffmpeg_utils.bed_mix_filter(): both inputs resampled to 48 kHz before the mix, a compensating volume multiply that cancels amix's normalization exactly (the weights ARE the absolute gains: bed 0.9, voice 1.1), and a transparent peak limiter for the rare summed peak that full-scale mixing makes possible. Measured A/B on the reporting user's job (bed vs bed-through-mix, silent voice): 57% -> 90% of original level, 24 kHz -> 48 kHz output. The remaining -0.9 dB is deliberate dialogue headroom, one constant to change if policy shifts. Tests: the export command must carry the resample + compensation + limiter (fails on the old strings), builder label-uniqueness for multi-track graphs, and the existing export suites unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(dub): amix renormalizes when a stream ends — disable normalization instead of compensating for it Greptile P1 on this PR, confirmed real by measurement: amix's normalization is DYNAMIC — it rescales the remaining inputs whenever one ends. The previous commit cancelled it with a constant post-mix multiply, which is exact only while both streams are active; once the (even marginally shorter) voice track ends, the bed's internal scale jumps to 1.0 and the fixed multiply BOOSTS the tail music into the limiter. Measured on the real job with a deliberately short voice: bed at 90% while the voice runs, 189% after it ends. The original A/B used equal-length streams, which is why this never showed. Fix: amix normalize=0 (a plain sum) with per-input volume gains — levels are exact for the whole timeline regardless of stream lifetimes. Same measurement now: 90% / 90%. normalize= arrived in ffmpeg 5.x, and system-ffmpeg users can be older, where an unknown option rejects the whole graph (= no export at all). The builder probes `ffmpeg -h filter=amix` once per process and falls back to the compensated form on legacy builds — its tail quirk is the lesser evil next to a failed export, and every bundled/imageio tier ships 7.x. Tests: both paths pinned (normalize=0 + per-input gains on modern; the compensation multiply on legacy), probe monkeypatched per test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(dub): anchor the amix monkeypatches to the call chain — module aliases miss under random order Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
87eb5ad078 |
feat(dub): audio-only dubbing mode (#119) (#150)
* feat(dub): audio-only dubbing mode (#119) Add an audio→audio dubbing path: upload an audio file, get dubbed audio out, with no video processing. The transcribe → translate → TTS core is unchanged; only the video-coupled stages are skipped. Backend: - dub_core /dub/upload: new `input_type` form field ("video"|"audio"). Audio mode validates the upload is a known audio container (else 400) and threads input_type into the ingest source dict. - dub_pipeline ingest: for audio input, skip scene detection + thumbnail ffmpeg passes (still emits scene_done count=0 so the prep SSE contract the frontend waits on is unchanged); stores input_type on the job. - dub_export /dub/download: for audio jobs, branch to an audio-only export (_build_audio_export_cmd) — no video input/map/codec/subtitle pass. Outputs dubbed_audio_{lang}_{stamp}.{wav|m4a|mp3|flac} via `out_format` (default m4a), optionally mixed with the separated background. Unknown formats fall back to AAC. Frontend: - dubSlice: dubInputType state + setter (default 'video'). - DubTab: auto-select audio-only mode when an audio file is dropped/picked. - dub.ts/useDubWorkflow: pass input_type on upload. Tests (11): _build_audio_export_cmd format/mix matrix; end-to-end audio-only export produces an audio file (no video mux); unknown-format fallback; upload rejects a video extension in audio mode. Closes #119. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * harden(#119): allowlist-sanitize lang_code in audio export path The track id is already constrained to an existing track key, but allowlist-sanitize it before it reaches the output path (same pattern as the existing safe_name) so a path component can never carry separators — clears the CodeQL path-injection flag on the new audio-export branch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * polish(#119): address Greptile P2s on audio-only dubbing - dub_pipeline: emit scene_start before scene_done(count=0) for audio so the prep SSE stage sequence is symmetric with the video path. - useDubWorkflow: 'Preparing audio…' pill for audio jobs (was always 'Preparing video…'). - DubTab: widen the drop-accept regex + file-input accept to the full supported audio set (aac/opus/wma) so it matches the input-type detection and the backend allowlist. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#119): drop unused dubInputType read in DubTab (CodeQL) Only setDubInputType is used; the value read was dead. Clears the CodeQL unused-variable alert. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |