Commit Graph
4 Commits
Author SHA1 Message Date
018cdcb47f fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store (#1149)
* fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store

Review round on #1148, both findings real:

- Greptile P1 "missing translations leave mixed text": the in-browser
  translations map can be PARTIAL (tracks generated before per-language
  persistence, partial regens); the non-destructive switch then left those
  rows in the previous language under a single-language preview. New
  GET /dub/segments-text/{job}?lang= exposes segments_i18n (the
  authoritative per-language map every generate rebuilds); the tab click
  hydrates only the gap rows, failure-silent, and skips stale responses if
  the user switched again mid-fetch.
- CodeRabbit "clear stale dialect": the dropdown paths each cleared a
  non-matching dubDialect by hand; the guard now lives inside
  switchDubLangCode so every caller (dropdown, multi-language loop, preview
  tabs, future ones) inherits it. Matching dialects survive.

Tests: endpoint (i18n map served, never-generated track -> empty map, legacy
job -> empty map), hydration (stored rows swap instantly, missing row
hydrates from the mock backend and is cached into translations), dialect
guard (cleared on mismatch, kept on match). Suites: dub sweep 262, frontend
1253, both green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(api): register /dub/segments-text in the route-inventory snapshot

The inventory guard caught the new endpoint exactly as designed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 02:11:43 +05:30
dc3527ab05 fix(dub): stereo, full-band music bed — separate the HQ extraction, pin the mix to stereo (#1138)
* fix(dub): stereo, full-band music bed — separate the HQ extraction, pin the mix to stereo

Owner asked for a channels/Hz/samples comparison of a dub against its
original to tune generation toward the source. The measurements found a
class, not a knob:

  L/R correlation: original 0.754, dub 1.000 (mono in a stereo container)
  stereo width (S/M): 0.375 vs 0.003
  LUFS: -17.8 vs -17.2 (already fine)

Two stacked causes:

1. INGEST: Demucs separated audio.wav — the 16 kHz MONO extraction made for
   ASR. The music bed therefore inherited mono AND an 8 kHz bandwidth
   ceiling at its source (Demucs upsamples to 44.1 kHz internally, so the
   stems LOOKED like 44.1k stereo files while carrying neither). Ingest now
   extracts a second full-quality file (44.1 kHz stereo, pcm_s16le) just for
   separation; ASR keeps its 16 kHz mono file; Demucs cost is ~unchanged
   (it resampled to 44.1 kHz internally either way). Best-effort: if the HQ
   extraction fails, separation falls back to the ASR file — exactly the old
   behavior. The stem-move path follows the input's basename.

2. MIX: amix negotiates ONE channel layout across inputs, and the
   synthesized voice is mono — so even a true-stereo bed was collapsed at
   the mix. bed_mix_filter now pins BOTH legs to stereo
   (aformat=channel_layouts=stereo); upmixing the mono voice duplicates it
   dead-center, which is where dubbed dialogue belongs anyway.

Verified with real ffmpeg: the new graph preserves a stereo bed's width
through the mix (and the ingest test pins that demucs receives audio_hq.wav
with -ac 2 -ar 44100 while ASR keeps -ac 1 -ar 16000). Both tests fail with
their half of the fix reverted. Full suite: 2989 passed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(dub): pre-HQ stem caches are not reused (review)

Greptile P1, real: the content-hash cache restores a previous job's stems for
the same video and skips Demucs — so every video processed BEFORE the
HQ-extraction change would keep its 16 kHz-mono-derived bed forever, and the
fix would never apply to exactly the videos users re-upload to hear the
difference. find_cached_job now requires the audio_hq.wav marker in the
cached job dir; older candidates are skipped with a log line and separation
reruns once at full quality. Regression test covers both directions of the
gate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 21:04:06 +05:30
a7efaa24f7 fix(dub): background bed no longer plays quiet and muffled — cancel amix normalization, mix at 48 kHz (#1136)
* fix(dub): background bed no longer plays quiet and muffled — cancel amix normalization, mix at 48 kHz

Reported live: "background music is so much not like the original." Two
stacked fidelity bugs in every bed-mix site, measured with real ffmpeg on a
real dub job:

1. LEVEL — ffmpeg's amix NORMALIZES its inputs, so the per-site weight
   strings meant "favor dialogue slightly" but actually played the music bed
   at ~57% of its original level (batch.py stacked an explicit volume=0.15
   under the same normalization, leaving its bed near 8%).
2. BANDWIDTH — the voice track is synthesized at 24 kHz and amix negotiates
   one common rate, so the 44.1 kHz bed was silently downsampled to 24 kHz:
   everything above 12 kHz (cymbals, air, brightness) vanished.

Six call sites carried six hand-rolled variants of the same filter string
(dub_export x5, batch x1) with inconsistent input ordering — the same
copy-divergence pattern that orphaned the clone-prompt cache (#1130). They now
share one builder, services.ffmpeg_utils.bed_mix_filter(): both inputs
resampled to 48 kHz before the mix, a compensating volume multiply that
cancels amix's normalization exactly (the weights ARE the absolute gains: bed
0.9, voice 1.1), and a transparent peak limiter for the rare summed peak that
full-scale mixing makes possible.

Measured A/B on the reporting user's job (bed vs bed-through-mix, silent
voice): 57% -> 90% of original level, 24 kHz -> 48 kHz output. The remaining
-0.9 dB is deliberate dialogue headroom, one constant to change if policy
shifts.

Tests: the export command must carry the resample + compensation + limiter
(fails on the old strings), builder label-uniqueness for multi-track graphs,
and the existing export suites unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(dub): amix renormalizes when a stream ends — disable normalization instead of compensating for it

Greptile P1 on this PR, confirmed real by measurement: amix's normalization is
DYNAMIC — it rescales the remaining inputs whenever one ends. The previous
commit cancelled it with a constant post-mix multiply, which is exact only
while both streams are active; once the (even marginally shorter) voice track
ends, the bed's internal scale jumps to 1.0 and the fixed multiply BOOSTS the
tail music into the limiter. Measured on the real job with a deliberately
short voice: bed at 90% while the voice runs, 189% after it ends. The original
A/B used equal-length streams, which is why this never showed.

Fix: amix normalize=0 (a plain sum) with per-input volume gains — levels are
exact for the whole timeline regardless of stream lifetimes. Same measurement
now: 90% / 90%.

normalize= arrived in ffmpeg 5.x, and system-ffmpeg users can be older, where
an unknown option rejects the whole graph (= no export at all). The builder
probes `ffmpeg -h filter=amix` once per process and falls back to the
compensated form on legacy builds — its tail quirk is the lesser evil next to
a failed export, and every bundled/imageio tier ships 7.x.

Tests: both paths pinned (normalize=0 + per-input gains on modern; the
compensation multiply on legacy), probe monkeypatched per test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(dub): anchor the amix monkeypatches to the call chain — module aliases miss under random order

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 16:28:14 +05:30
Palash DebnathandClaude Opus 4.8 87eb5ad078 feat(dub): audio-only dubbing mode (#119) (#150)
* feat(dub): audio-only dubbing mode (#119)

Add an audio→audio dubbing path: upload an audio file, get dubbed audio
out, with no video processing. The transcribe → translate → TTS core is
unchanged; only the video-coupled stages are skipped.

Backend:
- dub_core /dub/upload: new `input_type` form field ("video"|"audio").
  Audio mode validates the upload is a known audio container (else 400)
  and threads input_type into the ingest source dict.
- dub_pipeline ingest: for audio input, skip scene detection + thumbnail
  ffmpeg passes (still emits scene_done count=0 so the prep SSE contract
  the frontend waits on is unchanged); stores input_type on the job.
- dub_export /dub/download: for audio jobs, branch to an audio-only export
  (_build_audio_export_cmd) — no video input/map/codec/subtitle pass.
  Outputs dubbed_audio_{lang}_{stamp}.{wav|m4a|mp3|flac} via `out_format`
  (default m4a), optionally mixed with the separated background. Unknown
  formats fall back to AAC.

Frontend:
- dubSlice: dubInputType state + setter (default 'video').
- DubTab: auto-select audio-only mode when an audio file is dropped/picked.
- dub.ts/useDubWorkflow: pass input_type on upload.

Tests (11): _build_audio_export_cmd format/mix matrix; end-to-end audio-only
export produces an audio file (no video mux); unknown-format fallback;
upload rejects a video extension in audio mode.

Closes #119.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* harden(#119): allowlist-sanitize lang_code in audio export path

The track id is already constrained to an existing track key, but
allowlist-sanitize it before it reaches the output path (same pattern as
the existing safe_name) so a path component can never carry separators —
clears the CodeQL path-injection flag on the new audio-export branch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* polish(#119): address Greptile P2s on audio-only dubbing

- dub_pipeline: emit scene_start before scene_done(count=0) for audio so
  the prep SSE stage sequence is symmetric with the video path.
- useDubWorkflow: 'Preparing audio…' pill for audio jobs (was always
  'Preparing video…').
- DubTab: widen the drop-accept regex + file-input accept to the full
  supported audio set (aac/opus/wma) so it matches the input-type
  detection and the backend allowlist.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#119): drop unused dubInputType read in DubTab (CodeQL)

Only setDubInputType is used; the value read was dead. Clears the
CodeQL unused-variable alert.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 19:07:38 +05:30