d91beef0fd314250d8d9b94de86dfea019a8bd96
46
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4053397921 |
feat(dub): karaoke word-highlight caption burn-in (#1764)
Adds opt-in word-timed ASS karaoke captions while preserving the existing line-caption default.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com> |
||
|
|
8cc7c88692 |
fix(dub): restore audio when the media preview falls back (#1693)
Preserve audible, language-matched playback when WebView video decoding falls back, keep waveform timing aligned with the selected dub, prevent recursive recovery failures, and disable stale audio caching. Closes #1692. |
||
|
|
e1e3a477a7 | fix: normalize original-only dub exports | ||
|
|
a04c972b71 | fix: keep dub export paths trusted | ||
|
|
45ec840ead | fix(dubbing): resolve effective export track once | ||
|
|
f72439d6cf | fix(security): use trusted audio download labels | ||
|
|
109199e024 | fix(security): keep download labels out of path boundary | ||
|
|
5baf82bfb9 | fix(security): validate audio export labels | ||
|
|
e450a37d4b | fix(security): separate export labels from paths | ||
|
|
dd1aa3654d | fix(dubbing): default exports to dubbed audio | ||
|
|
41722afe3b |
refactor(launchpad): quieter, borderless design refresh (#1515)
* refactor(launchpad): quieter, borderless design refresh The launchpad carried decoration from an earlier direction: icon chips, corner-hung count badges, a permanently visible filled arrow, uppercase mono card titles, and a dotted stipple divider — plus a frame that had been invisible since the app-wide border tokens were zeroed. Rework it around what the borderless direction actually implies: - Feature tiles get a whisper-faint surface instead of a dead frame, and read as three bands (bare glyph + count / title + arrow / description). `--card-hue` is spent sparingly — the glyph at rest, the surface, count and arrow only once raised. Titles move to sans sentence case; counts are plain tabular numerals. Lift softened 4px -> 2px, coloured glow -> neutral shadow, plus an explicit focus ring and a staggered entrance. - Hero drops the boxed "646" pill and the filled A/B-Compare button for quiet type, with a hairline standing in for the separation. - Section labels trade the dotted stipple for a single fading hairline; rows are transparent until hover and reveal "Open" on hover/focus (it stays in the DOM, so AT and keyboard always reach it). - Hero, tiles, recent files, callout and project lists now share one 1180px column — previously only the top half was capped, so lists ran edge-to-edge on a wide display while the deck stayed centred. Two bugs found and fixed while doing it: - Buttons that had `border border-solid border-transparent` removed fell back to the UA default border and rendered a visible 1px outline. They now carry `border-0` explicitly. - `.lp-animate` used `animation-fill-mode: both`, so after the entrance it kept owning `transform` — and animation-origin declarations outrank normal ones, which silently killed the card hover lift. Now `backwards`, which still holds the from-state through the stagger delay. Also drops CSS the page has not rendered since #904: the cursor-spotlight layer, the breath ring, and the per-card waveform strip. Verified with headless renders at 1600/1280/940 and the empty state. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(dictation): decode Wayland portal signals and show the capture pill The GlobalShortcuts portal declares Activated/Deactivated as (o session, s shortcut_id, t timestamp, a{sv} options). We decoded the timestamp as u32, so zbus rejected every signal with Signature mismatch: got `(osta{sv})`, expected `(osua{sv})` and the press was dropped as an invalid signal. Registration succeeded and the desktop even reported the bound chord back, so the hotkey looked wired up while doing nothing at all — on every Wayland compositor, for the whole life of the feature (#1490). Decode the 64-bit timestamp, and keep the 32-bit spelling as a fallback so a non-conforming portal degrades to working rather than to silence. With presses arriving, the second half of the failure showed: nothing had shown the widget window since it became a hidden recorder host, so a capture ran with no pill on screen — and a mic or Accessibility failure rendered into a window nobody could see. Add show_dictation_pill, which bottom-centres the capsule on the monitor under the pointer and shows it without taking focus (Windows keeps SW_SHOWNOACTIVATE so paste still lands in the user's document), and call it from the widget for every state but idle. Wayland denies clients their own placement, so the compositor picks the spot there; the pill still appears. dispatch_dictation_capture now logs whether a press was emitted or queued — a press that reaches Rust and produces nothing was otherwise indistinguishable from one the compositor never delivered. Tests: portal signals decode at both timestamp widths (the 64-bit case fails before this change with the exact production error); pill placement centres, respects a second monitor's origin, and clamps rather than going off-screen; the widget shows for a state needing the user, stays hidden while idle, and never shows for a press that arrives while dictation is disabled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: sync in-progress workspace changes Uncommitted work already in the tree, checkpointed so the branch matches the local machine: - Remote GPU workers: join-from-the-app flow, one-time secrets, QR join codes, a Compute control in the status bar, and the device-list Workers panel (#1516) - Model Catalogue workspace, with Settings pointing at it - Settings sidebar search and keyboard navigation - Demo assets for dubbing, dictation and voice design, plus the scripts that render them - Backend: validation-error handling, ASR request-path degradation, and the accompanying tests - CHANGELOG entries for the above and for the Wayland dictation fix Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tests): follow Engines to the Model Catalogue, and green the sweep - test_supertonic3 asserted the license gate points at "Settings" while the engine now names Model Catalogue → Engines, which is where the accept button actually lives. The assertion follows the move; what it pins is unchanged — the hint must name a place the user can reach it. - Carries the CJK allowlist entries for the rendered dub bundle (#1517) and the regenerated route snapshot for /workers/agent (#1516), both of which this branch inherits from the workspace sync. - docs/install/linux.md: the dictation capsule is bottom-anchored everywhere except Wayland, where the protocol gives applications no say in their placement. Documented rather than left as a surprise (CodeRabbit). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: stop a flaky dependency fetch from failing green runs en-core-web-sm resolves to a direct GitHub release URL, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own three retries all land within the same few seconds and fail together, so the whole job dies on a dependency that has nothing to do with the change under test — it cost #1518 and #1517 an otherwise-green run tonight. Two changes: back off between whole `uv sync` attempts, which is what actually clears it, and pass --no-sync to the pytest steps. `uv run` re-resolves the environment before running, so every test step was a fresh chance to hit the same fetch even though the install step had already synced — that is exactly how #1518 failed, in the isolated backend/tests step, with all 5467 tests already passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: one retry seam for every uv sync, not just the job that failed last en-core-web-sm resolves to a direct GitHub *release* URL rather than a package index, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own retries all land inside the same ~10 seconds and fail together, so a job dies on a dependency unrelated to the change under test. Tonight that cost four otherwise-green runs across #1515, #1517 and #1518 — and the first fix only covered the Tests job, so the next failure simply moved to Smoke (Linux), which syncs separately. The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh backs off between whole attempts (15s, 45s, 90s) and every workflow that syncs now goes through it — ci.yml (tests + the platform matrix), release.yml, security.yml, evals.yml. It still fails loudly after four attempts, so a genuinely broken lockfile is not disguised as a flake. The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke matrix has always set, which is part of why it was the one that kept dying; it has them now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ci): pin the Intel-Mac contract by intent, not by command spelling test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted the literal line `run: uv sync --extra pockettts`, so routing every sync through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The contract it exists to protect is that the pockettts extra installs ONLY on backend_supported legs — which the regex now pins, while leaving how the sync is invoked free to change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: keep every uv run out of the resolver, and bound the retry budget CodeRabbit, #1517: - `uv run` re-resolves before running, so the smoke suite, the worker-artifact tests, the release test run and the eval run were each a fresh chance to hit the flaky direct-URL fetch outside the retry loop. All of them pass --no-sync now; the environment is already synced by the step that owns the retries. security.yml's `uv run --with pip-audit` is deliberately left alone — it layers an ephemeral package rather than running the project's own tests. - The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a 120 s timeout on the smoke matrix). Three attempts and 60 s of total backoff outlast the refusals actually observed while staying well inside the jobs' timeout-minutes. - The Intel-Mac contract test pinned the smoke command literally too, so --no-sync tripped it exactly like the sync line did. Same fix: assert the contract (smoke runs only on backend_supported legs), not its spelling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
40662aeb06 | fix: close log safety review gaps | ||
|
|
bdcd3923eb | fix: preserve fixed-shape dub security logs | ||
|
|
047e7d901e |
Merge remote-tracking branch 'origin/main' into fix/ghas-log-safety
# Conflicts: # CHANGELOG.md # backend/api/routers/dub_export.py # backend/api/routers/marketplace.py # backend/services/ffmpeg_utils.py # backend/services/sonitranslate.py |
||
|
|
9b71eb5e46 | fix: close remaining dub export path sinks | ||
|
|
360ddf2152 |
Merge remote-tracking branch 'origin/main' into fix/ghas-path-boundary
# Conflicts: # CHANGELOG.md |
||
|
|
321823bc52 | fix(security): bind relocated artifacts to their job | ||
|
|
c2b9951f28 |
Merge remote-tracking branch 'origin/main' into fix/ghas-path-boundary
# Conflicts: # CHANGELOG.md # backend/core/path_authorization.py # frontend/src-tauri/src/commands.rs # tests/test_network_share.py |
||
|
|
8d4a9d11c8 | fix(security): authorize native dub destinations | ||
|
|
854a637b83 | fix(security): make untrusted log values single-line | ||
|
|
3e6679c03e | fix(security): enforce filesystem trust boundaries | ||
|
|
73ebd6518d |
fix(errors): six failures that reached users as raw text (#1262, #1256, #1251, #1247, #1257, #1254) (#1264)
* fix(errors): four failures that reached users as raw OS text (#1262, #1256, #1251, #1252) #1262 — a voice profile named in any non-latin-1 script 500'd every download endpoint with "'latin-1' codec can't encode characters in position 22-25". `attachment; filename="` is exactly 22 characters, so those were the first four characters of the user's own name. The sanitisers in front of the header filtered with str.isalnum(), which is True for every alphabetic script — they stripped punctuation and passed exactly what breaks the header. Ten sites, one RFC 6266 builder, plus a guard so an eleventh can't be hand-written. #1256 — a synth died on FileNotFoundError: 'ffprobe' and was reported as "an error OmniVoice doesn't recognize", on a Mac where the app's own ffprobe was resolvable the whole time. Our call sites pass explicit paths; a dependency shelling out by bare name does not. The resolved directories are now published on PATH, and the failure is classified either way. #1251 — "The paging file is too small" reached the user as a bare 500. It was already counted as an OOM, but that remedy (close apps, lighter engine) is wrong on a 32 GB machine — the fix is a Windows setting, and the hint now says which. Matched on the code in both the Python and Rust spellings. #1252/#1253 — deleting a dub mid-import crashed it with `ingest: 'mgw39lx3'`: str(KeyError) is the repr of the key. The pipeline blind-subscripted a job that DELETE /dub/history/{id} had popped minutes earlier. It now stops quietly, and no exception whose str() is a bare value can present itself that way again. * fix(engines): Unload 400'd, a wrong language said nothing, DRM was retried by hand (#1247, #1257, #1254) #1247 — list_loaded() advertises in-process engines as `engine:<id>` with "unloadable": true, but unload() only ever handled tts/diarization/sidecars. The panel was rendering a button for ids the dispatcher rejected. The engines already implement unload(); only the routing was missing. The contract test written for it immediately found a second instance — `capture-asr`, listed the same way with no branch either — which is why it enumerates the listing rather than hard-coding ids. #1257 — MLXAudioBackend.supported_languages() returns ["multi"] on the stated assumption that "each engine silently ignores languages it doesn't know". It doesn't; the library raises. So the picker offers all 646 languages and the rejection arrived as a bare list of 23 codes, naming neither the engine nor the way out. Enumerating each model's real language set would be a brittle map that goes stale every engine update — name the engine and the fix instead. #1254 — reported as intermittent: the same URL failed as DRM-protected, then succeeded on retry. Real DRM doesn't lapse; the player client varies. That is the same shape as the 403 case which already escalates through _YT_PLAYER_CLIENTS, so DRM now routes into it. If every client still refuses, the failure is classified instead of arriving as a raw yt-dlp line. * fix(review): close the delete race, narrow the tool match, sanitize the fallback Greptile P1 + CodeRabbit Major — verified real, and mine: splitting merge from save left a window where a delete lands between them, so the pending save UPSERTs the row straight back and a dub the user deleted reappears. Now one atomic step under _dub_jobs_lock, with both delete endpoints purging rows and memory under that same lock. That also fixed DELETE /dub/history, which deleted every row but evicted nothing — an in-flight job survived 'clear history' outright and re-saved itself on completion. CodeRabbit Minor (#1256): the media-tool match accepted any message ending in 'ffmpeg'/'ffprobe', so a missing FILE at /tmp/ffmpeg got the 'repair your media engine' remedy. Now requires the name unquoted-and-unqualified. CodeRabbit Major (#1262): `fallback` reached the header verbatim whenever the real name folded away entirely, walking past every guard the name goes through. Folded like the name. CodeRabbit Major (#1256): the PATH log printed resolved directories, and a user-set FFMPEG_PATH sits under their home. Logs a count now. CodeRabbit Minor (#1262): the subtitle-route assertion also passed against the pre-fix header; it now asserts filename*= too. Skipped: 'Highlights bullets must end with (#N)'. CLAUDE.md scopes that to the ### subsections; none of the seven pre-existing highlights carry refs, and tests/test_changelog_style.py encodes the rule already. * fix(review): the remaining unlocked save paths, an over-broad signature, two weak tests Greptile P1 — the mid-pipeline put_job + save_job pairs were still unlocked, so a clear-history landing between them left a ghost row behind the purge. Both go through put_and_save_job now; only the final completion gate decides whether a withdrawn job's work is kept. CodeRabbit Major (#1257) — 'unsupported language' as a bare prefix also matches 'Unsupported language model configuration', handing a model/config failure engine-switch advice it has no use for. The loose wordings now require the rejected thing to be a code or to end there. CodeRabbit Major (#1257) — the OOM test asserted on SOURCE TEXT, which passes even if the call is unreachable or its result discarded; #1224 taught this same lesson on this codebase. Both it and the language rewrite now drive the real _run_backend_inference with a raising backend. CodeRabbit Minor (#1257) — 'or "engine" in message' always passed, since the production template contains the word. Asserts the resolved class name now. CodeRabbit Major (#1252) — the delete-race test deleted the job BEFORE the merge, which only re-tested the absent case and would pass with the two steps still split. It now interleaves a real second thread against a slow save. CodeRabbit Major (#1256) — a hardcoded /tmp literal trips Ruff S108; built from tmp_path instead. * fix(review): a withdrawal must survive the job's first write CodeRabbit Major — the concern is real, though its suggested fix (gate the checkpoint on the job already existing) would break creation: an ingest's FIRST persistence is what creates the entry, so that gate would never pass. The actual defect is that dict membership cannot express 'withdrawn'. An absent key means 'not written yet' for a new job and 'deleted' for an established one — two opposite instructions from one signal. So a clear-history arriving before the first checkpoint was silently undone by that checkpoint recreating the row, and the run then persisted its result into history the user had just cleared. Tombstone it explicitly: the ingest declares itself in flight, a purge marks any in-flight id withdrawn, and both write paths refuse a withdrawn id. Released in , so it's bounded by concurrent ingests and can't poison a later run that reuses the id. That also fixed clear-history properly: a job with no row yet appears in no id list, so only an in-flight sweep can catch it. CodeRabbit Minor — my race test waited on an event that could not be set while the save held the lock, so it burned its full 2s timeout every run and synchronised nothing. It now waits for the purge thread to REACH the purge. * test(dub): the race test was not testing the race Caught by verifying fail-before rather than trusting the test: splitting merge from save — the exact resurrection bug — passed all 22 tests. The assertions checked WHAT happened (the save ran, the row was deleted, the job left memory) but never WHEN. A save landing after the delete is indistinguishable from one landing before if you only assert that both occurred — and 'after' is precisely the resurrection. Now recorded and asserted as an order. With merge+save atomic the purge cannot start until the save finishes, so the sequence is always save-then-delete; split them and it fails with ['delete', 'save']. That is the second time this test needed rewriting: v1 deleted the job before the merge and only re-checked the absent case, v2 interleaved a real thread but asserted the wrong thing. Both looked like tests. Also documents why the DB write sits inside the lock (atomicity beats a rare 5 s sqlite busy-timeout stall) and that no locked region calls another, so the non-reentrant lock cannot deadlock — verified by walking every locked region. * fix(dub): gate the withdrawal at save_job, not at its callers Greptile P1 — and the same class I'd already fixed, unfixed elsewhere. The withdrawal check sat in the two ingest helpers, but eight direct save_job call sites across dub generate / translate / export / core bypass those entirely. Deleting a dub mid-RENDER therefore still resurrected it, which is at least as likely as deleting mid-import. Moved the gate into save_job itself: one choke point, every caller inherits it, and the ninth cannot forget. That needs a re-entrant lock, since the atomic helpers call save_job while already holding it — a plain Lock would deadlock the backend, so a test pins the lock type and another exercises the nested path. Verified fail-before: removing the gate fails the new test. * fix(dub): the withdrawal only covered ingests, so it covered almost nothing Caught by testing the reported scenario directly instead of trusting a green suite: CI passed, 26 tests passed, and a dub deleted during a RENDER was still resurrected. The tombstone was scoped to in-flight ingests. But a dub is imported once and rendered many times, so the realistic delete lands during a render — long after its ingest ended — and end_ingest was CLEARING the tombstone at exactly that point. The rare case was protected and the common one left open. Now scoped to deletions, not ingests. Kept in a bounded LRU rather than cleared on completion, because there is no moment at which a delete stops mattering: any operation still holding that job can persist it. Re-importing an id is the only thing that legitimately revives it. Verified fail-before: the previous scoping fails three of the new tests. * refactor: move the dub delete-resurrection fix to its own PR (#1270) The six fixes left here are independent error-message changes that needed no corrections. The dub concurrency change needed five rounds, each finding something real in work that was already reviewed, tested and CI-green — the last of them being that the fix did not fix the reported case at all. Riding a release on that record is a bad trade, so it ships separately as #1270. This branch keeps #1262, #1256, #1251, #1247, #1257 and #1254; the KeyError message half goes with the dub PR, since it is that issue's other half. |
||
|
|
63fd497caf |
feat: TTS-only first run, platform-curated ASR, guided OS permissions, parakeet-mlx
Only the TTS model (~2.4 GB) is required on first run; ASR models are per-platform curated picks (curated_on in models.yaml) installed on demand. Every transcription surface returns a typed asr_model_missing error with a one-click download CTA instead of silently pulling multi-GB Whisper weights. Settings -> Models is a grouped, platform-aware catalog. New guided permissions UX (wizard System Check + Settings -> Permissions + mic pre-flight) with native mic-state checks and OS settings deep-links. New parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated capture preference so multilingual dictation never regresses). Docs: expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ. Hardening: preflight fails open for custom model pins, ROCm curation no longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged consent key, CaptureWidget setup race fixed, offline-cache CI simulation fixes so empty-cache runners stay green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5b74627e78 |
fix(backend): stop discarding tracebacks in error-level exception logs (#1160 follow-up)
#1160 fixed one traceback-losing logger.error in dub_pipeline.save_job; this sweeps the remaining class. 19 sites across 10 files where a real, unexpected failure was summarized as "...: %s" at ERROR level — losing the stack trace that makes crash reports diagnosable — now use logger.exception (diarization/ASR crashes, dictation load/final failures, ffmpeg mixes that silently degrade output, Smart Fit retime fallbacks, RVC init/inference, models.yaml catalog load, gallery search/download 500s, dub-history JSON decode, MCP CLI fatal exit). Deliberately left alone: WARNING/INFO/DEBUG logs, expected classes with self-sufficient messages (GPU/ASR timeouts, request validation, cryptography-availability checks), sites that re-raise immediately (db migration, _ensure_mcp), subprocess returncode checks where stderr IS the diagnosis, and sites already logging exc_info/format_exc. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
155cebe911 |
fix(dub): diagnose export failures honestly; dodge the Windows 32K argv limit (#1152)
A Windows dub export died at ffmpeg *spawn* with [WinError 206] — the mux argv scales with tracks/segments and exceeded CreateProcess's 32,767- char limit — but the catch-all told the user their (8-char) filename was too long AND to check whether ffmpeg was installed. - explain_ffmpeg_failure() maps the three real failure modes to their own advice: argv-too-long (names the limit + the actual size, suggests fewer languages per export), ffmpeg-unlaunchable (the only case that suggests checking the install / FFMPEG_PATH), and ffmpeg-ran-and-failed (surfaces ffmpeg's stderr, no install advice). All three dub_export catch sites (video mux, audio export, MP3 encode) now use it. - run_ffmpeg() on Windows moves an oversized -filter_complex graph (the dominant argv consumer — the bed-mix/apad branches grow per track) into a -filter_complex_script temp file, so the spawn never hits the limit in the first place; the script file is removed after the run. Regression tests: tests/test_ffmpeg_failure_diagnosis.py. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
018cdcb47f |
fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store (#1149)
* fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store Review round on #1148, both findings real: - Greptile P1 "missing translations leave mixed text": the in-browser translations map can be PARTIAL (tracks generated before per-language persistence, partial regens); the non-destructive switch then left those rows in the previous language under a single-language preview. New GET /dub/segments-text/{job}?lang= exposes segments_i18n (the authoritative per-language map every generate rebuilds); the tab click hydrates only the gap rows, failure-silent, and skips stale responses if the user switched again mid-fetch. - CodeRabbit "clear stale dialect": the dropdown paths each cleared a non-matching dubDialect by hand; the guard now lives inside switchDubLangCode so every caller (dropdown, multi-language loop, preview tabs, future ones) inherits it. Matching dialects survive. Tests: endpoint (i18n map served, never-generated track -> empty map, legacy job -> empty map), hydration (stored rows swap instantly, missing row hydrates from the mock backend and is cached into translations), dialect guard (cleared on mismatch, kept on match). Suites: dub sweep 262, frontend 1253, both green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(api): register /dub/segments-text in the route-inventory snapshot The inventory guard caught the new endpoint exactly as designed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a7efaa24f7 |
fix(dub): background bed no longer plays quiet and muffled — cancel amix normalization, mix at 48 kHz (#1136)
* fix(dub): background bed no longer plays quiet and muffled — cancel amix normalization, mix at 48 kHz Reported live: "background music is so much not like the original." Two stacked fidelity bugs in every bed-mix site, measured with real ffmpeg on a real dub job: 1. LEVEL — ffmpeg's amix NORMALIZES its inputs, so the per-site weight strings meant "favor dialogue slightly" but actually played the music bed at ~57% of its original level (batch.py stacked an explicit volume=0.15 under the same normalization, leaving its bed near 8%). 2. BANDWIDTH — the voice track is synthesized at 24 kHz and amix negotiates one common rate, so the 44.1 kHz bed was silently downsampled to 24 kHz: everything above 12 kHz (cymbals, air, brightness) vanished. Six call sites carried six hand-rolled variants of the same filter string (dub_export x5, batch x1) with inconsistent input ordering — the same copy-divergence pattern that orphaned the clone-prompt cache (#1130). They now share one builder, services.ffmpeg_utils.bed_mix_filter(): both inputs resampled to 48 kHz before the mix, a compensating volume multiply that cancels amix's normalization exactly (the weights ARE the absolute gains: bed 0.9, voice 1.1), and a transparent peak limiter for the rare summed peak that full-scale mixing makes possible. Measured A/B on the reporting user's job (bed vs bed-through-mix, silent voice): 57% -> 90% of original level, 24 kHz -> 48 kHz output. The remaining -0.9 dB is deliberate dialogue headroom, one constant to change if policy shifts. Tests: the export command must carry the resample + compensation + limiter (fails on the old strings), builder label-uniqueness for multi-track graphs, and the existing export suites unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(dub): amix renormalizes when a stream ends — disable normalization instead of compensating for it Greptile P1 on this PR, confirmed real by measurement: amix's normalization is DYNAMIC — it rescales the remaining inputs whenever one ends. The previous commit cancelled it with a constant post-mix multiply, which is exact only while both streams are active; once the (even marginally shorter) voice track ends, the bed's internal scale jumps to 1.0 and the fixed multiply BOOSTS the tail music into the limiter. Measured on the real job with a deliberately short voice: bed at 90% while the voice runs, 189% after it ends. The original A/B used equal-length streams, which is why this never showed. Fix: amix normalize=0 (a plain sum) with per-input volume gains — levels are exact for the whole timeline regardless of stream lifetimes. Same measurement now: 90% / 90%. normalize= arrived in ffmpeg 5.x, and system-ffmpeg users can be older, where an unknown option rejects the whole graph (= no export at all). The builder probes `ffmpeg -h filter=amix` once per process and falls back to the compensated form on legacy builds — its tail quirk is the lesser evil next to a failed export, and every bundled/imageio tier ships 7.x. Tests: both paths pinned (normalize=0 + per-input gains on modern; the compensation multiply on legacy), probe monkeypatched per test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(dub): anchor the amix monkeypatches to the call chain — module aliases miss under random order Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d90cfde1bb |
feat(dub): per-language translations + per-track caches — switching languages stops destroying work (P1) (#958)
* feat(dub): per-language translations + per-track caches — switching languages stops destroying work (P1) Multi-language dubbing translated per language (#957) but stored everything in single-slot state, so tracks silently destroyed each other's work: P1.2 — per-language translation storage (additive): - Frontend keeps every translation in s.translations[langCode] alongside the legacy s.text slot (still = the shown language). Translate All writes both; the new store action switchDubLangCode swaps text through the map on a user-driven language switch (non-destructive; restore paths keep the plain setter); manual edits / restore-original update the current language's entry; merge joins per-language texts, split drops them. Rides project save/load inside dubSegments — legacy projects behave exactly as before. - Backend mirrors it as job["segments_i18n"] = {lang: {segKey: text}} (segKey = stable id, index for id-less legacy rows), written by _sync_job_segments; job["segments"] stays byte-identical for every existing consumer. /dub/srt|vtt?lang= and subtitle burn-in now emit THAT language's text when present — ExportModal's "all dubs" batch stops producing N identical files. Legacy jobs without the field fall back to today's output. P1.3 — per-track WAV cache + fingerprints: - Per-segment WAVs are language-keyed (seg_{lang}_{id}.wav). The partial-regen read path falls back to legacy seg_{id}.wav ONLY while the job has no other-language track — single-language jobs keep their whole on-disk cache; multi-track jobs stop splicing the last-generated language into the current track. Read-only endpoints (segment preview, clips zip) gained ?lang= with the permissive legacy fallback they always had. - Fingerprints include the track language (segment_fingerprint(track_lang=…), /tools/incremental lang=…) and live in job["seg_hashes_by_lang"]; the flat job["seg_hashes"] stays as the current track's mirror so the done event, history restore and older frontends read it unchanged. A legacy flat map is attributed to the job's last-generated language (dropped when unknown) and reads stale once — the safe direction. seg_wav_kind is per-track too. - The frontend stores fingerprints per language and judges "Regen N changed" against the ACTIVE track; project save/load and dub-history restore carry all tracks' hashes (segHashesByLang / seg_hashes_by_lang, additive). Tests: fail-before regression coverage — two-track regen never splices the other language's audio (sample-level assert on the mixed track), legacy single-track cache reuse + multi-track gate, per-lang seg_hashes with flat mirror + migration semantics, /dub/srt|vtt?lang= emitting different text per track with legacy fallbacks, per-lang burn-in, /tools/incremental lang scoping, and 14 frontend tests for translations round-trips, per-track fingerprints and legacy-project behaviour. Full backend + frontend suites, typecheck, lint and format:check green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add per-language storage + per-track caches under [Unreleased] (#958) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
252f0d4fac |
fix(asr): bound whole-file transcription so a stall isn't reported as "can't reach backend" (#656)
A Windows/CUDA user (Vietnam) hit "Can't reach the local backend" only when dubbing/transcribing. Their log proves the backend started fine — model loaded, preload complete, 25 models — and the log ends right after `whisperx transcribing …tmp.wav`. The backend was alive; the *transcription* stalled (large-v3 ASR contending with the resident TTS model for VRAM on an 8 GB-class GPU), which the UI surfaces as an unreachable backend. Root cause (class, not instance): the chunked dub pipeline already bounds each chunk (OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S), but the *whole-file* transcribe paths ran unbounded: - dub QC re-transcribe (dub_export) - dictation (capture) - OpenAI-compat /audio/transcriptions A slow/stuck transcribe on any of these hung the request AND held a GPU-pool worker — indistinguishable from a dead backend. Fix: add run_transcribe_guarded() in services/asr_backend.py — a shared asyncio.wait_for wrapper (ASRTimeoutError, a TimeoutError subclass) with a generous env-tunable bound (OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S, default 300 s). On timeout the request returns 504 with actionable guidance (backend is alive; free VRAM / pick a smaller ASR model / use CPU; restart to clear the stuck worker) instead of hanging forever. Wired into all three whole-file paths. Docs: new troubleshooting §14 — "Can't reach the local backend during transcription/dubbing" — explains it's ASR weight/VRAM pressure, not a network/ mirror problem, and corrects the misconception that a "Network → Restricted/Global mirror" Settings toggle exists (the Network control is LAN sharing). Serves the #602/#585/#567 "can't reach backend" cluster. Test: backend/tests/test_asr_transcribe_timeout.py — slow fn raises ASRTimeoutError with the actionable message, fast fn passes through, subclass-of-TimeoutError so the openai_compat broad catch still maps to 504. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d6562d6f30 |
feat(dub): Smart Fit phase B — per-segment video retime export, drift absorption, fitted subtitles (#350)
* feat(dub): Smart Fit phase B — per-segment video retime export, drift absorption, fitted subtitles Executes the video side of the Smart Fit plans persisted by Phase A (job["fit_plans"], #347) at export and preview time. Backend: - services/video_retime.py (new, clean-room): two-tier retime executor. ≤48 chunks → the proven single-pass split/trim/setpts/concat filter_complex; above → batches of 40 chunks rendered to intermediate slices (identical libx264 medium/crf20 params, keyframe at t=0) joined losslessly with the concat demuxer. Slices are CFR-resampled (fps=) because setpts leaves VFR-ish timestamps that broke tpad and drifted a frame per retimed chunk on ffmpeg 7.x. Temp slices cleaned on success AND failure/abort. - Drift absorption: fitted track longer than retimed video → freeze-frame tail (tpad=stop_mode=clone) predicted into the last slice / single-pass graph, with residual mux-side tpad; video longer → silence-pad the dub audio chain (apad=whole_dur). ±50 ms tolerance. - VFR guard: probe r_frame_rate vs avg_frame_rate; normalise with fps= before trim/setpts; probe failure degrades gracefully. - Plan resolution: _video_retime_plan_for spans legacy video_stretch_plans (byte-identical resolution + command construction) and fit_plans, gated on the track's own timing_strategy so stale plans never retime a track re-generated under another strategy. - Fitted subtitles: /dub/srt + /dub/vtt accept ?lang= and serve cue times from fitted_segments for Smart Fit tracks; _write_burn_srt does the same for burn-in. burn_subs+retime is now allowed for smart_fit (burn runs AFTER the retime graph); still rejected for legacy stretch_video. - /dub/preview-video resolves the same plan so in-app preview matches export. - Fallback ladder: batch encode failure/timeouts → un-retimed export with a structured core.failure warning (X-Dub-Export-Warning header + job["last_export_warning"]); concat join rejection → one single-pass retry while ≤96 chunks; abort → 409 + proc kill via run_ffmpeg job_id registration (/dub/abort reaches export encodes now) + temp cleanup. Frontend: - Export drawer passes ?lang= on subtitle exports and shows an i18n'd re-encode cost note (~0.5–2× video length on CPU) when a retiming strategy is active — translated in all 21 locales. Tests: tests/test_smart_fit_export.py — plan resolution, batch math, graph parity + new stages, fitted-cue SRT/VTT/burn selection, burn policy, VFR detection; ffmpeg-gated integration renders both executor tiers (batch size forced to 2) and the real /dub/download endpoint, ffprobing durations within ±50 ms across both pad branches. All existing dub export/subtitle/preview/timing tests pass unchanged. Refs docs/competitive-analysis.md Action 1 (dub-length fitting v2); completes Smart Fit (Phase A = #347). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): sanitize Smart Fit retime work paths at every sink (CodeQL py/path-injection) The job_id-derived retime work path (retimed_*.mp4 / preview_retimed_*.tmp.mp4) flowed unguarded from dub_export into prepare_smart_fit_video / render_retimed_video and their derived slice/concat paths and ffmpeg argv. Apply the repo's proven inline realpath+startswith containment pattern (helpers/commonpath are not recognized — see #309/#328/#329/#348): - dub_export.py: validate work_path against DUB_DIR at both construction sites (export + preview) and pass the validated realpath onward. - video_retime.py: make both entry points self-defending — realpath + DUB_DIR containment on out_path/work_path before any derivation, raising RetimeError(stage="plan") on escape; slices_dir/slice_path/list_path and RetimeDecision.file_path now all derive from the sanitized value. DUB_DIR is read via module attribute so test fixtures reloading core.config work. - ffmpeg_utils.py: document that all caller-assembled argv paths are realpath-validated upstream. - tests: sandbox DUB_DIR in the executor integration tests (tmp_path) so the new guard sees the test workspace. No behavior change for valid (server-built) paths — the guard only fires on traversal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(smart-fit): patch DUB_DIR on video_retime's own config ref — survives suite-wide reload The retime guard reads video_retime._config.DUB_DIR at call time; the sandbox fixture patched a fresh 'import core.config' instead. Another test reloads core.config in the full suite, so the two module refs diverged — the patch missed and the guard rejected the test's tmp paths (green in isolation, red in CI's full run). Patch the exact ref the guard dereferences. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): resolve DUB_DIR live at call time in retime guards — survive full-suite reload The path-containment guards bound DUB_DIR via a module-level 'from core import config as _config'. Other tests importlib.reload() core.config (sandboxing OMNIVOICE_DATA_DIR), after which the guard checked containment against a stale DUB_DIR while dub_export built the path under the reloaded one — every retime path then 'escaped the dub workspace' (green file-alone, red full-suite: the 5 integration failures CI hit). Re-import DUB_DIR locally in each guard so it always reads the current sys.modules value; simplify the sandbox fixture to patch the canonical module. Verified: full backend suite green on the Smart Fit tests (the 2 remaining settings_store failures are pre-existing on main, unrelated — local data-dir artifact). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): clear CodeQL alerts on Smart Fit export — job_id allowlist, proc-registry decouple - py/path-injection (8, video_retime.py): validate job_id with a strict inline regex allowlist (re.fullmatch [A-Za-z0-9_-]{1,64}) at the entry of dub_download and dub_preview_video, before it reaches any filesystem path or ffmpeg argv. The existing realpath containment guards stay as defense-in-depth; the regex barrier is the sanitizer CodeQL recognizes through the service-module call chain. - py/log-injection (4): newline-strip job_id inline at the logger calls in ffmpeg_utils.run_ffmpeg and the two retime-fallback logger.error sites in dub_export. - py/empty-except (3): best-effort cleanup os.remove handlers now log the OSError at debug instead of bare pass (video_retime + both dub_export mux finally blocks; _discard_tmp too for consistency). - py/cyclic-import (2): break the dub_pipeline <-> ffmpeg_utils cycle for real — the subprocess registry (register_proc/unregister_proc/ kill_job_procs/has_active_procs + state) moves to a new stdlib-only leaf module services/proc_registry.py. ffmpeg_utils now imports it at module top (no lazy import); dub_pipeline re-exports every name so dub_core aliases and tests keep working unchanged. No behavior change for valid inputs; invalid job ids now get a clean 400 instead of a 404/containment error. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): address #350 review — cancelled-vs-failed retime, logged best-effort excepts, redacted probe logs, narrowed test assert - rc<0 (killed by user cancel) now raises RetimeError(stage='aborted') instead of reporting an ordinary render failure (CodeRabbit) - best-effort cleanup/QC-event excepts log at debug instead of bare pass (CodeQL empty-except x3) - probe failure logs use basename, not full user paths (CodeRabbit/CodeQL) - test_render_cleans_slices_on_failure asserts RetimeError, not Exception Rebuttals (no change needed, see PR comment): fitted-cue subtitles track the fitted AUDIO timeline which is correct even on retime fallback; the planner only emits stretch ratios >1 so the early-exit guard is a true no-op check; '\'' is ffmpeg's own utility quoting for concat lists; has_active_procs is an intentional re-export (noqa'd). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
825f4f7ac6 |
feat(dub): regenerate subtitle timeline on the fitted timeline (Wave 3.1) (#371)
Smart Fit Phase A (planner) + the export-side video retime + audio stretch already shipped (#347 + dub_export stretch filter). The last piece of Spec 1 was the subtitle timeline: under stretch_video the dubbed audio plays at FITTED positions, but the standalone SRT/VTT export still used the original segment times — so external subtitles drifted against the dubbed video. - services/fitted_subtitles.py (pure, tested): map_time_to_fitted() + fitted_cues() remap original cue times onto the same per-chunk {orig→new, stretch_ratio} plan the video stretch uses, with a monotonicity guard. - dub_export SRT + VTT endpoints: when a job used stretch_video, cues are regenerated from the plan (subtitles track actual dub placement); no plan → original times, unchanged. New optional ?lang= selects the track. 7 pure tests (chunk-bound mapping, linear interpolation, unit-rate tail, fitted cues, monotonicity, empty-plan identity). Spec 1 (remaining) / parity program Wave 3.1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a12492af07 |
feat(dub): second-pass ASR QC — flag lines whose dub drifts from target (Wave 3.3) (#370)
After a dub is generated, re-recognize the synthetic audio and compare what
the ASR heard against what we asked the TTS to say. Lines that drift are
flagged for the user to re-listen / re-dub — turning subtitle timing and
pronunciation from trusted math into measured truth, and doubling as an
automatic dub-quality check.
Design delta from pyvideotrans (which lets recognized text REPLACE the
subtitles wholesale): we keep the generated text authoritative and use the
second pass only for MEASUREMENT — a per-line drift score + measured
start/end that feed the incremental re-dub loop, never silently overwriting
the translation.
- services/dub_qc.py (pure, tested): word_error_rate (normalized token edit
distance, case/punct-insensitive, script-agnostic) + score_dub (matches
recognized segments to dub segments by time overlap, concatenates the
hypothesis, scores drift, derives measured bounds).
- POST /dub/qc/{job_id}: runs the active ASR backend on the dubbed track in
the GPU pool, annotates each segment with qc_drift/qc_flagged/
qc_recognized/qc_measured_start-end (non-destructive — content untouched),
persists, emits a qc_done job event. Opt-in, never fatal.
- Frontend: dubQc() API fn + a red 'Verify' badge on flagged segment rows
(en.json keys; other locales fall back).
12 pure scoring tests (identical/substitution/empty/no-overlap/multi-segment
matching/measured-timing); endpoint validated in CI.
Spec 5 / parity program Wave 3.3.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
65fc5245dc |
feat(dub): timeline segment editor — drag, snap-to-onset, keyboard a11y (#280) (#348)
* feat(dub): full-track speech-onset detection + GET /dub/onsets/{job_id} (#280)
detect_speech_onsets() lists every speech rise across the track (frame RMS,
adaptive threshold, 150ms hysteresis) — powers the timeline editor's
snap-to-onset ticks. Route prefers the Demucs vocals stem, falls back to the
mix, and caches onsets.json per job (mtime-invalidated).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): timeline editor math core — windowing, snap, clamp, fingerprint-safe commit (#280)
Pure helpers for the segment track: binary-search windowing, snapTime with
deterministic ties, neighbour/min-duration clamps with Alt-overlap (<=200ms),
commitMoveResize with fingerprint parity (move touches only start/end; resize
sets speed exactly like the old Regions handler and DELETES the key at 1.0 so
_canon_value's missing-vs-1.0 hashing can't mark untouched segments stale),
and overlap detection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): SegmentTrack editing lane replaces the Regions plugin (#280)
Custom DOM segment boxes (6px edge handles, body-drag move, speaker colors,
stale/fresh tint, hatched overlap warning) virtualized by time over a single
{pxPerSec, scrollLeft} alignment source read off WaveSurfer's wrapper.
Snap-to-onset ticks on a viewport-sized canvas light up in snap range;
Ctrl/Cmd-wheel zooms centered on the cursor; double-click plays the slot via
playRange (timeupdate watcher pauses at slot end). Roving-tabindex listbox
keyboard model (arrows / Enter / Shift / Alt / Delete / S) with polite
aria-live announcements. WebKit fallback keeps a self-scrolling lane at a
fixed px/sec. timeline.* strings translated in all 21 locales.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): wire timeline editor — per-gesture undo, id fix, table selection sync (#280)
segmentMoveResize() pushes undo ONCE per gesture (drag commits on pointerup;
keyboard nudges coalesce per focus session) and matches by String(id) — the
old parseInt('seg-3_a') path edited the wrong segment after a split. Commits
go through commitMoveResize for fingerprint parity, and the existing
recomputeIncremental effect picks up every commit. Clicking a timeline box
scrolls + highlights its row in DubSegmentTable; 'preview dub here' parks
the player at the slot start, then synthesizes the line.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(dub): inline the onsets-cache containment guard — CodeQL can't track helpers
Same lesson as #328/#329: the realpath+startswith sanitizer must sit at
the sink, not behind a function return. Unused helper removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
853b9eefc7 |
fix(dub): burn translated subtitles, fix subtitle save JSON error (#309) (#328)
* fix(dub): burn translated subtitles, fix subtitle save JSON error (#309) Two symptoms, one root: the job kept the original-language ASR transcript while the editor only sent translated/edited text in the generate request. - dub_generate now persists the segments the dub was actually generated from back onto the job (metadata carried over by stable id, fallback index; text_original retained for dual-subtitle layouts) — SRT/VTT export and ffmpeg burn-in now render the dub language, not the source. - The SRT/VTT export endpoints honor the save_path query param the Tauri save dialog appends (like every other export) and return the standard JSON envelope — previously they ignored it and returned the raw body, so the frontend's JSON.parse choked on the SRT cue index ('Unexpected non-whitespace character after JSON'). - Frontend guards the save response content-type so any future raw-body response surfaces as a clear error. Fixes #309 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Uncontrolled data used in path expression' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(dub): use the file's established realpath+startswith containment idiom (CodeQL) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): write subtitle saves from the Tauri process, not the backend (#309) The backend save_path variant on /dub/srt and /dub/vtt routed a user-controlled destination through the loopback HTTP surface — six new CodeQL path-injection flows plus two log-injection flows. Subtitles are small text bodies, so the frontend now fetches them raw and writes the file via a new save_text_file Tauri command: the OS save dialog in the trusted process is the write authorization, and the backend never sees a destination path. Binary exports keep the established save_path flow. Also strips newlines from user-derived values in the two flagged log lines. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): leave _native_save byte-identical to main The newline-strip on the log line moved a path sink onto a changed line, which made CodeQL re-attribute the long-standing binary-export flow to this PR as a new alert. The subtitle endpoints no longer feed this function at all, so restore the exact original line — the baseline alert stays baseline, and hardening pre-existing flows belongs in its own PR. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> |
||
|
|
48ae4dae1d |
fix(dub): re-dub honors transcript edits — fingerprints canonicalised, preview cache-busted, atomic mux (#281) (#329)
* fix(dub): re-dub honors transcript edits — fingerprints canonicalised, preview cache-busted, mux made atomic (#281) Three symptoms, three causes: 1. Edited line, unchanged result: the dubbed preview-video URL was identical across re-dubs, so the WebView kept serving the previous dub. A generation nonce now cache-busts the preview after every completed generation. 2. Preview stuck loading forever: overlapping preview requests ran ffmpeg against the same output path and the mtime cache check saw the half-written file as valid. The mux now runs under a per-path lock, writes to a temp file, and os.replace()s into place. 3. One edit re-dubs all lines: server-side fingerprints were computed from pydantic-parsed segments (defaults filled in) but recomputed client-side from raw dicts (keys omitted), so every segment always looked stale and incremental degraded to a full re-dub. Values are now canonicalised on the backend and the frontend builds generation inputs through one shared helper (utils/segments.js) for both the generate request and the incremental plan. Fixes #281 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Uncontrolled data used in path expression' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(dub): realpath containment for job-derived preview paths (CodeQL) Request-supplied job_id/lang flowed into the preview mux output path. Both now pass a realpath containment guard against DUB_DIR (the file's existing per-segment pattern) and lang is allowlist-validated before it lands in a filename. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): inline the containment guard — CodeQL can't track it through a helper Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> |
||
|
|
47729057bd |
chore(lint): remove unused imports + variables (ruff F401/F841) (#210)
Autofixes the genuine lint behind the CodeQL py/unused-import and py/unused-local-variable note-level alerts — actually removing the dead code rather than dismissing it. 68 safe fixes via 'ruff check --select F401,F841 --fix' across 29 backend files (dead stdlib/symbol imports like io/sys/json/torch/typing.Optional and unused locals). Only ruff's safe fixes applied — the 9 'unsafe' fixes and the audio_dsp numpy availability import were left untouched. Not touched: empty-except (needs per-site judgement, not autofixable); frontend js/unused-local-variable (eslint no-unused-vars has no autofix); the loopback-low-risk path/log/stack-trace alerts (real, left visible). Verified: full tests/ suite unchanged at 601 passed (the 2 test_supertonic3 failures are pre-existing on main, local .venv state, green in CI). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7d9338bd2d |
fix(dub): key per-segment WAVs by stable id, not list index (#185) (#187)
* fix(dub): key per-segment WAVs by stable id, not list index (#185) Partial regeneration ('regenerate only changed segments') reloaded/wrote seg_{i}.wav by LIST INDEX while the regen allow-list + fingerprints were keyed by STABLE id. After a delete/merge/split (ids preserved, positions shifted), unchanged segments reused a different segment's audio → silently corrupted dub output on the default in-UI incremental path. - core.config.dub_seg_path(job_id, seg_id): per-segment path keyed by stable id, sanitized to a bare filename (defends against path traversal via crafted ids). A numeric index sanitizes to the legacy seg_{i}.wav, so old jobs resolve through the same helper. - dub_generate: write/reload per-segment WAVs by stable seg_id (deferred write, RVC write, regen reload) with a legacy seg_{i}.wav fallback; persist a job['seg_order'] manifest (index -> stable id) for index-keyed readers. - dub_export: preview + stems-zip resolve the file via seg_order (-> stable id) with legacy fallback, so they keep finding the right audio. - test: dub_seg_path id-naming, legacy-index equivalence, traversal sanitization. Back-compatible with in-flight jobs (legacy index files still resolve). Closes #185. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(dub): harden dub_seg_path with realpath containment; route all seg paths through it (CodeQL path-injection) Both job_id and seg_id are request-derived. Sanitise both and verify the resolved path stays inside DUB_DIR (realpath + startswith) — raises on escape. Route the legacy index fallbacks in dub_generate/dub_export through dub_seg_path so no raw os.path.join(DUB_DIR, job_id, ...) remains and a bare '..' component can't traverse. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(dub): assert path containment at export sinks (recognized CodeQL barrier) dub_seg_path already validates realpath containment, but CodeQL doesn't propagate the barrier across the call. Re-assert at the FileResponse / zf.write sinks (realpath + startswith on the value used) so the guard is recognized in-function — clears the py/path-injection false positives. * fix(dub): realpath+containment guard before any path sink in export/preview CodeQL flags os.path.exists/FileResponse/zf.write as path sinks and won't propagate dub_seg_path's internal guard. Resolve each candidate, realpath it, and containment-check (startswith DUB_DIR) BEFORE any filesystem access — the guard now dominates every sink in-function, clearing py/path-injection. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
87eb5ad078 |
feat(dub): audio-only dubbing mode (#119) (#150)
* feat(dub): audio-only dubbing mode (#119) Add an audio→audio dubbing path: upload an audio file, get dubbed audio out, with no video processing. The transcribe → translate → TTS core is unchanged; only the video-coupled stages are skipped. Backend: - dub_core /dub/upload: new `input_type` form field ("video"|"audio"). Audio mode validates the upload is a known audio container (else 400) and threads input_type into the ingest source dict. - dub_pipeline ingest: for audio input, skip scene detection + thumbnail ffmpeg passes (still emits scene_done count=0 so the prep SSE contract the frontend waits on is unchanged); stores input_type on the job. - dub_export /dub/download: for audio jobs, branch to an audio-only export (_build_audio_export_cmd) — no video input/map/codec/subtitle pass. Outputs dubbed_audio_{lang}_{stamp}.{wav|m4a|mp3|flac} via `out_format` (default m4a), optionally mixed with the separated background. Unknown formats fall back to AAC. Frontend: - dubSlice: dubInputType state + setter (default 'video'). - DubTab: auto-select audio-only mode when an audio file is dropped/picked. - dub.ts/useDubWorkflow: pass input_type on upload. Tests (11): _build_audio_export_cmd format/mix matrix; end-to-end audio-only export produces an audio file (no video mux); unknown-format fallback; upload rejects a video extension in audio mode. Closes #119. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * harden(#119): allowlist-sanitize lang_code in audio export path The track id is already constrained to an existing track key, but allowlist-sanitize it before it reaches the output path (same pattern as the existing safe_name) so a path component can never carry separators — clears the CodeQL path-injection flag on the new audio-export branch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * polish(#119): address Greptile P2s on audio-only dubbing - dub_pipeline: emit scene_start before scene_done(count=0) for audio so the prep SSE stage sequence is symmetric with the video path. - useDubWorkflow: 'Preparing audio…' pill for audio jobs (was always 'Preparing video…'). - DubTab: widen the drop-accept regex + file-input accept to the full supported audio set (aac/opus/wma) so it matches the input-type detection and the backend allowlist. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#119): drop unused dubInputType read in DubTab (CodeQL) Only setDubInputType is used; the value read was dead. Clears the CodeQL unused-variable alert. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8b00dc1f4f |
feat: onboarding demos, opt-in bug reporting, error-docs deeplinks + issue triage (#133)
* feat: onboarding demos, opt-in bug reporting, error-docs deeplinks + issue triage Working-tree snapshot bundling several in-flight workstreams (v0.3.0): - Onboarding/demo system: DemoPresetGrid, DictationDemo, DubbingDemo components + tests, render scripts (render_demos_omnivoice.py, build_demos.sh, build_dub_demo.sh), personalities preview URLs, alembic 0002 voice-profile demo fields. - Opt-in bug reporting: ReportBugButton (prefilled GitHub-issue URL path). - Error transparency UX: errorDocsMap deeplinks + BootstrapSplash/error wiring. - Dub workspace: DubSegmentRow/Table, WaveformTimeline, dubSlice tweaks. - Issue triage: .planning/issue-clusters/ (plan-01..05 root-cause masters, GH #128-#132). - CLAUDE.md: hard rule — everything ships on v0.3.0, no version bumps. KNOWN GAP (why this is a draft): the generated demo audio assets are NOT in this tree, and backend/assets/samples/demo_voice.wav is deleted. onboarding.py guards the missing file (skips seeding the demo profile with a warning), so no crash — but first-run Launchpad will be empty and /demo_audio/ preview URLs 404 until assets are regenerated via scripts/build_demos.sh. Do not merge before regenerating + committing the demo assets. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(dub): timing strategies — kill audio compression, add Concise + Stretch Video Replaces the current audio time-compression default (atempo squeeze to fit slot) that produced chipmunk/alien output on high-density target languages like Bengali. Two new user-selectable modes; legacy behaviour kept behind an explicit "Strict slot" choice. New `DubRequest.timing_strategy` enum (default "concise"): - "concise" Translator trims text to fit at natural rate; if it still overflows, hard-trim at slot with a fade so we never overlap the next speaker. Surface overflow_s per segment so the user can shorten the text. - "stretch_video" Audio plays at natural 1.0× rate. Backend computes a per-segment new timeline; persists a video_stretch_plan on the job. Mux step (dub_export) builds an ffmpeg trim+setpts+concat filter graph that stretches each segment's video portion to match the natural-rate dub audio. Gaps/pre-roll/tail pass through at 1.0×. Sub burn under stretch_video is skipped in one pass (cues would drift). - "strict_slot" Legacy atempo squeeze. Retained for back-compat. Director rate-bias side-effect (seg_speed *= bias) now gated on strict_slot only, so "urgent"/"slow" direction tokens keep their instruct effect in the new modes without chipmunking. Per-segment fit_status emitted in the SSE done event: {status: "fits" | "overflows" | "video_stretched", overflow_s?, stretch_ratio?} DubSegmentRow's "Sync: 100%" badge (which was lying — sync_ratio was always ~1.0 because the TTS loop pre-trimmed to slot) is replaced with a truthful "Fits / Overflows +Ns / Video 1.18×" label. Frontend: - prefsSlice.timingStrategy (persisted, store v3→v4 with safe migrate). - DubTab footer Segmented control: "Concise · Stretch Video · Strict slot". - useDubWorkflow passes timing_strategy on /dub/generate; consumes fit_status. Tests: tests/test_dub_timing_strategy.py — 13 cases covering schema defaults/validation, _build_video_stretch_filter_graph (pre-roll, gap, tail, empty-plan early return, post-subtitle chain-in), and _video_stretch_plan_for guards. 30/30 existing dub tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(waveform): surface missing source as "Source media missing" instead of code-4 black box When a project's underlying media file is gone (moved or deleted between save and reload) the <video> element fires MediaError code 4 and the companion audio fetch returns HTTP 404 — both were silently warned to the console while the user stared at an unresponsive black panel and an empty waveform. - WaveformTimeline now flips loadError when the video element rejects code 3 (decode) or 4 (src not supported), and tracks `sourceMissing` separately so the error UI can name the actual problem. - The audio decode fallback chain catches HTTP 404 specifically and treats it as source-missing instead of loading silent empty peaks — an empty waveform on a deleted source is more confusing than a clear "Re-upload the video to continue" message. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(tray): "Show OmniVoice" reloads when the webview is blank When the dev Vite server restarts (or the main window is created before the backend is ready), the webview load fails and the window is left with `<body></body>` plus a "Could not connect to the server" console error. Clicking "Show OmniVoice" from the tray menu just re-showed the broken window — there was no recovery path short of quit+relaunch. Now the show handler runs a tiny eval after `show()`/`set_focus()` that calls `location.reload()` only when `document.body.childElementCount === 0`. A healthy window doesn't blink (body is non-empty); a blank one self-recovers as soon as the user clicks Show. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(#133): bug-report diagnostics field mapping + drop unused imports Address PR #133 review: - ReportBugButton: /system/info exposes `platform` + `device`, not `os`/`torch_device`/`gpu` — those reads silently dropped OS/GPU from every bug report. Map to the real fields (CodeRabbit). Also remove the dead `home` local in stripHome (CodeQL unused-variable). - DictationDemo: drop unused `Loader` import (CodeQL unused-import). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a1ef66c321 |
Stability pass: DB leaks, App.jsx hooks refactor, desktop bootstrap (#49)
* fix: eliminate DB connection leaks, race conditions, and deprecated asyncio API ## DB Connection Leaks (P0) - Convert 38 raw get_db() calls to db_conn() context manager across 14 router files - Connections are now guaranteed to close even when exceptions are raised - profiles.py create_profile: clean up orphaned audio file if DB insert fails - profiles.py lock_profile: consolidate 3 separate conn.close() error paths ## Race Condition (P1) - Add _dub_jobs_lock (threading.Lock) to protect _dub_jobs dict in dub_pipeline.py - get_job/put_job now thread-safe for concurrent dub sessions ## asyncio Deprecation (P2) - Replace 23 asyncio.get_event_loop() calls with asyncio.get_running_loop() - Prevents DeprecationWarning on Python 3.12+ and future breakage on 3.14 ## Quick Fixes - gallery.py preview_voice: remove filesystem path from error response (P2) - dub_pipeline.py parse_vtt_segments: remove redundant `import re` inside loop (P3) - gallery.py _init_gallery_db: use db_conn() context manager (P2) * refactor: extract hooks, centralize isTauri, add pytest-cov ## Frontend - Extract useTTS hook (150 LOC) — TTS generation, streaming, audio ingestion - Extract useProfiles hook (219 LOC) — voice profile CRUD, lock/unlock, preview - Centralize isTauri detection: dialog.js, VoiceGallery.jsx, Settings.jsx now import from utils/media.js instead of 4 different detection patterns ## Backend - Add pytest-cov to dev dependencies - Baseline coverage: 39% across backend/ (214 tests pass) - Add .coverage to .gitignore * feat: add Vitest + checkJs, extract useDubWorkflow + useAppData hooks ## Frontend Testing (new) - Set up Vitest with jsdom environment + @testing-library/react - 11 tests: utils (isTauri, formatTime, constants) + Zustand store (mode, text, dubStep, pill) - Scripts: 'test' (vitest run), 'test:watch' (vitest), 'test:legacy' (node runner) ## App.jsx Decomposition (continued) - Extract useDubWorkflow hook (387 LOC) — upload, ingest, transcribe SSE, translate, generate SSE, abort, stop, cleanup - Extract useAppData hook (181 LOC) — data loading, localStorage persistence, WebSocket real-time updates, model-status pill management ## TypeScript checkJs - Enable checkJs: true in tsconfig.json for IDE-level type checking - 947 existing errors (informational, not blocking builds) - noImplicitAny remains false to avoid blocking * ci: add Vitest step, fix useProfiles duplicate state ## CI - Add 'Run Vitest (frontend)' step — runs 11 unit tests - Override --checkJs false in CI typecheck to avoid 947 pre-existing errors - Rename legacy test step for clarity ## Hooks - Fix useProfiles to accept loadProfiles from parent (useAppData) instead of managing its own duplicate profiles array * refactor: wire hooks into App.jsx — 2067 → 1129 LOC (-45%) App.jsx now delegates to extracted hooks instead of inline logic: - useAppData: data loading, localStorage, WebSocket, model pill - useProfiles: voice profile CRUD, lock/unlock, preview - useTTS: generation, streaming, audio ingestion - useDubWorkflow: upload, transcribe SSE, translate, generate SSE 988 lines removed. All handler logic lives in focused, independently testable hooks. Store selectors and render JSX stay in App.jsx as the shell. Verified: vite build clean, 11 frontend + 214 backend tests pass. * feat: show real-time percentage on model loading pill Backend: register hf_progress listener during _load_model_sync() so download/weight-loading tqdm events update _loading_detail with a progress percentage (0-99%). get_model_status() now includes a 'progress' field that the frontend polls. Frontend: useAppData reads msQuery.data.progress and calls setPillProgress() — the FloatingPill already renders the percentage text and progress bar width from this value. * fix: prevent FileNotFoundError in desktop bundle during model init transformers >=4.52 calls _can_set_experts_implementation() and _can_set_attn_implementation() during PreTrainedModel.__init__, which open the class source file via open(class_file). In a Tauri desktop bundle, module.__file__ points to a path that doesn't exist on disk, causing: FileNotFoundError: .../omnivoice/models/omnivoice.py Override both classmethods on OmniVoice to return static values without filesystem access. OmniVoice doesn't use MoE experts (return False), but does support flex/flash attn (return True). * fix: sync source dirs on every bootstrap, not just first run The Tauri bootstrap previously only copied omnivoice/ and backend/ to Application Support on the first run. Subsequent app updates kept using stale source files, preventing bug fixes from landing. Now ensure_venv_ready() always syncs both directories from the bundle resources before returning, even when the venv is healthy. This fixes the FileNotFoundError crash where the old omnivoice.py lacked the _can_set_experts_implementation override. * ui: premium setup wizard polish - Primary button: solid gradient fill with hover glow + lift + press - Stepper nav: connected pills with glow ring on active step - Welcome cards: glassmorphism with stagger-in animations, lucide icons, left-border accent strip, hover translate - Preflight panel: colored icon pill backgrounds, stagger-slide entrance - Step transitions: fade+slide animation via keyed wrapper - Footnote: shortened paths (~/ notation), Reveal in Finder button - Recommendation banner: gradient background with accent glow - Compact spacing throughout for denser, professional layout * fix: kill zombie backend on clean+retry bootstrap When clean_and_retry_bootstrap removes the project dir, any old uvicorn process still running from the deleted paths remains alive on port 3900. The subsequent retry_bootstrap sees the port is healthy and attaches to the zombie instead of re-bootstrapping. Now explicitly kill any process on the backend port after cleaning, before calling retry_bootstrap. * feat: integrate speaker clones into dubbing interface, sanitize system environment variables for subprocesses, and improve FFMPEG binary path resolution. * fix: restore docker compose default + drop dead setSeed call - deploy/docker-compose.yml: remove profiles: ["cpu"] from the default service so `docker compose up` matches the comment on line 5. With the profile present, no service auto-started. - frontend/src/App.jsx: drop the setSeed call in restoreHistory. The selector was never reintroduced after the App.jsx hooks split, and there is no seed state in the store — seeds are generated fresh per call in useTTS and only read from history items for display. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: address CodeRabbit review — async detection, dub stream, bootstrap fail-fast - backend/services/tts_backend.py: invert async-context detection in _ensure_loaded. The previous code unconditionally caught its own diagnostic RuntimeError and then called asyncio.run() inside a running loop, masking the intended error message. - frontend/src/hooks/useDubWorkflow.js: require a terminal `done` event before reporting dub success. Without this, a dropped stream after partial progress would flip the UI to `done`, refresh history, and play the completion ping as if generation finished. - frontend/src/hooks/useDubWorkflow.js: restore the previous step when tasksCancel() fails. The UI was getting stuck in `stopping` forever on cancel errors. - frontend/src-tauri/src/bootstrap.rs: fail-fast when source sync fails after the existing directory has already been removed. The previous warn-and-continue path could leave the install with no backend/ or omnivoice/ sources and defer the failure to backend startup with a cryptic error. - backend/api/routers/generation.py: add `from e` to the ValueError → HTTPException re-raise (Ruff B904). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: preserve % suffix in TTS generation timer The 100ms timer in useTTS was rewriting generationTime to a plain elapsed-seconds string, which immediately wiped the "(xx%)" download suffix written on the next iteration of the response-body loop. The real-time percentage was flickering on/off as a result. Read the previous value inside the setter and reattach any existing percent suffix. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
994c6cf065 |
feat(backend): setup wizard router, translation engines, export options, client-disconnect handling
- Add setup router (backend/api/routers/setup.py) for first-run wizard: system checks, engine probes, model downloads with progress - Add translation engines service with pluggable backends - Add utils/hf_progress for HuggingFace download progress streaming - Add PyInstaller runtime hooks (numpy compat, torch compiler disable) - Global exception handler short-circuits h11 LocalProtocolError and Starlette ClientDisconnect with HTTP 499 to silence noisy stack traces when users scrub or cancel video mid-stream - /dub/download-mp3 accepts bitrate query param (clamped 64–320kbps) - Refactor ASR/TTS backends, dub pipeline, engine management - Update backend.spec for PyInstaller packaging - Bump pyproject version to 0.2.0; refresh uv.lock Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
52d68d05dc | refactor: update backend architecture, expand frontend state management, and synchronize voice-pro research modules. | ||
|
|
6e89db0f77 | feat: add YouTube/URL ingestion support using yt-dlp and update UI with granular preparation progress tracking | ||
|
|
15ef65a686 |
refactor: extract App.jsx into pages/components/api, add studio design pass + NavRail
Frontend:
- Split App.jsx (2850 -> ~1300 lines) into pages/{Launchpad,CloneDesignTab,DubTab,Settings}.jsx
and components/{Header,Sidebar,NavRail,CompareModal,DubSegmentTable,DubSegmentRow}.jsx.
- Centralize every fetch through api/{client,dub,generate,profiles,projects,system,exports}.js
with consistent ApiError + JSON error detail extraction.
- Extract utils: constants (TAGS/CATEGORIES/PRESETS/POPULAR_*), languages (LANG_CODES),
format (formatTime/probeAudioDuration), consoleBuffer (ring for Settings > Logs > Frontend).
- Lazy-load AudioTrimmer, DubSegmentTable, Launchpad, CloneDesignTab, DubTab, Sidebar, CompareModal, Settings.
Initial bundle 438KB -> 220KB (-50%).
- Virtualize segment table (react-window) + React.memo row; dynamic row height when
original text row shown. Fix {proj.is_locked && ...} rendering literal "0" on falsy.
- Segment UX: multi-select + bulk voice/lang/delete, Ctrl+D split-at-cursor,
Ctrl+M merge-with-next, search/filter/speaker filter, char-budget warn
when translated text >1.3x source, preserve text_original across translations,
always stream segments via EventSource /dub/transcribe-stream.
- AudioTrimmer: pro-grade zoom/pan/scrub, rAF-throttled drag, peak precompute
+ async refine (keeps UI responsive on 1700s mp3), keyboard shortcuts,
click-drag = fresh selection, loop preview, Enter/Esc/Space/Home/End bindings.
Tests: 22 cases in tests/frontend/audioTrim.test.mjs cover encodeWav header,
peak min/max invariants, drag modes (start/end/region/pan/new), zoom math,
slice-to-mono, tick interval picker.
- NavRail: left/right vertical icon rail (VS Code style), side persisted to
localStorage. Removes tab group from Header.
- View-specific sidebar: hidden on Launchpad/Settings; dub gets 3 tabs; clone/design 2.
app-container grid updated with sidebar-hidden / rail-right variants.
Design pass (hand-drawn "cute" identity):
- Fraunces italic serif for headlines, Nunito rounded sans for body.
- Wobbly non-uniform border-radius on cards/buttons/inputs.
- Warm peach/rose/lime palette across Launchpad, Settings, Header, panels,
sidebar items, segment table.
- Header HQ: view breadcrumb (pulsing dot + kicker + accent-colored view label
+ active project), live mini-waveform reacting to model status.
- Settings: tabbed Models / Logs / About / Privacy with accent-colored pills;
Logs sub-tabbed Backend / Frontend / Tauri.
- Clone/Design: two columns, each split into two studio-panels (prompt vs lang/steps;
voice source vs overrides+synth). Rectangular corners + launchpad-warm gradient.
- Dub tab: migrate panels to studio-panel look.
- Sidebar item overhaul: kind pill + ago timestamp + hover-reveal pill actions,
accent left-edge bar on hover, click-whole-card = primary action.
Backend:
- Per-segment translate retry + auto-src fallback; surfaces errors per-segment
with logged type. Tests: 9 cases in tests/test_dub_translate.py (code coverage,
source-lang resolution, retry/auto fallback, empty text handling).
- Thumbnail extraction during dub upload (ffmpeg @ ~10% offset, scale 320px wide)
served via GET /dub/thumb/{job_id}.
- Preserve text_original during transcription so cross-language retranslations
re-run from pristine source, not compounding on prior translation.
- Settings endpoints: GET /system/info, GET /system/logs?tail=N,
GET /system/logs/tauri, POST /system/logs/clear. FK-safe profile_id NULL on
history insert to avoid "FOREIGN KEY constraint failed" on stale/preset ids.
- /dub/transcribe-stream SSE: chunked mlx_whisper / pytorch pipeline per 30s
window, diarization final pass, honors job.aborted.
Tests: 22 frontend trim + 9 backend translate all pass. Production build 220KB
gzipped (from 127KB main-only pre-split, which hid unshipped modules in main).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
508d4ba116 | feat: add audio trimming for reference clips, implement streaming transcription, and refactor ffmpeg utility handling | ||
|
|
67328d04fe |
refactor: split backend into api/core/services/schemas, harden security + fd pressure, add searchable language picker, fix segment fragmentation
Backend:
- Split monolithic main.py into backend/{api/routers,core,schemas,services}
- core/db.py: allowlist-gated migrations, db_conn context manager (kills SQL injection on ALTER)
- core/tasks.py: lock-guarded listener add/remove/push, snapshot-before-iterate
- services/ffmpeg_utils.py: run_ffmpeg helper with concurrency semaphore, EAGAIN retry, guaranteed reap
- services/segmentation.py: Bengali/CJK/Arabic punctuation, ultra-short tier, stitch_adjacent_shorts,
bounded-loop merge; public clean_up_segments API
- services/model_manager.py: robust lock.locked() handling
- api/routers/dub_core.py: job_id traversal guard, thread-safe _active_procs, timeouts on ffmpeg/demucs,
POST /dub/cleanup-segments endpoint
- api/routers/dub_export.py: guarded SSE listener remove, ffmpeg timeouts via run_ffmpeg
- api/routers/exports.py: destination_path validation, safe source resolver, subprocess list-form
- api/routers/generation.py: contextlib.suppress on tempfile cleanup, db_conn usage, safe output-path helper
- api/routers/system.py: try/finally tmp cleanup, subprocess timeouts
- schemas/requests.py: TranslateSegment.id int->str to match hex segment IDs
- main.py: threading.Lock around crash log writes
Frontend:
- components/SearchableSelect.jsx: popover combobox with search, keyboard nav, popular+recent pins, 200-item cap
- App.jsx: wire SearchableSelect for dub language / ISO code / voice-gen language; Clean Up segments button;
fix blob URL leak (object-shaped prev in setter, unmount cleanup via ref)
- components/WaveformTimeline.jsx: explicit <video> detach instead of innerHTML='' to release decoder
- index.css: ss-* combobox styles matching Gruvbox theme
Tests:
- tests/test_segmentation.py (26 cases), test_dub_transcribe.py, test_dub_export_unique.py, conftest.py
Chore:
- .gitignore: exclude omnivoice.zip, /research/ reference clones
- Remove tracked stray root test scripts + crash_log.txt
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|