Commit Graph
11 Commits
Author SHA1 Message Date
debpalash f70199db32 fix(security): make path containment explicit to analysis 2026-08-10 04:56:42 +00:00
debpalash 3e6679c03e fix(security): enforce filesystem trust boundaries 2026-08-09 21:16:55 +00:00
283ef36b13 feat(dub): Voice match toggle — per-line prosody vs one consistent reference per speaker (#1147)
* feat(dub): Voice match toggle — per-line prosody vs one consistent reference per speaker

Owner report: "still 4 segments different in voice as they are 4 times done
from each segment?" — Wave 3.2 clones each dub line from a reference cut from
its OWN source audio (great prosody match), but the voice IDENTITY drifts
line to line, and heuristic-diarized jobs have no pooled speaker clones to
anchor it. The precedence was hardcoded; now it's a per-dub-job setting.

DubRequest.voice_match:
- "per_line" (DEFAULT, unchanged): segment clip preferred, speaker clone
  fallback — byte-identical to the previous behaviour.
- "consistent": ONE reference per speaker for the whole dub. `auto:` bindings
  use the pooled speaker clone; when none exists (heuristic diarization skips
  extraction entirely — the key case) a deterministic pick among that
  speaker's segment clips (longest ≥3 s, tie-break lowest segment id) is
  reused for every line. Server-default self `auto-seg:` bindings join the
  pick (they're what prepare stamps on heuristic jobs — the Voice dropdown
  can't even render them, so no user choice is overridden); explicit CROSS
  auto-seg bindings still honour their clip. The shared pick is multi-use,
  so it stays warm in the clone-prompt cache (#1132 cache_ref semantics) at
  both the main generate and the OOM-retry call site.

voice_match is part of the segment fingerprint when non-default (mixed in
like track_lang, so all stored hashes keep their values): flipping the toggle
marks segments stale instead of letting "Regen changed" splice mixed-identity
voices (#281 class). The client sends the mode on both /tools/incremental
recompute paths.

UI: a compact Voice-match Segmented control next to the Timing picker in the
dub panel, persisted in the prefs slice; labels + tooltips in all 21 locales.

Tests: resolution through the real dub_generate path for both modes (incl.
the 4-segment heuristic job unifying on one ref — fail-before/pass-after),
pick determinism + tie-breaks, schema validation, fingerprint semantics, and
frontend store→request wiring.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(changelog): Voice match toggle entry under Unreleased (#1147)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 01:08:50 +05:30
d90cfde1bb feat(dub): per-language translations + per-track caches — switching languages stops destroying work (P1) (#958)
* feat(dub): per-language translations + per-track caches — switching languages stops destroying work (P1)

Multi-language dubbing translated per language (#957) but stored everything
in single-slot state, so tracks silently destroyed each other's work:

P1.2 — per-language translation storage (additive):
- Frontend keeps every translation in s.translations[langCode] alongside the
  legacy s.text slot (still = the shown language). Translate All writes both;
  the new store action switchDubLangCode swaps text through the map on a
  user-driven language switch (non-destructive; restore paths keep the plain
  setter); manual edits / restore-original update the current language's
  entry; merge joins per-language texts, split drops them. Rides project
  save/load inside dubSegments — legacy projects behave exactly as before.
- Backend mirrors it as job["segments_i18n"] = {lang: {segKey: text}}
  (segKey = stable id, index for id-less legacy rows), written by
  _sync_job_segments; job["segments"] stays byte-identical for every existing
  consumer. /dub/srt|vtt?lang= and subtitle burn-in now emit THAT language's
  text when present — ExportModal's "all dubs" batch stops producing N
  identical files. Legacy jobs without the field fall back to today's output.

P1.3 — per-track WAV cache + fingerprints:
- Per-segment WAVs are language-keyed (seg_{lang}_{id}.wav). The partial-regen
  read path falls back to legacy seg_{id}.wav ONLY while the job has no
  other-language track — single-language jobs keep their whole on-disk cache;
  multi-track jobs stop splicing the last-generated language into the current
  track. Read-only endpoints (segment preview, clips zip) gained ?lang= with
  the permissive legacy fallback they always had.
- Fingerprints include the track language (segment_fingerprint(track_lang=…),
  /tools/incremental lang=…) and live in job["seg_hashes_by_lang"]; the flat
  job["seg_hashes"] stays as the current track's mirror so the done event,
  history restore and older frontends read it unchanged. A legacy flat map is
  attributed to the job's last-generated language (dropped when unknown) and
  reads stale once — the safe direction. seg_wav_kind is per-track too.
- The frontend stores fingerprints per language and judges "Regen N changed"
  against the ACTIVE track; project save/load and dub-history restore carry
  all tracks' hashes (segHashesByLang / seg_hashes_by_lang, additive).

Tests: fail-before regression coverage — two-track regen never splices the
other language's audio (sample-level assert on the mixed track), legacy
single-track cache reuse + multi-track gate, per-lang seg_hashes with flat
mirror + migration semantics, /dub/srt|vtt?lang= emitting different text per
track with legacy fallbacks, per-lang burn-in, /tools/incremental lang
scoping, and 14 frontend tests for translations round-trips, per-track
fingerprints and legacy-project behaviour. Full backend + frontend suites,
typecheck, lint and format:check green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): add per-language storage + per-track caches under [Unreleased] (#958)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 03:05:40 +05:30
Palash DebnathandClaude Opus 4.8 47729057bd chore(lint): remove unused imports + variables (ruff F401/F841) (#210)
Autofixes the genuine lint behind the CodeQL py/unused-import and
py/unused-local-variable note-level alerts — actually removing the dead
code rather than dismissing it. 68 safe fixes via 'ruff check --select
F401,F841 --fix' across 29 backend files (dead stdlib/symbol imports like
io/sys/json/torch/typing.Optional and unused locals). Only ruff's safe
fixes applied — the 9 'unsafe' fixes and the audio_dsp numpy availability
import were left untouched.

Not touched: empty-except (needs per-site judgement, not autofixable);
frontend js/unused-local-variable (eslint no-unused-vars has no autofix);
the loopback-low-risk path/log/stack-trace alerts (real, left visible).

Verified: full tests/ suite unchanged at 601 passed (the 2 test_supertonic3
failures are pre-existing on main, local .venv state, green in CI).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 07:53:06 +05:30
Palash DebnathandClaude Opus 4.8 e69dcbb6b1 fix(win): subprocess spawns work under bun run dev (--reload) on Windows (#122) (#175)
issue #122 'Extract: Unknown Error' on Windows: 'bun run dev' fails, running
backend+frontend separately works. Root cause: dev:api launches uvicorn with
--reload, so use_subprocess=True, and uvicorn 0.42's asyncio_loop_factory
EXPLICITLY forces the SelectorEventLoop on Windows in that case (passed as
loop_factory to asyncio_run, overriding any policy). The SelectorEventLoop has
no subprocess support -> asyncio.create_subprocess_exec raises
NotImplementedError. 'python backend/main.py' (no reload) uses ProactorEventLoop
-> works. So an event-loop-policy fix is futile; the thread fallback is the fix.

The ffmpeg extract path already routed through _spawn_async's thread fallback
(landed in #157), but several other spawn sites used raw create_subprocess_exec
and stayed broken on the dev loop:
- add public spawn_subprocess() (drop-in for create_subprocess_exec) that routes
  through _spawn_with_retry -> _spawn_async (NotImplementedError -> thread
  fallback + EAGAIN retry); native asyncio path unchanged on supported loops.
- fix _spawn_thread_fallback to forward cwd/env/etc. to subprocess.Popen (was
  silently dropping them -- breaks sonitranslate's cwd= pip install).
- convert raw spawns: dub_generate atempo, tools ffprobe, gallery yt-dlp (x2),
  sonitranslate install (x4). translation_engines already had its own fallback.
- tests: NotImplementedError -> thread fallback; cwd forwarding; stdin input
  (atempo); native path unchanged.

No behavior change off the broken loop (macOS/Linux/Windows-prod): the native
asyncio subprocess is still used; the fallback only triggers on NotImplementedError.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 21:15:09 +05:30
debpalash 41c23f6b3a feat: enhance ASR performance and reliability with binary bundling, model warmup, sub-stage progress tracking, and optimized polling. 2026-04-30 07:46:35 +05:30
debpalash 2cd1ab4fb9 feat: batched TTS, cold start, audiobook editor, context-aware pipeline
Batched TTS:
- Profile-grouped segment processing for cache locality
- CPU/GPU pipelining (ref audio load overlaps TTS inference)
- ~25-40% throughput improvement over sequential loop
- SegmentSpec container + generate_segments_batched() async API

Cold Start Optimization:
- Deferred torch + OmniVoice imports in model_manager.py
- Server starts in ~0.03s (was ~4s) — health/status respond immediately
- _lazy_torch() / _lazy_omnivoice() wrappers with singleton caching
- All downstream refs updated (idle_worker, free_vram, offload, restore)

Stories / Audiobook Editor:
- StoriesEditor component — multi-track with per-character voice assignment
- 7 character slots (Narrator + 6 characters) with color-coded dots
- Inline TTS preview per line via /dub/preview-segment endpoint
- Add/remove/reorder tracks, Generate All workflow
- Character stats footer (lines, characters, est. duration)

Context-Aware Pipeline:
- Video frame extraction via ffmpeg at segment midpoints
- Frame analysis: brightness, mood, complexity via PIL image stats
- Per-segment and global context (VideoContext container)
- get_segment_context() → natural-language TTS instruct hints
  e.g. 'Speak with vibrant energy, dark atmosphere, fast-paced scene'
- POST /tools/video-context/{job_id} API endpoint

Roadmap: ALL items completed ✅
2026-04-28 12:00:02 +05:30
debpalash b054249be2 feat: plugin SDK, GPU sandbox, waveform v2, accessibility
Plugin SDK:
- Abstract TTSPlugin base class with register/discover pattern
- Built-in plugins: ElevenLabs (cloud) + Bark (local)
- Auto-discovery from backend/plugins/ directory
- GET /tools/plugins API for frontend engine picker

GPU Crash Sandbox:
- Subprocess isolation for GPU-intensive operations
- CUDA OOM / driver crash kills worker, not the server
- Async wrapper with configurable timeout
- Platform availability check

Waveform Timeline v2:
- Added MinimapPlugin (20px overview bar)
- Added TimelinePlugin (time labels)
- Keyboard shortcuts: J/K/L (rewind/play/forward), Space
- Full ARIA labels on all controls
- role=region, role=toolbar for assistive tech

Accessibility:
- ARIA labels on waveform controls, theme picker, capture button
- role=radiogroup on theme dots
- aria-checked state on theme selection
- Keyboard hint icon (J/K/L) in waveform toolbar

LLM Translation: already implemented (OpenAI provider in dub_translate)
Roadmap: cleaned up, only batched TTS + vision items remain
2026-04-28 11:52:53 +05:30
debpalash e2f576f59e feat: MCP server + audio effects chain
MCP Server:
- Full Model Context Protocol server (backend/mcp_server.py)
- 5 tools: generate_speech, list_voices, list_personalities,
  list_languages, check_health
- 2 resources: voice://{id}, history://recent
- stdio + SSE transports for Claude Desktop / Cursor / remote agents
- Example config: mcp.json

Audio Effects Chain:
- 6 presets: Broadcast, Cinematic, Podcast, Warm, Bright, Raw
- Configurable pipeline via apply_effects_chain() with pedalboard
- Effects: highpass, lowpass, compressor, reverb, noise_gate, eq, limiter
- GET /tools/effects API for frontend preset picker
- Graceful fallback when pedalboard isn't installed
2026-04-28 11:28:26 +05:30
debpalash 52d68d05dc refactor: update backend architecture, expand frontend state management, and synchronize voice-pro research modules. 2026-04-21 18:32:25 +05:30