Direct torchaudio.save(path, ...) writes bytes as the encoder produces
them. SIGKILL, OOM kill, or Tauri sidecar reap mid-write leaves the file
at `path` truncated. Downstream tools (ffmpeg in the dub mux, NLEs the
user imports the WAV into) happily read truncated RIFF — the header
appears first, then the data chunk gets cut short — and surface as
silently corrupt audio later in the pipeline. That's the shape of #48.
New helper services/audio_io.py:atomic_save_wav() writes to a sibling
temp file in the same directory then os.replace() into place. POSIX
rename(2) is atomic; os.replace() ports the same guarantee to Windows.
Either the target ends up with a complete WAV or it keeps its previous
contents (or never exists) — no third state.
Migrated three call sites in api/routers/dub_generate.py:
- L289: RVC per-segment write
- L328: deferred batch write of all segments
- L390: final mixed-track export
Left L508 alone — it writes to BytesIO (in-memory response body), no
atomicity needed.
Implementation note (recorded as a docstring in audio_io.py): the temp
file must end in `.wav`, not `.tmp`. torchaudio.save infers the output
format from the path suffix and ignores the `format=` kwarg with the
soundfile backend. A `.tmp` suffix raises "Unsupported format: tmp". The
leading dot + target-name prefix still marks the file as transient.
Tests in backend/tests/test_atomic_wav.py:
- success path: writes valid WAV, no temp leaks, overwrites cleanly
- atomicity: target unchanged when save raises (pre-existing target)
- atomicity: target absent when save raises (new target path)
- no temp leaks on failure
- temp file lives in target_dir (cross-fs renames are not atomic)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The streaming-ASR WebSocket accepted any local connection without an
origin check. Any process running as the same user (rogue extension,
malware, sibling Electron app) could open ws://127.0.0.1:3900/ws/transcribe
and exfiltrate the user's live microphone audio in real time.
HTTP routers gate sensitive endpoints with Depends(require_loopback) at
the router level (see backend/api/dependencies.py). FastAPI's WebSocket
dependency injection differs across versions, so the guard is inlined in
ws_transcribe before websocket.accept() — non-loopback origins receive a
1008 (Policy Violation) close and never see the open socket.
Three source-level tests in backend/tests/test_capture_ws.py guard
against regression — same shape as tests/test_bind_host.py:
- references _LOOPBACK_HOSTS in the handler
- closes non-loopback with code 1008
- close() appears before accept() in the source
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three P0 fixes bundled — foundation cleanup before v0.3.0 phase work. Closes release.yml drift (PR #51's tabs broke v0.3.0 tag releases), production bind exposure (Critic F1), and 9-endpoint LAN gap on /system/* (Critic F2+F3). 5 new tests; 243 full pass.
Community contribution from @fishandsheep.
Adds PYTHONPATH=/app/backend to docker-compose env for both omnivoice (CPU) and omnivoice-gpu service blocks, so the pre-built Docker image can import backend modules correctly on first boot.
Complements PR #74 (Docker GPU detection) — different sections of docker-compose.yml.
Thanks @fishandsheep!
Production deployment hardening, dub OOM recovery, new SoniTranslate sidecar engine, ASR backend expansion.
- Dub generation OOM recovery: backend/api/routers/dub_generate.py:163-209 adds OOM detection + one retry with reduced nstep
- New SoniTranslate sidecar engine: backend/api/routers/sonitranslate.py + backend/services/sonitranslate.py (subprocess-based dubbing pipeline, opt-in)
- ASR backends expansion: backend/services/asr_backend.py adds NeMo Parakeet TDT, Moonshine, additional Whisper variants; new GET /system/asr-backends endpoint
- Dub UI polish: tighter spacing in DubSegmentRow.css, DubTab.css
Issue #78 (speaker diarization mis-assignment) NOT addressed by this PR — the bundled diarization changes are in the new SoniTranslate sidecar, not the existing pyannote pipeline. Keeping #78 open.
No DB schema changes, no migration. Backward-compatible for existing user data.
Dictation pill widget no longer displays the idle "Ready — hold shortcut to speak" state by default. The widget now appears only when actively used (global shortcut press or tray "Start Dictation" click).
Two surgical edits to frontend/src-tauri/src/lib.rs:
1. Pill-mode setup: removed win.show() + win.set_focus() on the widget. Kept positioning so the first show appears at top-center without animation flicker.
2. Tray "dictate" handler: now positions + shows + focuses widget BEFORE emitting tray-dictate, mirroring the global-shortcut handler. Previously tray-initiated dictation would record silently with no visible UI.
Trade-off accepted: the original auto-show was intended to prevent a "looks-launch-failed" first-run experience for users without Accessibility permission. The tray icon + "OmniVoice Dictation" tooltip provide app-running signal; first-launch onboarding toast can be added later if support requests indicate confusion.
First v0.3.x release on the Phase 0 cross-platform CI baseline.
## Cross-platform bug fixes (375ea4e)
User-reported bugs from a Pinokio/Windows session:
- Docker `compose --profile gpu up` no longer port-conflicts on 3900 — restored `profiles: ["cpu"]` that #49 wrongly reverted on CodeRabbit's advice.
- Argos / pip install from the UI now works inside Docker — added `_in_virtualenv()` runtime check; `run_pip` injects `--system` automatically when on system Python.
- Speaker diarization warning toast — when pyannote silently falls back to the silence-gap heuristic (missing HF_TOKEN, license not accepted, network blocked), `_diarize()` now returns `(segments, warning)`; `useDubWorkflow` renders an 8-second toast.
## Dub editor UX (d5df454)
Six fixes per annotated screenshots:
- Editable segment start times (`m:ss.s` or raw seconds; Esc reverts, Enter commits; rejects overlap with end).
- Click a transcript row → seek the waveform/video (`WaveformTimeline` now forwardRef's `seekTo(time)`).
- Speaker is datalist-backed (pulls from detected speaker clones; free text still allowed).
- Scissors menu splits at cursor — uses live caret, then last caret, then sentence-boundary fallback.
- Mouse-wheel scrolls the waveform; Cmd/Ctrl left alone for browser pinch-zoom.
- Menu popover collision: added `avoidCollisions` + `collisionPadding=8` to Radix Content; removed `position: fixed` from `.ui-menu`.
## VRAM-aware GPU pool (73dbe18)
`_gpu_pool` was hardcoded `ThreadPoolExecutor(max_workers=1)` since introduction — every TTS forward serialized through one thread.
- CUDA / ROCm: `workers = clamp(1, free_GB // 2.5, 4)`. 16 GB card with ~14 GB free → 4 workers → ~4× throughput on multi-segment dubs.
- MPS / CPU / unknown: 1 worker.
- `OMNIVOICE_GPU_WORKERS` env var override (clamped 1..16).
- Module `__getattr__` preserves the public `_gpu_pool` symbol for existing callers.
## Stories tab — wire-up + UX (f6bbc7a)
The 264-line `StoriesEditor` component existed but was mounted nowhere. Now wired into NavRail + lazy-loaded on `mode === 'stories'`. Added Paste & Split panel (sentence-boundary chunking) and per-track `[pause 0.5s]` insertion.
## Stories — pauses + inline voice (edd3a1d)
`frontend/src/utils/storyTokens.js` — tokenizer for `[pause X.Ys]` and `[voice:X]…[voice:default]` markers. Voice switches are stateful (carry forward). 13 new vitest cases (vitest now 24/24).
## Verified
- 214 backend tests pass (3 skipped, 10 xfailed, 3 xpassed)
- 23 router-smoke tests pass
- 24/24 vitest cases pass (13 new)
- All 7 Phase 0 CI checks green (Tauri shell + Smoke on macOS/Windows/Linux + Tests)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Adds `request.client.host` allow-list check (`127.0.0.1` / `::1` / `localhost`) to `POST /system/set-env`. Non-loopback callers receive `403` instead of being able to mutate `os.environ` for HF_TOKEN / TRANSLATE_API_KEY.
Surfaced during security review of PR #66, which widens the pre-existing window by persisting these keys to disk via prefs.json. This fix closes the underlying vulnerability so PR #66's revision lands onto a clean base.
Defensive `request.client is None` branch handles ASGI middleware that strips client info. Three new tests cover non-loopback reject, loopback allow, and allow-list still validated on loopback.
Follow-up: `260518-ivy-deferred-items.md` enumerates 5 sibling POST routes in `system.py` that share the same gap — separate PR.
* docs: initialize OmniVoice stabilization milestone project
* chore: add project config (yolo + balanced)
* docs: domain research for stabilization milestone
* docs: define v1 requirements for stabilization milestone
* docs: add GGUF + singing engine spike requirements (Phase 4 new)
* docs: roadmap revision + CLAUDE.md (7 phases, 62 reqs, +GGUF/SING spikes)
* docs(phase-0): add Gates phase RESEARCH.md
Phase 0 research synthesizes the cross-platform CI matrix, frozen
omnivoice_data fixture, installer post-build smoke, SHA-256 checksum
publishing, and PR-template extension into copy-paste-ready YAML and
Python snippets composed entirely from existing in-repo patterns.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(phase-0): add Gates phase CONTEXT, PATTERNS, and PLAN
Phase 0 — Gates is the hard pre-condition for v0.3.x stabilization.
Lays cross-platform CI matrix (macos-14/windows-2022/ubuntu-22.04),
regression fixture (≤200 KB), installer smoke on tag push, SHA-256
checksums in release body + per-OS SHA256SUMS-*.txt assets, PR
template with RC cadence + fixture line, and the open-PR landing
for #51.
Plan covers GATE-01..06; structured into 7 slices (A–G) with explicit
Slice C → Slice G dependency reordering so the new smoke-matrix lands
on main before PR #51 (CONTEXT.md L86 interleave decision).
Plan-checker iteration 2: APPROVED — all 3 BLOCKERs + 3 MAJORs from
iteration 1 resolved (file truncation/Slice-G missing, GATE-06 sibling
PR verification, Slice C ordering, Truth #5 wording, macOS Tauri
WebView avoidance per Pitfall #5, Windows taskkill per Pitfall #2).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(00-gates): seed regression fixture (GATE-01)
- scripts/seed-test-fixture.py — deterministic builder for tests/fixtures/omnivoice_data/
- wipes + rebuilds; fixed created_at=1700000000.0; all-zero PCM for byte-deterministic diffs
- calls backend.core.db.init_db() directly (alembic versions/ is empty — see CONTEXT.md)
- checkpoints WAL → DELETE on close so no -shm/-wal sidecars pollute git status
- exits non-zero if fixture > 200 KB
- tests/fixtures/omnivoice_data/{omnivoice.db, README.md} — 8-table empty DB + 1 voice_profiles row
- tests/fixtures/omnivoice_data/voices/test-voice/{profile.json, sample.wav} — 1-sec 24 kHz mono silence
- .gitignore — explicit allow-list (!tests/fixtures/omnivoice_data/**) so the existing
omnivoice_data/, *.db, *.wav patterns don't hide the fixture from git
Verifies: du = 144 KB on disk; sqlite_master lists 8 init_db tables + sqlite_sequence;
voice_profiles has exactly 1 row id='test-voice'; 0 rows in generation_history.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(00-gates): add tests/smoke/test_boot_smoke.py (GATE-01)
- tests/smoke/__init__.py — package marker so pytest treats tests/smoke/ as a module
- tests/smoke/test_boot_smoke.py — 4 in-process FastAPI TestClient smoke tests:
* test_health_returns_ok — /health returns 200 + {status:ok, device:...}
* test_profiles_endpoint_lists_fixture_voice — /profiles surfaces the seeded
test-voice row (validates OMNIVOICE_DATA_DIR wiring → DB_PATH → init_db schema)
* test_system_info_includes_data_dir — /system/info resolves data_dir
* test_history_endpoint_empty — /history reaches DB and returns []
Test isolation env vars (OMNIVOICE_MODEL=test, OMNIVOICE_DISABLE_FILE_LOG=1)
set at module top BEFORE any backend import — pattern from tests/test_router_smoke.py.
Fixture is copied to a per-session temp dir so the test never mutates the
checked-in artifact (SQLite file-change counter + runtime subdirs like dub_jobs/
would otherwise dirty `git status` after every run).
Failure mode: if tests/fixtures/omnivoice_data/ is missing, pytest.fail at
import time with the regenerate command.
- .gitignore — tighten the GATE-01 allow-list to ONLY the seed-produced files
(README.md, omnivoice.db, voices/test-voice/profile.json, sample.wav).
Prevents future runtime subdirs the backend may create under the fixture
from being accidentally committed.
Verifies: `uv run pytest tests/smoke/ -q --tb=short` → 4 passed in 1.31 s
(target was < 30 s). `git status` clean after a test run.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(triage): record post-planning GitHub state — PR #62, new issues, OOS deferrals
- GATE-06: mark #53 + #61 merged (2026-05-16); add #62 (Wave 1 quick wins) to gate set
- INST-01: note PR #62 implements setuptools pin (closes#58)
- INST-04: note PR #62 lands README docs for #56 workaround
- INST-12: new requirement for #65 Windows Triton/torch.compile OOM (filed post-planning)
- Out of Scope: defer #67/PR #68 (audio effects), #64 (custom model dir),
PR #66 zh-CN (i18n milestone), #63 (empty-template bug)
PR #62 is the user's own Wave 1 work landed as a separate PR while
GSD planning ran in parallel. Merging it eliminates duplicate work
in Phase 1.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(00-gates): add cross-platform smoke matrix (GATE-02)
- New smoke-matrix job on macos-14, windows-2022, ubuntu-22.04
- needs: test, fail-fast: false, timeout-minutes: 10
- Pinned actions: checkout@v4, setup-python@v5, setup-uv@v3 (cache enabled)
- Per-OS ffmpeg + libsndfile install (brew/choco/apt via awalsh128 cache)
- UV_HTTP_TIMEOUT=120, UV_HTTP_RETRIES=5 for restricted-network resilience
- Narrow scope: uv run pytest tests/smoke/ -q --tb=short
- Existing `test` and `tauri-cross-platform` jobs untouched
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci: add workflow_dispatch to ci.yml so smoke-matrix can run on feature branches
* feat(00-gates): add --health-check CLI flag to backend entrypoint (GATE-03)
- argparse on __main__ block; --health-check boots uvicorn in a daemon
thread and polls http://127.0.0.1:3900/health every 5s for up to 60s.
- Prints 'OK — /health responded 200 after Ns' and exits 0 on first 200.
- Prints 'FAIL — /health did not respond 200 within 60s' to stderr and
exits 1 on timeout. Default invocation behavior unchanged.
- No new deps (stdlib argparse/threading/time/urllib.request/sys + uvicorn).
- Consumed by per-OS installer-smoke step in .github/workflows/release.yml.
Verified locally: exits 0 in 5s against tests/fixtures/omnivoice_data/.
* ci(00-gates): add per-OS installer smoke to release.yml (GATE-03)
Adds three matrix-leg-specific steps after 'Build + release (Tauri)',
each gated by runner.os with timeout-minutes: 5:
- macOS (macos-14): hdiutil attach DMG → locate bundled Python backend
inside *.app/Contents (NOT the Tauri WebView shell — RESEARCH Pitfall
#5: WebView hangs on headless runners) → invoke --health-check →
hdiutil detach. Falls back to *.app/Contents/Resources and hard-fails
with a directory listing if no backend binary found.
- Windows (windows-2022): msiexec /quiet install → find backend.exe
under 'C:/Program Files/OmniVoice Studio' → invoke --health-check in
background, wait, then taskkill //F //T //PID to cleanup orphaned
PyInstaller child processes on port 3900 (RESEARCH Pitfall #2).
- Linux (ubuntu-22.04): --appimage-extract (no FUSE on GH runners),
locate binary or AppRun, run under xvfb-run -a.
Bundle-only regressions (PyInstaller missing-module, Tauri sidecar
path mismatch) are invisible to ci.yml's in-process smoke matrix —
this step closes that gap before any release is published.
Verified: YAML parses; all three steps present; gating + timeout
correct; Pitfall #2/#5 mitigations preserved.
* ci(00-gates): publish SHA-256 checksums in release body + as asset (GATE-05)
- Add 'Compute SHA-256 checksums' step writing SHA256SUMS-<label>.txt
per matrix leg using native shasum/sha256sum (Git Bash on Windows).
- Add 'Append checksums to release + attach SHA256SUMS file' step using
softprops/action-gh-release@v2 with append_body: true so the hashes
land in the release body alongside tauri-action's content (not
replacing it) and the file is uploaded as a release asset for
'shasum -c SHA256SUMS-<label>.txt' verification.
- Both steps gated by 'github.event_name == push && refs/tags/v*' so
workflow_dispatch dry-runs do not attempt to attach to a non-existent
release (per CONTEXT.md L70 + RESEARCH Pitfall #7 deferral of any
aggregate cross-leg SHA256SUMS job).
- fail_on_unmatched_files: true to surface path-resolution errors loudly.
* docs(00-gates): document RC cadence + regression-fixture check in PR template (GATE-04)
* docs(setup): add HF token persistence guide for macOS/Windows/Linux (DOCS-05)
Covers two persistent paths:
- Method A — canonical ~/.cache/huggingface/token via huggingface-cli login
- Method B — shell env var (~/.zshrc / ~/.bashrc / Windows User scope)
Documents the v0.2.7 "session only" in-app behavior + notes that
Phase 1 AUTH-03 will make in-app pastes write to the canonical file.
Bundled with Phase 0 PR per user request. Strictly DOCS-05 scope —
zero code changes, no engine touches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* spec(auth): redesign HF token resolution as 3-source cascade with fallback (AUTH-01..06)
Replaces the env_store.py file-based design with a SQLite-backed app
store + cascade resolver that checks app → env var → ~/.cache/huggingface/token
in priority order, with automatic fallback to next source on HTTP 401.
User-explicit design decision:
- App-stored token (SQLite settings table, AES-GCM encrypted) wins
- Env var ($HF_TOKEN) second
- Global huggingface-cli login file third
- All three sources visible in Settings → API Keys with "Active" badge
- Save action populates BOTH app store AND canonical HF file (defense in depth)
New requirement:
- AUTH-06 — on 401, auto-retry next source in cascade before erroring
Also: traceability count corrected (62 → 74 — undercount at planning +
INST-12 + AUTH-06 added post-planning). All 74 v1 reqs mapped.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): backend recognizes HF token from canonical file, not just env var
Two call sites were only checking $HF_TOKEN env var, missing the canonical
~/.cache/huggingface/token file written by `huggingface-cli login` (or the
app's future Save action):
- system.py `/system/info` `has_hf_token` flag — UI showed "No HF token"
even when `huggingface-cli login` had populated the file.
- model_manager.get_diarization_pipeline — pyannote diarization silently
returned None when only the canonical file was set. This is the bug
behind issue #35 (speaker diarization setup failure).
Both fixes use the same pattern: env var > huggingface_hub.get_token()
(which reads the canonical file). Adds a local _has_hf_token() helper
to system.py with a comment marking it as prelude to the AUTH-01..06
cascade (Phase 1 token_resolver.py will layer SQLite app-store on top).
Closes#35 sub-issue (canonical token invisible to diarization).
Cross-cuts AUTH-02 + AUTH-06 design for Phase 1.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(dictation): make pill-widget mode reachable from GUI + scripts (INST-13)
The dictation widget infrastructure shipped in PR #40 but was only reachable
via the undocumented --pill CLI flag. Adds three discovery paths:
1. Tray menu: "Switch to Dictation Widget" (studio mode) — saves
launch_as_widget=true to config, relaunches with --pill, exits current.
Mirrors the existing "Open Studio" path in pill-mode tray.
2. Persistent config: AppConfig.launch_as_widget (bool, default false). Read
at startup via load_config_pre_app() (uses dirs-next, no AppHandle
required). CLI --pill still takes precedence when explicitly passed.
3. Tauri commands: get_launch_as_widget / set_launch_as_widget for the
Phase 2 Settings UI to bind a checkbox to.
4. Scripts: bun desktop-prod:pill / desktop-prod:run:pill — forward --pill
to the bundled app launch. macOS uses `open -n --args` to spawn fresh
instance with the flag.
Closes the GUI half of INST-13. Phase 2 closes the Settings UI half.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(dictation): show widget unconditionally on pill-mode launch + visible Suspense fallback
Before: pill mode set up correctly but the widget window stayed hidden
until ⌘⇧Space was pressed. New users saw absolutely nothing on launch
(no main window, no dock icon, hidden widget) and assumed the app
failed. If global-shortcut Accessibility permission wasn't granted,
they had no path to discover the widget at all.
Two changes:
1. lib.rs: in pill_mode_setup, explicitly show + position + focus the
widget window after hiding main. With per-call error logging so we
can diagnose failures (and a clear error log if widget window
wasn't created at all — points at tauri.conf.json regression).
2. main-app.jsx: Suspense fallback was `null`, which combined with
widget's transparent+decorations:false config made any lazy-import
delay or failure invisible. Now renders a dark pill saying
"Loading dictation…" so even if CaptureWidget lazy-import stalls,
the user sees the window exists.
Studio mode behavior unchanged — widget stays hidden until hotkey
or tray click triggers it (existing show() call in the shortcut/
menu handlers is preserved).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(dictation): create widget window programmatically; Tauri 2 silently dropped config-array creation
Root cause: declaring the widget window in tauri.conf.json's app.windows[]
silently failed in Tauri 2 — get_webview_window("widget") returned None
even though the config was syntactically valid. Probable culprit was the
transparent + decorations:false + visible:false combo, but Tauri offered
no error message either at startup or via webview_windows() enumeration.
Diagnosed by adding webview_windows() enumeration logging at setup start
(only ["main"] ever appeared) and a programmatic WebviewWindowBuilder
fallback that surfaces real Result errors.
Fix:
- tauri.conf.json: widget entry now has `create: false` to make the
config-vs-programmatic handoff explicit.
- lib.rs setup(): call WebviewWindowBuilder::new(app, "widget", ...).build()
with the exact same surface attributes the config used to declare.
- capabilities/default.json: include "widget" in windows array so the new
window inherits the same Tauri permissions as main.
- tauri.conf.json: remove the invalid `"url": "/?window=widget"` field —
WebviewUrl::App takes a path only, query strings aren't supported.
Both windows now load index.html.
- main-app.jsx: replace URL-query-based widget detection with
getCurrentWindow().label === 'widget' via @tauri-apps/api/window. This
is the Tauri 2-recommended pattern for multi-window apps and works
regardless of URL routing.
Closes the immediate UX bug behind the dictation widget being invisible.
Builds cleanly + manually verified: pill widget visible on screen at
top-center after `bun desktop-prod:pill`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Pin setuptools>=75.0 in pyproject.toml to ensure pkg_resources is
available on Python 3.12+ (fixes#58)
- Add Linux white-screen workaround for Fedora 44 / Ubuntu 24.04 (#56)
- Add firewall/Russia install guide for uv venv failures (#60, #57)
Closes#58
Closes#52. Users who already have correct, pre-synced subtitles can now
skip ASR entirely — they upload a video as normal and then hit "Import
.srt" instead of "Upload & Transcribe". The .srt cues populate the dub
segment list directly, so the rest of the pipeline (translate, dub,
export) just works.
Backend
- services/srt_parser.py: lenient SubRip parser. Tolerates BOM, CRLF,
missing index numbers, dot-vs-comma ms separator, and overlap (shifts
the later cue's start to the earlier's end rather than dropping). Skips
cues with non-positive duration or empty bodies; reports counts so the
UI can warn.
- dub_core.py: new POST /dub/import-srt/{job_id} accepts the .srt file,
parses it, clamps cues that run past the source media's duration, and
replaces job["segments"]. Tries UTF-8 with BOM first, falls back to
latin-1 for legacy Windows subs.
Frontend
- api/dub.ts: dubImportSrt helper with a typed response.
- hooks/useDubWorkflow.js: handleDubImportSrt — sets segments, flips
dubStep to 'editing', shows a toast with per-bucket counts (imported /
skipped / overlap-shifted / clamped) so the user sees what happened.
- pages/DubTab.jsx: "Import .srt" button next to "Upload & Transcribe"
once a job exists, plus a smaller "Import .srt instead" affordance in
the transcription-failure banner — the exact recovery path the
reporter asked for.
Tests
- tests/test_srt_parser.py: 12 cases covering well-formed input,
multi-line cues, dot-as-separator, BOM, CRLF, malformed cues, empty
bodies, overlap shift, overlap-becomes-zero-drop, missing indices,
empty input, and segment shape (sequential ids, speaker filler).
pytest is now 226 passed (was 214); vitest unchanged at 11.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: post-refactor cleanup — wire fingerprints, drop dead code, scope pytest
Follow-up to PR #49. Fixes residual issues from the App.jsx hooks split
and tightens repo hygiene so a bare `pytest` doesn't foot-gun.
Real bug
- frontend/src/hooks/useDubWorkflow.js: setLastGenFingerprints lives in
useSegmentEditing, not on the store. The previous code called
useAppStore.getState().setLastGenFingerprints?.(...) — the optional
chain swallowed the missing method, so the "N segments changed" badge
never updated after a fresh generate until a project save+reopen.
Thread setLastGenFingerprints in from App.jsx; useSegmentEditing()
now runs before useDubWorkflow() to make the setter available.
Dead code from the refactor
- frontend/src/App.jsx: drop unused `showAllProjects` useState and
`pushUndo` from the useSegmentEditing destructure.
- frontend/src/hooks/useDubWorkflow.js: drop 5 unused selectors
(preserveBg, defaultTrack, exportTracks, dualSubs, burnSubs) — the
dub-download logic that needs these lives in App.jsx, not the hook.
Repo hygiene
- backend/api/routers/setup.py.bak: delete 38 KB tracked-in-git backup.
The setup/ subpackage replacement has been in place for a while.
- pyproject.toml: add [tool.pytest.ini_options] with testpaths +
norecursedirs. Previously a bare `pytest` would INTERNALERROR walking
into research/ (1.2 GB of vendored upstream projects with their own
test_*.py files that call sys.exit at module level).
- .github/workflows/ci.yml: run backend/tests/ as a second pytest
invocation. The 23 tests there stub core.config in sys.modules to
avoid the heavy main app import chain — that pollutes import state
for other tests, so they need their own session. Previously these
tests existed in the repo but never ran on CI.
Net effect on lint: 60 → 52 problems (-8) from dead-code removal.
Test counts unchanged: pytest 214 + 23, vitest 11.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: log silent catch failures that mask real bugs
CodeRabbit nitpick on #50: empty catch on the incremental-plan fallback
swallows errors. Extending the fix to the catches in this area that have
the same problem (a real failure would be invisible) while leaving the
genuinely non-actionable cleanup catches alone (EventSource.close(),
localStorage.setItem, fire-and-forget UI promises).
Logged:
- useDubWorkflow.js:97 — transcribe SSE message handler
- useDubWorkflow.js:347 — incremental-plan fallback (the CR finding)
- useDubWorkflow.js:352 — dub generate SSE event dispatch
- App.jsx:552 — exportRecord on Tauri save path
- App.jsx:580 — exportRecord on browser download path
Left silent (cleanup / non-actionable):
- useDubWorkflow.js:68, 112 — evt.close() in SSE teardown
- useDubWorkflow.js:102 — SSE error-event payload parse fallback
- App.jsx:124 — localStorage.setItem (quota / privacy mode)
- App.jsx:789, 901 — fire-and-forget UI promise tails
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: eliminate DB connection leaks, race conditions, and deprecated asyncio API
## DB Connection Leaks (P0)
- Convert 38 raw get_db() calls to db_conn() context manager across 14 router files
- Connections are now guaranteed to close even when exceptions are raised
- profiles.py create_profile: clean up orphaned audio file if DB insert fails
- profiles.py lock_profile: consolidate 3 separate conn.close() error paths
## Race Condition (P1)
- Add _dub_jobs_lock (threading.Lock) to protect _dub_jobs dict in dub_pipeline.py
- get_job/put_job now thread-safe for concurrent dub sessions
## asyncio Deprecation (P2)
- Replace 23 asyncio.get_event_loop() calls with asyncio.get_running_loop()
- Prevents DeprecationWarning on Python 3.12+ and future breakage on 3.14
## Quick Fixes
- gallery.py preview_voice: remove filesystem path from error response (P2)
- dub_pipeline.py parse_vtt_segments: remove redundant `import re` inside loop (P3)
- gallery.py _init_gallery_db: use db_conn() context manager (P2)
* refactor: extract hooks, centralize isTauri, add pytest-cov
## Frontend
- Extract useTTS hook (150 LOC) — TTS generation, streaming, audio ingestion
- Extract useProfiles hook (219 LOC) — voice profile CRUD, lock/unlock, preview
- Centralize isTauri detection: dialog.js, VoiceGallery.jsx, Settings.jsx
now import from utils/media.js instead of 4 different detection patterns
## Backend
- Add pytest-cov to dev dependencies
- Baseline coverage: 39% across backend/ (214 tests pass)
- Add .coverage to .gitignore
* feat: add Vitest + checkJs, extract useDubWorkflow + useAppData hooks
## Frontend Testing (new)
- Set up Vitest with jsdom environment + @testing-library/react
- 11 tests: utils (isTauri, formatTime, constants) + Zustand store (mode, text, dubStep, pill)
- Scripts: 'test' (vitest run), 'test:watch' (vitest), 'test:legacy' (node runner)
## App.jsx Decomposition (continued)
- Extract useDubWorkflow hook (387 LOC) — upload, ingest, transcribe SSE,
translate, generate SSE, abort, stop, cleanup
- Extract useAppData hook (181 LOC) — data loading, localStorage persistence,
WebSocket real-time updates, model-status pill management
## TypeScript checkJs
- Enable checkJs: true in tsconfig.json for IDE-level type checking
- 947 existing errors (informational, not blocking builds)
- noImplicitAny remains false to avoid blocking
* ci: add Vitest step, fix useProfiles duplicate state
## CI
- Add 'Run Vitest (frontend)' step — runs 11 unit tests
- Override --checkJs false in CI typecheck to avoid 947 pre-existing errors
- Rename legacy test step for clarity
## Hooks
- Fix useProfiles to accept loadProfiles from parent (useAppData)
instead of managing its own duplicate profiles array
* refactor: wire hooks into App.jsx — 2067 → 1129 LOC (-45%)
App.jsx now delegates to extracted hooks instead of inline logic:
- useAppData: data loading, localStorage, WebSocket, model pill
- useProfiles: voice profile CRUD, lock/unlock, preview
- useTTS: generation, streaming, audio ingestion
- useDubWorkflow: upload, transcribe SSE, translate, generate SSE
988 lines removed. All handler logic lives in focused,
independently testable hooks. Store selectors and render
JSX stay in App.jsx as the shell.
Verified: vite build clean, 11 frontend + 214 backend tests pass.
* feat: show real-time percentage on model loading pill
Backend: register hf_progress listener during _load_model_sync()
so download/weight-loading tqdm events update _loading_detail with
a progress percentage (0-99%). get_model_status() now includes a
'progress' field that the frontend polls.
Frontend: useAppData reads msQuery.data.progress and calls
setPillProgress() — the FloatingPill already renders the percentage
text and progress bar width from this value.
* fix: prevent FileNotFoundError in desktop bundle during model init
transformers >=4.52 calls _can_set_experts_implementation() and
_can_set_attn_implementation() during PreTrainedModel.__init__,
which open the class source file via open(class_file). In a Tauri
desktop bundle, module.__file__ points to a path that doesn't
exist on disk, causing:
FileNotFoundError: .../omnivoice/models/omnivoice.py
Override both classmethods on OmniVoice to return static values
without filesystem access. OmniVoice doesn't use MoE experts
(return False), but does support flex/flash attn (return True).
* fix: sync source dirs on every bootstrap, not just first run
The Tauri bootstrap previously only copied omnivoice/ and backend/
to Application Support on the first run. Subsequent app updates
kept using stale source files, preventing bug fixes from landing.
Now ensure_venv_ready() always syncs both directories from the
bundle resources before returning, even when the venv is healthy.
This fixes the FileNotFoundError crash where the old omnivoice.py
lacked the _can_set_experts_implementation override.
* ui: premium setup wizard polish
- Primary button: solid gradient fill with hover glow + lift + press
- Stepper nav: connected pills with glow ring on active step
- Welcome cards: glassmorphism with stagger-in animations, lucide icons,
left-border accent strip, hover translate
- Preflight panel: colored icon pill backgrounds, stagger-slide entrance
- Step transitions: fade+slide animation via keyed wrapper
- Footnote: shortened paths (~/ notation), Reveal in Finder button
- Recommendation banner: gradient background with accent glow
- Compact spacing throughout for denser, professional layout
* fix: kill zombie backend on clean+retry bootstrap
When clean_and_retry_bootstrap removes the project dir, any old
uvicorn process still running from the deleted paths remains alive
on port 3900. The subsequent retry_bootstrap sees the port is
healthy and attaches to the zombie instead of re-bootstrapping.
Now explicitly kill any process on the backend port after cleaning,
before calling retry_bootstrap.
* feat: integrate speaker clones into dubbing interface, sanitize system environment variables for subprocesses, and improve FFMPEG binary path resolution.
* fix: restore docker compose default + drop dead setSeed call
- deploy/docker-compose.yml: remove profiles: ["cpu"] from the default
service so `docker compose up` matches the comment on line 5. With the
profile present, no service auto-started.
- frontend/src/App.jsx: drop the setSeed call in restoreHistory. The
selector was never reintroduced after the App.jsx hooks split, and
there is no seed state in the store — seeds are generated fresh per
call in useTTS and only read from history items for display.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: address CodeRabbit review — async detection, dub stream, bootstrap fail-fast
- backend/services/tts_backend.py: invert async-context detection in
_ensure_loaded. The previous code unconditionally caught its own
diagnostic RuntimeError and then called asyncio.run() inside a
running loop, masking the intended error message.
- frontend/src/hooks/useDubWorkflow.js: require a terminal `done` event
before reporting dub success. Without this, a dropped stream after
partial progress would flip the UI to `done`, refresh history, and
play the completion ping as if generation finished.
- frontend/src/hooks/useDubWorkflow.js: restore the previous step when
tasksCancel() fails. The UI was getting stuck in `stopping` forever
on cancel errors.
- frontend/src-tauri/src/bootstrap.rs: fail-fast when source sync fails
after the existing directory has already been removed. The previous
warn-and-continue path could leave the install with no backend/ or
omnivoice/ sources and defer the failure to backend startup with a
cryptic error.
- backend/api/routers/generation.py: add `from e` to the ValueError →
HTTPException re-raise (Ruff B904).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: preserve % suffix in TTS generation timer
The 100ms timer in useTTS was rewriting generationTime to a plain
elapsed-seconds string, which immediately wiped the "(xx%)" download
suffix written on the next iteration of the response-body loop. The
real-time percentage was flickering on/off as a result.
Read the previous value inside the setter and reattach any existing
percent suffix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- New .github/workflows/docker.yml publishes images to ghcr.io on tag push
- README Docker section now leads with 'docker pull' from GHCR
- docker-compose.yml defaults to GHCR image with build-from-source fallback
- Dockerfile: copy README.md for hatchling metadata resolution
* feat: implement frameless OS-level floating dictation widget
- Refactor CaptureButton into standalone CaptureWidget
- Add secondary transparent Tauri window configuration
- Map global hotkey to show/hide widget instead of focusing main app
- Implement auto-hide post-paste
- Add social preview image
* docs: up the game with enhanced README
- Use the high-quality social preview image as the hero image
- Bump download release links to v0.2.7
- Highlight the new Frameless Dictation Widget feature
* docs: complete README overhaul for maximum virality
- Add Highlights section with 2-column feature grid
- Move Quickstart to top with one-command install
- Collapse technical details into expandable sections
- Add 'Up Next' roadmap with concrete upcoming features
- Add star call-to-action banner
- Tighten navigation links and section hierarchy
* docs: add beta warning banner
* docs: add star request to beta banner
* docs: rewrite README with cognitive hooks, remove redundant CTAs
- Remove 2 premature star asks (beta banner + highlights)
- Rewrite highlights with loss-aversion framing
- Keep single earned CTA at the very bottom
- Use action-oriented headings that describe outcomes
* docs: rename section to 'Why OmniVoice Studio?'
* docs: rename 'What you get' to 'Features'
* docs: concise scannable features, remove duplicate section
- Each feature is one punchy emoji-led line
- No verbose paragraphs, no redundant collapsibles
- Removed duplicate Features section from merge
* docs: 3-column feature card grid for visual impact
Replaces flat bullet list with 4x3 HTML table grid.
Each feature gets its own visual cell with emoji header,
bold keywords, and 2-line description. Pops on dark mode.
* docs: fix feature grid vertical alignment
* chore: bump version to 0.2.7, add changelog entry
* fix: apply CodeRabbit auto-fixes
Fixed 1 file(s) based on 1 unresolved review comment.
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
---------
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
Replaces flat bullet list with 4x3 HTML table grid.
Each feature gets its own visual cell with emoji header,
bold keywords, and 2-line description. Pops on dark mode.
- Remove 2 premature star asks (beta banner + highlights)
- Rewrite highlights with loss-aversion framing
- Keep single earned CTA at the very bottom
- Use action-oriented headings that describe outcomes
- Add Highlights section with 2-column feature grid
- Move Quickstart to top with one-command install
- Collapse technical details into expandable sections
- Add 'Up Next' roadmap with concrete upcoming features
- Add star call-to-action banner
- Tighten navigation links and section hierarchy
- Use the high-quality social preview image as the hero image
- Bump download release links to v0.2.7
- Highlight the new Frameless Dictation Widget feature
WS dictation pipeline was producing exit-183 from ffmpeg on every
partial because MediaRecorder.start(250) ran before the WebSocket
handshake finished — the first chunk (WebM EBML header) was queued
only into chunksRef and never pushed to the WS, so concatenated
chunks 1..N decoded as malformed WebM. Fix:
- Construct the WebSocket BEFORE starting the recorder so wsRef is
set when the first ondataavailable fires.
- ondataavailable now queues every chunk through wsPendingRef when
the socket isn't OPEN; ws.onopen drains the queue.
- ws.onmessage('error'): fire HTTP fallback immediately instead of
waiting the full fallback-timeout window.
- ws.onclose without prior `final`: same — kick the HTTP path now
if the recorder has already stopped.
Mic permissions:
- New frontend/src-tauri/Info.plist with NSMicrophoneUsageDescription
+ NSCameraUsageDescription. Tauri 2 auto-merges the file at bundle
time (path is the same dir as tauri.conf.json — schema documents
this fallback). Without it, getUserMedia silently fails on macOS
10.14+ TCC.
- Mic-denial toast now includes platform-specific recovery (Settings
paths for macOS/Windows, audio-group check for Linux).
CI / release notes:
- release.yml extracts the matching `## [X.Y.Z]` section from
CHANGELOG.md and feeds it into tauri-action's releaseBody, so
v0.2.6+ tag pushes produce real release notes instead of the
placeholder "Auto-generated release. See commit log for changes."
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- LICENSE replaced with the canonical Functional Source License,
Version 1.1, ALv2 Future License (auto-converts to Apache 2.0 two
years after each release).
- Scope clarified: Studio (frontend + backend + tauri shell + scripts)
is FSL. Bundled `omnivoice/` Python TTS model package by Han Zhu
stays Apache-2.0 — not relicensed here.
- README license section + license badge updated to reflect FSL +
future-Apache; replaced "30-day free evaluation" copy with the FSL
Permitted Purposes wording.
- Enterprise page: drop hard-coded pricing tiers (Startup/Business/
Enterprise) since pricing is still being finalized. Replaced with a
"Pricing tiers coming soon — request a quote" panel. FAQ rewritten
around FSL semantics (internal use is permitted, source converts to
Apache 2.0 in 2yr). Drop now-unused TIERS const + TierCard
component + .ent-tier* CSS.
- CHANGELOG entry under 0.2.6 records the relicense.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Tray + lifecycle:
- tauri-plugin-single-instance — second launch focuses existing window
instead of racing for port 3900.
- Window close hides instead of destroying; backend shutdown moved to
RunEvent::ExitRequested so only the tray "Quit" item (or Cmd+Q on macOS)
actually exits.
- Tray icon flips to red-dot variant during dictation recording.
Hotkey customization:
- Settings → Capture tab. Records any modifier+key combo, persists to
app config, re-registers on launch.
- set_dictation_shortcut rolls back to the previous binding on register
failure so a bad combo never leaves the user with no shortcut.
Dictation latency / correctness:
- WS-final treated as source of truth; HTTP POST /transcribe runs only as
fallback (WS error / timeout / no-WS path). Audio transcribed once
instead of twice. Server accepts an "EOF" text frame (or empty binary
frame) so the socket stays open for `final` to be delivered before the
client closes.
- MediaRecorder chunks queued during the WS handshake are drained in
ws.onopen — the server's final transcript no longer drops the first
~250 ms of audio.
- Fallback timeout scales with recording length (max(15s, recordedMs+10s))
so long-form dictations don't trip duplicate transcription.
Donate page:
- Drop Patreon, Bitcoin / Ethereum / Solana cards. Drop qrcode.react.
- Move "Commercial License" CTA from page bottom to top-right header bar.
Docker hygiene:
- docker-compose binds 127.0.0.1 by default. README documents the LAN
exposure trade-off + recommends a reverse proxy with auth.
CI:
- New cross-platform `tauri-cross-platform` job runs `cargo check` against
the Tauri shell on macOS / Windows / Linux per PR. Catches platform
cfg-gate regressions without paying the full ~15min/platform bundle
cost (full bundling stays in release.yml on tag push).
Tests:
- tests/test_capture_ws.py (3 cases) covers EOF text-frame, empty-binary
EOF, and legacy disconnect-finalize paths.
Includes the user's previously-staged 0.2.5 polish: cross-platform
desktop-prod.sh, Dockerfile base-image fix, bun.lock churn.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Backend emits 'resolving' heartbeat every 2s during HF metadata
resolution so UI shows 'Resolving repo metadata...' instead of
being stuck on 'Connecting to HuggingFace…' indefinitely
- fmtBytes(0) now returns '0 B' instead of '—'
- desktop-prod.sh only tolerates signing errors, surfaces real
build failures with exit code
- Handle install_retry phase in frontend with attempt number
The --no-bundle flag caused the script to build only the raw binary
while launching the OLD stale .app bundle from a previous build.
Now builds the full bundle (tolerating the signing error which is
non-fatal) and deletes the old bundle first to prevent stale code.
shutil.disk_usage() throws on non-existent paths, causing the
preflight to report 0.0 GB free after a fresh wipe. Now resolves
up to the nearest existing ancestor directory so it probes the
actual volume free space correctly.
- 'Connecting to HuggingFace…' when no file events yet
- 'Resolving N files…' when tqdm init fired but total unknown
- Speed shows immediately from backend tqdm rate (no 2s warmup)
- '0 B / …' instead of '— / ?' for early progress
- 1s tick timer forces re-render so speed/ETA updates smoothly
- ETA shortened to ~3m instead of ~3m left for compactness