* chore: drop stray v0.4 references — everything ships on the v0.3.0 line
Per the project's versioning rule (no v0.4, no unprompted version chatter):
- backend/main.py + marketplace.py: the app reported version "0.4.0" (ahead of
even pyproject's 0.2.7 and referencing a forbidden version). Aligned to
"0.2.7" to match pyproject.toml / tauri.conf.json — a consistency fix, not a
bump.
- errorDocsMap.ts / indextts/bootstrap.py / _secret_key.py: reworded "v0.4"
deferral comments to version-agnostic "deferred / later hardening pass".
- docs/install/troubleshooting.md: the "tracked for v0.4" notarization line now
matches macos.md (signing is wired; activates on the Apple cert secrets).
Note: historical planning records under .planning/ still contain "defer to v0.4"
notes; left as-is (a record of superseded decisions) — CLAUDE.md + the
constitution are the live source of truth.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: set version to 0.3.0 across all sources (current dev line)
The current/upcoming version is v0.3.0 (0.2.7 is the prior stable). Bump every
version source so the codebase consistently reports 0.3.0 — the in-code dev
version; the git *tag* still happens later per the release cadence.
- pyproject.toml, frontend/src-tauri/Cargo.toml, tauri.conf.json,
frontend/package.json: 0.2.7 → 0.3.0
- backend/main.py (FastAPI) + marketplace.py export metadata → 0.3.0
(these had drifted to a phantom "0.4.0")
- CHANGELOG.md: "[0.2.7] — Unreleased" → "[0.3.0] — Unreleased"
- uv.lock + Cargo.lock reconciled (1-line each) so `--frozen` installs hold.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(version): read app version from package metadata (no more drift)
Greptile (#145): the FastAPI version + marketplace bundle metadata were bare
string literals — they'd go stale-wrong again at the next bump (the exact class
of bug this PR fixes; that's how "0.4.0" happened). Read once from
importlib.metadata.version("omnivoice") via core.version.APP_VERSION, with a
"0.3.0" fallback only for a non-installed source checkout. pyproject.toml is now
the single source of truth for the runtime version.
Tests: tests/test_app_version.py (semver + equals installed metadata). 2 pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Phase 3 Plan 03-01: Supertonic-3 engine on SubprocessBackend
Adds Supertonic-3 as a 7th opt-in TTS engine on the Phase 2
SubprocessBackend primitive. Closes TTS-01..06 (REQUIREMENTS.md):
* TTS-01 — _REGISTRY["supertonic3"] resolves to Supertonic3Backend,
a SubprocessBackend subclass.
* TTS-02 — `supertonic==1.3.1` lives under [project.optional-dependencies];
default `uv sync --no-dev` does NOT install it. Exactly one
`onnxruntime` row in `uv pip list` after `--extra supertonic`.
* TTS-03 — Model revision pinned by 40-char commit SHA
(724fb5abbf5502583fb520898d45929e62f02c0b — the "Initial
Supertonic 3 release" SHA, same as the SDK's own pin).
Resolver script for intentional bumps:
scripts/resolve_supertonic3_sha.py.
* TTS-04 — Honest CPU-only reporting. `is_available()` message says
"ready (CPU-only via onnxruntime)" and never mentions
"cuda" or "mps". `gpu_compat = ("cpu",)`.
* TTS-05 — License gate via settings_store helpers
(get/set_license_accepted) + Loopback-only
/api/settings/license endpoint + SupertonicLicenseDialog
frontend modal showing MIT (code) and OpenRAIL-M (model).
Wired into EngineCompatibilityMatrix as an "Accept license"
button on rows whose `reason` mentions "license not
accepted".
* TTS-06 — 3 langs (en/ja/ru) × 3 sec smoke test in
tests/test_supertonic3.py::test_smoke_3langs_3sec
(OMNIVOICE_SMOKE-gated; asserts no onnxruntime-gpu row
post-synthesize).
Package legitimacy gate (Task 1 in plan): supertonic on PyPI verified
to be published by Supertone Inc. (ato@supertone.ai), repo
github.com/supertone-inc/supertonic, wheel is pure-Python with no
postinstall scripts. Same publisher ships supertonic-js on npm under
the same maintainer email.
Test results:
* tests/test_supertonic3.py — 10 passed, 3 skipped (network-gated).
* tests/smoke/ — 4 passed.
* tests/ (full, --ignore=tests/manual) — 412 passed, 0 failed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(tests): uv sync --all-extras so optional-engine tests can import their package
Phase 3 added `supertonic` as an optional dependency. The CI Tests job
runs `uv sync` (no extras), so `test_cpu_only_honest` and `test_license_gate`
in tests/test_supertonic3.py hit the "supertonic package not installed"
fallback instead of the real import path, and fail.
Bare `uv sync` is the right default for users (engines are opt-in), but
the test environment should exercise the full surface. `--all-extras`
keeps the smoke job lean (still bare `uv sync`) while letting Tests
verify the integrated behavior of every optional engine.
Future-proofs against the same failure mode in Phase 4 (GGUF) and any
later optional engines.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(01-01): encrypted settings store + alembic migration (AUTH-02, T-01-01)
Adds the SQLite-backed encrypted settings store that Phase 1 token resolver
will read from. Closes the at-rest plaintext risk for HF tokens (T-01-01).
- backend/services/settings_store.py: get_hf_token / set_hf_token /
clear_hf_token using Fernet symmetric AEAD. Stored value column never
contains the literal "hf_" substring.
- backend/services/_secret_key.py: per-install Fernet key derived via
scrypt(machine-id + 16-byte random salt). machine-id resolution covers
macOS (ioreg IOPlatformUUID), Linux (/etc/machine-id and dbus fallback),
Windows (HKLM Cryptography MachineGuid via winreg). Final fallback to
hostname+user with a warn log.
- backend/migrations/versions/0001_phase1_settings_table.py: alembic
migration adding `settings(key, value, updated_at)`. Idempotent — checks
for an existing table so fresh installs (where _BASE_SCHEMA already
created it) and v0.2.7 upgrades both succeed.
- backend/core/db.py: _BASE_SCHEMA grows the settings table for fresh
installs; init_db() now runs `alembic upgrade head` after the CREATE.
- backend/migrations/env.py: honours an externally-set sqlalchemy.url so
tests can point alembic at a fixture DB; falls back to core.config
DB_PATH for production.
- pyproject.toml: cryptography>=41 added explicitly (RESEARCH.md
Assumption A1 was checked at execute-time and proved false; the dep was
not present transitively, so the install would fail without this).
Tests (10 cases, all green):
- Round-trip encryption + plaintext-leakage check (T-01-01 invariant)
- Salt persistence across clear/set cycles
- InvalidToken decrypt path returns None (Open Question #5 resolution)
- Concurrent reads consistent under sqlite WAL
- Alembic upgrade on a hand-built v0.2.7 fixture DB preserves all
existing tables + seeded rows (CLAUDE.md backward-compat constraint)
- Alembic downgrade -1 drops only the settings table
Refs #35.
* feat(01-01): 3-source HF token resolver + log redactor + 5 read sites patched
Closes the #35 bug class (bare os.environ.get('HF_TOKEN') reads) by routing
every backend HF-token consumer through one resolver, and mitigates
T-01-02 (info disclosure via logs) by stripping `hf_[A-Za-z0-9]{30,}`
substrings from every log record at the root logger.
backend/services/token_resolver.py:
- resolve(skip) — 3-source cascade (App → Env → HF-CLI), each source
validated via huggingface_hub.whoami(); first valid wins.
- on_401(active) — invalidate cache and re-resolve skipping the source
that just 401'd (AUTH-06).
- state() — three SourceState rows for the Settings UI: set,
masked preview (hf_…<last 3>), whoami_user, whoami_ok.
- save_app_token / clear_app_token — wraps settings_store + calls
huggingface_hub.login(add_to_git_credential=False) per Pitfall #2.
- 300-second whoami cache so repeated Settings-page renders don't hit
the HF API.
backend/core/logging_filter.py:
- HFTokenRedactor(logging.Filter) — regex `hf_[A-Za-z0-9]{30,}` so real
tokens are masked but `hf_hub` / `hf_token` literals survive.
- install_redaction_filter() — idempotent attach to root + every handler.
backend/main.py: install the redactor at startup, BEFORE the file
handler is added. Re-installed after the file handler attaches so the
handler-attached filter list includes it too.
Read-side call sites patched (per Pitfall #1 — every HF token read must
flow through token_resolver.resolve()):
- backend/api/routers/dub_core.py:540 (the original #35 site)
- backend/api/routers/system.py:38 (_has_hf_token notification)
- backend/services/model_manager.py:480 (diarization pipeline auth)
- backend/services/sonitranslate.py:143 (Popen env for SoniTranslate child)
- backend/services/sonitranslate.py:217 (gradio_client predict call)
New endpoint:
- GET /system/hf-token/state — returns the 3-source cascade state with
masked tokens for the Wave 2 Settings UI panel.
Grep gate confirmed clean: zero `os.environ.get("HF_TOKEN")` reads remain
outside token_resolver.py.
Tests (17 new cases, all green):
- tests/backend/services/test_token_resolver.py: priority cascade, 401
skip mid-resolve, on_401 fallback, state() shape, save+login
invariant (add_to_git_credential=False), HUGGING_FACE_HUB_TOKEN
alias acceptance.
- tests/backend/core/test_logging_filter.py: msg + args redaction,
multi-token redaction, non-string args pass-through, short-token
literals preserved, install_redaction_filter idempotence.
Refs #35.
* feat(01-01): Settings hf-token API endpoints + subprocess env injection (AUTH-03/04)
Backend half of the Wave 2 Settings → API Keys UI plus the AUTH-04
subprocess env-injection invariant.
backend/api/routers/settings.py:
- POST /api/settings/hf-token — body {token: str} → save_app_token
- DELETE /api/settings/hf-token — also_clear_hf_cli query → clear_app_token
- GET /api/settings/hf-token/state — same shape as token_resolver.state()
All three are gated by `Depends(require_loopback)` at the router level
(threat T-01-03 mitigation; non-loopback Host → 403).
backend/main.py: router mounted alongside existing API routers.
Subprocess env injection (AUTH-04, threat T-01-04 disposition=accept):
- backend/services/sonitranslate.py already updated in Task 2 to read
via token_resolver.resolve() and inject HF_TOKEN + YOUR_HF_TOKEN into
the SoniTranslate child env block.
- backend/services/gpu_sandbox.py: NOT patched — the GPU sandbox runs
in-process TTS generation that uses the parent's already-loaded HF
state. Adding env injection there is a no-op (parent and child share
state via multiprocessing.Pipe before any HF API call).
- backend/services/model_manager.py:480 (Task 2): resolves in-process,
no subprocess crosses here.
- backend/api/routers/exports.py: subprocess.Popen calls only spawn
`open` / `explorer` / `xdg-open` — file-manager launchers with no
HF needs. Skipped per Task 3 conservative-patching rule.
So the canonical AUTH-04 site for this milestone is sonitranslate.py.
Future SubprocessBackend work in Phase 2 will inherit the same pattern.
Tests (8 new cases, all green):
- tests/backend/test_engine_spawn_token.py
* POST /hf-token loopback → 200 + state.active == "app"
* POST /hf-token non-loopback → 403 ("loopback origin required")
* DELETE /hf-token clears settings_store + state.active == None
* GET /hf-token/state returns 3 source rows in priority order
* GET /hf-token/state non-loopback → 403
* env block contains HF_TOKEN + YOUR_HF_TOKEN when resolver returns one
* env block does NOT contain an injected empty HF_TOKEN when resolver
returns None
* source-level check that backend/services/sonitranslate.py still
reads via token_resolver.resolve() (regression guard against
silent reverts of the AUTH-04 wiring)
Full Wave 1 test suite: 35/35 green. Phase 0 smoke tests still green.
Refs #35.
* docs(01-01): SUMMARY + STATE update for Phase 1 Wave 1 completion
Records execution outcome of the 3-task plan: 10 files created, 9 modified,
35 new test cases, 5 read sites patched, grep gate clean. Documents the
two Rule-3/Rule-2 deviations applied (cryptography dep, env.py URL
override), the subprocess-launcher inventory for Phase 2, and the
known stray edit to the main repo's pyproject.toml that needs a one-
line user action to revert.
Updates STATE.md current-position table, progress bar, and open TODOs to
point at Wave 2 (Plan 01-02) and Wave 3 (Plan 01-03) as the next steps.
Production deployment hardening, dub OOM recovery, new SoniTranslate sidecar engine, ASR backend expansion.
- Dub generation OOM recovery: backend/api/routers/dub_generate.py:163-209 adds OOM detection + one retry with reduced nstep
- New SoniTranslate sidecar engine: backend/api/routers/sonitranslate.py + backend/services/sonitranslate.py (subprocess-based dubbing pipeline, opt-in)
- ASR backends expansion: backend/services/asr_backend.py adds NeMo Parakeet TDT, Moonshine, additional Whisper variants; new GET /system/asr-backends endpoint
- Dub UI polish: tighter spacing in DubSegmentRow.css, DubTab.css
Issue #78 (speaker diarization mis-assignment) NOT addressed by this PR — the bundled diarization changes are in the new SoniTranslate sidecar, not the existing pyannote pipeline. Keeping #78 open.
No DB schema changes, no migration. Backward-compatible for existing user data.
* fix: eliminate DB connection leaks, race conditions, and deprecated asyncio API
## DB Connection Leaks (P0)
- Convert 38 raw get_db() calls to db_conn() context manager across 14 router files
- Connections are now guaranteed to close even when exceptions are raised
- profiles.py create_profile: clean up orphaned audio file if DB insert fails
- profiles.py lock_profile: consolidate 3 separate conn.close() error paths
## Race Condition (P1)
- Add _dub_jobs_lock (threading.Lock) to protect _dub_jobs dict in dub_pipeline.py
- get_job/put_job now thread-safe for concurrent dub sessions
## asyncio Deprecation (P2)
- Replace 23 asyncio.get_event_loop() calls with asyncio.get_running_loop()
- Prevents DeprecationWarning on Python 3.12+ and future breakage on 3.14
## Quick Fixes
- gallery.py preview_voice: remove filesystem path from error response (P2)
- dub_pipeline.py parse_vtt_segments: remove redundant `import re` inside loop (P3)
- gallery.py _init_gallery_db: use db_conn() context manager (P2)
* refactor: extract hooks, centralize isTauri, add pytest-cov
## Frontend
- Extract useTTS hook (150 LOC) — TTS generation, streaming, audio ingestion
- Extract useProfiles hook (219 LOC) — voice profile CRUD, lock/unlock, preview
- Centralize isTauri detection: dialog.js, VoiceGallery.jsx, Settings.jsx
now import from utils/media.js instead of 4 different detection patterns
## Backend
- Add pytest-cov to dev dependencies
- Baseline coverage: 39% across backend/ (214 tests pass)
- Add .coverage to .gitignore
* feat: add Vitest + checkJs, extract useDubWorkflow + useAppData hooks
## Frontend Testing (new)
- Set up Vitest with jsdom environment + @testing-library/react
- 11 tests: utils (isTauri, formatTime, constants) + Zustand store (mode, text, dubStep, pill)
- Scripts: 'test' (vitest run), 'test:watch' (vitest), 'test:legacy' (node runner)
## App.jsx Decomposition (continued)
- Extract useDubWorkflow hook (387 LOC) — upload, ingest, transcribe SSE,
translate, generate SSE, abort, stop, cleanup
- Extract useAppData hook (181 LOC) — data loading, localStorage persistence,
WebSocket real-time updates, model-status pill management
## TypeScript checkJs
- Enable checkJs: true in tsconfig.json for IDE-level type checking
- 947 existing errors (informational, not blocking builds)
- noImplicitAny remains false to avoid blocking
* ci: add Vitest step, fix useProfiles duplicate state
## CI
- Add 'Run Vitest (frontend)' step — runs 11 unit tests
- Override --checkJs false in CI typecheck to avoid 947 pre-existing errors
- Rename legacy test step for clarity
## Hooks
- Fix useProfiles to accept loadProfiles from parent (useAppData)
instead of managing its own duplicate profiles array
* refactor: wire hooks into App.jsx — 2067 → 1129 LOC (-45%)
App.jsx now delegates to extracted hooks instead of inline logic:
- useAppData: data loading, localStorage, WebSocket, model pill
- useProfiles: voice profile CRUD, lock/unlock, preview
- useTTS: generation, streaming, audio ingestion
- useDubWorkflow: upload, transcribe SSE, translate, generate SSE
988 lines removed. All handler logic lives in focused,
independently testable hooks. Store selectors and render
JSX stay in App.jsx as the shell.
Verified: vite build clean, 11 frontend + 214 backend tests pass.
* feat: show real-time percentage on model loading pill
Backend: register hf_progress listener during _load_model_sync()
so download/weight-loading tqdm events update _loading_detail with
a progress percentage (0-99%). get_model_status() now includes a
'progress' field that the frontend polls.
Frontend: useAppData reads msQuery.data.progress and calls
setPillProgress() — the FloatingPill already renders the percentage
text and progress bar width from this value.
* fix: prevent FileNotFoundError in desktop bundle during model init
transformers >=4.52 calls _can_set_experts_implementation() and
_can_set_attn_implementation() during PreTrainedModel.__init__,
which open the class source file via open(class_file). In a Tauri
desktop bundle, module.__file__ points to a path that doesn't
exist on disk, causing:
FileNotFoundError: .../omnivoice/models/omnivoice.py
Override both classmethods on OmniVoice to return static values
without filesystem access. OmniVoice doesn't use MoE experts
(return False), but does support flex/flash attn (return True).
* fix: sync source dirs on every bootstrap, not just first run
The Tauri bootstrap previously only copied omnivoice/ and backend/
to Application Support on the first run. Subsequent app updates
kept using stale source files, preventing bug fixes from landing.
Now ensure_venv_ready() always syncs both directories from the
bundle resources before returning, even when the venv is healthy.
This fixes the FileNotFoundError crash where the old omnivoice.py
lacked the _can_set_experts_implementation override.
* ui: premium setup wizard polish
- Primary button: solid gradient fill with hover glow + lift + press
- Stepper nav: connected pills with glow ring on active step
- Welcome cards: glassmorphism with stagger-in animations, lucide icons,
left-border accent strip, hover translate
- Preflight panel: colored icon pill backgrounds, stagger-slide entrance
- Step transitions: fade+slide animation via keyed wrapper
- Footnote: shortened paths (~/ notation), Reveal in Finder button
- Recommendation banner: gradient background with accent glow
- Compact spacing throughout for denser, professional layout
* fix: kill zombie backend on clean+retry bootstrap
When clean_and_retry_bootstrap removes the project dir, any old
uvicorn process still running from the deleted paths remains alive
on port 3900. The subsequent retry_bootstrap sees the port is
healthy and attaches to the zombie instead of re-bootstrapping.
Now explicitly kill any process on the backend port after cleaning,
before calling retry_bootstrap.
* feat: integrate speaker clones into dubbing interface, sanitize system environment variables for subprocesses, and improve FFMPEG binary path resolution.
* fix: restore docker compose default + drop dead setSeed call
- deploy/docker-compose.yml: remove profiles: ["cpu"] from the default
service so `docker compose up` matches the comment on line 5. With the
profile present, no service auto-started.
- frontend/src/App.jsx: drop the setSeed call in restoreHistory. The
selector was never reintroduced after the App.jsx hooks split, and
there is no seed state in the store — seeds are generated fresh per
call in useTTS and only read from history items for display.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: address CodeRabbit review — async detection, dub stream, bootstrap fail-fast
- backend/services/tts_backend.py: invert async-context detection in
_ensure_loaded. The previous code unconditionally caught its own
diagnostic RuntimeError and then called asyncio.run() inside a
running loop, masking the intended error message.
- frontend/src/hooks/useDubWorkflow.js: require a terminal `done` event
before reporting dub success. Without this, a dropped stream after
partial progress would flip the UI to `done`, refresh history, and
play the completion ping as if generation finished.
- frontend/src/hooks/useDubWorkflow.js: restore the previous step when
tasksCancel() fails. The UI was getting stuck in `stopping` forever
on cancel errors.
- frontend/src-tauri/src/bootstrap.rs: fail-fast when source sync fails
after the existing directory has already been removed. The previous
warn-and-continue path could leave the install with no backend/ or
omnivoice/ sources and defer the failure to backend startup with a
cryptic error.
- backend/api/routers/generation.py: add `from e` to the ValueError →
HTTPException re-raise (Ruff B904).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: preserve % suffix in TTS generation timer
The 100ms timer in useTTS was rewriting generationTime to a plain
elapsed-seconds string, which immediately wiped the "(xx%)" download
suffix written on the next iteration of the response-body loop. The
real-time percentage was flickering on/off as a result.
Read the previous value inside the setter and reattach any existing
percent suffix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- New .github/workflows/docker.yml publishes images to ghcr.io on tag push
- README Docker section now leads with 'docker pull' from GHCR
- docker-compose.yml defaults to GHCR image with build-from-source fallback
- Dockerfile: copy README.md for hatchling metadata resolution
Three fixes for the Windows MSI first-launch failure:
1. **backend_log_path() was macOS-only** — used $HOME + Library/Logs
which doesn't exist on Windows. Now uses %LOCALAPPDATA% on Windows,
~/Library/Logs on macOS, and XDG_STATE_HOME on Linux. Without this,
stdout/stderr went to Stdio::null() and all backend crash output was
silently lost.
2. **Add TORCHDYNAMO_DISABLE=1 on Windows** — prevents PyTorch from
trying to download Triton (which has no Windows support), avoiding a
hang during first torch.compile() call (#26 workaround).
3. **Increase health timeout from 180s to 300s** — first-run PyTorch
import on Windows can take 120+ seconds for CUDA kernel JIT, plus
uv sync + torch + model loading. 3 min wasn't enough.
Fixes#30