fd7d20fe1e53d4e09b33fa59bb47760ba603b520
90
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
32103ad7b7 |
docs(readme): charm + organization overhaul (Opal-style) — collapsibles + OpenAI-compatible API section (#945)
* docs(readme): charm + organization overhaul (Opal-style) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): restore inventory-exact feature names (docs-drift guard) The charm pass sentence-cased five bold leads in the collapsed feature list; scripts/check-docs-drift.py greps for the inventory's exact title-case names. Restored: Vocal Isolation, Speaker Diarization, Batch Queue, AI Watermark, GPU Auto-Detect. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): dubbing screenshot shows a real completed dub (37 segs, EN→BN) Replaces the empty drop-zone shot with the populated editor — video + waveform + cast, 37 Bengali segment rows, DUB COMPLETE banner — captured live from the v0.3.9 app; caption updated to match. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0579ec91ad |
docs(readme): Opal-style restyle + fresh v0.3.9 screenshots (#937)
- Emoji section headers with explicit <a id> anchors. Emoji breaks GitHub's auto-generated heading slugs, so every in-page nav target keeps a stable explicit anchor (verified all href="#..." resolve). - Refresh the screenshot gallery. The prior set was from April, predating the launchpad / settings / dictation UI overhaul, so it misrepresented the app. Captured fresh at retina from the live v0.3.9 UI and led the gallery with the new Launchpad home: launchpad, studio, voice design, voice gallery, dubbing, engine-compatibility matrix, model store, embedded API reference (Scalar), and the in-app changelog reader. - Fix the stale engine count (11 -> 14 TTS engines) in the comparison table, FAQ, and roadmap to match the engine table + backend registry. - Use <kbd> keycaps for the dictation shortcut (Opal detail). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0d80ab2cb0 |
docs: backfill [0.3.9] batch bullets + add OSS sponsorship playbook (#931)
Bullets for #922 (release titles), #923+#924 (sponsors), #925 (contact), #927 (models), #928 (openapi), #930 (engines) — the agents kept off CHANGELOG.md during the merge chain. Plus a portable how-we-set-up- sponsorship playbook for reuse on other projects. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8f71c90f20 |
feat(settings): LLM Skills — per-feature enable/route control for every LLM call (#912)
New Settings → System → LLM Skills area: every LLM-powered capability
(Cinematic & Autofit translation, speech-rate slot fitting, glossary
auto-extract, direction parsing, dictation cleanup) becomes a "skill" the
user can toggle or route to a specific provider (local Ollama/LM Studio vs
a remote key) instead of everything riding the one global active provider.
Backend:
- services/llm_skills.py — skill registry + settings_store persistence
(llm_skill.<id>.enabled / .provider), resolution precedence
override > active > none, resolve_skill_client() (OpenAI-compat client
bound to the effective provider; None when disabled/unconfigured) and
skill_backend() (OffBackend when disabled — the exact no-LLM object every
caller already degrades on).
- All five consumption points wired through the registry; a disabled skill
degrades exactly like "no LLM configured" today (Fast translation
fallback, refinement pass-through, heuristic direction parse, no-llm slot
fit, 503 on glossary auto-extract). No new degradation modes; defaults
(enabled + no override) keep existing setups byte-identical.
- OpenAICompatBackend gains an optional bound provider (None = active, the
historical behavior).
- GET /api/settings/llm-skills + PUT /api/settings/llm-skills/{skill_id}
(404 unknown skill/provider); route snapshot updated.
Frontend:
- LLMSkillsPanel (Sparkles, next to LLM Providers): one row per skill —
i18n name/description, enable toggle, provider Select ("Use active
provider" + configured providers, local ones tagged), ready /
needs-setup badge linking to LLM Providers. All strings via t()
(settings.llmskills_*).
Tests: 30 backend (precedence, per-consumption-point disabled semantics,
endpoint round-trips, validation) + 4 panel render/PUT tests. Docs:
translation-engines.md gains an LLM Skills section.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
af6690840e |
fix(translate): run Cinematic/Autofit on every engine (incl. default Argos), bound the fit pass, scrub provider errors (#910)
P0 — Cinematic/Autofit silently no-op'd on argos/nllb/openai. Those three branches returned BEFORE _maybe_cinematic, so only the deep_translator fall-through reached the refine/fit pass. A user on the DEFAULT Argos engine who picked Cinematic/Autofit got plain Fast output with a success toast and no quality_used/cinematic_skipped/rate_ratio. All three now route through _maybe_cinematic. provider=openai is already an LLM translation, so it skips the reflect/adapt re-refine (new already_llm flag) but still stamps rate-ratio badges and runs the Autofit fit pass; the dialect it baked into its translate prompt is now reported applied. P1 — the Autofit fit pass ran one blocking adjust_for_slot per segment in the merge loop, OUTSIDE any budget (a 50-seg dub vs a slow provider spun ~50×timeout unbounded). New speech_rate.adjust_for_slot_many fans it out concurrently under a wall-clock deadline SHARED with the cinematic refine; segments still running at the deadline degrade to their literal with rate_error='fit-budget'. Also set max_retries=0 on the OpenAI clients used for translate/refine/fit so a 429 + Retry-After can't sleep through the budget. P2 — glossary auto-extract's no-LLM message now points at Settings → LLM Providers (was the stale TRANSLATE_BASE_URL/TRANSLATE_API_KEY). Provider error bodies on the glossary auto-extract, the OpenAI translate-segment path, and the DeepL/Microsoft translate-segment path are now scrubbed (core.scrub.scrub_provider_error) — they could echo the API key / a user_id. DubTab re-polls LLM availability on window focus / visibility so configuring a provider in Settings lifts the Cinematic gate without a remount. Documented LLM_DEFAULT_PROVIDER in docs/dubbing/translation-engines.md. Tests: fail-before/pass-after for argos+cinematic (refine runs), argos+cinematic no-LLM (cinematic_skipped), argos Fast (rate_ratio stamped), openai+autofit budget bound, and provider-error scrubbing on the translate + glossary paths. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
16294fed44 |
feat(updates): data-safe updates — pre-migration DB backups, guarded venv heal, release notes + changelog reader (#909)
Backend: - core/db_backup.py: WAL-safe SQLite snapshot to omnivoice.db.backup-<version>-<n> before pending alembic migrations run; keep newest 3, prune older; skip >500MB with a log line. Restore is never automatic. - core/db.py: _run_alembic_upgrade now plans the run (up_to_date / pending / unknown_revision), snapshots first when migrations will execute, and raises MigrationError on a mid-flight failure — startup stops with the backup path named instead of continuing on a half-migrated DB. The #552/#547 unknown-revision class stays non-fatal (warn + additive reconcile). - core/changelog.py + GET /api/settings/changelog: parse the shipped CHANGELOG.md (single-line and wrapped bullet styles) into structured releases. - GET /api/settings/db-backup: newest pre-migration backup for the panel. Rust (bootstrap.rs): - #314 heal guard: an exit-signature match alone can no longer delete the venv — venv_rebuild_justified requires a structural problem or a failed direct interpreter probe; a venv that probes healthy is kept and the real error surfaced. Drift/repair remains in-place `uv sync` (non-destructive). - CHANGELOG.md now ships as a bundle resource and is copied/refreshed into the project dir so the changelog endpoint works in packaged installs. Frontend (Settings → Updates): - Available update shows its actual release notes (updater metadata body) through a safe markdown-lite renderer (text nodes only, refs stay plain). - "Your data is backed up before every update" line with the latest backup timestamp from the new endpoint. - "What's new" changelog reader (accordion, newest expanded) over the shipped CHANGELOG.md; GitHub releases list reuses the same renderer. - One-time, non-blocking "What's new" footer pill after an update (persisted last-seen version; fresh installs baseline silently). - All strings via t() with en keys (other locales fall back to English). Tests: db backup/rotation/failure-path units, migration-safety units, changelog parser (both bullet styles + real CHANGELOG.md), endpoint tests, route inventory regenerated, Rust decision-logic + probe tests, vitest suites for renderer/viewer/panel/pill logic. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
72d137e1f3 |
fix(bootstrap): port cuDNN 8 (NVIDIA CUDA GPU) + VC++ redist to packaged installs (#869)
* fix(bootstrap): port cuDNN 8 (NVIDIA CUDA GPU) + VC++ redist install into ensure_venv_ready() * fix(bootstrap): address #869 review — drop dead VC++ half, cache negative CUDA probe, gate on ROCm, sync docs Per maintainer review on #869: 1. Drop the VC++ Redistributable half: LoadLibraryA("vcruntime140.dll") from the running Tauri exe is a tautology (the exe itself links the MSVC CRT, so the process wouldn't be running without it), and torch's real failure mode is msvcp140.dll inside the venv python process. Dead code removed; a comment records why for future readers. 2. Stop taxing every non-CUDA launch: a negative torch probe (CPU / Intel / AMD — most installs) is now cached in a .venv/.cudnn8_probe_negative marker, so the synchronous `import torch` runs at most once per venv lifetime. Invalidated on every path that can change the torch build (drift sync #307, repair sync, first-run sync, ROCm reinstall) and implicitly by a venv rebuild. A probe that fails to run cleanly is skipped WITHOUT caching so a transient error can't wedge a real CUDA machine. 3. Rewrite docs/install/troubleshooting.md §10 to the actual root cause: packaged installs never had the cudnn8_compat libs (so reinstalling never restored them); the bootstrap now installs them automatically on CUDA machines, with the manual uv pip command as the offline fallback and PyTorch Whisper as the sidestep. 4. Gate the ~700 MB nvidia-cudnn-cu12 download on the venv torch being a real CUDA build: the probe now reports 'hip' before checking cuda.is_available() (which HIP spoofs), so opt-in ROCm installs (#124) never fetch the CUDA wheel. Also reflow the CHANGELOG entry to house style (bold one-line lead, 1-3 lines of why, (#827, #869) refs) and extend the bootstrap unit tests: classify_cuda_probe verdict mapping and the marker write/invalidate round-trip (6 cuDNN tests total, 43 lib tests green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
83e71c5689 |
fix(asr): close the #730 residuals — chunked dub wedge shares the guarded reset; repeated timeouts recommend the crash-isolated engine (#895)
Residual A — the chunked dub-stream had a PARALLEL wedge mechanism (its own ping-loop timeout, its own _reset_pool_on_wedge, a dead-end "Try restarting the server" message). A wedged chunk now routes through the SAME run_transcribe_guarded bound+reset as the whole-file paths (#731/#851): the guard resets the poisoned pool once per wedged attempt (no double-reset on retry) and the user sees the actionable ASRTimeoutError. The reset logic is extracted to asr_backend.reset_pool_after_wedge — one shared mechanism, so the semantics can't drift again. run_transcribe_guarded also gains a timeout_env param so chunk errors name OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S instead of the whole-file knob. Residual B — the crash-isolated ASR sidecar (#393, faster-whisper-isolated) is wired as an explicit ESCAPE HATCH, not a default: - selectable end-to-end: Settings engine list gets an explanatory install_hint; honest gpu_compat ("cuda","cpu" — it wraps the same CTranslate2 engine as faster-whisper); get_active_asr_backend now hands back a process-wide singleton for subprocess-isolated backends (a fresh instance per request would leak atexit hooks and respawn the sidecar — reloading its model — on every transcribe). - on the SECOND consecutive guarded timeout-with-reset in one session (resets aren't recovering the hang; the wedged thread keeps its VRAM), the error the user sees + the log recommend switching to the isolated engine in Settings → Engines. Never auto-switched (owner rule: no silent behavior divergence); a completed transcribe resets the streak. Tests (fail-before/pass-after verified against origin/main): wedged-chunk SSE integration (reset count + actionable error + recommendation surfaces), consecutive-timeout streak (fires at 2, resets on success, suppressed when already on the isolated engine), timeout_env parametrization, shared-reset helper, isolated backend in list_backends with hint + honest availability, singleton caching, gpu_compat matrix entry. Docs: troubleshooting §14 gains the chunk knob + escape-hatch guidance. Closes the residuals tracked on #730. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
86f5213055 |
fix(splash): IPC-independent watchdog + recovery panel for dead Tauri IPC (#879) (#892)
After an unclean shutdown (Windows BSOD), the WebView2 profile cache (%LOCALAPPDATA%\com.debpalash.omnivoice-studio\EBWebView) can corrupt: Tauri's IPC custom protocol fails AND the postMessage fallback breaks, so invoke() hangs forever. useBootstrapStage's poll loop rode entirely on that IPC — a hung bootstrap_status call silently killed the loop and the splash sat at "preparing" forever, even with a fully healthy backend answering over plain HTTP. Class fix, three parts: - splashWatchdog.js: IPC-independent escape hatch. If no IPC signal arrives within 10s, poll GET /health over plain HTTP; healthy → proceed to the app as if 'ready' was received (console.warn breadcrumb so diagnostic bundles carry it). Any successful IPC response disarms it for good. - Recovery panel (stage 'ipc_lost'): if neither IPC nor HTTP succeed within 45s, show an actionable panel instead of the infinite spinner — "Open logs" (with an inline path fallback when IPC is dead) and, Windows-only and only in this error state, "Repair and restart". Health polling continues behind the panel so a slow first-run install with broken IPC still reaches the app. - clear_webview_cache_and_relaunch (Rust): writes a marker and relaunches; the fresh process deletes EBWebView at the top of run() before any webview exists (WebView2 holds locks while running), with a bounded retry while the old instance exits. Runtime cfg! guards keep the whole path compiling on every platform. Tauri 2 exposes no reliable flag for the postMessage-fallback mode (closure-local in its injected ipc.js), so the logged detector is the observable combination: zero IPC signals + working plain HTTP. Fail-before/pass-after regression tests: hung invoke + healthy HTTP → ready; hung invoke + dead backend → recovery panel, then auto-continue; working IPC → normal path untouched, zero HTTP polling. Plus watchdog state-machine unit tests and recovery-panel render/interaction tests (6/7 fail on the pre-fix component). Troubleshooting doc gains the matching section (docs-sync). Fixes #879 Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
be1ec3ade0 |
fix(platform): declare Intel-Mac local backend unsupported — honest first-run gate + docs (#889); Windows portable-install docs (#766 follow-up) (#891)
torch >=2.3 ships no macOS x86_64 wheels (transformers 5.x needs torch >=2.6), so `uv sync` can never resolve on an Intel Mac — per the platform-parity rule the honest option is declaring the platform unsupported, not letting first launch die in a raw resolver error: - bootstrap.rs: pre-check on macOS x86_64 before any venv create / uv sync (first-run AND repair paths) fails fast with an actionable message (remote-backend escape hatch + docs link); healthy pre-torch-bump venvs are deliberately untouched. Unit test pins the message's load-bearing phrases. - BootstrapSplash: routes the failure to a dedicated localized hint (bootstrap.hint_intel_mac, all 21 locales) and suppresses the useless Retry-oriented hints for it. - README + docs/install/macos.md (+ troubleshooting #9): every Intel-Mac support claim now says UI-installs-but-backend-cannot-run, including the from-source path (also broken); remote backend documented as the only use. - release.yml: #889 note on the macos-15-intel leg — artifact is UI-only; keep-or-drop is an owner call, deliberately not changed here. - docs/install/windows.md: new "Portable install (Windows)" section promised in #766 — custom MSI wizard folder / msiexec INSTALLDIR=..., what lives in OmniVoiceStudio-Data next to the exe, and the Program-Files-greyed-out why. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4eed552153 |
fix(engines): Confucius4-TTS validated E2E — clone sys.path import, 22.05 kHz, real install docs (#590) (#872)
Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3d0705fdb7 |
feat(engines): Confucius4-TTS — finalized (API-validated + unit-tested; opt-in, GPU run pending) (#590) (#637)
* feat(engines): Confucius4-TTS scaffold (opt-in, needs hardware validation) (#590) Plumbing for netease-youdao's Confucius4-TTS — LLM-based 14-language cross-lingual zero-shot voice cloning, Apache-2.0 — mirroring the opt-in subprocess-venv pattern of dots.tts / MOSS-TTS-v1.5: - engines/confucius4/__init__.py: Confucius4Backend(SubprocessBackend), CUDA-only (gpu_compat=("cuda",)), language passthrough, ref_audio→prompt_wav. is_available reports a clear reason and stays unavailable without a clone. - bootstrap.py: dedicated Python 3.10 venv resolution (user clone-level venv → package venv → uv bootstrap), import-probed on `confuciustts`. - main.py: sidecar speaking the same length-prefixed JSON-over-stdio protocol as the other engines, calling ConfuciusTTS(config_path, device).generate(text, lang, prompt_wav). - Registered lazily in _LAZY_REGISTRY; docs/engines/confucius4-tts.md. Gated behind OMNIVOICE_CONFUCIUS4_TTS_DIR — inert on every default install, never imports the upstream package unless opted in. The sidecar's synthesis API is derived from the upstream README and is NOT yet validated on a CUDA box; the module, docs, and CHANGELOG all flag this. 4 tests pin registration + inert-by-default. No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#590): register Confucius4 in install-hints + docs inventory (CI gates) Registering the engine tripped two completeness gates: every backend needs an install_hint (test_issue_fixes) and every registry engine must appear in the tts_engines docs inventory + README (check-docs-drift). Add the install_hint, the docs/features.yaml entry, and the README engine-table row (with the scaffold caveat). Docs-drift clean; gates pass. No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(confucius4): finalize — validate API vs upstream, add 22 sidecar unit tests, document external deps (Amphion/w2v-bert/weights) The synthesis API (ConfuciusTTS(config_path, device) → generate(text, lang, prompt_wav) → tensor, model.sample_rate) is confirmed against the netease-youdao/Confucius4-TTS repo. Added runnable unit tests for the sidecar's pure logic (language norm, tensor→PCM mono/stereo/clip, config resolution, wire framing, synthesize dispatch with the model mocked) — 22 cases, all green. Docs now list the external deps (Amphion/MaskGCT codec, facebook/w2v-bert-2.0, ~2-4GB HF checkpoint) and CUDA 12.6. Softened the scaffold warnings to reflect API-validated + unit-tested status; a one-time CUDA GPU run is still needed to confirm live inference + true sample rate. --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
29269b9cf0 |
feat: LLM Providers page + Autofit translation quality (fit-to-segment-time) (#838) (#854)
* feat(llm): multi-provider LLM registry + encrypted key storage + settings API (v0.3.8, phase 1)
Foundation for the LLM Providers settings page and timing-aware (Autofit)
translation. Every provider in the shipped .env is OpenAI-compatible, so one
client drives all of them via a registry instead of a class-per-provider.
- llm_providers.py: registry of 16 providers (OpenAI, OpenRouter, Groq,
Cerebras, Google AI, Mistral, Cohere, NVIDIA, GitHub Models, Cloudflare,
HuggingFace, SambaNova, SiliconFlow, + local Ollama/LM Studio + Custom).
Field resolution precedence env → encrypted store → default; active-provider
selection (LLM_DEFAULT_PROVIDER → stored → first keyed remote; local requires
explicit pick so we never assume a local server is up). Legacy TRANSLATE_*
maps to the Custom provider (keyless-with-base_url preserved).
- settings_store.py: generic ENCRYPTED secrets (get/set/clear_secret,
list_secret_names) reusing the HF-token Fernet path; get_text/set_text now
refuse the secret namespace (no ciphertext leak).
- llm_backend.py: OpenAICompatBackend resolves the active provider's
base_url/key/model from the registry. Backward-compatible.
- settings API: GET /llm-providers, PUT /llm-providers/{id} (encrypted key +
overrides), POST /llm-providers/active, POST /llm-providers/{id}/test.
Loopback-gated; never returns key material.
- 11 registry tests; existing llm-endpoint/openai-available tests still green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(settings): LLM Providers page — configure any provider's key/URL/model + Test + set active (v0.3.8, phase 2)
New Settings → System → LLM Providers pane (Brain icon, searchable). Lists all
16 registry providers; pick one to configure its encrypted API key, base URL,
model (and Cloudflare account id), Test the connection with one round-trip, and
'Save & use for translation' to make it the active provider for Cinematic/
Autofit. Keys are write-only from the UI (masked placeholder, never echoed);
env-set keys show as read-only. Local providers (Ollama/LM Studio) need no key.
- LLMProvidersPanel.jsx: provider selector + per-provider config + Test/activate,
following the LLMEndpointPanel pattern (apiJson/apiFetch/apiPost, SettingsSection
primitives).
- settingsCategories.jsx: new 'llm-providers' category under System + Brain icon.
- Settings.jsx: route the category to the panel.
- en.json: settings.llm_providers label.
- Frontend build passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(translate): Autofit quality style + one-click LLM setup from the dub menu (v0.3.8, phases 3-4)
Autofit = Cinematic + a strict 'never exceed the segment time' fit. The LLM
rewrites each translated line so its target-language reading time fits within
the slot, preserving the video timing without harsh audio time-stretch.
Backend:
- speech_rate.adjust_for_slot(strict=): strict caps the accepted upper ratio at
1.0 (fit within slot) vs Cinematic's 1.08; best-effort, degrades gracefully
with no LLM.
- dub_translate: quality='autofit' takes the LLM refine path and runs the fit
pass with strict=True; reports quality_used accurately.
- TranslateRequest.quality doc note.
Frontend:
- 'autofit' added to the quality control (Settings Translation + dub menu) and
the TranslateQuality type.
- Dub menu: picking Cinematic/Autofit with no LLM no longer dead-ends on a toast
— it offers a one-click 'Set up' that routes to Settings → LLM Providers,
with copy about fitting translations to segment time (#838).
- 4 strict-fit tests; frontend build green; i18n keys added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(translate): document Autofit quality + the LLM Providers page (v0.3.8, phase 5)
- CHANGELOG [0.3.8] Added: Autofit style + LLM Providers page.
- docs/dubbing/translation-engines.md: Fast/Autofit/Cinematic quality section
and an LLM Providers setup section (16 providers, encrypted keys, offline
Ollama/LM Studio, env overrides).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(api): add /api/settings/llm-providers routes to the route-inventory snapshot
Regenerated tests/fixtures/api_routes.txt for the 4 new LLM-providers endpoints
so test_route_inventory_matches_snapshot passes (keep-main-green).
* style(frontend): oxfmt the LLM Providers panel + dub quality control (format:check green)
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
e347f99542 |
fix(tts): bound + reset the GPU pool on a hung generate so it can't brick the backend (#730 class) (#851)
* fix(tts): bound + reset the GPU pool on a hung generate so it can't brick the backend (#730 class) A GPU job that wedges on some Windows+CUDA setups occupies its worker forever — run_in_executor can't cancel the thread — so on the 1–2 worker pools we ship, one stuck job starves every other request and the next action surfaces as the misleading "Can't reach the local backend" even though the process is alive. ASR/dub/model-load already bound+reset the pool on hang (#730). The TTS **generate** paths (generation.py, tts_stream.py) were the last unguarded GPU dispatch — and the residual on-main reports (#850 #802 #755 #723 #721, plus the 0.3.7 generate cohort) all fail on generate:start (audio). - model_manager: add run_on_gpu_pool_guarded() + GpuJobTimeoutError, a generalized version of the ASR guard so every GPU dispatch shares one bound+reset recovery path. Env-tunable via OMNIVOICE_GENERATE_TIMEOUT_S (default 300s). - generation.py: route both inference branches + the reference-clip transcribe through the guard; map a timeout to an actionable 503. - tts_stream.py: same guard on the streaming path (timeout → error frame). - test_generate_timeout_730: fail-before/pass-after regression (timeout resets pool + restores capacity, happy path, env override, no-reset exec). - docs + CHANGELOG: extend troubleshooting §14 to cover generate; document the new env var. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(tts): extend the GPU-pool hang guard to batch/dub/archetype/openai-compat generate (#730 class) The generate-hang class wasn't only in Studio + streaming: batch generate, the dub per-segment + preview generate, archetype preview render, and the OpenAI-compat /v1/audio/speech path all dispatched the TTS model to the GPU pool with no wall-clock bound either. Any one of them wedging on a Windows+CUDA hang starves the pool and bricks the backend the same way. Route all of them through run_on_gpu_pool_guarded so the whole class is closed — a hung generate anywhere resets the pool and returns an actionable timeout instead of a dead backend. Batch/dub recover per-segment on a fresh worker; drop the now-dead loop/_gpu_pool/asyncio locals ruff flagged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
522bbddccf |
feat(translate): highlighted Install affordance for uninstalled engines + dismissable/auto-clearing error banner (#847)
Two related Dub-tab translation-flow fixes, one PR. TASK 1 — proactive, highlighted Install affordance in the translate engine selector (replaces "find out only via a translate-time 400"): - FROM-SOURCE lane (activeEngineUnavailable && !enginesSandboxed): the muted install chip is promoted to a HIGHLIGHTED brand-accent Install button, still wired to handleInstallEngine(translateProvider) with the installing/disabled state. Selecting any uninstalled engine surfaces it immediately. - FROZEN lane (enginesSandboxed): pip install is impossible in the read-only, signed packaged env, so the disabled "needs dev install" span becomes an equally highlighted button opening a popover with (1) the exact install command + copy-to-clipboard, (2) one-click "Switch to Argos (bundled, offline)" — the guaranteed importable escape hatch, and (3) a Docs link via the existing Tauri shell.open path. Gated on the existing `sandboxed` flag, not platform. - Single-source install command: new translation_engines.install_command() is the one source of truth; list_engines() stamps `install_command` per engine and BOTH the argos + deep_translator translate-time 400 messages build their command from it, so the proactive button and the 400 can't drift. engines.ts gains `install_command: string | null`. TASK 2 — the translation error banner now dismisses and clears (class fix): - Root cause: handleTranslateAll never cleared dubError, so a stale 400 survived even a successful retry. It now clears at the start of every attempt. - Corrective-action clears (whole class): changing the engine and installing the package both clear dubError (wrapped setTranslateProvider + handleInstallEngine in DubTab). - DubFooter's banner gains a × dismiss and a guarded auto-timeout (skipped while generating so live per-segment errors persist). i18n: 8 new dub.* keys translated across all 21 locales. Docs: new docs/dubbing/translation-engines.md (from-source vs packaged build) linked from the popover Docs button + a troubleshooting cross-reference. Tests: FE regression for both lanes + never-installs-when-sandboxed + banner dismiss/auto-clear; BE regression that list_engines() install_command is embedded verbatim in the dub_translate 400s. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cb70c2b1af |
feat(ui): back Input/Select/Textarea/Slider with shadcn (prop APIs preserved) (#798)
P1 of the shadcn/ui primitive migration (docs/shadcn-migration.md): route the
OmniVoice form/data primitives through the shadcn components in
src/components/ui/* while keeping their exact exports and prop APIs, so no call
site changes.
- input.tsx: export `inputBaseClass` (the shell) with no behaviour change —
ShadcnInput baseline stays byte-identical.
- New shadcn components: textarea.tsx, select.tsx (+@radix-ui/react-select),
slider.tsx, table.tsx.
- src/ui/Input.jsx (Input/Textarea/Select/Field): Input/Textarea now render the
shadcn components; a small `fieldSizeVariants` cva (named palette utilities,
tailwind-merge-clean) restores the OmniVoice padding-based sm/md/lg scale +
filled bg-bg-elev-2 over the shell. Select stays a NATIVE <select> wearing the
same shell — DubSegmentTable/CompareModal/GeneralTab depend on
onChange={(e) => …e.target.value}, which Radix's value-only Select would break;
the Radix select.tsx is added for new call sites only.
- src/ui/Slider.jsx: wraps the shadcn Slider, keeping the number-based onChange +
label/value-bubble chrome; track/thumb sized via the data-slot selectors.
- Table deliberately NOT rerouted: ui/Table.jsx is a flex-<div> chrome wrapper
whose .ui-table*/.segment-table global classes are a SHARED CONTRACT used
directly by ModelsTable/DubSegmentTable/EngineCompatibilityMatrix (virtualised
lists needing the div/flex layout, not a semantic <table>). table.tsx is
provided for new tabular data; Table.jsx + its globals are untouched. Its
toolbar inherits the shadcn-backed Input/Button for free.
Verification: only the 3 Input-* visual baselines moved (palette-coherent across
default/midnight/catppuccin); Slider/Table stayed within tolerance. vitest 641
green, oxlint 0 errors, oxfmt --check clean, vite build green, root bun.lock
regenerated and bun install --frozen-lockfile in sync (Docker).
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
200a559183 |
feat(ui): shadcn/ui foundation + OmniVoice palette token bridge (Button/Input proof) (#797)
Lay the foundation for migrating OmniVoice's UI to clean Tailwind v4 + shadcn/ui WITHOUT changing the look: shadcn primitives inherit the existing OmniVoice palette (Gruvbox-pink default + every [data-theme] variant) through a semantic token bridge. Foundation only — no existing component is replaced. What landed: - shadcn init for Tailwind v4 + Vite + React 19: frontend/components.json (new-york, rsc:false, tsx:true), src/lib/utils.ts (cn = clsx + tailwind-merge), and a @/* -> src/* alias in vite.config.js + tsconfig.json so future `npx shadcn add` resolves. - Token bridge in src/index.css: a single `@theme inline` block maps shadcn's semantic vocab (--color-background/-foreground/-card/-popover/-primary/ -secondary/-muted/-muted-foreground/-accent-foreground/-destructive/-input/ -ring + --radius) onto the existing OmniVoice --color-* tokens. Because those tokens are re-declared per theme in ui/themes.css, theme switching recolors shadcn components automatically — no per-theme shadcn block. Existing --color-accent/--color-border and the --radius-* scale are left intact. - Two proof components: src/components/ui/button.tsx + input.tsx (verbatim shadcn new-york), rendered across default/midnight/catppuccin in the visual harness with committed baselines (brand-pink / purple / lavender confirmed). - New deps: class-variance-authority, clsx, tailwind-merge, tw-animate-css, @radix-ui/react-slot. Root bun.lock regenerated; `bun install --frozen-lockfile` verified in sync (Docker-green). - Migration plan at docs/shadcn-migration.md (bridge table, primitive->shadcn mapping, prop-compat wrapper strategy, staged waves, honest risk/effort). Verified: vite build, typecheck:ci, oxlint (0 errors), oxfmt --check, vitest (641 pass), test:visual (48 pass incl. 6 new baselines), frozen lockfile in sync. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
721cb34a9d |
docs: add the CSS → Tailwind v4 migration plan (#772)
Phased, bounded migration plan (not a big-bang): convert the mechanical ~80% (flex/grid/gap/padding/typography/simple color) to Tailwind v4 utilities, deliberately keep ~15-25% as CSS (glass/backdrop-filter, @keyframes, ::before/::after, :has(), !important). Realistic end state ~10-12k of 16.6k CSS lines removed across ~5-7 weeks of small PRs. Key gates the plan establishes before any conversion starts (P0): - A Playwright screenshot baseline (default + dark + light) — the className-diff trick used for the page refactors is useless here since class names change. - Fix the @theme ↔ tokens.css token drift (single source + a parity test). - Rewrite the CONTRIBUTING.md "no Tailwind" line (docs-sync rule). Companion to docs/maintenance-pages-modularization.md. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f33bdc731d |
refactor(settings): modularize Settings page (1969→399 lines, all files under 500) (#758)
* refactor(settings): extract Settings.jsx tabs into components/settings (1969→602 lines) Settings.jsx had grown to 1969 lines — every edit reloaded the whole file into context and risked unrelated breakage. This finishes the migration the existing components/settings/*Panel.jsx pattern started: the page is now a thin orchestrator and each heavy tab lives in its own file. Extracted (logic byte-for-byte identical; only import paths adjusted + the shared isTauri/askConfirm moved to components/settings/native.js): - GeneralTab, ModelStoreTab, EnginesTab, HotkeyTab, CredentialsTab - native.js — shared isTauri() wrapper + askConfirm() Tauri-dialog helper Also establishes the standard so files can't silently regrow: - CONTRIBUTING.md: frontend file-structure & size limits (soft 300 / hard 500) - eslint.config.js: warn-only max-lines:500 guardrail (CI stays green) - docs/maintenance-pages-modularization.md: the phased refactor plan Verified: vite build passes (all imports resolve); 18/18 settings tests pass; no new lint errors introduced (the pruned imports were the only regressions). Follow-ups (tracked in the plan doc): ModelStoreTab.jsx is 836 lines and Settings.jsx 602 — both still over the 500 cap (warn-only); split next. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(settings): split ModelStoreTab + Settings.jsx under the 500-line cap Follow-up to the tab extraction: bring the two remaining over-cap files into compliance with the new standard. Pure-mechanical, no behavior change. Settings.jsx 602 → 399: - Extract AboutTab, PrivacyTab, LogsTab into components/settings/ - Move the shared Row helper to components/settings/Row.jsx - LogsTab keeps its state in Settings() (lower-risk); About/Privacy take props ModelStoreTab.jsx 836 → 439, split into components/settings/models/: - format.js (fmtBytes/orgColor), runtime.js (computeRowRuntime) - columns.jsx exposes makeModelColumns(...) — a factory so the TanStack cell closures keep working; called with the same useMemo dep array as before - ModelsTable.jsx (virtualized table view), RecoBanner.jsx Every settings file is now under 500 lines. Verified: vite build passes; 18/18 settings tests pass; no new lint errors (the 4 remaining in Settings.jsx are pre-existing — refreshInfo no-op, a catch(e), two set-state-in-effect). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
de3d83f14b |
docs: add rust as prerequisite for from-source builds (#704)
Adds Rust/Cargo as a from-source build prerequisite across the linux/macos/windows install docs. Thanks @Deepakv2104. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
25105605f1 |
docs(specs): ElevenLabs-parity roadmap + Tier-1 implementation specs (#684)
Add the implementation-ready spec set mapping OmniVoice to ElevenLabs parity while preserving local-first: - 00-roadmap-elevenlabs-parity.md — gap analysis, prioritized tiers, sequencing, prior-art reconciliation, and the deliberate "won't build" list. - 01-expressive-tts.md — engine-agnostic emotion/style intent lowered onto each TTS engine's real mechanism (degrade-visibly) + a DB-backed pronunciation dict. - 02-conversational-agent.md — fully-offline full-duplex voice agent (/ws/converse, Silero-VAD barge-in on AEC-cleaned mic) composing existing streaming STT/TTS + LLM. - 03-longform-studio-editor.md — per-segment edit/regenerate across dub/audiobook/ stories, extending the existing content-addressed cache to longform. Reconcile prior planning docs: banner the superseded parity/studio docs pointing here; keep distinct-scope docs untouched (classification table in 00). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
022a3bd6b9 |
feat(dictation): live local dictation via sherpa-onnx + Voice settings panel (#683)
* feat(dictation): live local dictation via sherpa-onnx + Voice settings panel Add a sherpa-onnx ASR engine alongside the existing Whisper/NeMo dictation path, powering a genuinely live experience: as you speak, words type straight into the focused field (streaming partials via a new simulate_type command, self-correcting with backspaces) and commit per pause. Backend: - SherpaDictationBackend + sherpa_dictation registry of the 7 models (Parakeet TDT v3/v2, streaming Zipformer EN/ZH/bilingual, Paraformer bilingual, Whisper Tiny) from csukuangfj/* int8 HF repos; CPU provider for cross-platform parity. - /dictation/models + /dictation/prefs router; get_capture_asr_backend() honors the selected dictation model. get_active_asr_backend() (dub transcription) and the legacy WebM/Opus capture path are untouched. - True streaming over /ws/transcribe (OnlineRecognizer: live partials + per-endpoint finals); offline models surface partials via short re-decode. Frontend: - New "Voice" settings panel (enable, Toggle/Hold mode, model picker with offline/streaming/recommended badges + per-model download/delete). - Live word-by-word typing via simulate_type (enigo) with prefix-diff delta and backspace correction; paste fallback retained, no double-insertion. Deps: sherpa-onnx>=1.13.3 (+ sherpa-onnx-core); uv.lock regenerated, Docker frozen-install verified. API route-inventory snapshot updated. 40+ new tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dictation): register sherpa-onnx-asr engine in README + features inventory Fixes the docs-drift CI guard: the new sherpa-onnx-asr ASR engine existed in the registry but not in docs/features.yaml or README. Adds the live-dictation engine row to the ASR Engines table, bumps the engine counts (8→9), and adds the inventory entry. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): fold live-dictation into the [0.3.8] section main is 0.3.8 (untagged), so the dictation feature belongs in that release, not a separate [Unreleased] block. Merge the two Added lists under one [0.3.8], refresh the headline to lead with live dictation, and correct the capture description to reflect live word-by-word typing (not paste-on-pause). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2339cc8e85 |
fix(model): actionable "reinstall transformers" hint on a corrupted-install model-load error (#676)
A model load failed with `[Errno 2] No such file or directory:
'…/site-packages/transformers/models/qwen3/modeling_qwen3.py'` — the user's
transformers install was incomplete (the file is missing while a correct 5.3.0
install has it; an interrupted `uv sync` / antivirus / partial update drops it).
The System Check showed the raw path + "Check logs and try restarting", which is
useless — restarting can't restore a missing file.
Two fixes:
1. core.failure.classify(): recognize this corrupted-install variant. It's a
FileNotFoundError, not an ImportError, so the existing TRANSFORMERS_IMPORT
match ("could not import module"/"AutoFeatureExtractor") missed it. Now also
matches a "no such file"/"errno 2" + "transformers" + "site-packages" signal
(substrings checked separately so it works on POSIX `/` and Windows `\`
paths). An unrelated package's missing file is NOT mislabelled.
2. model_manager._load(): build the /model/status error via build_failure so it
carries the classified hint AND strips the home dir, instead of storing the
raw str(exc). The System Check now shows "Your transformers install is
incomplete. Reinstall it (uv pip install --reinstall transformers) or switch
ASR to faster-whisper" — the existing TRANSFORMERS_IMPORT hint.
Docs: troubleshooting §1a documents the error + the reinstall fix.
Test: test_failure_classify.py pins the POSIX + Windows path forms classify as
TRANSFORMERS_IMPORT with a "reinstall" hint, and that an unrelated package's
missing file does not.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
b7cecde57e |
feat(setup): faster downloads by default + prominent, encouraged HF-token entry (#669)
Two changes that make first-run downloads faster and easier to speed up further.
1. Segmented (multi-connection) downloader is now ON by default. The app forces
the legacy-LFS path (HF_HUB_DISABLE_XET=1) for clear progress, but that path
is single-stream and slow — which is why downloads felt sluggish. The built-in
IDM/uGet-style segmented accelerator (parallel byte-ranges, live speed/ETA)
was already implemented but defaulted OFF. Flip it ON: it only engages when
Xet is inactive (the default), and ANY failure falls back to snapshot_download
("can never compromise a correct install"). Pure-httpx, cross-platform,
auth-safe (token never forwarded to a CDN). Override with
OMNIVOICE_SEGMENTED_DOWNLOAD=0.
2. The Hugging Face token field is now a prominent, always-visible card right
above Continue — was a collapsed "advanced" fold almost nobody opened. A free
token gives authenticated downloads (higher rate limits, fewer stalls), so it
pairs with change #1 to keep the parallel fetch from getting throttled. The
card leads with the speed benefit, shows a saved-state, and adds a one-click
"Get one free →" link to huggingface.co/settings/tokens.
Docs: downloading-models.md updated — the legacy-LFS section now documents the
default-on segmented accelerator + the HF-token speed tip, and the tuning table
reflects OMNIVOICE_SEGMENTED_DOWNLOAD=0 as the disable knob (docs-sync).
Test: test_segmented_download_default.py pins the new default ON and that the
env override still disables it; existing FDL-08 behavior tests stay green.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
252f0d4fac |
fix(asr): bound whole-file transcription so a stall isn't reported as "can't reach backend" (#656)
A Windows/CUDA user (Vietnam) hit "Can't reach the local backend" only when dubbing/transcribing. Their log proves the backend started fine — model loaded, preload complete, 25 models — and the log ends right after `whisperx transcribing …tmp.wav`. The backend was alive; the *transcription* stalled (large-v3 ASR contending with the resident TTS model for VRAM on an 8 GB-class GPU), which the UI surfaces as an unreachable backend. Root cause (class, not instance): the chunked dub pipeline already bounds each chunk (OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S), but the *whole-file* transcribe paths ran unbounded: - dub QC re-transcribe (dub_export) - dictation (capture) - OpenAI-compat /audio/transcriptions A slow/stuck transcribe on any of these hung the request AND held a GPU-pool worker — indistinguishable from a dead backend. Fix: add run_transcribe_guarded() in services/asr_backend.py — a shared asyncio.wait_for wrapper (ASRTimeoutError, a TimeoutError subclass) with a generous env-tunable bound (OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S, default 300 s). On timeout the request returns 504 with actionable guidance (backend is alive; free VRAM / pick a smaller ASR model / use CPU; restart to clear the stuck worker) instead of hanging forever. Wired into all three whole-file paths. Docs: new troubleshooting §14 — "Can't reach the local backend during transcription/dubbing" — explains it's ASR weight/VRAM pressure, not a network/ mirror problem, and corrects the misconception that a "Network → Restricted/Global mirror" Settings toggle exists (the Network control is LAN sharing). Serves the #602/#585/#567 "can't reach backend" cluster. Test: backend/tests/test_asr_transcribe_timeout.py — slow fn raises ASRTimeoutError with the actionable message, fast fn passes through, subclass-of-TimeoutError so the openai_compat broad catch still maps to 504. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
79f3e35682 |
docs(troubleshooting): add stuck-download / incomplete-cache recovery (#622) (#643)
The 'stuck on the download page, model folder has only refs/ no weights' case (a connection dropping/blocking mid-pull) is a recurring support report but wasn't in the install troubleshooting guide. Add section 13 with the recovery steps + antivirus/VPN/mirror escalation + a huggingface-cli manual fallback. Docs-only; no version bump. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0e17caa52a |
fix(install): actionable torch-wheel-download failure + local-wheel recovery (#569) (#574)
#569: on a restricted network the first-run install fails downloading the ~2.5 GB cu128 PyTorch wheel from download.pytorch.org, and the app won't launch. Two problems: the error told users to "set UV_DEFAULT_INDEX to a mirror" — which CANNOT redirect torch, because it comes from a *named, explicit* uv index (uv 0.11 rejects index-name override values and `--frozen` pins the exact wheel URLs); and there was no way to supply a manually-downloaded wheel. - Detect a torch/pytorch-host `uv sync` failure and emit torch-specific guidance (Clean & Retry → VPN → drop the wheel locally) instead of the wrong mirror advice. - Add a local wheel-drop dir `<env_root>/wheels` (survives Clean & Retry) wired via `UV_FIND_LINKS`. On a frozen-sync torch-download failure WITH wheels present, retry NON-frozen with find-links so uv re-resolves from the local wheels. Verified empirically: a non-frozen find-links sync installs from a local wheel fully offline, while a `--frozen` sync ignores find-links — so the retry is the only mechanism that can consume a dropped wheel. Best-effort: if it can't satisfy, it fails identically to before and the actionable error still fires. - docs/install/troubleshooting.md: new "#12 CUDA PyTorch wheel download fails" entry (docs-sync) — the offline wheel-drop path + why a PyPI mirror can't fix this index. Note: an automatic mirror redirect for the cu128 index is intentionally NOT shipped — uv provides no working override for a named explicit index, so it couldn't be verified; the offline wheel path is the reliable escape hatch. Test: sync_failure_is_torch_download host/keyword detection + negative guard. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3777d3a62c |
feat(tts): add MOSS-TTS-v1.5 (8B) and dots.tts (2B) as opt-in engines (#498)
Adds two zero-shot voice-cloning TTS engines requested in #498, both opt-in and subprocess-isolated with their own dedicated venv — the same pattern as IndexTTS-2. The dedicated venv is forced, not just chosen: each upstream pins a transformers version that conflicts with the parent's >=5.3 (MOSS-TTS-v1.5 ==5.0.0, dots.tts ==4.57.0), so they cannot share the parent interpreter. Because they use the clone+venv bootstrap (env var -> clone -> uv venv), this touches no pyproject.toml / uv.lock / bun.lock — `uv sync --all-extras` and Docker's `bun install --frozen-lockfile` are unchanged, so main's CI/Docker matrix stays green. Engines: - moss-tts-v15: 8B, 31 langs, ~16 GB weights, 24 kHz. AutoModel/ AutoProcessor via trust_remote_code. gpu_compat=(cuda,cpu) — MPS is undocumented/untested upstream so it is never claimed; on a Mac it runs on CPU. Apache-2.0, no license gate. - dots-tts: 2B, 24 langs, ~9 GB weights, 48 kHz. DotsTtsRuntime; continuation cloning (prompt_audio_path+prompt_text). Upstream is Linux/macOS-only, so is_available() gates it off cleanly on Windows (cross-platform parity rule — it is opt-in, never a broken default). Wiring: registered in _LAZY_REGISTRY + _INSTALL_HINTS. list_backends() surfaces both as subprocess/[cuda,cpu]/available-until-installed; the data-driven Settings engine picker needs no frontend change. Tests (19, fail-before/pass-after): registry resolution, subprocess marker, no-MPS gpu_compat, the Windows gate, not-installed honesty, and the parent-side generate() kwarg arbitration. Existing engine suite still 55 passed / 5 skipped. Sidecar inference follows the upstream-documented APIs but, like IndexTTS/Supertonic, can't be executed in CI without the multi-GB model clones. Docs (same-PR per docs-sync rule): README + README_CN engine tables, new docs/engines/moss-tts-v15.md + dots-tts.md, disk-usage.md (torch-dedup note), CHANGELOG. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e58106d552 |
fix(updater): preview channel builds nightly from main + prerelease/parity guards (#500)
The Preview update channel was effectively dead: its only build trigger was a manual workflow_dispatch, so "preview = main" was never enforced — the live preview manifest was stuck at 0.3.5-41 (June 7) while main moved to 0.3.7. It also shipped two latent hazards: the `preview` GitHub release had drifted to isPrerelease=false (a non-prerelease `preview` is eligible to become GitHub's "Latest" — the exact URL the *stable* updater reads, so it could hijack the Stable channel), and its updater manifest dropped darwin-x86_64 (Intel-Mac preview users silently got no updates — a cross-platform-parity breach). Changes (release.yml): - Add a nightly `schedule` (07:00 UTC) that rebuilds the rolling `preview` prerelease from main. A new `preview-gate` job no-ops the 4-platform matrix on nights when main didn't move, so idle days cost only a ~30s gate job. - Centralize the preview-vs-stable decision in `preview-gate.outputs.is_preview` (schedule OR workflow_dispatch+publish_preview), consumed by the stamp step, tauri-action, and preview-notes — replacing the repeated inline conditions. - Harden the prerelease flag: preview-notes' `gh release edit` now re-asserts `--prerelease` every run, and a new post-publish step fails the run if the preview release isn't a prerelease or its manifest is missing any platform stable ships (catches the Intel-Mac regression in CI). Docs (docs-sync): update docs/update-channels.md — previews are no longer "manual / no scheduled spend"; they build nightly from main (+ on demand). The live `preview` release was re-flagged prerelease out-of-band to close the hazard immediately; this makes it recurrence-proof. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
af3da58584 |
fix(bootstrap): force-reinstall setuptools so pkg_resources repair actually works (#248) (#495)
The auto-repair ran `uv pip install setuptools>=75,<80`, which `uv` treats as "already satisfied" (no-op, "Checked 1 package in 5ms") whenever setuptools' *metadata* is present but its `pkg_resources` files are gone — the common cause being Windows Defender quarantining `pkg_resources/`, or a partial extract on a restricted network. So the repair never restored the files, the post-check failed, and users hit the #248 dead-end. The error message *also* told them to run the same no-op command, so the suggested manual fix didn't work either (reported on Discord, Win11 + RTX 5070 Ti). Fix: both repair sites in bootstrap.rs now use `--reinstall` (the flag already used for the ROCm torch repair), which force re-extracts pkg_resources even when uv thinks setuptools is satisfied. The fail() message and the failure.py hint now suggest `uv pip install --reinstall 'setuptools>=75,<80'` + an antivirus-exclusion note, and docs/install/troubleshooting.md (#pkg_resources-missing) is updated with the real cause (metadata-present/files-missing) + AV guidance. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3656f0a4ef |
fix(profiles): decouple design-profile save from TTS render (#476) (#488)
* fix(profiles): decouple design-profile save from TTS render (#476) Saving a design voice profile forced a full TTS model load + inference to render a deterministic identity sample. On a fresh model-less image (Docker first-run) that 503'd, so the save failed. A secondary guard also rejected an all-Auto design (empty instruct) with a 422. Saving a design profile is now a pure persistence operation: - The seed-42 identity sample render is attempted opportunistically but is non-fatal — if the engine isn't ready the row is persisted with ref_audio_path=NULL (sample pending). The row's vd_states + instruct already make the voice fully usable (generation.py falls back to instruct-only conditioning for design profiles with no ref audio). - The sample is rendered lazily + cached on the first GET /profiles/{id}/audio request; if the engine is still unavailable that path returns a precise "model not ready — finish setup / download a model" 503. - The all-Auto (empty-instruct) design is now saveable (vd_states still required). Adds tests/test_profile_design_save_decouple.py (top-level tests/, asyncio.run per test) covering: design save with model unavailable creates the row instead of 503-ing; all-Auto design is saveable; the pending sample materializes on first /audio request. Updates the unification spec (docs-sync). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): contain profile-audio paths under VOICES_DIR (CodeQL CWE-22) The lazy design-sample path was built as `os.path.join(VOICES_DIR, f"{profile_id}.wav")` / `os.path.join(VOICES_DIR, audio_file)` where profile_id is the request path param — CodeQL flagged 5 high-severity path-injection alerts (profiles.py + the taint flowing into archetypes.py's torchaudio save). Add `_safe_voice_path()` (basename + safe-char sanitise + realpath containment, mirroring core.config.dub_seg_path) and route both the read and lazy-render sites through it; a traversal id now 404s instead of escaping VOICES_DIR. Regression test covers the containment guard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): use CodeQL-recognized path-injection guards (CWE-22) The previous `_safe_voice_path()` helper was correct (basename + realpath containment) but CodeQL's taint tracking didn't propagate the barrier through the function return, so the 5 path-injection alerts persisted. Switch to guards CodeQL recognizes, inline at each file-op site: - validate `profile_id` against the generated-id charset (`[A-Za-z0-9_-]{1,64}`) with `re.fullmatch` and 404 on mismatch (covers the `f"{profile_id}.wav"` render path); - read only `os.path.join(VOICES_DIR, os.path.basename(name))` so a stored/derived filename is always a direct child of VOICES_DIR (covers the read + the taint flowing into archetypes.py's torchaudio save). Drop the helper. Test now asserts a traversal/separator/NUL profile_id 404s at the guard. Same security property, recognized by CodeQL. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): inline realpath+commonpath containment for CodeQL (CWE-22) CodeQL didn't recognize the earlier sanitizers — neither the helper (barrier hidden behind a function return) nor os.path.basename / a cross-function regex guard cleared the 5 path-injection alerts. Use the canonical, CodeQL-recognized form INLINE at each file-op site: resolve the path with os.path.realpath (which collapses any `..`) and confirm os.path.commonpath((base, path)) == base before the read / the render, returning 404 / raising on escape. Same property the helper had, now in a shape CodeQL's taint tracking follows. Keeps the profile_id charset guard as defense-in-depth. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): route design-sample path through shared _voices_path guard (#476) The inline realpath+commonpath containment in get_profile_audio and _materialize_design_sample wasn't recognized by CodeQL as a path-injection sanitizer (5 new high-severity py/path-injection alerts at the file-op sites, incl. archetypes.py mkdir via the rendered Path). Both now reuse the existing _voices_path() helper, which applies the os.path.basename() barrier plus symlink-resolved containment — the same guard the consent endpoint uses and that CodeQL already accepts. Behavior is unchanged: the DB columns only ever hold bare {profile_id}.wav filenames, so basename() is a no-op here. Tests: tests/test_profile_design_save_decouple, test_profile_unification, test_profile_consent, test_archetype_blank_guard — 25 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d0e1c19e88 |
docs(persona): document the .ovsvoice portable format (#29 slice D) (#464)
- docs/persona-format.md: export (privacy/include-reference, watermarked preview), import (consent/verification non-forgeability rule), the ZIP layout table, SPDX-license semantics (metadata only), and the local-first / zero- network guarantee. Notes legacy .omnivoice compatibility. - CHANGELOG.md: [Unreleased] → Added entry for portable personas. Satisfies the docs-sync hard rule for the new bundle format. |
||
|
|
e0b59f3984 |
docs(longform): implementation specs for the 14 roadmap tasks (#21–#34) (#429)
* docs(longform): implementation specs for the 14 roadmap/integration tasks (#21–#34) Per-task implementation specs under docs/specs/longform/ for the remaining longform + #346-roadmap work: GPU compat matrix, shared VoiceSelector, Transcriptions import, Story⇄Audiobook export, inline Create Voice, gallery handoff, parser unification, two-pass ACX, .ovsvoice format, Dub→Stories, unified LongformProject store, phone calls, cue-sheet, runtime-verify. Authored by a draft + iterative-refinement workflow (codebase-grounded: exact file:line anchors, API/data shapes, test plans, constraints, deps, risk, PR slices). NOTE: the 10-round refinement was cut to ~rounds 4–5 by an account session limit; rounds 5–10 (incl. the final de-bloat/polish pass) are pending — the specs carry per-round revision-note preambles that the polish round trims. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(longform): strip accreted (this-revision) note preambles from specs Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
4cc55ab852 |
Fast model downloads: Xet fast path + accurate progress (FDL W0–W2 + W4) (#424)
* feat(downloads): Xet fast path + accurate progress (FDL W0–W2)
Make model downloads fast and show accurate downloaded/remaining/speed.
Research confirmed hf-xet already implements the IDM/uGet technique
(content-defined chunking, parallel byte-range gets, dedup, resume), and
the spike found all 25 catalog repos are Xet-backed — so the win is
driving Xet well + accurate progress, not a custom downloader.
W1 — maximize + guarantee Xet:
- pin huggingface_hub>=1.7 + hf-xet>=1.1 (was transitive); no hf_transfer
- drive snapshot_download with explicit tqdm_class + max_workers + endpoint
- opt-in HF_XET_HIGH_PERFORMANCE / HDD sequential-write knobs (default off)
- /system/info reports fast_download {xet_enabled, xet_version, high_perf}
W2 — accurate progress:
- dry_run preflight -> install_plan event (exact total/cached/remaining)
- utils/download_aggregator.py: one overall bar; byte bars (by id) vs the
"Fetching N files" count bar; windowed rate; emits one 'aggregate' event
- frontend overall bar (speed/remaining/ETA), cached-skip, ⚡ fast badge
Known limit (verified live): under Xet+hf_hub 1.7.2 per-file byte bars
never advance/close via tqdm, so mid-download the bar is file-granular and
bytes flush to the exact total on completion. Classic-LFS/mirror repos get
true byte progress (W4).
Drive-by: download.py used os.walk without importing os (latent NameError
in _validate_snapshot_has_weights on every install) — fixed.
Tests: tests/backend/setup/test_download_preflight.py (10). Spike + plan
under .planning/quick/260613-fdl-fast-model-downloads/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in mirror + cancel + docs (FDL W4)
- mirror (FDL-10): snapshot_download(endpoint=) honours prefs hf_endpoint /
env HF_ENDPOINT on preflight + download (per-call, no process-wide env).
Documented as the classic-LFS path (no Xet) for restricted networks.
- cancel (FDL-11): POST /models/install/cancel {repo_id} stops further
retries at the next boundary, emits install_cancelled, clears the cooldown
(cancel is intent, not failure). Frontend treats it as a terminator.
- docs (FDL-12): docs/downloading-models.md (Xet fast path, progress
semantics + byte-speed limitation, opt-in tuning, mirror, cancel,
troubleshooting) + README pointer. Docs-sync rule satisfied.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(planning): model-management v2 cleanup plan (mm2)
GSD plan for cleaning the model-management subsystem: registry unload-on-
switch + per-engine unload() (fixes VRAM leak), model_lifecycle facade,
unified idle/timeout config, bounded cooldowns, sidecar VRAM self-report,
cache-fallback logging. Planning artifact only — no code.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(downloads): reconcile with main's HF_HUB_DISABLE_XET; honest status
Rebasing onto main surfaced that main forces HF_HUB_DISABLE_XET=1 (classic
LFS) because Xet progress bypasses the tqdm hook — the same limitation found
here. Reconcile instead of fight:
- /system/info fast_download now reports runtime truth: xet_installed +
xet_active (installed AND not HF_HUB_DISABLE_XET) + xet_enabled alias. The
⚡ badge only shows when Xet actually runs; startup log says
"downloads: Xet disabled → legacy LFS".
- complete(): clear the rate window before the final flush so crediting the
full size in one step can't emit an absurd instantaneous rate.
- docs/downloading-models.md rewritten: default is legacy LFS for accurate
progress; Xet is opt-in via HF_HUB_DISABLE_XET=0. hf-xet pin stays (ready
for a future Xet progress hook).
W2 (preflight total/remaining + aggregate bar + exact completion) is the
value on either path; W1's "maximize Xet" is dormant by main's design.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in segmented multi-connection accelerator (FDL W3)
Since main forces Xet off (HF_HUB_DISABLE_XET=1), the default path is
single-stream legacy LFS — so a segmented downloader is the way to get BOTH
parallel speed and live byte progress.
- services/segmented_download.py: async multi-connection Range downloader for
one file — parallel byte-ranges, resume (.part + manifest), per-segment
short-read truncation guard, optional sha256/etag verify, cancel, and a
single-stream fallback when the server won't range. Auth-safe: the HF
Authorization header is sent only to huggingface.co/hf.co and never
forwarded to a CDN host on redirect (unit-tested).
- dispatch (download.py): opt-in via prefs segmented_downloader / env
OMNIVOICE_SEGMENTED_DOWNLOAD (default off). When on and Xet inactive,
fetches each file into the HF cache mirroring hf_hub_download (blobs +
snapshot symlinks + refs/main), feeding real bytes to the aggregator. Any
failure falls back to snapshot_download — never breaks a correct install.
- fix: complete() was adding a full total on top of accumulated segmented
bytes (2x); now replaces byte bars so the sum is exactly total.
Verified live (accelerator on): real byte progress to ~16.6 MB/s, final
bytes==total, /models installed=True, delete frees correctly.
Tests: test_segmented_download.py (7) + aggregator double-count regression.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(downloads): relocate FDL tests to top-level; loop-isolate segmented test
CI runs the full suite, which exposed a pre-existing test-isolation leak:
several tests/backend/** fixtures purge core.*/services.* from sys.modules
under a temp OMNIVOICE_DATA_DIR and never restore, leaving core.config/core.db
bound to a dead temp dir. It only bites when collection order puts a purging
test ahead of a real-DB reader (test_longform_jobs). Adding tests under
tests/backend/setup/ reordered collection and tripped it.
Fix without touching the shared (fragile) fixtures or risking class-identity
breakage from a blanket sys.modules restore:
- move the two FDL test files to top-level tests/ (tests/test_fdl_*.py) so
tests/backend/** collection order is identical to main — longform passes.
- rewrite the segmented test to run each case under asyncio.run() (fresh loop)
instead of asyncio.get_event_loop(), which an earlier async test can leave
closed in the full suite.
Full suite green locally: 1364 passed, 0 failed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
e9481ef307 |
feat(longform): shared render core — loudness, metadata, cover art (PR 1/8) (#408)
First slice of the Stories+Audiobook convergence (spec: docs/specs/2026-06-13-stories-audiobook-maturity.md). Both features will compile to one server-side chapterized renderer; this lands the shared pure builders and wires them behind Audiobook. New `backend/services/longform_render.py` (all pure, unit-tested without ffmpeg/torch): - build_ffmetadata(chapters, global_meta) — FFMETADATA1 with an optional global tag block (title/author→artist/narrator→composer/year→date/genre/description→ comment) + chapter table. - build_loudnorm_filter(preset) — `-af loudnorm` for ACX (~-19 LUFS, -3 dBTP) or podcast (-16 LUFS); off/unknown → None. Opt-in, so default behavior stays platform-identical. - validate_cover_image — jpg/png + 8 MB cap guard. - build_render_cmd — generalizes the m4b mux: m4b|mp3, optional cover (attached_pic) + loudness, bitrate validated. - build_concat_list — moved here. `services/audiobook.py`: build_chapter_ffmetadata / build_m4b_cmd / build_concat_list are now backward-compatible wrappers over the core (existing imports + tests unchanged). `POST /audiobook`: now accepts optional `format` (m4b|mp3), `loudness`, `cover_path`, and `metadata` and passes them through — backend-complete; the UI for these lands in PR 2. Tests: tests/test_longform_render.py (28) + existing test_audiobook.py (11) green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6704d062fc |
fix(persona): preserve design kind + vd_states across share/import (Wave 5 §R3) (#405)
The persona-gallery surface already exists (VoiceGallery Community zone + community.py manifest + marketplace .omnivoice bundles). The blocker for §R3's 'synthetic-only' gate was data integrity: a *designed* persona lost its kind='design' (and vd_states) when imported from the community gallery or round-tripped through a bundle — silently demoting it to a clone. - community.py /use: a 'preset' (rendered from instruct) imports as kind='design'; a 'voice' (real reference clip) as 'clone'. - marketplace.py: extract a pure _bundle_metadata() (dedupes export+publish) that captures kind + vd_states; import restores them. Old bundles without the keys import as 'clone' (backward-compatible). This makes 'accept only designed/synthetic voices' enforceable instead of everything defaulting to clone. No new persona-gallery feature was built — that would duplicate the existing community/marketplace surface. 4 torch-free tests (isolated DB): _bundle_metadata captures design + defaults to clone; import round-trip preserves design kind+vd_states; legacy bundle → clone. docs §R3 status updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
151f73f794 |
feat(audiobook): Audiobook tab — script → plan → m4b (Wave 5 UI) (#404)
Frontend for the audiobook backend (#402/#403): a dedicated Audiobook tab. - pages/AudiobookTab.jsx: script textarea + default-voice picker (reuses the app's profiles), 'Preview plan' (POST /audiobook/plan → chapter list) and 'Create' (POST /audiobook → reads the SSE stream, shows per-chapter progress + assembling, then an <audio> player + m4b download via the /audio mount). - api/audiobook.ts: typed plan() + generate() (returns the raw streaming Response). - utils/sseParse.js: pure splitSSEBuffer/parseSSELine helpers for reading the POST event-stream (EventSource is GET-only) — unit-tested (the buffer/line handling is the easy thing to get subtly wrong). - NavRail + App.jsx wiring (lazy tab, hideSidebar); i18n keys in en.json. All strings via i18n (CJK gate green). 7 new SSE tests; full vitest 326 + vite build green. Runtime-unverifiable here (Tauri webview) — wants an in-app pass. docs §R3 updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9441274ab6 |
feat(audiobook): synth job → chapterized m4b, SSE progress (Wave 5) (#403)
Completes the audiobook backend: POST /audiobook renders each chapter through the active TTS engine (synthesize_chapter + chunked_tts), writes per-chapter WAVs, then muxes a chapterized m4b (FFMETADATA1 chapters via build_m4b_cmd + concat demuxer). Progress streams as SSE (started/chapter/assembling/done/ error), recorded to job_store. ffmpeg-gated — emits an error event and stops when ffmpeg is absent (m4b is the only output). - services/audiobook.build_concat_list: pure ffmpeg concat-list builder with proper single-quote escaping (no arg injection). Unit-tested. - router: voice resolution (compact form of generation.py's locked/design/ clone cases) cached per id; OmniVoice native model path + generic TTSBackend path; chapter synthesis runs on the GPU pool, ffmpeg via run_ffmpeg. Reuses the tested building blocks from #402 (parser, synthesize_chapter, FFMETADATA + m4b argv builders) — the new router glue is thin and import-checked by CI. Deferred: epub/pdf ingest, ACX loudnorm mastering, crash-resume, UI. 15 audiobook tests (added concat-list); docs §R3 updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
34b47282af |
feat(audiobook): chapterized audiobook core + plan preview (Wave 5) (#402)
* feat(audiobook): chapterized audiobook core + plan preview (Wave 5) First cut of the long-form vertical (parity §R3). Engine-agnostic core in services/audiobook.py: - parse_audiobook_script: pure parser. Markdown '# H1' headings → chapters; inline [voice:NAME] switches the narrator; [pause …] is delegated to the shared omnivoice.utils.text.parse_pause_markers so audiobooks and single-shot synthesis keep one pause dialect. Returns a chapter/span plan. - synthesize_chapter: orchestration via an injected synth(text, voice) callable (reuses chunked_tts split + crossfade, stitches inter-span silence) — so it's unit-testable with a stub backend, no model/GPU. - build_chapter_ffmetadata + build_m4b_cmd: pure FFMETADATA1 [CHAPTER] builder and faststart-m4b concat-demux argv (bitrate-validated, no injection). POST /audiobook/plan returns the parsed plan (no TTS/ffmpeg, no side effects). Deferred (follow-ups): the streaming synth job + chapterized-m4b run, epub/pdf ingest (new dep), ACX loudnorm mastering, crash-resume, UI. 14 tests: parser (chapters/voice/pause/intro/empties/to_dict), FFMETADATA offsets+escaping, m4b argv + bitrate guard, and stub-synth orchestration (span+silence stitching, voice threading). docs §R3 status updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): linear-time regexes (CodeQL ReDoS) CodeQL flagged polynomial backtracking on user-provided input in three regexes reachable from the new POST /audiobook/plan endpoint: - _VOICE_RE: \s*(...)\s* → single [^\]]* class, stripped in code. - _HEADING_RE: trailing [ \t]* removed; title captured greedily + stripped. - _PAUSE_RE (omnivoice/utils/text.py): the numeric spec is now an atomic group (?>…) so its leading \s+ can't backtrack against the trailing \s*. Behavior-preserving (Python >=3.11 already required); 14 pause tests + 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): require non-space heading title start (CodeQL ReDoS) The previous _HEADING_RE '[ \t]+(.+)' still let the leading whitespace class and the title '.+' both match the same tab run (overlap → polynomial). Anchor the title capture with \S so the two can't overlap. 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): exclude '[' from voice-tag content (CodeQL ReDoS) [^\]]* still matched '[', so a run of nested [voice: prefixes produced overlapping finditer match attempts → O(n^2). Excluding both brackets ([^\]\[]) makes matches non-overlapping and linear. A voice name never contains a bracket. 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
34c8ab2409 |
feat(engines): idle-reap subprocess-engine sidecars to free VRAM (Wave 13) (#401)
Parity Action 13 (dynamic load/unload), subprocess-engine half. A subprocess engine's sidecar holds a process — and, for GPU engines, VRAM — for the life of the backend, even after the user switches engines. The default in-process OmniVoice model already idle-unloads (model_manager.idle_worker); this gives the subprocess engine class the same treatment. subprocess_backend gains a background reaper (lazy daemon thread, started on first spawn) that shuts down sidecars idle past OMNIVOICE_SIDECAR_IDLE_TIMEOUT_S (default 300 s; <= 0 disables). The next request transparently respawns one via the existing dead-process relaunch. Safety: the reaper only acts while holding the per-backend lock acquired NON-blockingly, so it can never run mid-op — if an op holds the lock it skips that backend this round. Reuses the idempotent shutdown() (which doesn't take the lock, so no re-entrancy). Each backend tracks last-use and registers in a weak live-set. Scope: subprocess engines only (the heavy, VRAM-holding, process-isolated class). In-process non-default engines and cross-engine VRAM preemption remain TODO — get_active_tts_backend returns a fresh instance per call, so those need an instance-tracking refactor. 6 reaper tests via the stdlib echo sidecar (no torch): kills idle, respawns, skips busy (lock held), recent-use kept, disabled at <=0, ignores dead. The 3 subprocess suites pass together (24). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7d7d07c8fc |
feat(dictation): wire AEC end-to-end in the frontend (Wave 8, opt-in) (#400)
Completes Action 8: dictate-over-playback echo cancellation now works
end-to-end, gated behind a new off-by-default 'aecEnabled' pref so the
standard dictation + playback paths are untouched when off.
- utils/aec/{pcm,farEndBus,micCapture,playbackTap}.js + public/aec-worklet.js:
AudioWorklet captures the mic as raw int16 PCM; a player tap routes playback
output through Web Audio to a singleton far-end bus. Pure framing/encode
helpers are unit-tested.
- CaptureWidget: when aecEnabled, opens /ws/transcribe?aec=1, streams tagged
PCM (0x00 mic / 0x01 far-end) instead of MediaRecorder/WebM. Default path
unchanged; no POST fallback in AEC mode (the WS is the sole channel).
- WaveformPlayer: while actually playing AND aecEnabled, taps its decoded
output as the echo reference. Gated on isPlaying so only the one active
player holds an AudioContext (well under the browser cap); audio stays
audible (source always reconnected to destination).
- Settings → Capture: AecPanel toggle. prefsSlice: aecEnabled (persisted).
Runtime-unverifiable here (jsdom has no Web Audio); needs in-app testing in
the Tauri shell. 7 new pure-helper tests; full vitest (319) + vite build green.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
e8705a106d |
feat(dictation): opt-in NLMS AEC for dictate-over-playback (Wave 8b) (#399)
* feat(dictation): opt-in NLMS AEC for dictate-over-playback (Wave 8b) Dictating while OmniVoice plays audio (TTS preview, dub, video) leaks the loudspeaker signal into the mic, and the streaming ASR transcribes that bleed. Browser echoCancellation varies per platform/webview — it can't be a cross-platform default — so this adds a server-side canceller that behaves identically everywhere. services/aec.py ports Patter's NlmsEchoCanceller (MIT): a time-domain NLMS adaptive filter with a Geigel double-talk detector, warm-up step ramp, and far-end staleness pass-through. /ws/transcribe gains an opt-in '?aec=1[&sr=]' mode: frames are raw int16 mono PCM tagged with a 1-byte prefix (0x00 mic, 0x01 playback reference); the mic is cleaned against the reference before buffering, and the cleaned PCM is muxed via stdlib wave (not ffmpeg). Without the param the protocol and behaviour are byte-for-byte unchanged. Backend ships dark (no new deps — numpy already pinned); frontend far-end streaming is a follow-up. Tests cover echo attenuation, double-talk preservation, cold/stale pass-through, param validation, and the framing helpers — all pure-numpy/stdlib so they skip the torch ASR stack. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(capture_ws): stubs accept the new pcm_sr kwarg _transcribe_buffer/_transcribe_buffer_full gained an optional pcm_sr kwarg for the AEC PCM path; the protocol-test stubs had fixed signatures and raised TypeError on it, so the handler sent 'error' instead of 'final'. Accept **kw in the stubs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a330c9774c |
docs(spec): Voice Console 10/10 polish spec (#394)
* docs(spec): Voice Console 10/10 — pinned action bar, two-kicker hierarchy, unified presets, identity-first right rail Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(spec): ASCII '+' in wireframes — clears the CJK gate Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b60cceb3e6 |
docs(engines): uv dedupe + sidecar torch-pin disk-usage policy (Wave 4.5) (#392)
Explain why dedicated-venv engines (IndexTTS2) add disk (a second torch + CUDA libs: Linux cu128 ~0.83 GiB, Windows ~3.2 GiB), and how uv's link-mode dedup (clone on macOS/Linux, hardlink on Windows) shares identical wheels for free — provided UV_CACHE_DIR and the venvs are on the same filesystem. Key policy: pin the same torch build as the parent whenever the engine allows, since only identical wheels dedupe; UV_LINK_MODE=hardlink on Linux ext4. Linked from the IndexTTS engine doc. Spec §R4(a) / parity program Wave 4.5. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6306f6edae |
docs: official Docker Hub image palashdeb/omnivoice-studio (#388)
Link the published Docker Hub repo (https://hub.docker.com/r/palashdeb/ omnivoice-studio) as an official image alongside GHCR in the README install list and docker.md header. Same images, same tags. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
851ca4e012 |
feat(studio): workspace UX overhaul — right-side panels, shared waveform player, dub pipeline UX, setup polish (#374)
* feat(studio): workspace UX overhaul — right-side panels, shared waveform player, dub pipeline UX, setup polish, UI-wide fixes Voice workspace (specs: docs/specs/voice-studio-unification.md, workspace-connectivity.md): - Right-side panels replace the left sidebar for clone/design and dub: WorkspaceVoices (saved profiles), WorkspaceHistory (scoped history with All/Clone/Design filters), WorkspaceProjects (dub projects) - Prompt restacked over Voice Source in one definition column (spec §1) - Gallery "Use voice" now hands off via pendingProfileId and lands in clone - Shared <WaveformPlayer> (wavesurfer + in-DOM media element for Tauri WebKit, blob routing via preview endpoint, 404 -> "audio file missing") replaces every bare <audio controls>; lazy-mounted via IntersectionObserver Dub: - Pipeline stepper (Upload -> Prepare -> Transcribe -> Edit -> Generate -> Export) - Multi-language preview switcher pills (Original + per-track, ElevenLabs-style) - Batch multi-language generation via langOverride loop - FloatingPill: bottom-center, suppressed on its homeMode tab (no dup progress) - Transcript skeleton shimmer (no fake data), progress overlays the video, exports demoted behind Generate, empty right-panels collapse Chrome/layout: - Nav rail is full-window-height; content yields to the logs footer via padding-bottom; footer joins the rail edge (no overlap at any UI scale) - UI scale 60–175% slider with zoom-compensated container sizing - LogsFooter: merged single Logs tab when collapsed, per-source tabs on expand; Updates chip lives with the logs tabs - Gallery: three independently scrollable filter lanes, uniform 26px controls - Font picker as live-preview grid; double-click titlebar maximize fixed (single mousedown detail-2 handler) First-run: - Setup wizard: pinned action row + scrollable content at every window size, one-line head-ellipsized paths, height budget for short windows, library rows back to one-line grammar, raw i18n key + duplicate host fixed Performance/i18n/consistency sweep (10-agent scan, 47 fixes): - i18n locales lazy-loaded per language (i18n chunk 1.84 MB -> 76 kB) - Undefined CSS vars replaced with real tokens across 8 stylesheets; hardcoded hexes tokenized; emoji swept to lucide icons app-wide - Poll throttling (sysinfo subscription scoped to Header, logs 45s when collapsed, rAF only during playback), hardcoded strings moved to t() Build clean; 312/312 tests pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(studio): re-flow clone/design columns (grid rows collapsed in restack) + strip placeholder emoji across locales The base .studio-column grid (minmax(0,1fr) rows) collapsed to 0 height inside the new auto-height definition column, overlapping every panel in design mode — found via Playwright visual pass. Columns now re-flow as natural-height flex stacks. Also removed the leftover pencil emoji from clone.prompt_placeholder in all 21 locales. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(design): compact the design control stack — 2-up facet selects, scrollable tag row, tighter rhythm English accent + Chinese dialect dropdowns share one row (full-width on narrow), insertable tag chips collapse from three wrapped rows to one scrollable line, and describe/personality spacing tightens — the whole design stack now fits a single viewport. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(spec): unification migration renumbered 0004 — upstream 0003 is voice-profile consent Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): clear hardcoded-CJK gate — ASCII '+' in spec wireframes, reword voiceIcons comment Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(spec): migration is 0005 — 0004 taken by mcp bindings upstream Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1561ff4428 |
ci(docker): also publish to Docker Hub palashdeb/omnivoice-studio (#375)
Push the same images (same tag set: :latest rolling main, :stable/:X.Y.Z releases, :sha-) to docker.io/palashdeb/omnivoice-studio alongside GHCR. Gated on DOCKERHUB_USERNAME/DOCKERHUB_TOKEN secrets — without them the build still publishes to GHCR only. Docs-sync: docker.md mirror note. Requires repo secrets: DOCKERHUB_USERNAME, DOCKERHUB_TOKEN. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d6562d6f30 |
feat(dub): Smart Fit phase B — per-segment video retime export, drift absorption, fitted subtitles (#350)
* feat(dub): Smart Fit phase B — per-segment video retime export, drift absorption, fitted subtitles Executes the video side of the Smart Fit plans persisted by Phase A (job["fit_plans"], #347) at export and preview time. Backend: - services/video_retime.py (new, clean-room): two-tier retime executor. ≤48 chunks → the proven single-pass split/trim/setpts/concat filter_complex; above → batches of 40 chunks rendered to intermediate slices (identical libx264 medium/crf20 params, keyframe at t=0) joined losslessly with the concat demuxer. Slices are CFR-resampled (fps=) because setpts leaves VFR-ish timestamps that broke tpad and drifted a frame per retimed chunk on ffmpeg 7.x. Temp slices cleaned on success AND failure/abort. - Drift absorption: fitted track longer than retimed video → freeze-frame tail (tpad=stop_mode=clone) predicted into the last slice / single-pass graph, with residual mux-side tpad; video longer → silence-pad the dub audio chain (apad=whole_dur). ±50 ms tolerance. - VFR guard: probe r_frame_rate vs avg_frame_rate; normalise with fps= before trim/setpts; probe failure degrades gracefully. - Plan resolution: _video_retime_plan_for spans legacy video_stretch_plans (byte-identical resolution + command construction) and fit_plans, gated on the track's own timing_strategy so stale plans never retime a track re-generated under another strategy. - Fitted subtitles: /dub/srt + /dub/vtt accept ?lang= and serve cue times from fitted_segments for Smart Fit tracks; _write_burn_srt does the same for burn-in. burn_subs+retime is now allowed for smart_fit (burn runs AFTER the retime graph); still rejected for legacy stretch_video. - /dub/preview-video resolves the same plan so in-app preview matches export. - Fallback ladder: batch encode failure/timeouts → un-retimed export with a structured core.failure warning (X-Dub-Export-Warning header + job["last_export_warning"]); concat join rejection → one single-pass retry while ≤96 chunks; abort → 409 + proc kill via run_ffmpeg job_id registration (/dub/abort reaches export encodes now) + temp cleanup. Frontend: - Export drawer passes ?lang= on subtitle exports and shows an i18n'd re-encode cost note (~0.5–2× video length on CPU) when a retiming strategy is active — translated in all 21 locales. Tests: tests/test_smart_fit_export.py — plan resolution, batch math, graph parity + new stages, fitted-cue SRT/VTT/burn selection, burn policy, VFR detection; ffmpeg-gated integration renders both executor tiers (batch size forced to 2) and the real /dub/download endpoint, ffprobing durations within ±50 ms across both pad branches. All existing dub export/subtitle/preview/timing tests pass unchanged. Refs docs/competitive-analysis.md Action 1 (dub-length fitting v2); completes Smart Fit (Phase A = #347). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): sanitize Smart Fit retime work paths at every sink (CodeQL py/path-injection) The job_id-derived retime work path (retimed_*.mp4 / preview_retimed_*.tmp.mp4) flowed unguarded from dub_export into prepare_smart_fit_video / render_retimed_video and their derived slice/concat paths and ffmpeg argv. Apply the repo's proven inline realpath+startswith containment pattern (helpers/commonpath are not recognized — see #309/#328/#329/#348): - dub_export.py: validate work_path against DUB_DIR at both construction sites (export + preview) and pass the validated realpath onward. - video_retime.py: make both entry points self-defending — realpath + DUB_DIR containment on out_path/work_path before any derivation, raising RetimeError(stage="plan") on escape; slices_dir/slice_path/list_path and RetimeDecision.file_path now all derive from the sanitized value. DUB_DIR is read via module attribute so test fixtures reloading core.config work. - ffmpeg_utils.py: document that all caller-assembled argv paths are realpath-validated upstream. - tests: sandbox DUB_DIR in the executor integration tests (tmp_path) so the new guard sees the test workspace. No behavior change for valid (server-built) paths — the guard only fires on traversal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(smart-fit): patch DUB_DIR on video_retime's own config ref — survives suite-wide reload The retime guard reads video_retime._config.DUB_DIR at call time; the sandbox fixture patched a fresh 'import core.config' instead. Another test reloads core.config in the full suite, so the two module refs diverged — the patch missed and the guard rejected the test's tmp paths (green in isolation, red in CI's full run). Patch the exact ref the guard dereferences. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): resolve DUB_DIR live at call time in retime guards — survive full-suite reload The path-containment guards bound DUB_DIR via a module-level 'from core import config as _config'. Other tests importlib.reload() core.config (sandboxing OMNIVOICE_DATA_DIR), after which the guard checked containment against a stale DUB_DIR while dub_export built the path under the reloaded one — every retime path then 'escaped the dub workspace' (green file-alone, red full-suite: the 5 integration failures CI hit). Re-import DUB_DIR locally in each guard so it always reads the current sys.modules value; simplify the sandbox fixture to patch the canonical module. Verified: full backend suite green on the Smart Fit tests (the 2 remaining settings_store failures are pre-existing on main, unrelated — local data-dir artifact). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): clear CodeQL alerts on Smart Fit export — job_id allowlist, proc-registry decouple - py/path-injection (8, video_retime.py): validate job_id with a strict inline regex allowlist (re.fullmatch [A-Za-z0-9_-]{1,64}) at the entry of dub_download and dub_preview_video, before it reaches any filesystem path or ffmpeg argv. The existing realpath containment guards stay as defense-in-depth; the regex barrier is the sanitizer CodeQL recognizes through the service-module call chain. - py/log-injection (4): newline-strip job_id inline at the logger calls in ffmpeg_utils.run_ffmpeg and the two retime-fallback logger.error sites in dub_export. - py/empty-except (3): best-effort cleanup os.remove handlers now log the OSError at debug instead of bare pass (video_retime + both dub_export mux finally blocks; _discard_tmp too for consistency). - py/cyclic-import (2): break the dub_pipeline <-> ffmpeg_utils cycle for real — the subprocess registry (register_proc/unregister_proc/ kill_job_procs/has_active_procs + state) moves to a new stdlib-only leaf module services/proc_registry.py. ffmpeg_utils now imports it at module top (no lazy import); dub_pipeline re-exports every name so dub_core aliases and tests keep working unchanged. No behavior change for valid inputs; invalid job ids now get a clean 400 instead of a 404/containment error. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): address #350 review — cancelled-vs-failed retime, logged best-effort excepts, redacted probe logs, narrowed test assert - rc<0 (killed by user cancel) now raises RetimeError(stage='aborted') instead of reporting an ordinary render failure (CodeRabbit) - best-effort cleanup/QC-event excepts log at debug instead of bare pass (CodeQL empty-except x3) - probe failure logs use basename, not full user paths (CodeRabbit/CodeQL) - test_render_cleans_slices_on_failure asserts RetimeError, not Exception Rebuttals (no change needed, see PR comment): fitted-cue subtitles track the fitted AUDIO timeline which is correct even on retime fallback; the planner only emits stretch ratios >1 so the early-exit guard is a true no-op check; '\'' is ffmpeg's own utility quoting for concat lists; has_active_procs is an intentional re-export (noqa'd). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
99357e8c5b |
feat(mcp): MCP server v1 — mount on /mcp, per-agent voice binding, stdio shim (Wave 2.2) (#368)
* feat(mcp): MCP server v1 — mount on /mcp, per-agent voice binding, stdio shim (Wave 2.2) The FastMCP server (previously dead code, never mounted) is now mounted on the main FastAPI app at /mcp via Streamable HTTP, with its session manager composed into the app lifespan through an AsyncExitStack (best-effort: a missing mcp package or OMNIVOICE_MCP_DISABLE=1 never breaks startup). streamable_http_path set to '/' so the sub-mount lands at /mcp, not /mcp/mcp. Adds the 'mcp' dependency (1.27.x). Per-agent voice binding (Spec 2 headline): each MCP client sends an X-OmniVoice-Client-Id header; generate_speech resolves the voice as explicit arg > the client's binding > global default > app default. New mcp_client_bindings table (alembic 0004 + _BASE_SCHEMA, additive/idempotent), services/mcp_bindings.py (CRUD + resolve_voice + best-effort last_seen), and a loopback-gated REST router (/api/mcp/bindings) the Settings panel drives. New transcribe tool (base64 audio in, 200 MB cap). Stdio shim (backend/mcp_shim, httpx-only, ported from voicebox MIT) proxies stdio clients to the mounted endpoint and forwards OMNIVOICE_CLIENT_ID as the binding header. Settings → Sharing gains an MCP bindings panel. Docs: docs/mcp.md (both connection modes + binding REST) and docs/mcp.json updated to the shim form. Tests: bindings service + resolution precedence + migration up/down (pure, run locally); REST CRUD + mount-not-404 + disable-flag (main-importing, validated in CI). MCP build + mount + initialize handshake verified out-of-band (no torch). Spec: docs/competitive-analysis.md Spec 2 / parity program Wave 2.2. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(mcp): assert /mcp mount via app.routes, not a lifespan client The two main-importing mount tests ran the app lifespan, which now starts the FastMCP session manager and binds asyncio queues to the test loop — contaminating later lifespan-running tests ('bound to a different event loop'). The mount happens at import time, so inspecting app.routes for the /mcp Mount is the correct loop-free assertion. Same fix shape as the Wave 0.2 consent tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(mcp): stop reload-main poisoning across the MCP test files Root cause of the CI failure: the bindings REST fixture set OMNIVOICE_MCP_DISABLE=1 and reloaded main but never restored it, so a later 'from main import app' in test_mcp_mount saw /mcp un-mounted ({'/audio','/voice_audio'}). Reloading main mutates the shared module for every subsequent test. - REST fixture: drop the disable flag (the mount is harmless without a lifespan), yield the client, and restore main (+ core.config/db) to the default data dir in teardown so the global module is clean again. - test_main_mounts_mcp_route: reload main with the disable flag cleared so the assertion is independent of any earlier reload. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9b6d1d0863 |
docs(agentic): OmniVoice as a TTS/STT provider for pipecat/LiveKit (Wave 2.5) (#366)
Agentic v1: OmniVoice is a provider, not the orchestrator. Its existing OpenAI-compatible API already serves everything pipecat/LiveKit need (POST /v1/audio/speech with pcm/wav, voice-profile id, speed; default 24 kHz output matching pipecat's OpenAITTSService) — so this is docs + an example + a contract test, no new endpoint. - docs/agentic-voice.md: the provider recipe for pipecat (base_url to :3900/v1) and LiveKit, the remote-backend note (bearer from 2.3), the consent-locked-voice nudge (0.2), and an explicit telephony-is-deferred scope box. - examples/agentic/pipecat_minimal.py: lazy-import skeleton wiring the OmniVoice STT/TTS services (importable without pipecat installed). - tests/test_agentic_provider_contract.py: pins the /v1/audio/speech request shape pipecat sends (pcm + wav formats, voice-profile passthrough, speed) so the documented recipe can't silently break. Validated in CI. Spec: Action 15 / §R1 v1 / parity program Wave 2.5. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |