d91beef0fd314250d8d9b94de86dfea019a8bd96
312
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5ca9f54e47 |
docs: migration guide for stranded Real-Time-Voice-Cloning users + sharper local-first promise (#1085)
RTVC (CorentinJ's 50k-star SV2TTS repo) is archived; its users need a maintained home. New docs/migration/real-time-voice-cloning.md maps every RTVC concept to its OmniVoice equivalent (encoder+utterance → reference clip, toolbox → app, vocoder choice → Settings → Engines, demo_cli.py → REST API/CLI/MCP), is honest about what RTVC did that we don't (research toolbox, three-stage training, MIT license, smaller footprint), and walks the first clone with verified UI labels only. Wired into docs/features.yaml's existence-checked docs list and linked from the README Quickstart. README tagline now states the local-first promise verbatim at the very top: "No accounts. No API keys. No cloud." — everything else on the front page is unchanged. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ff56865cf7 |
feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines (#1083)
* feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines
IndexTTS-2 required four manual terminal steps (git clone, uv venv,
uv pip install -e ., export OMNIVOICE_INDEXTTS_DIR). This turns that into
a guided in-app install:
- backend/services/sidecar_install.py — parametrized sidecar provisioner
(SidecarSpec/SPECS so future sidecar engines are one entry, not another
installer). Resumable background job with step-by-step status: disk-space
preflight (needs-X/have-Y message), source fetch (git clone --depth 1
primary, GitHub tarball fallback when git is absent/fails), dedicated
venv via uv (OMNIVOICE_BUNDLED_UV → PATH resolution; transformers<5
isolation preserved — the parent env is never touched), import-probe
verification, IndexTeam/IndexTTS-2 weights into <checkout>/checkpoints
(where the sidecar actually loads from) via snapshot_download with the
auto-selected/configured HF endpoint + token — no hardcoded
huggingface.co — and persistence of OMNIVOICE_INDEXTTS_DIR (os.environ
for immediate use, prefs.json env.* for the next launch). Idempotent:
partial installs repair, downloads resume, healthy installs (incl. a
user's own clone) report already_installed and are never touched.
- API: POST /engines/{id}/install starts the job, GET
/engines/{id}/install/status polls it, DELETE /engines/{id}/install
removes an app-managed install (loopback-gated; refuses user-managed
clones). list_backends() gains one_click_install.
- Frontend: Settings → Engines shows an Install button on the IndexTTS2
row with per-step progress, live log tail, weight-download %, and
error+remediation; the manual setup snippet is demoted to a collapsed
"Manual install" fallback. All strings via i18n (en.json).
- OMNIVOICE_INDEXTTS_DIR joins the Settings env-var allowlist
(single-sourced from the installer SPECS).
- Docs: docs/engines/indextts.md leads with the one-click flow; manual
steps become the fallback section. CHANGELOG Unreleased entry added.
- Tests: tests/test_sidecar_install.py (24 cases — happy path, disk-space
fail, git-absent/git-failing tarball fallback, partial-install repair,
already-installed/running gating, uninstall safety, spec↔bootstrap
contract, router wiring) + 6 new EngineCompatibilityMatrix RTL cases.
API route snapshot regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): harden the sidecar installer — review findings
- Route namespace: /engines/sidecar/{id}/install — a dynamic
/engines/{id}/install would shadow the literal
POST /engines/sonitranslate/install (engines router registers first);
regression-guarded by test_sidecar_routes_never_shadow_literal_engine_routes.
- Weights completion marker: a killed-mid-download multi-shard weights dir
(config.yaml + plausible shards) no longer passes for healthy; the marker
is written only after snapshot_download returns, so re-runs resume.
- _run_logged: drain thread + proc.wait(timeout) + POSIX process-group kill
— a grandchild holding the stdout pipe can no longer hang the step past
its timeout.
- Job log lock: the status poll's list(deque) copy no longer races the
worker's appends (RuntimeError under active logging).
- Self-heal: a healthy managed install whose env var was lost (prefs wiped)
is re-pointed by start_install instead of reported already_installed
while the engine stays unavailable; legacy bootstrap installs (Probe-2
venv) are trusted via the engine's own probe.
- Single-sourced uv/venv-layout resolution: engines.indextts.bootstrap now
delegates _locate_uv/_venv_python_path to services.sidecar_install.
- Frontend: stable poll interval (keyed on the running-id set, not the
status map), reload on a job that finishes before the first poll,
re-attach to an in-flight job on remount, i18n'd Install aria-label,
manual-install <details> auto-opens on failure, snippet block hoisted
out of the JSX IIFE.
- list_backends: sidecar-installable set hoisted out of the per-engine
loop; exhaustive-shape registry test updated for one_click_install.
- Tests rebind the live services.sidecar_install module per test (other
suites purge sys.modules["services"], which made router tests
order-dependent).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): fill in the PR ref (#1083)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): validated tarball fallback + scanner-clean installer
- The pre-filter= extractall fallback (Python < 3.11.4) now extracts
member-by-member behind the same guards extractall(filter="data")
enforces — regular files/dirs only, no absolute paths, no ../ escapes,
resolved-path containment. Kills the new CodeQL py/tarslip (high) and
Bandit B202 (error) alerts; regression-tested with a malicious tarball
(test_safe_extract_members_blocks_tar_slip).
- snapshot_download tracks the weights repo's default branch on purpose
(same policy as every other model download; artifacts are
checksum-verified by hf_hub) — documented + B615 waived at the call.
- Explanatory comments on the intentional empty-except blocks
(CodeQL py/empty-except notes).
Verified locally: bandit -ll -ii on the module reports 0 MEDIUM+ findings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): address Greptile review — Windows tree kill, prefs write race, poll robustness
- _kill_tree: Windows now uses taskkill /F /T so a git/uv helper spawned by
the timed-out child can't keep writing into the checkout (POSIX already
killed the process group). Unit-tested with os.name patched to nt.
- core/prefs: mutations (set_/delete) are serialized behind a module lock —
the installer worker persisting its env.* key concurrently with a Settings
write could previously drop whichever key saved first (whole-class fix:
every threaded prefs writer, not just the installer). Fail-before/
pass-after: tests/test_prefs_thread_safety.py.
- Matrix polling: at most one in-flight status request per engine (an old
'running' response can no longer land after a newer 'succeeded' and
restart the poller), and four consecutive poll failures drop the stale
snapshot instead of showing "Installing…" and hammering a dead backend
forever.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9c81e3389d |
feat(network): automatic Hugging Face endpoint selection — probe, pick, remember (#1082)
* feat(network): automatic Hugging Face endpoint selection — probe, pick, remember Restricted-network first-runs (the #984 class: huggingface.co unreachable, user dead-ends before discovering the mirror setting) now self-heal by default, while explicit endpoint choices are never second-guessed. - New backend/services/endpoint_race.py: parallel HTTPS reachability + latency probes of huggingface.co and the hf-mirror.com community mirror (3s timeouts). Probes are the only signal — no geo-IP, no third-party calls. Reachable beats unreachable; with both reachable the official endpoint wins unless the mirror is decisively faster (anti-flap hysteresis). The pick is cached in prefs and re-raced only on first run, a network-classified download failure, staleness (>7 days), or an explicit "Test again". - Manual mode is sacred: HF_ENDPOINT env, an hf_endpoint pref, or any explicit Settings pick disables auto-switching entirely; OMNIVOICE_HF_ENDPOINT_MODE=manual is a hard opt-out. - Wiring: the wizard preflight races endpoints when nothing is configured (honest copy when the mirror wins; warn-not-block when nothing is reachable); Model Store installs and the model-cache auto-repair resolve their per-call endpoint= through the cached decision, and a network-classified failure re-races once per repo per process and retries on the new winner (same guard pattern as the cache-recovery ladder). - Settings → Models → Hugging Face mirror gains "Auto (recommended)": shows the current pick, measured latency, last-checked time, and a "Test again" button (POST /api/settings/hf-mirror/test). Existing explicit configs surface as the matching manual mode. Panel notes that hf_hub checksums every download regardless of endpoint. - Tests: policy/cache/failover matrices in tests/test_endpoint_race.py, preflight + settings + repair-failover integration with mocked probers, HFMirrorPanel mode tests, and a suite-wide conftest guard that pins the probers so no test can hit the real network. - Docs: downloading-models.md and install/troubleshooting.md describe the automatic default and both opt-outs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): Unreleased entry for automatic HF endpoint selection Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint-probe pin uses an isolated MonkeyPatch and clears the decision cache; dtype guard tolerates stubbed torch The autouse probe pin requested the shared monkeypatch fixture, hoisting its setup earlier for every test and reordering teardown against the fp16 guard — which then ran torch.get_default_dtype() on test_torch_compile_gate's SimpleNamespace stub. The pin now uses its own MonkeyPatch context and also clears the prefs-cached endpoint decision per test (one test's auto pick leaked into other tests' preflight labels on CI ordering). The dtype guard additionally skips non-module torch stubs outright. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint env vars can no longer leak out of the mirror-settings suite set_hf_mirror writes os.environ[HF_ENDPOINT] during the test, and monkeypatch.delenv(raising=False) on an absent var records nothing to undo — so the write leaked process-wide and flipped later suites' preflight network checks into the explicit-endpoint branch (the CI-order failures). Guaranteed save/restore autouse fixture at the source, plus defensive env shedding in the preflight suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
254f071b45 |
fix(docker): ship alembic.ini in the image, add an image-level HEALTHCHECK (#1080)
Migrations in Docker fell back to the additive-column self-heal because alembic.ini was never copied; the real migration chain now runs. The HEALTHCHECK covers plain docker-run (compose files keep their own), with a start period sized for first-boot schema creation. Docs example tag freshened. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5562aa16a7 |
feat(setup): media tools become invisible — bundled by default, controllable in Settings (#1071)
Most users should never learn what ffmpeg is. The Setup Wizard's SYSTEM
PREFLIGHT stops listing FFmpeg / FFprobe / yt-dlp as user-installed
requirements ("brew install ffmpeg…"): they are internal dependencies the
app provisions for itself. Genuine user facts (OS, RAM, disk, GPU,
network, Python) are untouched.
Backend
- New services/media_tools.py: per-tool status {version, path, origin:
sidecar|bundled|system|custom}; background acquisition of a pinned,
SHA-256-verified static ffmpeg+ffprobe build (immutable-commit fetch
from the same upstream the static-ffmpeg pip package uses — that
package itself was audited and rejected: mutable raw/main URL, no
checksums, writes into site-packages); binaries are `-version`-probed
via the existing _binary_runs before being trusted, installed under
DATA_DIR (update-surviving, frozen-build-safe), zero new Python deps.
- ffmpeg_utils resolution chain gains the acquired-bundled tier — and
ffprobe finally has a bundled tier at all (imageio-ffmpeg ships none),
closing the source-install gap.
- New /media-tools router (loopback-gated, same contract as
/system/set-env): status, acquire, {tool}/custom-path | use-system |
restore, ytdlp/update | restore. Overrides persist via the existing
env.FFMPEG_PATH / env.FFPROBE_PATH prefs convention — one store, no
competing controls.
- yt-dlp updates: audited in-venv pip/uv upgrade and rejected (venv is
uv-managed with no pip; yt-dlp is a locked dep, so the updater's
--inexact drift sync would revert it). Instead the newest wheel —
verified against PyPI's own sha256 — lands in a DATA_DIR overlay
prepended to sys.path at startup: survives app updates, works in
frozen builds, and "Restore tested version" is just deleting the
overlay. Gallery now runs yt-dlp via `python -m yt_dlp` (module, not
PATH) so the CLI can never be a user-install task either.
- /setup/preflight drops the three tool rows, carries a media_tools
verdict, and self-heals: kicks the bundled download in the background
when no tier resolves (never re-fires after a failure — the wizard's
card owns Retry). diagnose + the ffmpeg-missing notification now point
at Settings → Audio tools instead of package managers.
Frontend
- Wizard: new MediaEngineCard — renders NOTHING when the engine is ready,
a one-line progress while acquiring, and only on failure an actionable
card (Retry / Use a system copy / Choose file…).
- Settings → Audio tools (new category, System group): FFmpeg + FFprobe
rows with version, path, origin badge, Use system copy / Choose file… /
Restore bundled, header-level "Update bundled build"; yt-dlp row with
Update + Restore tested version (+ restart affordance). Package-manager
commands appear only as copyable prose, never executed.
- The FFmpeg-path override moved out of Settings → Network (pointer row
deep-links to Audio tools; no second writer of env.FFMPEG_PATH).
Notifications gain a settings-tab action type.
- All strings i18n (en + defaultValue), a11y labels on every control.
Tests: 29 new backend (origin classification, checksum/size/probe
rejection, override persistence, overlay update/restore, router gating +
route-shadowing) + preflight contract tests (tool rows gone, verdict
present, auto-acquire fires once); 14 new frontend (wizard hide/progress/
failure-card, Audio tools rows/badges/actions). Route snapshot
regenerated. Docs (macos/linux install, troubleshooting §7b) describe the
new reality in the same commit.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
fd083749c6 |
fix(settings): network & privacy panels — clearable proxy after reload, HF-mirror panel never vanishes, guarded remote-backend save, honest privacy claims (#1063)
Settings → Network / Models / Sharing / Privacy / OpenAPI fixes:
- NetworkTab: a proxy persisted in a previous session can now be cleared —
the Clear button and "Set" badge derive from the backend-persisted value
(sysInfo.proxy_url), not only from a save in the current session. Proxy row
copy now matches its real semantics ("Applies now" badge; desc/toast no
longer claim a restart is needed or leak yt-dlp jargon — reworded in all
21 locales). FFmpeg path placeholder is platform-appropriate instead of
Windows-only on every OS.
- HFMirrorPanel: the panel no longer disappears when the initial GET fails —
the section shell always renders, with a loading state and an error +
Retry affordance. Saving now toasts, the active preset is marked
(aria-pressed), and the custom-URL row is labelled "Custom mirror URL"
instead of raw HF_ENDPOINT jargon (env var moved to the row note).
- RemoteBackendPanel: full i18n (was 100% hardcoded English); Save & reload
now validates the URL (http/https, parseable) and asks for confirmation
before saving a URL that hasn't passed a connection test — a typo'd base
no longer bricks every API call after reload. Dropped the contradictory
"Restart required" badge (saving reloads the app itself; description says
so). docs/remote-gpu.md updated to match (docs-sync).
- PrivacyTab: the "Network calls" row no longer shows the green "Offline
translator" assurance when the backend is down or reports 'unknown' —
green is reserved for confirmed-offline providers (nllb/argos/
libretranslate), everything unconfirmed shows a neutral "Unknown" badge.
The online-translator warning now deep-links to Translation settings.
- OpenApiPanel: a failed clipboard copy toasts an error instead of silence.
- a11y: all five text inputs across these panels now carry accessible names
(aria-label), previously announced only by their vanishing placeholders.
Tests: new colocated suites for NetworkTab, HFMirrorPanel,
RemoteBackendPanel, PrivacyTab; OpenApiPanel suite extended with copy
success/failure. Frontend suite 140 files / 1061 tests green; i18n parity
probes green (new keys en-only with defaultValue, reworded keys updated in
every locale).
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
5ce9d0e51d |
feat(dub): predict segment fit before synthesis — tight/impossible badges + opt-in shorter rewrites (#1051)
New pure planning layer (services/duration_planner.py) runs after translation, before TTS: estimates each translated line's natural speech duration (self- calibrating from the job's already-synthesized segments, static per-language rates as cold-start fallback) and classifies it fits/tight/impossible against slot + capped gap borrow, with thresholds derived from fit_planner's own caps so "impossible" means "would be trimmed". Verdicts ride the /dub/translate response and badge the segment table; an opt-in (default OFF) LLM pass attaches one-click shorter-rewrite suggestions for impossible lines. Never blocks generation — informs before GPU time is burned. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7dbb95fa15 |
feat(dub): LLM translations keep terms consistent and sound spoken — auto-glossary brief + reflect pass (#1050)
One up-front LLM pass over the full transcript extracts a theme summary + terminology map, merges it under the user's manual glossary (user entries always win), caches it on the dub job per target language (job_data blob, no schema change), and injects the brief into every per-segment prompt. A new reflect pass then critiques each segment's direct translation for wordiness / stiff register and rewrites it as natural spoken dialogue — any failure or divergence silently keeps the direct translation. Both stages have Dub-tab toggles (default ON for the LLM engine, persisted; MT engines unaffected), with i18n strings across all 21 locales and docs updated. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
df870e9ed3 |
docs(agents): add verified Tesla T4 (16GB) inference notes (#1014)
* docs: add AGENTS.md with verified Tesla T4 (16GB) inference notes
Documents two things found while verifying inference on a real T4:
1. Cold-cache first /v1/audio/speech call can hit the 300s
OMNIVOICE_GENERATE_TIMEOUT_S because the checkpoint download happens
inside that budget — workaround via existing POST /models/install or
raising the timeout, no code change needed.
2. The OpenAI-compatible endpoint silently ignores num_step/guidance_scale
(schema doesn't declare them) — use native /generate for those.
Also documents the T4 acceleration checklist (dtype/attention/int8/CUDA
graphs) and measured VRAM (peak 2.05GB). No code changes.
* fix(docs): make /models/install workaround command actually executable
Addresses Greptile review: the instruction omitted the required
repo_id body field (InstallModelRequest rejects an empty body).
* fix(docs): correct port in /models/install example (3900, not 8000)
The app serves on port 3900 (confirmed: /health returns 200 there,
connection refused on 8000). Verified the exact corrected curl command
returns 200 {"status":"install_started",...}.
* move T4 notes to docs/hardware-notes-tesla-t4.md — AGENTS.md is the auto-loaded agent-instructions filename
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
80f10289fe |
feat(asr): ASR engines get the same Settings picker TTS has (env var still wins) (#1026)
Settings → Engines now stacks one pinned Engine Compatibility Matrix per family (TTS, ASR, LLM) instead of a single TTS-titled table with the other families tucked behind a low-discoverability tab. The backend select/prefs path (family="asr" → prefs.asr_backend, env > prefs > auto-detect) already worked but was unexercised and undocumented — it's now locked by API and resolution-order tests, and README + the openai-compat-asr doc stop promising a picker that didn't exist / denying one that now does. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1808a373a1 |
fix(linux): AppRun workaround detection reads the BUNDLED WebKitGTK version, not the host's (#961 follow-up) (#1024)
The launcher decided whether to export WEBKIT_DISABLE_COMPOSITING_MODE by asking the host's pkg-config — but LD_LIBRARY_PATH makes the BUNDLED libwebkit2gtk the one that actually runs, so on any machine where the two diverge the detection read the wrong number. This was the second bug identified during #961's investigation (the reporter built from source, so their dev packages answered pkg-config with a healthy version while the shipped bundle ran an older lib) and was explicitly deferred in #1007 as not-safely-fixable at runtime. The fix makes it knowable by construction instead: inject-apprun.sh runs at bundle time ON the build host whose libwebkit2gtk gets bundled, so it stamps that version into .bundled-webkitgtk-version inside the AppDir. AppRun reads the stamp first and only falls back to host pkg-config for bundles predating it. Empty/unreadable stamp fails safe (workaround on), same philosophy as the missing-pkg-config path. Tests: 3 new cases in AppRun.test.sh — marker-beats-host in both directions (broken-marker/healthy-host and the #961 inversion, healthy-marker/broken-host) plus empty-marker fail-safe. Also wires AppRun.test.sh into pytest (tests/test_apprun_launcher.py) — it was previously run by NO CI job, so the launcher could regress silently. Also documents Windows install-to-another-drive behavior in docs/install/windows.md (#938): local drives work via the wizard's directory picker, mapped network drives are a Windows Installer limitation, and the data directory moves independently of the app. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
da4bef8e42 |
docs: correct troubleshooting §16 — mic bug was a missing entitlement, not an upstream limitation; changelog for #1016/#1020/#1021 (#1022)
troubleshooting.md §16 claimed the macOS microphone-permission bug was an unresolved upstream Tauri/wry limitation with no available fix. That was wrong: @MahdiHedhli read the wry/tauri sources more carefully and found the real cause — Tauri's Hardened Runtime default blocks mic hardware access without com.apple.security.device.audio-input in the bundle's entitlements, which also explains why TCC never listed the app. Their fix (#1016) is merged; §16 now documents the real mechanism, credits the correction, and keeps the record-elsewhere workaround for users on ≤0.3.12 builds. Also brings CHANGELOG [Unreleased] current for the three merges that lacked entries: #1016 (mic fix), #1020 (shutdown wait 3s→20s), #1021 (CI flaky-trio root cause + guard). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7f8a42ce51 |
fix(tts): mlx-audio CSM cloning drops ref_text, breaking every clone attempt (#1012, #1013) (#1017)
MLXAudioBackend.generate() reads voice/ref_audio/language/speed from its kwargs but never extracted ref_text — it was built, then silently never passed through to self._model.generate(). CSM (sesame.py) only builds its cloning context when BOTH ref_audio AND ref_text are present; with ref_text missing, the context list stays empty and indexing into it raises "IndexError: list index out of range" deep inside mlx-audio, instead of the clone ever being attempted. Voice cloning on the CSM engine could never have worked as shipped. generation.py already threads ref_text all the way through — even auto-transcribing it via the GPU pool when the caller supplies ref_audio without one (~line 780) — so the value was always available in kwargs; it just never survived the crossing into this specific backend. Reported with the precise root cause and a working fix (community member independently diagnosed and patched it locally, confirmed working on MPS/0.3.12). Two-line fix: extract ref_text and pass it through when both ref_audio and ref_text are present (guards against passing an orphaned ref_text with no accompanying audio to engines that don't expect it). Tests: tests/test_engines.py — ref_text is passed through when paired with ref_audio, omitted when ref_audio is absent. Also documents the second bug from the same report (#1013): macOS microphone permission never prompts, so OmniVoice never appears in System Settings to grant access. Root-caused to an unresolved upstream Tauri/WebKit limitation (WKWebView's requestMediaCapturePermissionFor delegate — wry#1195, tauri#11951, fix wry#1196 still open/unmerged, no released version to bump to) — not something fixable here without an unverified native Rust/WKWebView hack this session has no way to test. Documented in docs/install/troubleshooting.md with the confirmed workaround (record elsewhere, upload the file). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5a7d9cc05c |
feat(asr): generic OpenAI-compatible transcription backend (#877) (#1003)
First slice of the community's two-track proposal for #877: a generic OpenAI-compatible ASR backend that works TODAY, without waiting on transformers to ship a direct Qwen3-ASR integration (tracked separately, still blocked upstream). Points OmniVoice's transcription at any server exposing POST /v1/audio/transcriptions — a self-hosted Qwen3-ASR/ FunASR/SenseVoice server, or OpenAI's own API. - New OpenAICompatASRBackend (backend/services/asr_backend.py): a pure network client, no local model, no install. Prefers response_format=verbose_json for real per-segment timestamps, degrades to plain text (matching MoonshineASRBackend's shape) when a minimal server rejects that format. Never leaks a raw SDK/httpx exception to the caller (#977 convention) — wraps network/auth failures in a clean, actionable RuntimeError naming the server. - Settings persist via the same encrypted-secret convention as services/llm_providers.py (settings_store.set_secret for the API key — Fernet-encrypted, never a .env row, never echoed back; get_text/ set_text for base_url/model). New GET/PUT /api/settings/ asr-openai-compat, loopback-gated like every other settings route. - Frontend: a small settings panel (Settings → Models) mirroring HFMirrorPanel's exact structure. No ASR engine picker exists yet for ANY ASR backend (only TTS has one) — activating this engine still needs OMNIVOICE_ASR_BACKEND=openai-compat-asr; documented plainly rather than pretending otherwise. - README's ASR Engines table (9 → 10 engines) and docs/features.yaml's drift-checker inventory updated; the '9 engines, all fully local' claim corrected since this one genuinely isn't. - docs/engines/openai-compatible-asr.md: setup steps + an explicit privacy note (unlike every other ASR engine, audio leaves the machine to whatever server is configured). Regression tests: tests/test_asr_openai_compat_877.py (12 tests) — is_available() gating, verbose_json + plain-text response adaptation, network-failure error hygiene, SDK retry disabling, and the settings endpoints' persist/mask/clear-vs-unchanged semantics. Fixed two real full-suite-only failures found during verification (not brushed aside): the API route inventory snapshot needed regenerating for the two new routes, and this file's own tests had a module- staleness bug — a collection-time settings_store import went stale relative to a test-time-fresh fixture when another test elsewhere in the ~2400-test suite reimports the module — fixed by making settings_store itself a fixture resolved at test-run time, same lesson already applied to tests/test_mm2_lifecycle.py earlier this session. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
93aab6dadc |
docs(linux): mention yt-dlp as an optional prerequisite (#973) (#997)
The preflight system check already warns in-app when yt-dlp is missing (Voice Gallery/Dub YouTube downloads fail without it), but the install docs never mentioned it — a user has to hit the in-app warning first instead of seeing it up front alongside the other optional prereqs. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
637020b82b |
fix(engines): nemo-parakeet install hint stops recommending a shared-venv-breaking pip install (#974) (#991)
The Engines page told users to run `pip install nemo_toolkit[asr]` for the NeMo Parakeet ASR engine. nemo_toolkit[asr]==2.7.3 hard-pins transformers>=4.57,<4.58, which is unsatisfiable alongside OmniVoice's own transformers>=5.3 requirement (needed by omnivoice/models/omnivoice.py for HiggsAudioV2TokenizerModel). A user who followed the hint ended up with a backend that wouldn't start (ImportError: cannot import name 'HiggsAudioV2TokenizerModel'). _INSTALL_HINTS["nemo-parakeet"] in backend/services/asr_backend.py now states plainly that installing into the shared venv will break the backend, names the transformers conflict, and tells users to use a separate/dedicated Python environment instead — without implying a safe one-line fix or an isolated-venv env var exists (unlike dots-tts/moss-tts-v15/confucius4-tts, nemo-parakeet has no isolated venv option yet; that's a separate, larger follow-up). Also adds one sentence to docs/install/troubleshooting.md's existing "engine venv clash" section (#11) pointing at the same class of issue on the ASR side, and a regression test asserting the hint never again contains the literal bare `pip install nemo_toolkit[asr]` string. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
dab6456581 |
docs(linux): stop advertising a .deb package that isn't published (#990)
README's Quickstart badges linked a 'Download Debian .deb' button straight to the releases page — but .deb bundling was deliberately dropped from release.yml (tauri-cli bug, 'Failed to create control scripts') and no release has ever shipped one. A community member investigating #961 confirmed this by checking the actual release assets. Users clicking that badge got a broken promise, not a package. Removed the badge; docs/install/linux.md's '## Install (.deb)' section now honestly states it's unavailable pending a tauri-cli fix, points to the AppImage as the supported path, and keeps the historical pre-v0.3 .deb upgrade note (ffprobe conflict) since that's still relevant to existing installs. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1a03c59f82 |
fix(install): AMD ROCm torch reinstall targets rocm6.4, not rocm6.2 (#988)
Community-diagnosed (issue #972, Kaihui-AMD): pyproject.toml pins torch==2.8.0, but the rocm6.2 wheel index only ever published up to 2.5.1 — the reinstall silently failed to resolve and fell back to the default CUDA build, which runs on CPU on an AMD GPU. The failure was correctly logged (bootstrap.rs's emit_log warning), just never actioned because the index itself couldn't succeed. rocm6.4 carries a matching torch==2.8.0 build. Docs updated with the corrected index plus a repo.radeon.com find-links path for users who want a driver-matched ROCm 7.2.x build the PyTorch index doesn't carry (OMNIVOICE_TORCH_INDEX only accepts a PEP 503 index, not find-links, so that's documented as a manual step). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7f77f4d7bf |
fix(audio): remove hidden reverb from the mastering pre-stage — reverb is preset-declared only (#986)
* fix(audio): remove hidden reverb from the mastering pre-stage — reverb is preset-declared only (#TBD) Field report (Discord): baked-in echo/reverb on some voices. apply_mastering() hardcoded a Reverb that ran on every non-raw synthesis before the user's preset chain — broadcast shipped reverb it never declared, podcast broke its "no reverb" promise, cinematic/warm got doubled reverb. The mastering pre-stage is now data-driven (MASTERING_CHAIN: highpass + compressor, same params as before) and reverb-free; cinematic/warm keep their user-chosen reverb. Regression tests pin the contract, incl. a burst-then- silence echo-tail check and pedalboard-missing passthrough. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): hidden mastering reverb entry (#986) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
087309259b |
fix(setup): first-run network check is mirror-aware and never hard-blocks (#984)
* fix(setup): first-run network check is mirror-aware and never hard-blocks Field report (Discord, China): the Launchpad preflight probed hardcoded huggingface.co:443 and any failure disabled Continue outright — users behind the GFW were stuck on the very first screen, before Settings (and its HF mirror quick-pick) was even reachable. - The probe now targets the HF endpoint actually in effect (HF_ENDPOINT / hf_endpoint pref via configured_hf_mirror), with the real port. - An unreachable endpoint is a WARNING, not a blocker: local-first — cached models work offline, and downloads surface their own actionable errors. - When huggingface.co is blocked but hf-mirror.com answers, the fix text says exactly that, and the wizard shows an inline mirror quick-pick (presets + custom URL) that applies via PUT /hf-mirror — effective immediately for downloads — then re-checks. - Docs updated (downloading-models, install troubleshooting); regression tests cover warn-not-fail, mirror-host probing, and the mirror suggestion. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): open [Unreleased] with the preflight mirror fix (#984) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
620321a9cc |
feat(diagnostics): backend crashes become self-documenting — exit code + stderr tail surfaced and attached to bug reports (#969)
* feat(diagnostics): backend crashes become self-documenting — exit code + stderr tail surfaced and attached to bug reports When the backend PROCESS died (native CUDA abort, OOM kill, DLL crash) the user saw only "Can't reach the local OmniVoice backend" and the evidence died with the process — every #941-class report needed a logs-please round-trip nobody answers. The v0.3.9 guard fixed HANGS; this fixes the class of invisible DEATHS: - Rust (crash.rs): every unexpected child exit — detected by the startup health poll and the post-Ready supervisor — writes a rotating (last 3) JSON crash marker next to the backend logs: ts, exit code/signal, backend version, uptime, ~40-line stderr tail. Intentional shutdowns never forensicate: app-quit raises the quitting flag first (now also on macOS Cmd+Q via ExitRequested), and retry/clean-retry kills set a BACKEND_KILL_INTENDED flag cleared when the fresh child is tracked. - Tauri commands get_last_backend_crash / acknowledge_backend_crash; ack is a persisted watermark, never a delete — bug reports still get the evidence after the user viewed it. - Crash-loop escalation: the supervisor budget goes 5-in-60s → 3-in-10min so slow crash loops stop respawning and land on the Failed screen with the last exit code + stderr tail. - Frontend: apiFetch's transport-failure path swaps the vague message for "the backend crashed (exit code X) N s ago…" when an unacknowledged marker exists, and BackendCrashNotice (banner + details dialog, i18n'd, ack-on-view) surfaces it even with no request in flight. - Bug-report prefill gains a "Last backend crash" section (exit code + home-path-scrubbed stderr tail via the existing scrubText), so the next report arrives WITH the evidence. Tests: cargo --lib 57 pass (marker rotation write-4-keep-3, ack semantics, store IO, ExitStatus decomposition, 3-in-10min policy); vitest 909 pass incl. crash-notice branch, client crash-message branch, bug-report enrichment; legacy node:test 41 pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add backend crash forensics under [Unreleased] (#969) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4aa9abe22a |
docs+scripts: install fixes — desktop-prod tauri resolution, Ubuntu white-screen guidance, honest GPU/prereq docs (#960 #961 #962) (#964)
* docs+scripts: install fixes — desktop-prod tauri resolution, Ubuntu white-screen guidance, honest GPU/prereq docs (#960 #961 #962) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add the install-fixes batch under [Unreleased] (#964) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
32103ad7b7 |
docs(readme): charm + organization overhaul (Opal-style) — collapsibles + OpenAI-compatible API section (#945)
* docs(readme): charm + organization overhaul (Opal-style) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): restore inventory-exact feature names (docs-drift guard) The charm pass sentence-cased five bold leads in the collapsed feature list; scripts/check-docs-drift.py greps for the inventory's exact title-case names. Restored: Vocal Isolation, Speaker Diarization, Batch Queue, AI Watermark, GPU Auto-Detect. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): dubbing screenshot shows a real completed dub (37 segs, EN→BN) Replaces the empty drop-zone shot with the populated editor — video + waveform + cast, 37 Bengali segment rows, DUB COMPLETE banner — captured live from the v0.3.9 app; caption updated to match. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0579ec91ad |
docs(readme): Opal-style restyle + fresh v0.3.9 screenshots (#937)
- Emoji section headers with explicit <a id> anchors. Emoji breaks GitHub's auto-generated heading slugs, so every in-page nav target keeps a stable explicit anchor (verified all href="#..." resolve). - Refresh the screenshot gallery. The prior set was from April, predating the launchpad / settings / dictation UI overhaul, so it misrepresented the app. Captured fresh at retina from the live v0.3.9 UI and led the gallery with the new Launchpad home: launchpad, studio, voice design, voice gallery, dubbing, engine-compatibility matrix, model store, embedded API reference (Scalar), and the in-app changelog reader. - Fix the stale engine count (11 -> 14 TTS engines) in the comparison table, FAQ, and roadmap to match the engine table + backend registry. - Use <kbd> keycaps for the dictation shortcut (Opal detail). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0d80ab2cb0 |
docs: backfill [0.3.9] batch bullets + add OSS sponsorship playbook (#931)
Bullets for #922 (release titles), #923+#924 (sponsors), #925 (contact), #927 (models), #928 (openapi), #930 (engines) — the agents kept off CHANGELOG.md during the merge chain. Plus a portable how-we-set-up- sponsorship playbook for reuse on other projects. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8f71c90f20 |
feat(settings): LLM Skills — per-feature enable/route control for every LLM call (#912)
New Settings → System → LLM Skills area: every LLM-powered capability
(Cinematic & Autofit translation, speech-rate slot fitting, glossary
auto-extract, direction parsing, dictation cleanup) becomes a "skill" the
user can toggle or route to a specific provider (local Ollama/LM Studio vs
a remote key) instead of everything riding the one global active provider.
Backend:
- services/llm_skills.py — skill registry + settings_store persistence
(llm_skill.<id>.enabled / .provider), resolution precedence
override > active > none, resolve_skill_client() (OpenAI-compat client
bound to the effective provider; None when disabled/unconfigured) and
skill_backend() (OffBackend when disabled — the exact no-LLM object every
caller already degrades on).
- All five consumption points wired through the registry; a disabled skill
degrades exactly like "no LLM configured" today (Fast translation
fallback, refinement pass-through, heuristic direction parse, no-llm slot
fit, 503 on glossary auto-extract). No new degradation modes; defaults
(enabled + no override) keep existing setups byte-identical.
- OpenAICompatBackend gains an optional bound provider (None = active, the
historical behavior).
- GET /api/settings/llm-skills + PUT /api/settings/llm-skills/{skill_id}
(404 unknown skill/provider); route snapshot updated.
Frontend:
- LLMSkillsPanel (Sparkles, next to LLM Providers): one row per skill —
i18n name/description, enable toggle, provider Select ("Use active
provider" + configured providers, local ones tagged), ready /
needs-setup badge linking to LLM Providers. All strings via t()
(settings.llmskills_*).
Tests: 30 backend (precedence, per-consumption-point disabled semantics,
endpoint round-trips, validation) + 4 panel render/PUT tests. Docs:
translation-engines.md gains an LLM Skills section.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
af6690840e |
fix(translate): run Cinematic/Autofit on every engine (incl. default Argos), bound the fit pass, scrub provider errors (#910)
P0 — Cinematic/Autofit silently no-op'd on argos/nllb/openai. Those three branches returned BEFORE _maybe_cinematic, so only the deep_translator fall-through reached the refine/fit pass. A user on the DEFAULT Argos engine who picked Cinematic/Autofit got plain Fast output with a success toast and no quality_used/cinematic_skipped/rate_ratio. All three now route through _maybe_cinematic. provider=openai is already an LLM translation, so it skips the reflect/adapt re-refine (new already_llm flag) but still stamps rate-ratio badges and runs the Autofit fit pass; the dialect it baked into its translate prompt is now reported applied. P1 — the Autofit fit pass ran one blocking adjust_for_slot per segment in the merge loop, OUTSIDE any budget (a 50-seg dub vs a slow provider spun ~50×timeout unbounded). New speech_rate.adjust_for_slot_many fans it out concurrently under a wall-clock deadline SHARED with the cinematic refine; segments still running at the deadline degrade to their literal with rate_error='fit-budget'. Also set max_retries=0 on the OpenAI clients used for translate/refine/fit so a 429 + Retry-After can't sleep through the budget. P2 — glossary auto-extract's no-LLM message now points at Settings → LLM Providers (was the stale TRANSLATE_BASE_URL/TRANSLATE_API_KEY). Provider error bodies on the glossary auto-extract, the OpenAI translate-segment path, and the DeepL/Microsoft translate-segment path are now scrubbed (core.scrub.scrub_provider_error) — they could echo the API key / a user_id. DubTab re-polls LLM availability on window focus / visibility so configuring a provider in Settings lifts the Cinematic gate without a remount. Documented LLM_DEFAULT_PROVIDER in docs/dubbing/translation-engines.md. Tests: fail-before/pass-after for argos+cinematic (refine runs), argos+cinematic no-LLM (cinematic_skipped), argos Fast (rate_ratio stamped), openai+autofit budget bound, and provider-error scrubbing on the translate + glossary paths. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
16294fed44 |
feat(updates): data-safe updates — pre-migration DB backups, guarded venv heal, release notes + changelog reader (#909)
Backend: - core/db_backup.py: WAL-safe SQLite snapshot to omnivoice.db.backup-<version>-<n> before pending alembic migrations run; keep newest 3, prune older; skip >500MB with a log line. Restore is never automatic. - core/db.py: _run_alembic_upgrade now plans the run (up_to_date / pending / unknown_revision), snapshots first when migrations will execute, and raises MigrationError on a mid-flight failure — startup stops with the backup path named instead of continuing on a half-migrated DB. The #552/#547 unknown-revision class stays non-fatal (warn + additive reconcile). - core/changelog.py + GET /api/settings/changelog: parse the shipped CHANGELOG.md (single-line and wrapped bullet styles) into structured releases. - GET /api/settings/db-backup: newest pre-migration backup for the panel. Rust (bootstrap.rs): - #314 heal guard: an exit-signature match alone can no longer delete the venv — venv_rebuild_justified requires a structural problem or a failed direct interpreter probe; a venv that probes healthy is kept and the real error surfaced. Drift/repair remains in-place `uv sync` (non-destructive). - CHANGELOG.md now ships as a bundle resource and is copied/refreshed into the project dir so the changelog endpoint works in packaged installs. Frontend (Settings → Updates): - Available update shows its actual release notes (updater metadata body) through a safe markdown-lite renderer (text nodes only, refs stay plain). - "Your data is backed up before every update" line with the latest backup timestamp from the new endpoint. - "What's new" changelog reader (accordion, newest expanded) over the shipped CHANGELOG.md; GitHub releases list reuses the same renderer. - One-time, non-blocking "What's new" footer pill after an update (persisted last-seen version; fresh installs baseline silently). - All strings via t() with en keys (other locales fall back to English). Tests: db backup/rotation/failure-path units, migration-safety units, changelog parser (both bullet styles + real CHANGELOG.md), endpoint tests, route inventory regenerated, Rust decision-logic + probe tests, vitest suites for renderer/viewer/panel/pill logic. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
72d137e1f3 |
fix(bootstrap): port cuDNN 8 (NVIDIA CUDA GPU) + VC++ redist to packaged installs (#869)
* fix(bootstrap): port cuDNN 8 (NVIDIA CUDA GPU) + VC++ redist install into ensure_venv_ready() * fix(bootstrap): address #869 review — drop dead VC++ half, cache negative CUDA probe, gate on ROCm, sync docs Per maintainer review on #869: 1. Drop the VC++ Redistributable half: LoadLibraryA("vcruntime140.dll") from the running Tauri exe is a tautology (the exe itself links the MSVC CRT, so the process wouldn't be running without it), and torch's real failure mode is msvcp140.dll inside the venv python process. Dead code removed; a comment records why for future readers. 2. Stop taxing every non-CUDA launch: a negative torch probe (CPU / Intel / AMD — most installs) is now cached in a .venv/.cudnn8_probe_negative marker, so the synchronous `import torch` runs at most once per venv lifetime. Invalidated on every path that can change the torch build (drift sync #307, repair sync, first-run sync, ROCm reinstall) and implicitly by a venv rebuild. A probe that fails to run cleanly is skipped WITHOUT caching so a transient error can't wedge a real CUDA machine. 3. Rewrite docs/install/troubleshooting.md §10 to the actual root cause: packaged installs never had the cudnn8_compat libs (so reinstalling never restored them); the bootstrap now installs them automatically on CUDA machines, with the manual uv pip command as the offline fallback and PyTorch Whisper as the sidestep. 4. Gate the ~700 MB nvidia-cudnn-cu12 download on the venv torch being a real CUDA build: the probe now reports 'hip' before checking cuda.is_available() (which HIP spoofs), so opt-in ROCm installs (#124) never fetch the CUDA wheel. Also reflow the CHANGELOG entry to house style (bold one-line lead, 1-3 lines of why, (#827, #869) refs) and extend the bootstrap unit tests: classify_cuda_probe verdict mapping and the marker write/invalidate round-trip (6 cuDNN tests total, 43 lib tests green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
83e71c5689 |
fix(asr): close the #730 residuals — chunked dub wedge shares the guarded reset; repeated timeouts recommend the crash-isolated engine (#895)
Residual A — the chunked dub-stream had a PARALLEL wedge mechanism (its own ping-loop timeout, its own _reset_pool_on_wedge, a dead-end "Try restarting the server" message). A wedged chunk now routes through the SAME run_transcribe_guarded bound+reset as the whole-file paths (#731/#851): the guard resets the poisoned pool once per wedged attempt (no double-reset on retry) and the user sees the actionable ASRTimeoutError. The reset logic is extracted to asr_backend.reset_pool_after_wedge — one shared mechanism, so the semantics can't drift again. run_transcribe_guarded also gains a timeout_env param so chunk errors name OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S instead of the whole-file knob. Residual B — the crash-isolated ASR sidecar (#393, faster-whisper-isolated) is wired as an explicit ESCAPE HATCH, not a default: - selectable end-to-end: Settings engine list gets an explanatory install_hint; honest gpu_compat ("cuda","cpu" — it wraps the same CTranslate2 engine as faster-whisper); get_active_asr_backend now hands back a process-wide singleton for subprocess-isolated backends (a fresh instance per request would leak atexit hooks and respawn the sidecar — reloading its model — on every transcribe). - on the SECOND consecutive guarded timeout-with-reset in one session (resets aren't recovering the hang; the wedged thread keeps its VRAM), the error the user sees + the log recommend switching to the isolated engine in Settings → Engines. Never auto-switched (owner rule: no silent behavior divergence); a completed transcribe resets the streak. Tests (fail-before/pass-after verified against origin/main): wedged-chunk SSE integration (reset count + actionable error + recommendation surfaces), consecutive-timeout streak (fires at 2, resets on success, suppressed when already on the isolated engine), timeout_env parametrization, shared-reset helper, isolated backend in list_backends with hint + honest availability, singleton caching, gpu_compat matrix entry. Docs: troubleshooting §14 gains the chunk knob + escape-hatch guidance. Closes the residuals tracked on #730. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
86f5213055 |
fix(splash): IPC-independent watchdog + recovery panel for dead Tauri IPC (#879) (#892)
After an unclean shutdown (Windows BSOD), the WebView2 profile cache (%LOCALAPPDATA%\com.debpalash.omnivoice-studio\EBWebView) can corrupt: Tauri's IPC custom protocol fails AND the postMessage fallback breaks, so invoke() hangs forever. useBootstrapStage's poll loop rode entirely on that IPC — a hung bootstrap_status call silently killed the loop and the splash sat at "preparing" forever, even with a fully healthy backend answering over plain HTTP. Class fix, three parts: - splashWatchdog.js: IPC-independent escape hatch. If no IPC signal arrives within 10s, poll GET /health over plain HTTP; healthy → proceed to the app as if 'ready' was received (console.warn breadcrumb so diagnostic bundles carry it). Any successful IPC response disarms it for good. - Recovery panel (stage 'ipc_lost'): if neither IPC nor HTTP succeed within 45s, show an actionable panel instead of the infinite spinner — "Open logs" (with an inline path fallback when IPC is dead) and, Windows-only and only in this error state, "Repair and restart". Health polling continues behind the panel so a slow first-run install with broken IPC still reaches the app. - clear_webview_cache_and_relaunch (Rust): writes a marker and relaunches; the fresh process deletes EBWebView at the top of run() before any webview exists (WebView2 holds locks while running), with a bounded retry while the old instance exits. Runtime cfg! guards keep the whole path compiling on every platform. Tauri 2 exposes no reliable flag for the postMessage-fallback mode (closure-local in its injected ipc.js), so the logged detector is the observable combination: zero IPC signals + working plain HTTP. Fail-before/pass-after regression tests: hung invoke + healthy HTTP → ready; hung invoke + dead backend → recovery panel, then auto-continue; working IPC → normal path untouched, zero HTTP polling. Plus watchdog state-machine unit tests and recovery-panel render/interaction tests (6/7 fail on the pre-fix component). Troubleshooting doc gains the matching section (docs-sync). Fixes #879 Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
be1ec3ade0 |
fix(platform): declare Intel-Mac local backend unsupported — honest first-run gate + docs (#889); Windows portable-install docs (#766 follow-up) (#891)
torch >=2.3 ships no macOS x86_64 wheels (transformers 5.x needs torch >=2.6), so `uv sync` can never resolve on an Intel Mac — per the platform-parity rule the honest option is declaring the platform unsupported, not letting first launch die in a raw resolver error: - bootstrap.rs: pre-check on macOS x86_64 before any venv create / uv sync (first-run AND repair paths) fails fast with an actionable message (remote-backend escape hatch + docs link); healthy pre-torch-bump venvs are deliberately untouched. Unit test pins the message's load-bearing phrases. - BootstrapSplash: routes the failure to a dedicated localized hint (bootstrap.hint_intel_mac, all 21 locales) and suppresses the useless Retry-oriented hints for it. - README + docs/install/macos.md (+ troubleshooting #9): every Intel-Mac support claim now says UI-installs-but-backend-cannot-run, including the from-source path (also broken); remote backend documented as the only use. - release.yml: #889 note on the macos-15-intel leg — artifact is UI-only; keep-or-drop is an owner call, deliberately not changed here. - docs/install/windows.md: new "Portable install (Windows)" section promised in #766 — custom MSI wizard folder / msiexec INSTALLDIR=..., what lives in OmniVoiceStudio-Data next to the exe, and the Program-Files-greyed-out why. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4eed552153 |
fix(engines): Confucius4-TTS validated E2E — clone sys.path import, 22.05 kHz, real install docs (#590) (#872)
Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3d0705fdb7 |
feat(engines): Confucius4-TTS — finalized (API-validated + unit-tested; opt-in, GPU run pending) (#590) (#637)
* feat(engines): Confucius4-TTS scaffold (opt-in, needs hardware validation) (#590) Plumbing for netease-youdao's Confucius4-TTS — LLM-based 14-language cross-lingual zero-shot voice cloning, Apache-2.0 — mirroring the opt-in subprocess-venv pattern of dots.tts / MOSS-TTS-v1.5: - engines/confucius4/__init__.py: Confucius4Backend(SubprocessBackend), CUDA-only (gpu_compat=("cuda",)), language passthrough, ref_audio→prompt_wav. is_available reports a clear reason and stays unavailable without a clone. - bootstrap.py: dedicated Python 3.10 venv resolution (user clone-level venv → package venv → uv bootstrap), import-probed on `confuciustts`. - main.py: sidecar speaking the same length-prefixed JSON-over-stdio protocol as the other engines, calling ConfuciusTTS(config_path, device).generate(text, lang, prompt_wav). - Registered lazily in _LAZY_REGISTRY; docs/engines/confucius4-tts.md. Gated behind OMNIVOICE_CONFUCIUS4_TTS_DIR — inert on every default install, never imports the upstream package unless opted in. The sidecar's synthesis API is derived from the upstream README and is NOT yet validated on a CUDA box; the module, docs, and CHANGELOG all flag this. 4 tests pin registration + inert-by-default. No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#590): register Confucius4 in install-hints + docs inventory (CI gates) Registering the engine tripped two completeness gates: every backend needs an install_hint (test_issue_fixes) and every registry engine must appear in the tts_engines docs inventory + README (check-docs-drift). Add the install_hint, the docs/features.yaml entry, and the README engine-table row (with the scaffold caveat). Docs-drift clean; gates pass. No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(confucius4): finalize — validate API vs upstream, add 22 sidecar unit tests, document external deps (Amphion/w2v-bert/weights) The synthesis API (ConfuciusTTS(config_path, device) → generate(text, lang, prompt_wav) → tensor, model.sample_rate) is confirmed against the netease-youdao/Confucius4-TTS repo. Added runnable unit tests for the sidecar's pure logic (language norm, tensor→PCM mono/stereo/clip, config resolution, wire framing, synthesize dispatch with the model mocked) — 22 cases, all green. Docs now list the external deps (Amphion/MaskGCT codec, facebook/w2v-bert-2.0, ~2-4GB HF checkpoint) and CUDA 12.6. Softened the scaffold warnings to reflect API-validated + unit-tested status; a one-time CUDA GPU run is still needed to confirm live inference + true sample rate. --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
29269b9cf0 |
feat: LLM Providers page + Autofit translation quality (fit-to-segment-time) (#838) (#854)
* feat(llm): multi-provider LLM registry + encrypted key storage + settings API (v0.3.8, phase 1)
Foundation for the LLM Providers settings page and timing-aware (Autofit)
translation. Every provider in the shipped .env is OpenAI-compatible, so one
client drives all of them via a registry instead of a class-per-provider.
- llm_providers.py: registry of 16 providers (OpenAI, OpenRouter, Groq,
Cerebras, Google AI, Mistral, Cohere, NVIDIA, GitHub Models, Cloudflare,
HuggingFace, SambaNova, SiliconFlow, + local Ollama/LM Studio + Custom).
Field resolution precedence env → encrypted store → default; active-provider
selection (LLM_DEFAULT_PROVIDER → stored → first keyed remote; local requires
explicit pick so we never assume a local server is up). Legacy TRANSLATE_*
maps to the Custom provider (keyless-with-base_url preserved).
- settings_store.py: generic ENCRYPTED secrets (get/set/clear_secret,
list_secret_names) reusing the HF-token Fernet path; get_text/set_text now
refuse the secret namespace (no ciphertext leak).
- llm_backend.py: OpenAICompatBackend resolves the active provider's
base_url/key/model from the registry. Backward-compatible.
- settings API: GET /llm-providers, PUT /llm-providers/{id} (encrypted key +
overrides), POST /llm-providers/active, POST /llm-providers/{id}/test.
Loopback-gated; never returns key material.
- 11 registry tests; existing llm-endpoint/openai-available tests still green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(settings): LLM Providers page — configure any provider's key/URL/model + Test + set active (v0.3.8, phase 2)
New Settings → System → LLM Providers pane (Brain icon, searchable). Lists all
16 registry providers; pick one to configure its encrypted API key, base URL,
model (and Cloudflare account id), Test the connection with one round-trip, and
'Save & use for translation' to make it the active provider for Cinematic/
Autofit. Keys are write-only from the UI (masked placeholder, never echoed);
env-set keys show as read-only. Local providers (Ollama/LM Studio) need no key.
- LLMProvidersPanel.jsx: provider selector + per-provider config + Test/activate,
following the LLMEndpointPanel pattern (apiJson/apiFetch/apiPost, SettingsSection
primitives).
- settingsCategories.jsx: new 'llm-providers' category under System + Brain icon.
- Settings.jsx: route the category to the panel.
- en.json: settings.llm_providers label.
- Frontend build passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(translate): Autofit quality style + one-click LLM setup from the dub menu (v0.3.8, phases 3-4)
Autofit = Cinematic + a strict 'never exceed the segment time' fit. The LLM
rewrites each translated line so its target-language reading time fits within
the slot, preserving the video timing without harsh audio time-stretch.
Backend:
- speech_rate.adjust_for_slot(strict=): strict caps the accepted upper ratio at
1.0 (fit within slot) vs Cinematic's 1.08; best-effort, degrades gracefully
with no LLM.
- dub_translate: quality='autofit' takes the LLM refine path and runs the fit
pass with strict=True; reports quality_used accurately.
- TranslateRequest.quality doc note.
Frontend:
- 'autofit' added to the quality control (Settings Translation + dub menu) and
the TranslateQuality type.
- Dub menu: picking Cinematic/Autofit with no LLM no longer dead-ends on a toast
— it offers a one-click 'Set up' that routes to Settings → LLM Providers,
with copy about fitting translations to segment time (#838).
- 4 strict-fit tests; frontend build green; i18n keys added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(translate): document Autofit quality + the LLM Providers page (v0.3.8, phase 5)
- CHANGELOG [0.3.8] Added: Autofit style + LLM Providers page.
- docs/dubbing/translation-engines.md: Fast/Autofit/Cinematic quality section
and an LLM Providers setup section (16 providers, encrypted keys, offline
Ollama/LM Studio, env overrides).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(api): add /api/settings/llm-providers routes to the route-inventory snapshot
Regenerated tests/fixtures/api_routes.txt for the 4 new LLM-providers endpoints
so test_route_inventory_matches_snapshot passes (keep-main-green).
* style(frontend): oxfmt the LLM Providers panel + dub quality control (format:check green)
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
e347f99542 |
fix(tts): bound + reset the GPU pool on a hung generate so it can't brick the backend (#730 class) (#851)
* fix(tts): bound + reset the GPU pool on a hung generate so it can't brick the backend (#730 class) A GPU job that wedges on some Windows+CUDA setups occupies its worker forever — run_in_executor can't cancel the thread — so on the 1–2 worker pools we ship, one stuck job starves every other request and the next action surfaces as the misleading "Can't reach the local backend" even though the process is alive. ASR/dub/model-load already bound+reset the pool on hang (#730). The TTS **generate** paths (generation.py, tts_stream.py) were the last unguarded GPU dispatch — and the residual on-main reports (#850 #802 #755 #723 #721, plus the 0.3.7 generate cohort) all fail on generate:start (audio). - model_manager: add run_on_gpu_pool_guarded() + GpuJobTimeoutError, a generalized version of the ASR guard so every GPU dispatch shares one bound+reset recovery path. Env-tunable via OMNIVOICE_GENERATE_TIMEOUT_S (default 300s). - generation.py: route both inference branches + the reference-clip transcribe through the guard; map a timeout to an actionable 503. - tts_stream.py: same guard on the streaming path (timeout → error frame). - test_generate_timeout_730: fail-before/pass-after regression (timeout resets pool + restores capacity, happy path, env override, no-reset exec). - docs + CHANGELOG: extend troubleshooting §14 to cover generate; document the new env var. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(tts): extend the GPU-pool hang guard to batch/dub/archetype/openai-compat generate (#730 class) The generate-hang class wasn't only in Studio + streaming: batch generate, the dub per-segment + preview generate, archetype preview render, and the OpenAI-compat /v1/audio/speech path all dispatched the TTS model to the GPU pool with no wall-clock bound either. Any one of them wedging on a Windows+CUDA hang starves the pool and bricks the backend the same way. Route all of them through run_on_gpu_pool_guarded so the whole class is closed — a hung generate anywhere resets the pool and returns an actionable timeout instead of a dead backend. Batch/dub recover per-segment on a fresh worker; drop the now-dead loop/_gpu_pool/asyncio locals ruff flagged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
522bbddccf |
feat(translate): highlighted Install affordance for uninstalled engines + dismissable/auto-clearing error banner (#847)
Two related Dub-tab translation-flow fixes, one PR. TASK 1 — proactive, highlighted Install affordance in the translate engine selector (replaces "find out only via a translate-time 400"): - FROM-SOURCE lane (activeEngineUnavailable && !enginesSandboxed): the muted install chip is promoted to a HIGHLIGHTED brand-accent Install button, still wired to handleInstallEngine(translateProvider) with the installing/disabled state. Selecting any uninstalled engine surfaces it immediately. - FROZEN lane (enginesSandboxed): pip install is impossible in the read-only, signed packaged env, so the disabled "needs dev install" span becomes an equally highlighted button opening a popover with (1) the exact install command + copy-to-clipboard, (2) one-click "Switch to Argos (bundled, offline)" — the guaranteed importable escape hatch, and (3) a Docs link via the existing Tauri shell.open path. Gated on the existing `sandboxed` flag, not platform. - Single-source install command: new translation_engines.install_command() is the one source of truth; list_engines() stamps `install_command` per engine and BOTH the argos + deep_translator translate-time 400 messages build their command from it, so the proactive button and the 400 can't drift. engines.ts gains `install_command: string | null`. TASK 2 — the translation error banner now dismisses and clears (class fix): - Root cause: handleTranslateAll never cleared dubError, so a stale 400 survived even a successful retry. It now clears at the start of every attempt. - Corrective-action clears (whole class): changing the engine and installing the package both clear dubError (wrapped setTranslateProvider + handleInstallEngine in DubTab). - DubFooter's banner gains a × dismiss and a guarded auto-timeout (skipped while generating so live per-segment errors persist). i18n: 8 new dub.* keys translated across all 21 locales. Docs: new docs/dubbing/translation-engines.md (from-source vs packaged build) linked from the popover Docs button + a troubleshooting cross-reference. Tests: FE regression for both lanes + never-installs-when-sandboxed + banner dismiss/auto-clear; BE regression that list_engines() install_command is embedded verbatim in the dub_translate 400s. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cb70c2b1af |
feat(ui): back Input/Select/Textarea/Slider with shadcn (prop APIs preserved) (#798)
P1 of the shadcn/ui primitive migration (docs/shadcn-migration.md): route the
OmniVoice form/data primitives through the shadcn components in
src/components/ui/* while keeping their exact exports and prop APIs, so no call
site changes.
- input.tsx: export `inputBaseClass` (the shell) with no behaviour change —
ShadcnInput baseline stays byte-identical.
- New shadcn components: textarea.tsx, select.tsx (+@radix-ui/react-select),
slider.tsx, table.tsx.
- src/ui/Input.jsx (Input/Textarea/Select/Field): Input/Textarea now render the
shadcn components; a small `fieldSizeVariants` cva (named palette utilities,
tailwind-merge-clean) restores the OmniVoice padding-based sm/md/lg scale +
filled bg-bg-elev-2 over the shell. Select stays a NATIVE <select> wearing the
same shell — DubSegmentTable/CompareModal/GeneralTab depend on
onChange={(e) => …e.target.value}, which Radix's value-only Select would break;
the Radix select.tsx is added for new call sites only.
- src/ui/Slider.jsx: wraps the shadcn Slider, keeping the number-based onChange +
label/value-bubble chrome; track/thumb sized via the data-slot selectors.
- Table deliberately NOT rerouted: ui/Table.jsx is a flex-<div> chrome wrapper
whose .ui-table*/.segment-table global classes are a SHARED CONTRACT used
directly by ModelsTable/DubSegmentTable/EngineCompatibilityMatrix (virtualised
lists needing the div/flex layout, not a semantic <table>). table.tsx is
provided for new tabular data; Table.jsx + its globals are untouched. Its
toolbar inherits the shadcn-backed Input/Button for free.
Verification: only the 3 Input-* visual baselines moved (palette-coherent across
default/midnight/catppuccin); Slider/Table stayed within tolerance. vitest 641
green, oxlint 0 errors, oxfmt --check clean, vite build green, root bun.lock
regenerated and bun install --frozen-lockfile in sync (Docker).
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
200a559183 |
feat(ui): shadcn/ui foundation + OmniVoice palette token bridge (Button/Input proof) (#797)
Lay the foundation for migrating OmniVoice's UI to clean Tailwind v4 + shadcn/ui WITHOUT changing the look: shadcn primitives inherit the existing OmniVoice palette (Gruvbox-pink default + every [data-theme] variant) through a semantic token bridge. Foundation only — no existing component is replaced. What landed: - shadcn init for Tailwind v4 + Vite + React 19: frontend/components.json (new-york, rsc:false, tsx:true), src/lib/utils.ts (cn = clsx + tailwind-merge), and a @/* -> src/* alias in vite.config.js + tsconfig.json so future `npx shadcn add` resolves. - Token bridge in src/index.css: a single `@theme inline` block maps shadcn's semantic vocab (--color-background/-foreground/-card/-popover/-primary/ -secondary/-muted/-muted-foreground/-accent-foreground/-destructive/-input/ -ring + --radius) onto the existing OmniVoice --color-* tokens. Because those tokens are re-declared per theme in ui/themes.css, theme switching recolors shadcn components automatically — no per-theme shadcn block. Existing --color-accent/--color-border and the --radius-* scale are left intact. - Two proof components: src/components/ui/button.tsx + input.tsx (verbatim shadcn new-york), rendered across default/midnight/catppuccin in the visual harness with committed baselines (brand-pink / purple / lavender confirmed). - New deps: class-variance-authority, clsx, tailwind-merge, tw-animate-css, @radix-ui/react-slot. Root bun.lock regenerated; `bun install --frozen-lockfile` verified in sync (Docker-green). - Migration plan at docs/shadcn-migration.md (bridge table, primitive->shadcn mapping, prop-compat wrapper strategy, staged waves, honest risk/effort). Verified: vite build, typecheck:ci, oxlint (0 errors), oxfmt --check, vitest (641 pass), test:visual (48 pass incl. 6 new baselines), frozen lockfile in sync. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
721cb34a9d |
docs: add the CSS → Tailwind v4 migration plan (#772)
Phased, bounded migration plan (not a big-bang): convert the mechanical ~80% (flex/grid/gap/padding/typography/simple color) to Tailwind v4 utilities, deliberately keep ~15-25% as CSS (glass/backdrop-filter, @keyframes, ::before/::after, :has(), !important). Realistic end state ~10-12k of 16.6k CSS lines removed across ~5-7 weeks of small PRs. Key gates the plan establishes before any conversion starts (P0): - A Playwright screenshot baseline (default + dark + light) — the className-diff trick used for the page refactors is useless here since class names change. - Fix the @theme ↔ tokens.css token drift (single source + a parity test). - Rewrite the CONTRIBUTING.md "no Tailwind" line (docs-sync rule). Companion to docs/maintenance-pages-modularization.md. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f33bdc731d |
refactor(settings): modularize Settings page (1969→399 lines, all files under 500) (#758)
* refactor(settings): extract Settings.jsx tabs into components/settings (1969→602 lines) Settings.jsx had grown to 1969 lines — every edit reloaded the whole file into context and risked unrelated breakage. This finishes the migration the existing components/settings/*Panel.jsx pattern started: the page is now a thin orchestrator and each heavy tab lives in its own file. Extracted (logic byte-for-byte identical; only import paths adjusted + the shared isTauri/askConfirm moved to components/settings/native.js): - GeneralTab, ModelStoreTab, EnginesTab, HotkeyTab, CredentialsTab - native.js — shared isTauri() wrapper + askConfirm() Tauri-dialog helper Also establishes the standard so files can't silently regrow: - CONTRIBUTING.md: frontend file-structure & size limits (soft 300 / hard 500) - eslint.config.js: warn-only max-lines:500 guardrail (CI stays green) - docs/maintenance-pages-modularization.md: the phased refactor plan Verified: vite build passes (all imports resolve); 18/18 settings tests pass; no new lint errors introduced (the pruned imports were the only regressions). Follow-ups (tracked in the plan doc): ModelStoreTab.jsx is 836 lines and Settings.jsx 602 — both still over the 500 cap (warn-only); split next. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(settings): split ModelStoreTab + Settings.jsx under the 500-line cap Follow-up to the tab extraction: bring the two remaining over-cap files into compliance with the new standard. Pure-mechanical, no behavior change. Settings.jsx 602 → 399: - Extract AboutTab, PrivacyTab, LogsTab into components/settings/ - Move the shared Row helper to components/settings/Row.jsx - LogsTab keeps its state in Settings() (lower-risk); About/Privacy take props ModelStoreTab.jsx 836 → 439, split into components/settings/models/: - format.js (fmtBytes/orgColor), runtime.js (computeRowRuntime) - columns.jsx exposes makeModelColumns(...) — a factory so the TanStack cell closures keep working; called with the same useMemo dep array as before - ModelsTable.jsx (virtualized table view), RecoBanner.jsx Every settings file is now under 500 lines. Verified: vite build passes; 18/18 settings tests pass; no new lint errors (the 4 remaining in Settings.jsx are pre-existing — refreshInfo no-op, a catch(e), two set-state-in-effect). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
de3d83f14b |
docs: add rust as prerequisite for from-source builds (#704)
Adds Rust/Cargo as a from-source build prerequisite across the linux/macos/windows install docs. Thanks @Deepakv2104. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
25105605f1 |
docs(specs): ElevenLabs-parity roadmap + Tier-1 implementation specs (#684)
Add the implementation-ready spec set mapping OmniVoice to ElevenLabs parity while preserving local-first: - 00-roadmap-elevenlabs-parity.md — gap analysis, prioritized tiers, sequencing, prior-art reconciliation, and the deliberate "won't build" list. - 01-expressive-tts.md — engine-agnostic emotion/style intent lowered onto each TTS engine's real mechanism (degrade-visibly) + a DB-backed pronunciation dict. - 02-conversational-agent.md — fully-offline full-duplex voice agent (/ws/converse, Silero-VAD barge-in on AEC-cleaned mic) composing existing streaming STT/TTS + LLM. - 03-longform-studio-editor.md — per-segment edit/regenerate across dub/audiobook/ stories, extending the existing content-addressed cache to longform. Reconcile prior planning docs: banner the superseded parity/studio docs pointing here; keep distinct-scope docs untouched (classification table in 00). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
022a3bd6b9 |
feat(dictation): live local dictation via sherpa-onnx + Voice settings panel (#683)
* feat(dictation): live local dictation via sherpa-onnx + Voice settings panel Add a sherpa-onnx ASR engine alongside the existing Whisper/NeMo dictation path, powering a genuinely live experience: as you speak, words type straight into the focused field (streaming partials via a new simulate_type command, self-correcting with backspaces) and commit per pause. Backend: - SherpaDictationBackend + sherpa_dictation registry of the 7 models (Parakeet TDT v3/v2, streaming Zipformer EN/ZH/bilingual, Paraformer bilingual, Whisper Tiny) from csukuangfj/* int8 HF repos; CPU provider for cross-platform parity. - /dictation/models + /dictation/prefs router; get_capture_asr_backend() honors the selected dictation model. get_active_asr_backend() (dub transcription) and the legacy WebM/Opus capture path are untouched. - True streaming over /ws/transcribe (OnlineRecognizer: live partials + per-endpoint finals); offline models surface partials via short re-decode. Frontend: - New "Voice" settings panel (enable, Toggle/Hold mode, model picker with offline/streaming/recommended badges + per-model download/delete). - Live word-by-word typing via simulate_type (enigo) with prefix-diff delta and backspace correction; paste fallback retained, no double-insertion. Deps: sherpa-onnx>=1.13.3 (+ sherpa-onnx-core); uv.lock regenerated, Docker frozen-install verified. API route-inventory snapshot updated. 40+ new tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dictation): register sherpa-onnx-asr engine in README + features inventory Fixes the docs-drift CI guard: the new sherpa-onnx-asr ASR engine existed in the registry but not in docs/features.yaml or README. Adds the live-dictation engine row to the ASR Engines table, bumps the engine counts (8→9), and adds the inventory entry. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): fold live-dictation into the [0.3.8] section main is 0.3.8 (untagged), so the dictation feature belongs in that release, not a separate [Unreleased] block. Merge the two Added lists under one [0.3.8], refresh the headline to lead with live dictation, and correct the capture description to reflect live word-by-word typing (not paste-on-pause). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2339cc8e85 |
fix(model): actionable "reinstall transformers" hint on a corrupted-install model-load error (#676)
A model load failed with `[Errno 2] No such file or directory:
'…/site-packages/transformers/models/qwen3/modeling_qwen3.py'` — the user's
transformers install was incomplete (the file is missing while a correct 5.3.0
install has it; an interrupted `uv sync` / antivirus / partial update drops it).
The System Check showed the raw path + "Check logs and try restarting", which is
useless — restarting can't restore a missing file.
Two fixes:
1. core.failure.classify(): recognize this corrupted-install variant. It's a
FileNotFoundError, not an ImportError, so the existing TRANSFORMERS_IMPORT
match ("could not import module"/"AutoFeatureExtractor") missed it. Now also
matches a "no such file"/"errno 2" + "transformers" + "site-packages" signal
(substrings checked separately so it works on POSIX `/` and Windows `\`
paths). An unrelated package's missing file is NOT mislabelled.
2. model_manager._load(): build the /model/status error via build_failure so it
carries the classified hint AND strips the home dir, instead of storing the
raw str(exc). The System Check now shows "Your transformers install is
incomplete. Reinstall it (uv pip install --reinstall transformers) or switch
ASR to faster-whisper" — the existing TRANSFORMERS_IMPORT hint.
Docs: troubleshooting §1a documents the error + the reinstall fix.
Test: test_failure_classify.py pins the POSIX + Windows path forms classify as
TRANSFORMERS_IMPORT with a "reinstall" hint, and that an unrelated package's
missing file does not.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
b7cecde57e |
feat(setup): faster downloads by default + prominent, encouraged HF-token entry (#669)
Two changes that make first-run downloads faster and easier to speed up further.
1. Segmented (multi-connection) downloader is now ON by default. The app forces
the legacy-LFS path (HF_HUB_DISABLE_XET=1) for clear progress, but that path
is single-stream and slow — which is why downloads felt sluggish. The built-in
IDM/uGet-style segmented accelerator (parallel byte-ranges, live speed/ETA)
was already implemented but defaulted OFF. Flip it ON: it only engages when
Xet is inactive (the default), and ANY failure falls back to snapshot_download
("can never compromise a correct install"). Pure-httpx, cross-platform,
auth-safe (token never forwarded to a CDN). Override with
OMNIVOICE_SEGMENTED_DOWNLOAD=0.
2. The Hugging Face token field is now a prominent, always-visible card right
above Continue — was a collapsed "advanced" fold almost nobody opened. A free
token gives authenticated downloads (higher rate limits, fewer stalls), so it
pairs with change #1 to keep the parallel fetch from getting throttled. The
card leads with the speed benefit, shows a saved-state, and adds a one-click
"Get one free →" link to huggingface.co/settings/tokens.
Docs: downloading-models.md updated — the legacy-LFS section now documents the
default-on segmented accelerator + the HF-token speed tip, and the tuning table
reflects OMNIVOICE_SEGMENTED_DOWNLOAD=0 as the disable knob (docs-sync).
Test: test_segmented_download_default.py pins the new default ON and that the
env override still disables it; existing FDL-08 behavior tests stay green.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
252f0d4fac |
fix(asr): bound whole-file transcription so a stall isn't reported as "can't reach backend" (#656)
A Windows/CUDA user (Vietnam) hit "Can't reach the local backend" only when dubbing/transcribing. Their log proves the backend started fine — model loaded, preload complete, 25 models — and the log ends right after `whisperx transcribing …tmp.wav`. The backend was alive; the *transcription* stalled (large-v3 ASR contending with the resident TTS model for VRAM on an 8 GB-class GPU), which the UI surfaces as an unreachable backend. Root cause (class, not instance): the chunked dub pipeline already bounds each chunk (OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S), but the *whole-file* transcribe paths ran unbounded: - dub QC re-transcribe (dub_export) - dictation (capture) - OpenAI-compat /audio/transcriptions A slow/stuck transcribe on any of these hung the request AND held a GPU-pool worker — indistinguishable from a dead backend. Fix: add run_transcribe_guarded() in services/asr_backend.py — a shared asyncio.wait_for wrapper (ASRTimeoutError, a TimeoutError subclass) with a generous env-tunable bound (OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S, default 300 s). On timeout the request returns 504 with actionable guidance (backend is alive; free VRAM / pick a smaller ASR model / use CPU; restart to clear the stuck worker) instead of hanging forever. Wired into all three whole-file paths. Docs: new troubleshooting §14 — "Can't reach the local backend during transcription/dubbing" — explains it's ASR weight/VRAM pressure, not a network/ mirror problem, and corrects the misconception that a "Network → Restricted/Global mirror" Settings toggle exists (the Network control is LAN sharing). Serves the #602/#585/#567 "can't reach backend" cluster. Test: backend/tests/test_asr_transcribe_timeout.py — slow fn raises ASRTimeoutError with the actionable message, fast fn passes through, subclass-of-TimeoutError so the openai_compat broad catch still maps to 504. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
79f3e35682 |
docs(troubleshooting): add stuck-download / incomplete-cache recovery (#622) (#643)
The 'stuck on the download page, model folder has only refs/ no weights' case (a connection dropping/blocking mid-pull) is a recurring support report but wasn't in the install troubleshooting guide. Add section 13 with the recovery steps + antivirus/VPN/mirror escalation + a huggingface-cli manual fallback. Docs-only; no version bump. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0e17caa52a |
fix(install): actionable torch-wheel-download failure + local-wheel recovery (#569) (#574)
#569: on a restricted network the first-run install fails downloading the ~2.5 GB cu128 PyTorch wheel from download.pytorch.org, and the app won't launch. Two problems: the error told users to "set UV_DEFAULT_INDEX to a mirror" — which CANNOT redirect torch, because it comes from a *named, explicit* uv index (uv 0.11 rejects index-name override values and `--frozen` pins the exact wheel URLs); and there was no way to supply a manually-downloaded wheel. - Detect a torch/pytorch-host `uv sync` failure and emit torch-specific guidance (Clean & Retry → VPN → drop the wheel locally) instead of the wrong mirror advice. - Add a local wheel-drop dir `<env_root>/wheels` (survives Clean & Retry) wired via `UV_FIND_LINKS`. On a frozen-sync torch-download failure WITH wheels present, retry NON-frozen with find-links so uv re-resolves from the local wheels. Verified empirically: a non-frozen find-links sync installs from a local wheel fully offline, while a `--frozen` sync ignores find-links — so the retry is the only mechanism that can consume a dropped wheel. Best-effort: if it can't satisfy, it fails identically to before and the actionable error still fires. - docs/install/troubleshooting.md: new "#12 CUDA PyTorch wheel download fails" entry (docs-sync) — the offline wheel-drop path + why a PyPI mirror can't fix this index. Note: an automatic mirror redirect for the cu128 index is intentionally NOT shipped — uv provides no working override for a named explicit index, so it couldn't be verified; the offline wheel path is the reliable escape hatch. Test: sync_failure_is_torch_download host/keyword detection + negative guard. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3777d3a62c |
feat(tts): add MOSS-TTS-v1.5 (8B) and dots.tts (2B) as opt-in engines (#498)
Adds two zero-shot voice-cloning TTS engines requested in #498, both opt-in and subprocess-isolated with their own dedicated venv — the same pattern as IndexTTS-2. The dedicated venv is forced, not just chosen: each upstream pins a transformers version that conflicts with the parent's >=5.3 (MOSS-TTS-v1.5 ==5.0.0, dots.tts ==4.57.0), so they cannot share the parent interpreter. Because they use the clone+venv bootstrap (env var -> clone -> uv venv), this touches no pyproject.toml / uv.lock / bun.lock — `uv sync --all-extras` and Docker's `bun install --frozen-lockfile` are unchanged, so main's CI/Docker matrix stays green. Engines: - moss-tts-v15: 8B, 31 langs, ~16 GB weights, 24 kHz. AutoModel/ AutoProcessor via trust_remote_code. gpu_compat=(cuda,cpu) — MPS is undocumented/untested upstream so it is never claimed; on a Mac it runs on CPU. Apache-2.0, no license gate. - dots-tts: 2B, 24 langs, ~9 GB weights, 48 kHz. DotsTtsRuntime; continuation cloning (prompt_audio_path+prompt_text). Upstream is Linux/macOS-only, so is_available() gates it off cleanly on Windows (cross-platform parity rule — it is opt-in, never a broken default). Wiring: registered in _LAZY_REGISTRY + _INSTALL_HINTS. list_backends() surfaces both as subprocess/[cuda,cpu]/available-until-installed; the data-driven Settings engine picker needs no frontend change. Tests (19, fail-before/pass-after): registry resolution, subprocess marker, no-MPS gpu_compat, the Windows gate, not-installed honesty, and the parent-side generate() kwarg arbitration. Existing engine suite still 55 passed / 5 skipped. Sidecar inference follows the upstream-documented APIs but, like IndexTTS/Supertonic, can't be executed in CI without the multi-GB model clones. Docs (same-PR per docs-sync rule): README + README_CN engine tables, new docs/engines/moss-tts-v15.md + dots-tts.md, disk-usage.md (torch-dedup note), CHANGELOG. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |