d23e56a5ecee53da0460ea9da2cbe83dbb4eae03
68
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d23e56a5ec |
fix(errors): the transformers-import advice names torchvision now (#1376) (#1377)
The TRANSFORMERS_IMPORT hint and the ASR pipeline error told users to reinstall torch + torchaudio + transformers. torchvision — the package whose ABI mismatch actually produces this exact lazy-import wording (#1357's torchvision::nms, wrapped into "Could not import module 'AutoFeatureExtractor'") — was the one package the advice omitted. Following it to the letter left the broken package untouched (#1376). Both surfaces now name the mismatch as a cause and prescribe the pinned reinstall with literal versions (desktop installs ship no deploy/, so the constraint-file form fails there) targeting the venv explicitly. A lockstep test asserts the exact command on every advice surface against deploy/torch-constraints.txt, so a pin bump stays red until the advice matches. Docs gain the same-wording-different-cause section (1a-bis). |
||
|
|
810b598739 |
fix(appimage): let the host GStreamer win, and stop sharing its registry (#1333) (#1354)
* fix(appimage): let the host GStreamer win, and stop sharing its registry (#1333) Recording from the AppImage failed with "No microphone found" on a Debian 13 host whose audio stack the reporter verified healthy (pactl, wpctl, gst-launch with both pulsesrc and pipewiresrc), while the same build`s raw binary recorded fine. GST_DEBUG=2 named it: WARN GST_REGISTRY gst_registry_binary_check_magic: Binary registry magic version is different : 1.23.90 != 1.3.0 GStreamer element appsink not found. Please install it. linuxdeploy bundles libgstreamer-1.0 because WebKit links it, but not the plugins: those are dlopen`d, so nothing static can see them to copy. The bundled core falls back to the host plugin directory, whose plugins were built against the host core, the version check rejects them, and the scan yields nothing. appsink is one of the casualties and it is the element WebKit hands a capture stream to, so getUserMedia() rejects NotFoundError. Same class as #1258 (frozen bundled library against a host that moved on) in a different library, which is why OMNIVOICE_PREFER_SYSTEM_WEBKIT=1 did nothing for the reporter. Since we ship no plugins, the host core is the only one that can agree with the plugins that will load — so prefer it, with OMNIVOICE_PREFER_SYSTEM_GSTREAMER=0 as the escape hatch. Also isolate the registry cache. GStreamer keys ~/.cache/gstreamer-1.0/ registry.<arch>.bin by architecture alone, so two cores of different versions clobber each other`s file: that makes the failure depend on which app ran last, and the AppImage corrupts the cache for every other GStreamer app on the machine. Both directions go away with a private path. AppRun.test.sh covers host-present, host-absent and opt-out; all three fail against the previous AppRun. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(appimage): cover the ldconfig discovery path; docs fixes CodeRabbit, all three valid: - every GStreamer case forced ldconfig to fail, so the runtime-only-host fallback (no -dev package, hence no .pc file) was never exercised. The cases now select their discovery path, and the new ldconfig one fails if that branch is removed. - MD040: the GST_DEBUG fence had no language tag. - the registry cache path follows XDG_CACHE_HOME when set; ~/.cache is only the default. Documented, along with WHY the shared file is a problem. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(appimage): compose LD_LIBRARY_PATH once; host WebKit stays first CI caught a real regression, not a flaky test. The GStreamer block prepended its own directory, which put it AHEAD of the host WebKit dir — and "host WebKit first" is the invariant #1258 turns on. On a host where the two libraries live in different directories that silently changes which WebKit resolves. It only showed on Linux because the WebKit ldpath cases do not stub away a real host GStreamer, so the runner had one to find and macOS did not. Reproduced locally with an ldconfig shim, and confirmed the ordering is what fixes it: with the old order the suite is 19/2, with this one 21/0. Both decisions now compose one path in one place — host WebKit, host GStreamer, bundle, inherited — so neither preference is weakened and the ordering is stated where it is applied rather than implied by two independent prepends. Same directory for both (the common case) is not listed twice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(appimage): preload the host GStreamer instead of hoisting its libdir greptile P1, valid. The host GStreamer lives in a general system library directory (/usr/lib/x86_64-linux-gnu on Debian), so putting that directory ahead of ${HERE}/usr/lib replaced EVERY other bundled library with the host copy — loader symbol errors, startup crashes, or a blank window on a distro we never built against. One library needs to come from the host and the mechanism has to be that narrow. LD_PRELOAD names exactly that library and leaves the search path alone, so the WebKit ordering from #1258 is untouched too (and this removes the composed-LD_LIBRARY_PATH block that only existed to keep the two prepends from fighting). The preload is inherited by the Python backend, where nothing links GStreamer and it is inert — the accepted cost. Tests now assert both halves: the library IS preloaded, and the libdir is NOT hoisted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(appimage): verify the host GStreamer loads before preloading it greptile P1, valid. The host core links GLib and the bundle ships GLib too, resolved bundle-first — so a host GStreamer built against newer GLib than we bundle fails its relocations and the app does not start at all. That is strictly worse than the broken microphone this PR fixes. Taking host GLib as well is not an option either: GLib is what WebKit is built against, so pulling it from the host reopens #961/#1258. Rather than predict the pairing, test it. The loader processes LD_PRELOAD for any binary, so running `true` under the exact environment the app will get is a complete check of whether the library loads there — a missing dependency or an unresolved version tag ("version GLIB_2.84 not found") fails it and nothing else runs. On failure the preload is skipped, the app starts on the bundled core, and a warning names the mismatch so the user has a thread to pull rather than a silent half-fix. OMNIVOICE_APPRUN_PRELOAD_PROBE lets the suite choose the outcome, matching the existing OMNIVOICE_APPRUN_WK_MARKER precedent; the new case fails if the guard is removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
168e8e5c61 |
fix(macos): declare the floor the app actually delivers (13.3, not 12) (#1314)
* fix(macos): declare the floor the app actually delivers (13.3, not 12) The app declared minimumSystemVersion 12.0 and the docs promised Monterey, while the frontend required Safari 16.4 in three independent places: Vite's default build target (baseline-widely-available = safari16.4), Tailwind v4's own documented floor, and `@property` throughout its generated utilities. On Monterey's WKWebView 15.6 the focus ring and accent surfaces resolve invalid, and a bundled dependency ships a RegExp lookbehind that is a PARSE-time SyntaxError no polyfill can reach. Option B — actually supporting 15.6 — means setting build.target back, replacing 64 color-mix() calls, dropping Tailwind v4 and replacing that dependency, indefinitely, for an OS that stopped receiving security updates in late 2024. The council was unanimous on A, and the precedent is uniform (Chrome 117, Electron 27, VS Code, Firefox 116). minimumSystemVersion is also the guard: macOS itself refuses to launch a bundle below it, so a Monterey user gets an explicit OS refusal rather than an app that opens to a blank window — which matters because the Tauri updater has no per-OS gating of its own. Docs updated in the same change (README support table, docs/install/macos.md) and the webCompat floor assertion re-derived to 16.4, so the post-floor API list must be revisited the next time the floor moves. Closes #1268 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(macos): raise the floor in the macOS overlay too, and assert it Greptile P1, and correct: Tauri merges tauri.macos.conf.json OVER the base config for a macOS build, and that file carried its own minimumSystemVersion: 12.0. Changing the base config alone decided nothing — the shipped bundle would have stayed Monterey-installable while the base config, the README and the install docs all said 13.3. Worse, the guard I added read only the base config, so it would have gone on passing. A test that validates the wrong file is not a guard; it now asserts both, with a comment saying why the overlay is the one that ships. Also per review: the changelog entry was an editorial paragraph rather than a one-line entry, and the section was missing ### Docs. Both fixed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(webcompat): the module header still described the old 12.0 floor The floor moved to 13.3/Safari 16.4 in this PR and the test was re-derived, but webCompat.js still told the next reader the oldest supported WebView was 15.6 — which would make every fill here look mandatory instead of retained for Linux's unpinnable WebKitGTK (CodeRabbit). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
15afc6611d |
fix(linux): AppImage blank window on Mesa 26.1+ hosts (#1258, #1244) (#1265)
* fix(linux): AppImage blank window on Mesa 26.1+ hosts (#1258, #1244) The AppImage bundles an Ubuntu-built WebKitGTK but ships no libEGL, so that bundled WebKit runs against the HOST's Mesa. On Mesa >= 26.1 it calls eglGetPlatformDisplay() in a way the newer driver rejects and the app dies before it paints: Could not create default EGL display: EGL_BAD_PARAMETER. Aborting... No environment variable helps, because the failure is in EGL display creation — before WebKit consults any rendering-path flag. #1258 confirmed WEBKIT_DISABLE_DMABUF_RENDERER, WEBKIT_DMABUF_RENDERER_FORCE_SHM, WEBKIT_SKIA_ENABLE_CPU_RENDERING, EGL_PLATFORM=surfaceless and MESA_LOADER_DRIVER_OVERRIDE=swrast all fail identically. Chasing the build runner's WebKit (#961 bumped 22.04 -> 24.04) cannot fix this class: what we bundle is frozen and host Mesa keeps moving. So when the host has a WebKitGTK at least as new as ours, let it win — the bundle still fills every gap, and a host without WebKitGTK is untouched. That is exactly why building from source works on the hardware where the AppImage does not. The compositing workaround is re-decided against whichever library ends up running, and AppRun.test.sh — which had never been wired into CI — now runs there, so this logic stops being a regression test nothing executes. * fix(review): the ordering change was a no-op; name the host libdir explicitly CodeRabbit Major — correct, and it made the whole fix inert. LD_LIBRARY_PATH is searched AHEAD of the linker's default paths no matter where in that variable a directory sits, so on a normal launch (empty LD_LIBRARY_PATH) the bundle remained the only explicit search directory and still won. Merely appending it changed nothing. The host's WebKit libdir is now named explicitly, ahead of ours. The new tests fail 3/3 against the previous version. Greptile P1 — a host with the runtime but no -dev package has no .pc file, so pkg-config can't answer and the check rejected a perfectly good system WebKit. The libdir probe now falls back to ldconfig, and OMNIVOICE_PREFER_SYSTEM_WEBKIT gives those users an explicit opt-in (=0 opts out) rather than gambling on an unverified version, which would risk the #961 regression. CodeRabbit — my changelog script had also inserted the CI entry into the published 0.4.0 section. Removed; it belongs only under Unreleased. CodeRabbit — the docs' source-build fallback used 'cd frontend', not the repo-root flow the rest of the page documents. Fixed. |
||
|
|
9736fd4859 |
release: v0.4.1 (#1239)
* release: v0.4.1 Seven user-reported issues fixed since v0.4.0 (#1221–#1229). Version bumped across the single source of truth (frontend/package.json) and its three toolchain mirrors; [Unreleased] renamed to the release section that release.yml extracts verbatim as the GitHub Release body. Docker tag examples in docs/install/docker.md and deploy/dockerhub-overview.md updated to 0.4.1 (docs-sync rule). * release: #1239 review — sync Cargo.lock to 0.4.1 Greptile: the manifest said 0.4.1 while Cargo.lock still recorded 0.4.0, so a `cargo build --locked` (and the Tauri bundler's own locked build) would fail on the mismatch. Regenerating locally updated it but it was never staged. * release: re-sync [0.4.1] after the fix merges, date it 2026-07-27 Picks up everything merged since the section was first written: the first-run wizard chrome (#1241), the MCP host allowlist (#1249), the macOS 12 startup crash (#1245), and the six error-message fixes (#1247, #1251, #1254, #1256, #1257, #1262). Deliberately NOT included: - the dub delete-resurrection fix (#1252, #1253) — split to #1270 after it needed six rounds of correction, the last two finding that the fix did not close the reported case and that its own bound reintroduced it; - the Linux AppImage WebKit fix (#1258, #1244) — held on #1265 pending confirmation on a Mesa 26.1 host, which nobody has run. |
||
|
|
d800f77e26 |
fix(rocm): #1228 review — a remap is only a fix if the build ships the target
Two P1s, both correct: - Being IN the override map was treated as proof of compatibility. If the wheel ships neither the native arch nor the remap target, setting HSA_OVERRIDE_GFX_VERSION only changes WHICH kernel is missing — gfx1151 with a gfx1030-only build was routed to the GPU and would fail at launch. Both arch_unsupported() and _configure_rocm_if_needed() now require the target to be present, and fall back to CPU otherwise. - An EMPTY arch list means the build's metadata is unavailable, not that the GPU is unsupported. The remap branch read that unknown state as a confirmed mismatch and would push a natively-supported gfx1151 onto foreign gfx1100 kernels. Now fails open and changes nothing, matching the fail-open contract the rest of the probe follows. ROCM_GFX_OVERRIDES values are now the target gfx NAME rather than the HSA version string, so the "is the target present?" check is a direct membership test; hsa_override_for() derives the env-var form, covered by a test that every entry in the map converts cleanly. Also: CHANGELOG entries shortened with refs last, and the MD028 blank line between the two docker.md blockquotes (CodeRabbit). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
791dae69db |
fix(rocm): stop force-routing every AMD GPU to the CPU (#1228)
The GPU compatibility gate built a CUDA-namespace tag from `get_device_capability()` (`sm_115` on a gfx1151 Strix Halo) and looked for it in `torch.cuda.get_arch_list()` — which on a ROCm wheel returns gfx *names* (`gfx1100`, `gfx1151`, …). The two namespaces can never intersect, so `check_device_compatibility()` returned False on every ROCm build and `get_best_device()` silently returned "cpu": torch saw the GPU, `torch.cuda.is_available()` was True, and the app ran on the CPU anyway. The comparison was copy-pasted in three places, all with the same bug, so it now lives once in `core.device_caps.arch_unsupported()` and branches on the build (gfx names on ROCm, sm_/compute_ tags on CUDA): - `model_manager.check_device_compatibility` — the CPU force-route, plus a ROCm-specific remedy instead of telling AMD users to install a cu128 wheel - `device_caps._probe` — the kernel-risk note that downgraded the routing badge - `engine_env._cuda_arch_supported_for_compile` — torch.compile off on all AMD `_configure_rocm_if_needed` also applied `HSA_OVERRIDE_GFX_VERSION` from a static map without checking whether the GPU needed it, remapping cards the installed build supports natively onto foreign kernels. It now applies only when the native GFX ID is genuinely absent from the arch list, and knows gfx1150/gfx1151 (Strix Point/Halo). Regression test: tests/test_rocm_arch_gate.py pins the reporter's host resolving to "cuda", a genuine ROCm mismatch still being caught, the CUDA path (#756 Blackwell fallback) unchanged, and the narrowed override. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f92e4223db |
docs(docker): refresh registry tag examples 0.3.22 → 0.4.0
Update the exact-version / minor / ROCm pin examples in the Docker Hub overview (deploy/dockerhub-overview.md — source of the hub.docker.com page, re-synced on this main push) and docs/install/docker.md to the v0.4.0 release. GHCR's package page inherits the repo README and the current org.opencontainers.image.description label, both already accurate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
433d4bc659 |
Merge remote-tracking branch 'origin/main' into fix/1177-surface-backend-diagnosis
# Conflicts: # CHANGELOG.md |
||
|
|
fc7fbf1227 |
fix(gpu-pool): bound execution, not queue wait (#1190, #1202)
A job queued behind a busy 1-worker pool burned its whole 300s budget without executing an instruction, then reported "too heavy for the available compute". The clock now starts when a worker picks the job up; queue wait has its own generous bound and surfaces as a retryable saturation error. Also: reset() no longer cancels innocent queued peers; the timeout message stops claiming capacity was restored (the abandoned job keeps the device until it drains); every GPU dispatch uses the shared length-scaled budget; watermark embeds move off the GPU pool; /v1/audio/speech gets 429/503 + Retry-After; a timed-out batch segment fails the job instead of shipping a silent gap. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
aab138ea57 |
fix(backend): surface the shell's start-failure diagnosis instead of "can't reach the backend" (#1177)
The reporter's string is apiFetch's LAST fallback, reached only when no crash
marker exists AND the shell's lifecycle stage is 'failed' or 'unknown'. The
'failed' half was the bug: `BootstrapStage::Failed { message }` carries the
whole diagnosis — exit code plus a ~30-line stderr tail, or the precise reason
`ensure_venv_ready` refused (Intel Mac, a failed `uv sync`, a blocked GitHub) —
and `backendLifecycleStage()` returned only the stage tag, throwing the message
away. Every backend-start failure mode collapsed into one generic, evidence-free
sentence that was also factually wrong: it is not starting, and it will not
recover on its own.
- backendLifecycleStage() returns `{ stage, message }`; a `failed` stage gets
its own branch in apiFetch that surfaces the shell's diagnosis, scrubbed.
- BackendStartFailureNotice renders it after the splash is gone, reusing the
splash's `detectHints` matcher (shared, so the two can't drift) and the
existing bug-report affordance. No Retry advice for unrecoverable failures.
- Rust retains the last `Failed { message }` past a later stage transition
(Retry sets Checking, the supervisor sets StartingBackend) so a respawn can't
erase the first diagnosis; exposed as `last_bootstrap_failure`.
- `bun desktop` prints the exit code and where to look instead of exiting
silently — the from-source twin of the same class ("builds but won't launch").
- Scrub primitives extracted to utils/scrub.js so the transport layer can scrub
without a bugReport -> client import cycle.
Non-Tauri deployments are untouched: there is no shell to fail this way, so the
stage stays 'unknown' and #1164's deployment-specific message still stands.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
83f943bead |
fix: bot-review harvest (16 findings) + deterministic style/locale CI + reviewer configs
Harvested and verified every CodeRabbit/Greptile finding from PRs #1175, #1189, #1192, #1195: 16 real ones fixed (fallback ASR preflight bypass, VRAM release on stream exit, typed 409 parity, uv env independence, path-privacy in errors, MCP clone_voice hardening, CaptureWidget WS guard, test hygiene), 4 refuted with evidence, rest documented as deliberate design or deferred. Deterministic CI replaces hand-enforcement: tests/test_changelog_style.py (quiet one-liner format) and tests/test_locale_parity.py (21-locale key/placeholder lockstep with a ratchet baseline) — the latter surfaced and fixes 151 already-broken locale strings. CodeRabbit/Greptile carry the house rules via .coderabbit.yaml + greptile.json; CLAUDE.md gains the harvest-before-merge and never-accept-as-is rules. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
933743e336 |
fix: resolve open issue batch — #1172 #1173 #1174 #1185 #1186 #1188
- validate managed binaries before exec; 0-byte GGUF placeholders fail actionably instead of "Exec format error" (#1172) - KittenTTS: tokenizer-measured chunking to the ONNX 512-token cap; clear 400 for unspeakable input (#1173) - clean SIGTERM during weight load: shutdown-aware loader, benign cancelled-load classification, lifespan hardening, scoped log silencers (transformers load + alembic fileConfig) (#1174) - broken ASR deep-imports (lightning_fabric) mark the engine unavailable with a repair hint and fall through (#1185) - uv cache + managed Python follow the chosen install drive on Windows (cherry-picked cross-drive class fix + spaces/D: tests) (#1186) - adaptive silence-removal ladder for quiet clone references; localized actionable error for truly silent clips, all 21 locales (#1188) - CHANGELOG: consolidated Unreleased into the quiet one-liner style Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
63fd497caf |
feat: TTS-only first run, platform-curated ASR, guided OS permissions, parakeet-mlx
Only the TTS model (~2.4 GB) is required on first run; ASR models are per-platform curated picks (curated_on in models.yaml) installed on demand. Every transcription surface returns a typed asr_model_missing error with a one-click download CTA instead of silently pulling multi-GB Whisper weights. Settings -> Models is a grouped, platform-aware catalog. New guided permissions UX (wizard System Check + Settings -> Permissions + mic pre-flight) with native mic-state checks and OS settings deep-links. New parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated capture preference so multilingual dictation never regresses). Docs: expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ. Hardening: preflight fails open for custom model pins, ROCm curation no longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged consent key, CaptureWidget setup race fixed, offline-cache CI simulation fixes so empty-cache runners stay green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
08e19e0c1c |
feat(uninstall): opt-in, content-free app_uninstalled ping in the uninstall scripts
Before deleting anything, uninstall.sh / uninstall.ps1 now send ONE
best-effort `app_uninstalled` event — but only when the user opted in:
consent is read from the same prefs.json the app writes, AND the ping needs
the backend-written analytics_info.json (present only while analytics is
enabled — consent + build token — and removed on opt-out), because the
generic scripts ship no token of their own. Payload is content-free: app
version, OS name, random per-install id. 2-second timeout, silent failure,
one honest console line ('Sending anonymous uninstall ping (you opted in to
analytics).'); not opted in => nothing sent, nothing printed, and the
dry-run never sends either way.
Tested by exercising the bash script for real (fake $HOME + a curl shim
recording argv: consented/not-consented/dry-run/missing-info flows, plus
data-still-deleted-after-ping) and by static contract checks on the ps1
(consent gate, TimeoutSec 2, try/catch, no baked token). docs/install/
uninstall.md documents the behavior (docs-sync).
|
||
|
|
e6b1179001 |
feat(dev): loud backend exit banner for bun run dev + docs + changelog (#1164)
In dev there is no supervisor: concurrently's --kill-others-on-fail tears the whole stack down the moment uvicorn exits, the cause scrolls away with the terminal, and the browser tab just says it can't reach the backend — which is exactly how #1164 arrived with zero diagnostics. - scripts/dev-backend.mjs: dev:api now runs uvicorn through a wrapper (command args byte-identical, stdio inherited). On a non-Ctrl+C, non-zero exit it prints a boxed banner: exit code/signal, the last 20 lines of omnivoice.log (data dir resolved exactly like backend/core/config.py), an OOM hint (SIGKILL/137 + the Linux journalctl -k check), and a pointer to the crash notice the run sentinel raises on the next backend start. Exits with the child's own code so --kill-others-on-fail still works. Verified live: started the dev backend, SIGKILLed it, banner printed with the real log tail and exit code 137. - docs-sync: troubleshooting.md gains §14c (browser/dev/Docker crash forensics: the mode-aware error, the dev banner, run_sentinel.json / last_run_crash.json / GET /system/last-run-crash, cap+ack+version-gate semantics) and §14's crash-notice blockquote no longer implies the notice is desktop-only; CONTRIBUTING.md documents the dev:api wrapper. - CHANGELOG.md: [Unreleased] entry for the #1164 class fix. Tests: tests/frontend/devBackend.test.mjs (5) — the uvicorn args are pinned byte-identical, data-dir resolution mirrors config.py, tail/banner content incl. the OOM shapes. |
||
|
|
b17df5c7a3 |
docs(docker): refresh the Docker Hub overview and docker guide
- Add a what-you-need line (RAM/disk/GPU from the README requirements table, compressed pull sizes measured from the registry) so homelab users can size the deployment before pulling. - Update stale version examples (:0.3.6 / :0.3.17 -> :0.3.22). - Fix the 'main is always one patch ahead' claim — with AUTO_VERSION_BUMP off, main can equal the released version; say 'at or ahead of the last release', which is true in both modes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
99e01610bb |
feat(docker): publish ROCm/AMD GPU image variant (#1165) (#1166)
The Docker image was CUDA-only, so AMD GPUs (e.g. RX 7900 XTX under Podman) silently ran on CPU. Every preview and release now also ships a ROCm variant built from the same Dockerfile: - deploy/Dockerfile: parameterize the runtime base with a BASE_IMAGE build-arg (default unchanged: pytorch/pytorch 2.8.0 CUDA). Add PIP/UV_BREAK_SYSTEM_PACKAGES for the ROCm base's PEP-668-marked Ubuntu 24.04 Python (no-op on the conda CUDA base), and a build-time GPU_FLAVOR guard asserting the dependency install did not clobber the base image's GPU torch/torchaudio — a future dep bump that forces a torch reinstall now fails the build instead of shipping a CPU-only "ROCm" image. - .github/workflows/docker.yml: new build-and-push-rocm job (separate job for runner disk — the ROCm base is ~25 GB unpacked, so it frees the preinstalled toolchains first). Tags mirror the CUDA semantics with a -rocm suffix (:rocm rolling preview, :stable-rocm, :X.Y.Z-rocm, :X.Y-rocm, :sha-xxxx-rocm) on both GHCR and Docker Hub, same secret gating. flavor latest=false so release tags can't clobber :latest. No cache-to: the ROCm layers would blow the 10 GB GHA cache budget. - deploy/docker-compose.yml: new opt-in 'rocm' profile passing the GPU through via /dev/kfd + /dev/dri, with HSA_OVERRIDE_GFX_VERSION=11.0.0 documented (user-set, not baked in — backend auto-sets it for known consumer GFX IDs). - Docs-sync: docker.md (ROCm quick start incl. Podman/Quadlet, tag table, troubleshooting), dockerhub-overview.md, README AMD note, linux.md ROCm section cross-link, CHANGELOG [Unreleased]. Base image: rocm/pytorch:rocm7.2.4_ubuntu24.04_py3.12_pytorch_release_2.8.0 — torch 2.8.0 exactly matches the CUDA image (identical resolution, so uv keeps it), py3.12 satisfies requires-python >=3.11 (the ubuntu22.04 variants are py3.10 and do not). Closes #1165 Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9ecb810946 |
fix(shell): version-gate crash markers, pin the WebView repair contract, run the Rust suite in CI (#1145)
* fix(shell): version-gate crash markers, pin the WebView repair contract, run the Rust suite in CI Two deferred items from the recurrence audit, plus the CI gap that made them possible: - crash.rs: a persisted "backend crashed" marker now only surfaces for the release that wrote it. After an upgrade, markers from the previous version (quite possibly the build whose crash the upgrade fixed) are ignored and pruned on read instead of resurfacing unacknowledged as if the new build had crashed. backend_version gains #[serde(default)] so legacy version-less markers still deserialize — as "", which the gate treats as stale by design. Preview stamps (X.Y.Z-N) count as their release. - commands.rs: the #879 WebView2 cache repair's filesystem half is extracted into clear_webview_cache_at() (paths + retry policy as parameters, zero behavior change) and its contract is pinned by tests: no marker → nothing touched; marker consumed first, unconditionally (one-shot — a failing repair can never loop across launches); missing cache is success; a locked cache is retried then abandoned with a log, never bricking startup. - ci.yml: the Tauri shell check only ran `cargo check`, which neither compiles nor runs #[cfg(test)] code — so the shell's ~90 unit tests (crash.rs, reset.rs, bootstrap.rs, …) never executed anywhere in CI. `cargo test --lib` now runs them natively on all three OSes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(shell): crash-notice read path is strictly read-only — a prune-save there could destroy a fresh marker Greptile's P1 is real, and hotter than stated: get_last_backend_crash is not just a startup check — streamDropError (#1119) polls it every second for 8 s after a stream drops, which is exactly when the death watcher is inside record_crash's load→push→save. The previous commit's read path did load→prune→save when stale-version markers existed (the post-upgrade state), so a poll could load the pre-crash snapshot, lose the race, and save over the freshly recorded marker — silently deleting the only evidence of the crash it was being polled to find. Smallest fix: reads never write. The read path (extracted as read_notice_from(path, version) so the contract is testable) filters stale-version markers in memory only; disk pruning stays on the write paths (record_crash, acknowledge_backend_crash), where load-modify-save already existed pre-PR and is paced by a crash or a user click rather than a 1 Hz poll. Regression test pins the file as byte-identical across reads, stale markers filtered and current ones surfacing as before. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4e5d795832 |
fix(uninstall,storage): remove the saved-env leftover; count sidecar engines in disk usage (#1108)
Two recon findings from the reset work, fixed properly (whole class + tests +
docs), plus the destructive reset path is now exercised end-to-end.
1. ~/.config/omnivoice/env survived every uninstall. The app persists the
model-cache location (and a possible HF_TOKEN) there via
backend/core/user_env.py, but the in-app "Remove all data" (uninstall.rs),
uninstall.sh, and uninstall.ps1 all walked past it — so a reinstall silently
inherited the old file and redirected downloads to a maybe-deleted location.
All three now remove it. It's the same expanduser("~/.config/omnivoice/env")
path on every OS, so the Windows script uses %USERPROFILE%\.config\omnivoice.
is_recognizably_ours accepts it (contains "omnivoice"); docs tables updated.
2. Disk usage measured the wrong engines dir. storage_report.default_engines_dir()
returned backend/engines (built-in engine *modules*, no venvs), while sidecar
installs live in DATA_DIR/engines/<id>. So a multi-GB IndexTTS-2 install was
invisible in the engine-venv category and rolled into data/"other". Now points
at DATA_DIR/engines and sizes the WHOLE install (venv + checkout + weights),
with the data category claiming that subtree so it isn't double-counted.
Reset hardening: extracted purge_scopes() as a pure fs function (no AppHandle),
so the actual delete loop runs in tests against a real on-disk install tree —
"everything" wipes the install but spares the venv/foreign temp/sibling folders,
a settings reset keeps content+config+models, and a poisoned data_dir="$HOME"
deletes NOTHING. This is the live drive-through of the destructive path, minus
the GUI.
Also: gitignore the node_modules symlink form (the directory rule node_modules/
never matched a worktree symlink, so it kept slipping into commits).
Tests: Rust 78 (6 new), storage_report 20 (2 new incl. once-not-twice count +
default-dir guard), frontend 1207, i18n probe green, format+lint clean.
Co-authored-by: mergetest <nizam4103@gmail.com>
|
||
|
|
94093605eb |
feat(settings): factory reset gets scopes — preferences, settings, assets, everything (#1100)
* feat(settings): factory reset gets scopes — preferences, settings, assets, everything Factory reset did exactly one thing: clear localStorage. The only other option was "Remove all data", which deletes the Python env and quits. Between "forget my theme" and "wipe the machine" sat every reset a user actually needs — drop a corrupt model download, remove a wedged sidecar engine, put the settings back without losing a single voice — and none of them existed. Settings → Storage → "Reset & remove" now offers four tiers (UI preferences / all settings / downloaded assets & models / everything OmniVoice did) plus a per-scope checklist. Every scope shows its real on-disk size, and the number on the confirm button is exactly what gets freed. Why the shell and not the backend: a loaded model memory-maps its weights out of the HF cache (locked on Windows while mapped), and ensure_dirs() runs at import, so a backend cannot delete voices/ or outputs/ and still write to them. reset.rs stops the backend, deletes, and starts it again — and that restart is also the repair: the fresh process re-runs ensure_dirs() and alembic, so a removed database comes back empty rather than missing. retry_bootstrap's respawn path is extracted to bootstrap::respawn_backend so both callers share one implementation. Deliberate scope choices: - "Everything" stops short of the managed Python env, so a reset hands back a working app on the first-run screen. The env is the uninstaller's business. - A settings reset keeps the storage locations (config.json, the user env file). Clearing the model-cache pointer would strand gigabytes at a path the app no longer looks in — install shape is not a preference. - content deletes the DB with the media: rows without files is how you get a library full of broken entries. - The shared HF cache is flagged as shared only when it IS — computed, so Windows and portable installs (app-private cache) get no caveat they don't need. Safety: nothing is removed unless it sits inside a validated root — one carrying an OmniVoice-owned path component OR holding an OmniVoice signature file, which is what lets a custom data dir on an external volume be cleared while a mis-set data_dir: "/" is refused. Voices/projects/audio need the word typed. 9 Rust tests (guard, scope composition, shared-cache computation) + 14 frontend (planning purity, typed confirm, disk-vs-frontend split, shared warning). Border utilities follow the design guard (tests/test_no_literal_borders.py). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(settings): give the Storage panels a design — proportional sizes, live totals, real tokens "Remove all data" listed four folders as a flat run of text: a 7.5 GB model cache and a 391-byte config file rendered at identical visual weight, so the one thing worth seeing — where the space actually went — was the one thing you couldn't. And the 391 B folder said "0 KB", which reads as "nothing here". - New shared StorageTargetRow, used by BOTH destructive panels so they read as one system: icon, label, dimmed path (truncated, full text on hover), size, and a proportional bar showing that row's share of what will be freed. Unticked rows claim none of the bar — the bars must sum to what the button promises. - The shared HF cache moves OUT of the confirm dialog into its own "Optional" row with the checkbox and the caveat in the list. Ticking it now moves the running total in front of the user, instead of springing a different number on them at the point of no return. The dialog lists exactly what is going. - One byte formatter for both panels (settings/bytes.js). models/format.fmtBytes floors at kilobytes, hence "0 KB"; it stays where it is for the model store. Real fix underneath: three of the tokens these panels styled with DO NOT EXIST (--chrome-fg-subtle, --chrome-bg-raised, --color-warning). An undefined var() makes the declaration invalid, the browser drops it, and the element silently inherits — which is why the paths that were meant to recede rendered at full body weight. That is a whole class of bug that fails invisibly, so it gets a guard: src/test/cssTokens.test.js fails on any var(--token) in JSX not defined in a stylesheet, with runtime-injected tokens (Radix, inline-style hues) allowlisted by reason. Six pre-existing offenders elsewhere in the app are recorded as known-broken and ratcheted so the list can only shrink — they are real bugs, but each is a visual change that wants its own review. Frontend suite 1196 → 1205 (6 UninstallPanel component tests incl. the live total and the bar proportions; 3 token-guard tests, verified fail-before). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: green CI + finish the token sweep + snapshot the panels Three things on top of the redesign: 1. CI was red on tests/probe/test_probe_i18n.py — removing the eight dead `factory_reset_*` keys from en.json orphaned them in all 20 other locales (the probe forbids a non-en key absent from en). Removed them everywhere. This guard scans locales at pytest time; a frontend-only run never sees it. 2. Finished the undefined-token sweep instead of grandfathering it. Six bare `var(--token)` references resolved to nothing; the only genuinely undefined, fallback-less one in shipping panels was `--chrome-input-bg` (input fields AND progress-bar tracks AND skeletons across StoragePanel, StorageUsagePanel, HistoryRetentionPanel, ModelStoreTab — tracks were rendering with no background at all). Repointed to --chrome-hover-bg. The rest (--chrome-menu-bg, --chrome-bg-inset, --border, --input-bg, --muted) already carry `var(--x, fallback)`, which is valid CSS. So cssTokens.test.js now checks only the BARE form and ships with zero exceptions — no known-broken ratchet, because there is nothing left broken. 3. Registered both Storage panels in the visual-regression harness (a Tauri `invoke` stub added to providers.jsx alongside the existing fetch stub) and committed baselines across all three themes. This is how I actually looked at the redesign: the bars render proportional (the 720 KB voices row fills, the 391 B row is a sliver), the shared-cache row sits in its own Optional group, and every token now resolves in default/midnight/catppuccin. `_forceAdvanced` on ResetPanel opens the checklist for the snapshot; no effect on the toggle. Full backend suite 2897 passed (incl. the i18n probe). Frontend 1205. * style: oxfmt the new panels and specs Format-check is a CI gate (oxfmt --check); the new files weren't run through oxfmt --write. No behavior change. * chore: stop tracking the node_modules symlink A worktree-local symlink slipped past .gitignore (which lists node_modules/ — the directory form — so it never matched the symlink file). Removed from the index; the symlink stays on disk for local test runs. --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0421be966e |
feat(settings): in-app uninstall — Settings → Storage → Remove all data (#1089) (#1099)
* feat(settings): in-app uninstall — Settings → Storage → "Remove all data" The v0.3.19 uninstaller was a SCRIPT, which never reaches the people who need it: anyone who installed the .dmg / .msi / AppImage has no repo to run scripts/uninstall.sh from — exactly the reporter in #1089, an AppImage user. "Where is uninstall in the app?" had no answer. Now it does. New Tauri commands (uninstall.rs): - uninstall_scan — every folder this install owns, with real sizes, resolved through the same setup.rs helpers the app itself uses, so custom + portable locations are cleaned instead of the defaults being assumed. - uninstall_purge — stops the backend (marking the kill intentional so the #567 supervisor doesn't respawn one into the directories being deleted), removes the folders, and lets the UI quit the app: the Python env it runs on is gone, so there is nothing to return to. This lives in the Rust shell, not the backend, because the biggest thing to remove is the managed Python environment and the backend is RUNNING FROM IT — a process can't delete its own interpreter (and Windows locks the files). Safety: every path must pass is_recognizably_ours() before any remove_dir_all — absolute, not `/` or $HOME, and carrying an OmniVoice-owned component (unit tested both ways). The shared Hugging Face cache is reported separately and is OPT-IN behind its own checkbox with the caveat spelled out: it's the standard HF cache other ML tools share, so sweeping it up silently would delete models this app never downloaded. Deleting voices/projects is irreversible, so the confirm requires TYPING the word, not just a click. Also fixes a real bug in what shipped in v0.3.19: the scripts and docs missed where the BACKEND writes its logs — ~/.local/state/OmniVoice on Linux and %LOCALAPPDATA%\OmniVoice\Logs on Windows (backend_log_path(), backend.rs) — so every Linux/Windows uninstall left a stray log dir behind. Covered now in the scripts, the docs, and the in-app scan. And the scripts now ship as release assets, so cleanup is possible without launching the app at all. Rust: 2 new guard tests. Frontend: 6 new tests (the size on the confirm button must equal what actually gets deleted); suite 1182 passed. Docs synced. Refs #1089 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(settings): drop token border utilities from UninstallPanel (design guard) tests/test_no_literal_borders.py::test_no_token_border_utilities_in_jsx is a backend guard that scans JSX — so a frontend-only test run misses it. It forbids `border-[var(--chrome-border)]` structural utilities: the app-wide border removal converted every panel/row frame away from them, and they render a stray hairline the moment the token doesn't resolve transparent. Row dividers → spacing + an alternating `--chrome-hover-bg` tint; the opt-in checkbox card → a background tint; the confirm input → the sanctioned arbitrary `[border:1px_solid_var(--chrome-border)]` property form the other settings inputs already use (explicitly not flagged by the guard). Guard green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b1a7ddc374 |
feat(install): clean uninstaller + a straight answer to "where is my data?" (#1097)
A Linux AppImage user asked which folders to delete to remove OmniVoice and whether an uninstaller exists (#1089) — they had to guess. They shouldn't have to: the app is fully local, so uninstalling IS just deleting the folders it wrote, and we never documented them. - scripts/uninstall.sh (macOS/Linux) + scripts/uninstall.ps1 (Windows): find every OmniVoice folder — app data, the multi-GB managed Python env, config, logs — plus, listed SEPARATELY because it is a shared cache, the Hugging Face model cache. Print each with its size as a DRY RUN and stop; delete only on --yes (--models / -Models to include the shared cache). They honor the same env overrides the app reads (OMNIVOICE_DATA_DIR, OMNIVOICE_CACHE_DIR, HF_HOME, HF_HUB_CACHE), and never touch the app binary or anything outside the paths they list. - docs/install/uninstall.md: the complete per-platform path table (what each folder holds and how big it is), the shared-HF-cache caveat, custom/portable locations, per-platform steps to remove the app itself, and what to keep if you plan to reinstall. - Linked from the README FAQ, SUPPORT.md, and install troubleshooting. Paths mirror backend/core/config.py + frontend/src-tauri/src/setup.rs. Verified on macOS: dry-run lists the real dirs; sandboxed HOME runs confirm --yes removes app folders while KEEPING the shared cache, --models removes it, and the env overrides retarget correctly. Closes #1089 Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
23367cccaf |
fix: first-run wizard version + mirror-unreachable rescue + lifecycle-aware backend reachability (#1094)
Three fixes from the same first-run session report:
- SetupWizard shows v{APP_VERSION} in its masthead (same identity mark as
the install splash footer), so setup screenshots identify the build.
- A dead configured HF mirror no longer strands the wizard: the
install_error SSE now carries docs_topic (core.failure.classify), and
WizardLibrary renders the MirrorRescue quick-pick (extracted from
SetupWizard, now including the official preset) next to the failed row,
retrying it the moment a new endpoint is applied. PUT /hf-mirror clears
the install cooldowns (no 429 on the immediate retry) and clearing to
official also drops the legacy hf_endpoint pref that silently kept the
dead mirror in effect. The hint's false "applied when the app starts"
claim is corrected: downloads resolve the endpoint per call, retry
first, restart only if it still fails.
- "Can't reach the local OmniVoice backend" stops firing during real
start/restart windows: a respawn takes 10-20+s (venv spawn + torch
import) but the transport cascade gave up at ~2.9s. apiFetch now asks
the shell (bootstrap_status via utils/backendLifecycle) whether a
start/restart is in progress and keeps retrying while it is (capped at
120s, matching the supervisor's respawn budget); the new
BackendRestartBanner finally implements the reconnecting banner the
#567 supervisor has emitted events for all along. Truly dead backends
(or non-Tauri deploys) still error promptly.
Regression tests for all three layers; docs synced
(downloading-models.md, troubleshooting.md §14b); CHANGELOG [Unreleased].
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9c81e3389d |
feat(network): automatic Hugging Face endpoint selection — probe, pick, remember (#1082)
* feat(network): automatic Hugging Face endpoint selection — probe, pick, remember Restricted-network first-runs (the #984 class: huggingface.co unreachable, user dead-ends before discovering the mirror setting) now self-heal by default, while explicit endpoint choices are never second-guessed. - New backend/services/endpoint_race.py: parallel HTTPS reachability + latency probes of huggingface.co and the hf-mirror.com community mirror (3s timeouts). Probes are the only signal — no geo-IP, no third-party calls. Reachable beats unreachable; with both reachable the official endpoint wins unless the mirror is decisively faster (anti-flap hysteresis). The pick is cached in prefs and re-raced only on first run, a network-classified download failure, staleness (>7 days), or an explicit "Test again". - Manual mode is sacred: HF_ENDPOINT env, an hf_endpoint pref, or any explicit Settings pick disables auto-switching entirely; OMNIVOICE_HF_ENDPOINT_MODE=manual is a hard opt-out. - Wiring: the wizard preflight races endpoints when nothing is configured (honest copy when the mirror wins; warn-not-block when nothing is reachable); Model Store installs and the model-cache auto-repair resolve their per-call endpoint= through the cached decision, and a network-classified failure re-races once per repo per process and retries on the new winner (same guard pattern as the cache-recovery ladder). - Settings → Models → Hugging Face mirror gains "Auto (recommended)": shows the current pick, measured latency, last-checked time, and a "Test again" button (POST /api/settings/hf-mirror/test). Existing explicit configs surface as the matching manual mode. Panel notes that hf_hub checksums every download regardless of endpoint. - Tests: policy/cache/failover matrices in tests/test_endpoint_race.py, preflight + settings + repair-failover integration with mocked probers, HFMirrorPanel mode tests, and a suite-wide conftest guard that pins the probers so no test can hit the real network. - Docs: downloading-models.md and install/troubleshooting.md describe the automatic default and both opt-outs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): Unreleased entry for automatic HF endpoint selection Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint-probe pin uses an isolated MonkeyPatch and clears the decision cache; dtype guard tolerates stubbed torch The autouse probe pin requested the shared monkeypatch fixture, hoisting its setup earlier for every test and reordering teardown against the fp16 guard — which then ran torch.get_default_dtype() on test_torch_compile_gate's SimpleNamespace stub. The pin now uses its own MonkeyPatch context and also clears the prefs-cached endpoint decision per test (one test's auto pick leaked into other tests' preflight labels on CI ordering). The dtype guard additionally skips non-module torch stubs outright. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint env vars can no longer leak out of the mirror-settings suite set_hf_mirror writes os.environ[HF_ENDPOINT] during the test, and monkeypatch.delenv(raising=False) on an absent var records nothing to undo — so the write leaked process-wide and flipped later suites' preflight network checks into the explicit-endpoint branch (the CI-order failures). Guaranteed save/restore autouse fixture at the source, plus defensive env shedding in the preflight suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
254f071b45 |
fix(docker): ship alembic.ini in the image, add an image-level HEALTHCHECK (#1080)
Migrations in Docker fell back to the additive-column self-heal because alembic.ini was never copied; the real migration chain now runs. The HEALTHCHECK covers plain docker-run (compose files keep their own), with a start period sized for first-boot schema creation. Docs example tag freshened. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5562aa16a7 |
feat(setup): media tools become invisible — bundled by default, controllable in Settings (#1071)
Most users should never learn what ffmpeg is. The Setup Wizard's SYSTEM
PREFLIGHT stops listing FFmpeg / FFprobe / yt-dlp as user-installed
requirements ("brew install ffmpeg…"): they are internal dependencies the
app provisions for itself. Genuine user facts (OS, RAM, disk, GPU,
network, Python) are untouched.
Backend
- New services/media_tools.py: per-tool status {version, path, origin:
sidecar|bundled|system|custom}; background acquisition of a pinned,
SHA-256-verified static ffmpeg+ffprobe build (immutable-commit fetch
from the same upstream the static-ffmpeg pip package uses — that
package itself was audited and rejected: mutable raw/main URL, no
checksums, writes into site-packages); binaries are `-version`-probed
via the existing _binary_runs before being trusted, installed under
DATA_DIR (update-surviving, frozen-build-safe), zero new Python deps.
- ffmpeg_utils resolution chain gains the acquired-bundled tier — and
ffprobe finally has a bundled tier at all (imageio-ffmpeg ships none),
closing the source-install gap.
- New /media-tools router (loopback-gated, same contract as
/system/set-env): status, acquire, {tool}/custom-path | use-system |
restore, ytdlp/update | restore. Overrides persist via the existing
env.FFMPEG_PATH / env.FFPROBE_PATH prefs convention — one store, no
competing controls.
- yt-dlp updates: audited in-venv pip/uv upgrade and rejected (venv is
uv-managed with no pip; yt-dlp is a locked dep, so the updater's
--inexact drift sync would revert it). Instead the newest wheel —
verified against PyPI's own sha256 — lands in a DATA_DIR overlay
prepended to sys.path at startup: survives app updates, works in
frozen builds, and "Restore tested version" is just deleting the
overlay. Gallery now runs yt-dlp via `python -m yt_dlp` (module, not
PATH) so the CLI can never be a user-install task either.
- /setup/preflight drops the three tool rows, carries a media_tools
verdict, and self-heals: kicks the bundled download in the background
when no tier resolves (never re-fires after a failure — the wizard's
card owns Retry). diagnose + the ffmpeg-missing notification now point
at Settings → Audio tools instead of package managers.
Frontend
- Wizard: new MediaEngineCard — renders NOTHING when the engine is ready,
a one-line progress while acquiring, and only on failure an actionable
card (Retry / Use a system copy / Choose file…).
- Settings → Audio tools (new category, System group): FFmpeg + FFprobe
rows with version, path, origin badge, Use system copy / Choose file… /
Restore bundled, header-level "Update bundled build"; yt-dlp row with
Update + Restore tested version (+ restart affordance). Package-manager
commands appear only as copyable prose, never executed.
- The FFmpeg-path override moved out of Settings → Network (pointer row
deep-links to Audio tools; no second writer of env.FFMPEG_PATH).
Notifications gain a settings-tab action type.
- All strings i18n (en + defaultValue), a11y labels on every control.
Tests: 29 new backend (origin classification, checksum/size/probe
rejection, override persistence, overlay update/restore, router gating +
route-shadowing) + preflight contract tests (tool rows gone, verdict
present, auto-acquire fires once); 14 new frontend (wizard hide/progress/
failure-card, Audio tools rows/badges/actions). Route snapshot
regenerated. Docs (macos/linux install, troubleshooting §7b) describe the
new reality in the same commit.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
1808a373a1 |
fix(linux): AppRun workaround detection reads the BUNDLED WebKitGTK version, not the host's (#961 follow-up) (#1024)
The launcher decided whether to export WEBKIT_DISABLE_COMPOSITING_MODE by asking the host's pkg-config — but LD_LIBRARY_PATH makes the BUNDLED libwebkit2gtk the one that actually runs, so on any machine where the two diverge the detection read the wrong number. This was the second bug identified during #961's investigation (the reporter built from source, so their dev packages answered pkg-config with a healthy version while the shipped bundle ran an older lib) and was explicitly deferred in #1007 as not-safely-fixable at runtime. The fix makes it knowable by construction instead: inject-apprun.sh runs at bundle time ON the build host whose libwebkit2gtk gets bundled, so it stamps that version into .bundled-webkitgtk-version inside the AppDir. AppRun reads the stamp first and only falls back to host pkg-config for bundles predating it. Empty/unreadable stamp fails safe (workaround on), same philosophy as the missing-pkg-config path. Tests: 3 new cases in AppRun.test.sh — marker-beats-host in both directions (broken-marker/healthy-host and the #961 inversion, healthy-marker/broken-host) plus empty-marker fail-safe. Also wires AppRun.test.sh into pytest (tests/test_apprun_launcher.py) — it was previously run by NO CI job, so the launcher could regress silently. Also documents Windows install-to-another-drive behavior in docs/install/windows.md (#938): local drives work via the wizard's directory picker, mapped network drives are a Windows Installer limitation, and the data directory moves independently of the app. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
da4bef8e42 |
docs: correct troubleshooting §16 — mic bug was a missing entitlement, not an upstream limitation; changelog for #1016/#1020/#1021 (#1022)
troubleshooting.md §16 claimed the macOS microphone-permission bug was an unresolved upstream Tauri/wry limitation with no available fix. That was wrong: @MahdiHedhli read the wry/tauri sources more carefully and found the real cause — Tauri's Hardened Runtime default blocks mic hardware access without com.apple.security.device.audio-input in the bundle's entitlements, which also explains why TCC never listed the app. Their fix (#1016) is merged; §16 now documents the real mechanism, credits the correction, and keeps the record-elsewhere workaround for users on ≤0.3.12 builds. Also brings CHANGELOG [Unreleased] current for the three merges that lacked entries: #1016 (mic fix), #1020 (shutdown wait 3s→20s), #1021 (CI flaky-trio root cause + guard). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7f8a42ce51 |
fix(tts): mlx-audio CSM cloning drops ref_text, breaking every clone attempt (#1012, #1013) (#1017)
MLXAudioBackend.generate() reads voice/ref_audio/language/speed from its kwargs but never extracted ref_text — it was built, then silently never passed through to self._model.generate(). CSM (sesame.py) only builds its cloning context when BOTH ref_audio AND ref_text are present; with ref_text missing, the context list stays empty and indexing into it raises "IndexError: list index out of range" deep inside mlx-audio, instead of the clone ever being attempted. Voice cloning on the CSM engine could never have worked as shipped. generation.py already threads ref_text all the way through — even auto-transcribing it via the GPU pool when the caller supplies ref_audio without one (~line 780) — so the value was always available in kwargs; it just never survived the crossing into this specific backend. Reported with the precise root cause and a working fix (community member independently diagnosed and patched it locally, confirmed working on MPS/0.3.12). Two-line fix: extract ref_text and pass it through when both ref_audio and ref_text are present (guards against passing an orphaned ref_text with no accompanying audio to engines that don't expect it). Tests: tests/test_engines.py — ref_text is passed through when paired with ref_audio, omitted when ref_audio is absent. Also documents the second bug from the same report (#1013): macOS microphone permission never prompts, so OmniVoice never appears in System Settings to grant access. Root-caused to an unresolved upstream Tauri/WebKit limitation (WKWebView's requestMediaCapturePermissionFor delegate — wry#1195, tauri#11951, fix wry#1196 still open/unmerged, no released version to bump to) — not something fixable here without an unverified native Rust/WKWebView hack this session has no way to test. Documented in docs/install/troubleshooting.md with the confirmed workaround (record elsewhere, upload the file). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
93aab6dadc |
docs(linux): mention yt-dlp as an optional prerequisite (#973) (#997)
The preflight system check already warns in-app when yt-dlp is missing (Voice Gallery/Dub YouTube downloads fail without it), but the install docs never mentioned it — a user has to hit the in-app warning first instead of seeing it up front alongside the other optional prereqs. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
637020b82b |
fix(engines): nemo-parakeet install hint stops recommending a shared-venv-breaking pip install (#974) (#991)
The Engines page told users to run `pip install nemo_toolkit[asr]` for the NeMo Parakeet ASR engine. nemo_toolkit[asr]==2.7.3 hard-pins transformers>=4.57,<4.58, which is unsatisfiable alongside OmniVoice's own transformers>=5.3 requirement (needed by omnivoice/models/omnivoice.py for HiggsAudioV2TokenizerModel). A user who followed the hint ended up with a backend that wouldn't start (ImportError: cannot import name 'HiggsAudioV2TokenizerModel'). _INSTALL_HINTS["nemo-parakeet"] in backend/services/asr_backend.py now states plainly that installing into the shared venv will break the backend, names the transformers conflict, and tells users to use a separate/dedicated Python environment instead — without implying a safe one-line fix or an isolated-venv env var exists (unlike dots-tts/moss-tts-v15/confucius4-tts, nemo-parakeet has no isolated venv option yet; that's a separate, larger follow-up). Also adds one sentence to docs/install/troubleshooting.md's existing "engine venv clash" section (#11) pointing at the same class of issue on the ASR side, and a regression test asserting the hint never again contains the literal bare `pip install nemo_toolkit[asr]` string. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
dab6456581 |
docs(linux): stop advertising a .deb package that isn't published (#990)
README's Quickstart badges linked a 'Download Debian .deb' button straight to the releases page — but .deb bundling was deliberately dropped from release.yml (tauri-cli bug, 'Failed to create control scripts') and no release has ever shipped one. A community member investigating #961 confirmed this by checking the actual release assets. Users clicking that badge got a broken promise, not a package. Removed the badge; docs/install/linux.md's '## Install (.deb)' section now honestly states it's unavailable pending a tauri-cli fix, points to the AppImage as the supported path, and keeps the historical pre-v0.3 .deb upgrade note (ffprobe conflict) since that's still relevant to existing installs. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1a03c59f82 |
fix(install): AMD ROCm torch reinstall targets rocm6.4, not rocm6.2 (#988)
Community-diagnosed (issue #972, Kaihui-AMD): pyproject.toml pins torch==2.8.0, but the rocm6.2 wheel index only ever published up to 2.5.1 — the reinstall silently failed to resolve and fell back to the default CUDA build, which runs on CPU on an AMD GPU. The failure was correctly logged (bootstrap.rs's emit_log warning), just never actioned because the index itself couldn't succeed. rocm6.4 carries a matching torch==2.8.0 build. Docs updated with the corrected index plus a repo.radeon.com find-links path for users who want a driver-matched ROCm 7.2.x build the PyTorch index doesn't carry (OMNIVOICE_TORCH_INDEX only accepts a PEP 503 index, not find-links, so that's documented as a manual step). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
087309259b |
fix(setup): first-run network check is mirror-aware and never hard-blocks (#984)
* fix(setup): first-run network check is mirror-aware and never hard-blocks Field report (Discord, China): the Launchpad preflight probed hardcoded huggingface.co:443 and any failure disabled Continue outright — users behind the GFW were stuck on the very first screen, before Settings (and its HF mirror quick-pick) was even reachable. - The probe now targets the HF endpoint actually in effect (HF_ENDPOINT / hf_endpoint pref via configured_hf_mirror), with the real port. - An unreachable endpoint is a WARNING, not a blocker: local-first — cached models work offline, and downloads surface their own actionable errors. - When huggingface.co is blocked but hf-mirror.com answers, the fix text says exactly that, and the wizard shows an inline mirror quick-pick (presets + custom URL) that applies via PUT /hf-mirror — effective immediately for downloads — then re-checks. - Docs updated (downloading-models, install troubleshooting); regression tests cover warn-not-fail, mirror-host probing, and the mirror suggestion. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): open [Unreleased] with the preflight mirror fix (#984) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
620321a9cc |
feat(diagnostics): backend crashes become self-documenting — exit code + stderr tail surfaced and attached to bug reports (#969)
* feat(diagnostics): backend crashes become self-documenting — exit code + stderr tail surfaced and attached to bug reports When the backend PROCESS died (native CUDA abort, OOM kill, DLL crash) the user saw only "Can't reach the local OmniVoice backend" and the evidence died with the process — every #941-class report needed a logs-please round-trip nobody answers. The v0.3.9 guard fixed HANGS; this fixes the class of invisible DEATHS: - Rust (crash.rs): every unexpected child exit — detected by the startup health poll and the post-Ready supervisor — writes a rotating (last 3) JSON crash marker next to the backend logs: ts, exit code/signal, backend version, uptime, ~40-line stderr tail. Intentional shutdowns never forensicate: app-quit raises the quitting flag first (now also on macOS Cmd+Q via ExitRequested), and retry/clean-retry kills set a BACKEND_KILL_INTENDED flag cleared when the fresh child is tracked. - Tauri commands get_last_backend_crash / acknowledge_backend_crash; ack is a persisted watermark, never a delete — bug reports still get the evidence after the user viewed it. - Crash-loop escalation: the supervisor budget goes 5-in-60s → 3-in-10min so slow crash loops stop respawning and land on the Failed screen with the last exit code + stderr tail. - Frontend: apiFetch's transport-failure path swaps the vague message for "the backend crashed (exit code X) N s ago…" when an unacknowledged marker exists, and BackendCrashNotice (banner + details dialog, i18n'd, ack-on-view) surfaces it even with no request in flight. - Bug-report prefill gains a "Last backend crash" section (exit code + home-path-scrubbed stderr tail via the existing scrubText), so the next report arrives WITH the evidence. Tests: cargo --lib 57 pass (marker rotation write-4-keep-3, ack semantics, store IO, ExitStatus decomposition, 3-in-10min policy); vitest 909 pass incl. crash-notice branch, client crash-message branch, bug-report enrichment; legacy node:test 41 pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add backend crash forensics under [Unreleased] (#969) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4aa9abe22a |
docs+scripts: install fixes — desktop-prod tauri resolution, Ubuntu white-screen guidance, honest GPU/prereq docs (#960 #961 #962) (#964)
* docs+scripts: install fixes — desktop-prod tauri resolution, Ubuntu white-screen guidance, honest GPU/prereq docs (#960 #961 #962) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add the install-fixes batch under [Unreleased] (#964) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
72d137e1f3 |
fix(bootstrap): port cuDNN 8 (NVIDIA CUDA GPU) + VC++ redist to packaged installs (#869)
* fix(bootstrap): port cuDNN 8 (NVIDIA CUDA GPU) + VC++ redist install into ensure_venv_ready() * fix(bootstrap): address #869 review — drop dead VC++ half, cache negative CUDA probe, gate on ROCm, sync docs Per maintainer review on #869: 1. Drop the VC++ Redistributable half: LoadLibraryA("vcruntime140.dll") from the running Tauri exe is a tautology (the exe itself links the MSVC CRT, so the process wouldn't be running without it), and torch's real failure mode is msvcp140.dll inside the venv python process. Dead code removed; a comment records why for future readers. 2. Stop taxing every non-CUDA launch: a negative torch probe (CPU / Intel / AMD — most installs) is now cached in a .venv/.cudnn8_probe_negative marker, so the synchronous `import torch` runs at most once per venv lifetime. Invalidated on every path that can change the torch build (drift sync #307, repair sync, first-run sync, ROCm reinstall) and implicitly by a venv rebuild. A probe that fails to run cleanly is skipped WITHOUT caching so a transient error can't wedge a real CUDA machine. 3. Rewrite docs/install/troubleshooting.md §10 to the actual root cause: packaged installs never had the cudnn8_compat libs (so reinstalling never restored them); the bootstrap now installs them automatically on CUDA machines, with the manual uv pip command as the offline fallback and PyTorch Whisper as the sidestep. 4. Gate the ~700 MB nvidia-cudnn-cu12 download on the venv torch being a real CUDA build: the probe now reports 'hip' before checking cuda.is_available() (which HIP spoofs), so opt-in ROCm installs (#124) never fetch the CUDA wheel. Also reflow the CHANGELOG entry to house style (bold one-line lead, 1-3 lines of why, (#827, #869) refs) and extend the bootstrap unit tests: classify_cuda_probe verdict mapping and the marker write/invalidate round-trip (6 cuDNN tests total, 43 lib tests green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
83e71c5689 |
fix(asr): close the #730 residuals — chunked dub wedge shares the guarded reset; repeated timeouts recommend the crash-isolated engine (#895)
Residual A — the chunked dub-stream had a PARALLEL wedge mechanism (its own ping-loop timeout, its own _reset_pool_on_wedge, a dead-end "Try restarting the server" message). A wedged chunk now routes through the SAME run_transcribe_guarded bound+reset as the whole-file paths (#731/#851): the guard resets the poisoned pool once per wedged attempt (no double-reset on retry) and the user sees the actionable ASRTimeoutError. The reset logic is extracted to asr_backend.reset_pool_after_wedge — one shared mechanism, so the semantics can't drift again. run_transcribe_guarded also gains a timeout_env param so chunk errors name OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S instead of the whole-file knob. Residual B — the crash-isolated ASR sidecar (#393, faster-whisper-isolated) is wired as an explicit ESCAPE HATCH, not a default: - selectable end-to-end: Settings engine list gets an explanatory install_hint; honest gpu_compat ("cuda","cpu" — it wraps the same CTranslate2 engine as faster-whisper); get_active_asr_backend now hands back a process-wide singleton for subprocess-isolated backends (a fresh instance per request would leak atexit hooks and respawn the sidecar — reloading its model — on every transcribe). - on the SECOND consecutive guarded timeout-with-reset in one session (resets aren't recovering the hang; the wedged thread keeps its VRAM), the error the user sees + the log recommend switching to the isolated engine in Settings → Engines. Never auto-switched (owner rule: no silent behavior divergence); a completed transcribe resets the streak. Tests (fail-before/pass-after verified against origin/main): wedged-chunk SSE integration (reset count + actionable error + recommendation surfaces), consecutive-timeout streak (fires at 2, resets on success, suppressed when already on the isolated engine), timeout_env parametrization, shared-reset helper, isolated backend in list_backends with hint + honest availability, singleton caching, gpu_compat matrix entry. Docs: troubleshooting §14 gains the chunk knob + escape-hatch guidance. Closes the residuals tracked on #730. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
86f5213055 |
fix(splash): IPC-independent watchdog + recovery panel for dead Tauri IPC (#879) (#892)
After an unclean shutdown (Windows BSOD), the WebView2 profile cache (%LOCALAPPDATA%\com.debpalash.omnivoice-studio\EBWebView) can corrupt: Tauri's IPC custom protocol fails AND the postMessage fallback breaks, so invoke() hangs forever. useBootstrapStage's poll loop rode entirely on that IPC — a hung bootstrap_status call silently killed the loop and the splash sat at "preparing" forever, even with a fully healthy backend answering over plain HTTP. Class fix, three parts: - splashWatchdog.js: IPC-independent escape hatch. If no IPC signal arrives within 10s, poll GET /health over plain HTTP; healthy → proceed to the app as if 'ready' was received (console.warn breadcrumb so diagnostic bundles carry it). Any successful IPC response disarms it for good. - Recovery panel (stage 'ipc_lost'): if neither IPC nor HTTP succeed within 45s, show an actionable panel instead of the infinite spinner — "Open logs" (with an inline path fallback when IPC is dead) and, Windows-only and only in this error state, "Repair and restart". Health polling continues behind the panel so a slow first-run install with broken IPC still reaches the app. - clear_webview_cache_and_relaunch (Rust): writes a marker and relaunches; the fresh process deletes EBWebView at the top of run() before any webview exists (WebView2 holds locks while running), with a bounded retry while the old instance exits. Runtime cfg! guards keep the whole path compiling on every platform. Tauri 2 exposes no reliable flag for the postMessage-fallback mode (closure-local in its injected ipc.js), so the logged detector is the observable combination: zero IPC signals + working plain HTTP. Fail-before/pass-after regression tests: hung invoke + healthy HTTP → ready; hung invoke + dead backend → recovery panel, then auto-continue; working IPC → normal path untouched, zero HTTP polling. Plus watchdog state-machine unit tests and recovery-panel render/interaction tests (6/7 fail on the pre-fix component). Troubleshooting doc gains the matching section (docs-sync). Fixes #879 Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
be1ec3ade0 |
fix(platform): declare Intel-Mac local backend unsupported — honest first-run gate + docs (#889); Windows portable-install docs (#766 follow-up) (#891)
torch >=2.3 ships no macOS x86_64 wheels (transformers 5.x needs torch >=2.6), so `uv sync` can never resolve on an Intel Mac — per the platform-parity rule the honest option is declaring the platform unsupported, not letting first launch die in a raw resolver error: - bootstrap.rs: pre-check on macOS x86_64 before any venv create / uv sync (first-run AND repair paths) fails fast with an actionable message (remote-backend escape hatch + docs link); healthy pre-torch-bump venvs are deliberately untouched. Unit test pins the message's load-bearing phrases. - BootstrapSplash: routes the failure to a dedicated localized hint (bootstrap.hint_intel_mac, all 21 locales) and suppresses the useless Retry-oriented hints for it. - README + docs/install/macos.md (+ troubleshooting #9): every Intel-Mac support claim now says UI-installs-but-backend-cannot-run, including the from-source path (also broken); remote backend documented as the only use. - release.yml: #889 note on the macos-15-intel leg — artifact is UI-only; keep-or-drop is an owner call, deliberately not changed here. - docs/install/windows.md: new "Portable install (Windows)" section promised in #766 — custom MSI wizard folder / msiexec INSTALLDIR=..., what lives in OmniVoiceStudio-Data next to the exe, and the Program-Files-greyed-out why. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e347f99542 |
fix(tts): bound + reset the GPU pool on a hung generate so it can't brick the backend (#730 class) (#851)
* fix(tts): bound + reset the GPU pool on a hung generate so it can't brick the backend (#730 class) A GPU job that wedges on some Windows+CUDA setups occupies its worker forever — run_in_executor can't cancel the thread — so on the 1–2 worker pools we ship, one stuck job starves every other request and the next action surfaces as the misleading "Can't reach the local backend" even though the process is alive. ASR/dub/model-load already bound+reset the pool on hang (#730). The TTS **generate** paths (generation.py, tts_stream.py) were the last unguarded GPU dispatch — and the residual on-main reports (#850 #802 #755 #723 #721, plus the 0.3.7 generate cohort) all fail on generate:start (audio). - model_manager: add run_on_gpu_pool_guarded() + GpuJobTimeoutError, a generalized version of the ASR guard so every GPU dispatch shares one bound+reset recovery path. Env-tunable via OMNIVOICE_GENERATE_TIMEOUT_S (default 300s). - generation.py: route both inference branches + the reference-clip transcribe through the guard; map a timeout to an actionable 503. - tts_stream.py: same guard on the streaming path (timeout → error frame). - test_generate_timeout_730: fail-before/pass-after regression (timeout resets pool + restores capacity, happy path, env override, no-reset exec). - docs + CHANGELOG: extend troubleshooting §14 to cover generate; document the new env var. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(tts): extend the GPU-pool hang guard to batch/dub/archetype/openai-compat generate (#730 class) The generate-hang class wasn't only in Studio + streaming: batch generate, the dub per-segment + preview generate, archetype preview render, and the OpenAI-compat /v1/audio/speech path all dispatched the TTS model to the GPU pool with no wall-clock bound either. Any one of them wedging on a Windows+CUDA hang starves the pool and bricks the backend the same way. Route all of them through run_on_gpu_pool_guarded so the whole class is closed — a hung generate anywhere resets the pool and returns an actionable timeout instead of a dead backend. Batch/dub recover per-segment on a fresh worker; drop the now-dead loop/_gpu_pool/asyncio locals ruff flagged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
522bbddccf |
feat(translate): highlighted Install affordance for uninstalled engines + dismissable/auto-clearing error banner (#847)
Two related Dub-tab translation-flow fixes, one PR. TASK 1 — proactive, highlighted Install affordance in the translate engine selector (replaces "find out only via a translate-time 400"): - FROM-SOURCE lane (activeEngineUnavailable && !enginesSandboxed): the muted install chip is promoted to a HIGHLIGHTED brand-accent Install button, still wired to handleInstallEngine(translateProvider) with the installing/disabled state. Selecting any uninstalled engine surfaces it immediately. - FROZEN lane (enginesSandboxed): pip install is impossible in the read-only, signed packaged env, so the disabled "needs dev install" span becomes an equally highlighted button opening a popover with (1) the exact install command + copy-to-clipboard, (2) one-click "Switch to Argos (bundled, offline)" — the guaranteed importable escape hatch, and (3) a Docs link via the existing Tauri shell.open path. Gated on the existing `sandboxed` flag, not platform. - Single-source install command: new translation_engines.install_command() is the one source of truth; list_engines() stamps `install_command` per engine and BOTH the argos + deep_translator translate-time 400 messages build their command from it, so the proactive button and the 400 can't drift. engines.ts gains `install_command: string | null`. TASK 2 — the translation error banner now dismisses and clears (class fix): - Root cause: handleTranslateAll never cleared dubError, so a stale 400 survived even a successful retry. It now clears at the start of every attempt. - Corrective-action clears (whole class): changing the engine and installing the package both clear dubError (wrapped setTranslateProvider + handleInstallEngine in DubTab). - DubFooter's banner gains a × dismiss and a guarded auto-timeout (skipped while generating so live per-segment errors persist). i18n: 8 new dub.* keys translated across all 21 locales. Docs: new docs/dubbing/translation-engines.md (from-source vs packaged build) linked from the popover Docs button + a troubleshooting cross-reference. Tests: FE regression for both lanes + never-installs-when-sandboxed + banner dismiss/auto-clear; BE regression that list_engines() install_command is embedded verbatim in the dub_translate 400s. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
de3d83f14b |
docs: add rust as prerequisite for from-source builds (#704)
Adds Rust/Cargo as a from-source build prerequisite across the linux/macos/windows install docs. Thanks @Deepakv2104. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2339cc8e85 |
fix(model): actionable "reinstall transformers" hint on a corrupted-install model-load error (#676)
A model load failed with `[Errno 2] No such file or directory:
'…/site-packages/transformers/models/qwen3/modeling_qwen3.py'` — the user's
transformers install was incomplete (the file is missing while a correct 5.3.0
install has it; an interrupted `uv sync` / antivirus / partial update drops it).
The System Check showed the raw path + "Check logs and try restarting", which is
useless — restarting can't restore a missing file.
Two fixes:
1. core.failure.classify(): recognize this corrupted-install variant. It's a
FileNotFoundError, not an ImportError, so the existing TRANSFORMERS_IMPORT
match ("could not import module"/"AutoFeatureExtractor") missed it. Now also
matches a "no such file"/"errno 2" + "transformers" + "site-packages" signal
(substrings checked separately so it works on POSIX `/` and Windows `\`
paths). An unrelated package's missing file is NOT mislabelled.
2. model_manager._load(): build the /model/status error via build_failure so it
carries the classified hint AND strips the home dir, instead of storing the
raw str(exc). The System Check now shows "Your transformers install is
incomplete. Reinstall it (uv pip install --reinstall transformers) or switch
ASR to faster-whisper" — the existing TRANSFORMERS_IMPORT hint.
Docs: troubleshooting §1a documents the error + the reinstall fix.
Test: test_failure_classify.py pins the POSIX + Windows path forms classify as
TRANSFORMERS_IMPORT with a "reinstall" hint, and that an unrelated package's
missing file does not.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
252f0d4fac |
fix(asr): bound whole-file transcription so a stall isn't reported as "can't reach backend" (#656)
A Windows/CUDA user (Vietnam) hit "Can't reach the local backend" only when dubbing/transcribing. Their log proves the backend started fine — model loaded, preload complete, 25 models — and the log ends right after `whisperx transcribing …tmp.wav`. The backend was alive; the *transcription* stalled (large-v3 ASR contending with the resident TTS model for VRAM on an 8 GB-class GPU), which the UI surfaces as an unreachable backend. Root cause (class, not instance): the chunked dub pipeline already bounds each chunk (OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S), but the *whole-file* transcribe paths ran unbounded: - dub QC re-transcribe (dub_export) - dictation (capture) - OpenAI-compat /audio/transcriptions A slow/stuck transcribe on any of these hung the request AND held a GPU-pool worker — indistinguishable from a dead backend. Fix: add run_transcribe_guarded() in services/asr_backend.py — a shared asyncio.wait_for wrapper (ASRTimeoutError, a TimeoutError subclass) with a generous env-tunable bound (OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S, default 300 s). On timeout the request returns 504 with actionable guidance (backend is alive; free VRAM / pick a smaller ASR model / use CPU; restart to clear the stuck worker) instead of hanging forever. Wired into all three whole-file paths. Docs: new troubleshooting §14 — "Can't reach the local backend during transcription/dubbing" — explains it's ASR weight/VRAM pressure, not a network/ mirror problem, and corrects the misconception that a "Network → Restricted/Global mirror" Settings toggle exists (the Network control is LAN sharing). Serves the #602/#585/#567 "can't reach backend" cluster. Test: backend/tests/test_asr_transcribe_timeout.py — slow fn raises ASRTimeoutError with the actionable message, fast fn passes through, subclass-of-TimeoutError so the openai_compat broad catch still maps to 504. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
79f3e35682 |
docs(troubleshooting): add stuck-download / incomplete-cache recovery (#622) (#643)
The 'stuck on the download page, model folder has only refs/ no weights' case (a connection dropping/blocking mid-pull) is a recurring support report but wasn't in the install troubleshooting guide. Add section 13 with the recovery steps + antivirus/VPN/mirror escalation + a huggingface-cli manual fallback. Docs-only; no version bump. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0e17caa52a |
fix(install): actionable torch-wheel-download failure + local-wheel recovery (#569) (#574)
#569: on a restricted network the first-run install fails downloading the ~2.5 GB cu128 PyTorch wheel from download.pytorch.org, and the app won't launch. Two problems: the error told users to "set UV_DEFAULT_INDEX to a mirror" — which CANNOT redirect torch, because it comes from a *named, explicit* uv index (uv 0.11 rejects index-name override values and `--frozen` pins the exact wheel URLs); and there was no way to supply a manually-downloaded wheel. - Detect a torch/pytorch-host `uv sync` failure and emit torch-specific guidance (Clean & Retry → VPN → drop the wheel locally) instead of the wrong mirror advice. - Add a local wheel-drop dir `<env_root>/wheels` (survives Clean & Retry) wired via `UV_FIND_LINKS`. On a frozen-sync torch-download failure WITH wheels present, retry NON-frozen with find-links so uv re-resolves from the local wheels. Verified empirically: a non-frozen find-links sync installs from a local wheel fully offline, while a `--frozen` sync ignores find-links — so the retry is the only mechanism that can consume a dropped wheel. Best-effort: if it can't satisfy, it fails identically to before and the actionable error still fires. - docs/install/troubleshooting.md: new "#12 CUDA PyTorch wheel download fails" entry (docs-sync) — the offline wheel-drop path + why a PyPI mirror can't fix this index. Note: an automatic mirror redirect for the cu128 index is intentionally NOT shipped — uv provides no working override for a named explicit index, so it couldn't be verified; the offline wheel path is the reliable escape hatch. Test: sync_failure_is_torch_download host/keyword detection + negative guard. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
af3da58584 |
fix(bootstrap): force-reinstall setuptools so pkg_resources repair actually works (#248) (#495)
The auto-repair ran `uv pip install setuptools>=75,<80`, which `uv` treats as "already satisfied" (no-op, "Checked 1 package in 5ms") whenever setuptools' *metadata* is present but its `pkg_resources` files are gone — the common cause being Windows Defender quarantining `pkg_resources/`, or a partial extract on a restricted network. So the repair never restored the files, the post-check failed, and users hit the #248 dead-end. The error message *also* told them to run the same no-op command, so the suggested manual fix didn't work either (reported on Discord, Win11 + RTX 5070 Ti). Fix: both repair sites in bootstrap.rs now use `--reinstall` (the flag already used for the ROCm torch repair), which force re-extracts pkg_resources even when uv thinks setuptools is satisfied. The fail() message and the failure.py hint now suggest `uv pip install --reinstall 'setuptools>=75,<80'` + an antivirus-exclusion note, and docs/install/troubleshooting.md (#pkg_resources-missing) is updated with the real cause (metadata-present/files-missing) + AV guidance. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6306f6edae |
docs: official Docker Hub image palashdeb/omnivoice-studio (#388)
Link the published Docker Hub repo (https://hub.docker.com/r/palashdeb/ omnivoice-studio) as an official image alongside GHCR in the README install list and docker.md header. Same images, same tags. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |