cb581f38df4e11ea34dc287c485b71c27aaf2c31
79
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
41c098e009 |
feat(demos): ship the demo audio and video the app already advertises (#1517)
* feat(demos): ship the demo audio and video the app already advertises Every demo asset in the app was a dead link on anything but a Mac. `personalities.py` has carried a `preview_url` for each of the seven voice-design presets since they were added; DictationDemo.jsx posts three bundled WAVs to /transcribe so the feature can be shown without microphone permission; the Dub workspace reads a manifest and plays a source video plus four dubbed languages. None of those files were committed, because the tooling that renders them (scripts/build_demos.sh, scripts/build_dub_demo.sh) hard- requires macOS `say` — it even carries a `TODO: add espeak-ng path for Linux contributors`. So the presets returned 404, the replay buttons did nothing, and the dubbing demo never loaded. Rendered with VoiceStudio's own engine, which runs wherever the app does: - 7 voice-design previews (2.2 MB) - 3 dictation replay clips (1.1 MB) — verified by transcribing them back: the conversational and French clips round-trip exactly - dubbing demo: source + 4 dubbed videos with subtitles and manifest (9.6 MB) Tooling fixes this turned up: - build_dub_demo.sh wrote to backend/assets/demo/dubbing, but main.py mounts backend/assets/samples at /demo_audio — so the frontend's /demo_audio/demo/dubbing/manifest.json could never have resolved even after a successful Mac build. Output moved under the mount. - `say` is now the fallback rather than the requirement: the new scripts/render_dub_demo_audio.py renders the five tracks with the engine and the shell script picks them up. - The five demo paragraphs lived in two files. They are now one JSON both read — two copies is one edit away from a video whose subtitles disagree with it. - render_demos_omnivoice.py peak-normalized, which a single-sample transient defeats: the Helpdesk preset landed at -30 dB RMS against -17 dB for its neighbours, so the preview row played at wildly different volumes. Now EBU R128 at -18 LUFS with a -1.5 dBTP ceiling. - …and pinning the output rate, because loudnorm resamples to 192 kHz internally and writes there unless told otherwise, which turned 2.1 MB of previews into 17.5 MB of identical-sounding audio. - update_manifest() looked for a manifest at a path nothing writes, so it always printed "not found" and did nothing. - Dictation is rendered here now too. It was excluded on the grounds that `say` was good enough and engine TTS was overkill — true only on macOS. tests/test_demo_assets_exist.py resolves every advertised URL against the directory main.py actually mounts, and checks each dubbing subtitle matches the script its manifest entry claims. A missing static file is not an import error and not a failing request; nothing would have caught this otherwise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(changelog): stamp the demo-asset entries with their PR ref Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(demos): watermark rendered demo audio, and harden the render scripts Review findings on #1517: - Greptile P1: the renderers wrote engine output straight to disk, so a re-render shipped demo audio with no provenance mark. These clips play back to users as VoiceStudio output — they are synthetic audio leaving the app like any other, and now go through mark_synthetic (#1169), the one chokepoint every producing route uses. It runs on the file AFTER loudnorm, since loudnorm re-encodes what it is handed, and says so loudly when marking is unavailable rather than committing an unmarked asset. The dubbing renderer shares the same helper. - CodeRabbit: build_dub_demo.sh checked only source.src.wav before deciding it could run without macOS `say`, so a Linux or Windows run with four of five tracks present reached a missing one, called `say`, and left a half-built bundle. It now requires all five. - CodeRabbit: shutil.move over an existing path delegates to os.rename, which raises FileExistsError on Windows — os.replace overwrites atomically everywhere. - CodeRabbit: the preview test discovered presets in a parametrize argument, importing app code at collection time and leaving core.personalities in sys.modules for later tests. Discovery moved into the test body. CI: the rendered dub bundle's zh/ja subtitles, its manifest and the script source are dubbing CONTENT, not UI strings — allowlisted in test_no_hardcoded_cjk.py with that justification. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(demos): a render that cannot be watermarked fails instead of warning CodeRabbit and Greptile, #1517: mark_synthetic degrades rather than raising — correct for generation, wrong for a render script, whose whole job is to produce files a human then commits. A printed warning on a scrolling console is not a gate, so both scripts exited 0 with unmarked assets sitting on disk ready to commit. They now raise, with the reason and the fix; OMNIVOICE_DEMO_ALLOW_UNMARKED=1 stays for a local listen. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: stop a flaky dependency fetch from failing green runs en-core-web-sm resolves to a direct GitHub release URL, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own three retries all land within the same few seconds and fail together, so the whole job dies on a dependency that has nothing to do with the change under test — it cost #1518 and #1517 an otherwise-green run tonight. Two changes: back off between whole `uv sync` attempts, which is what actually clears it, and pass --no-sync to the pytest steps. `uv run` re-resolves the environment before running, so every test step was a fresh chance to hit the same fetch even though the install step had already synced — that is exactly how #1518 failed, in the isolated backend/tests step, with all 5467 tests already passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: one retry seam for every uv sync, not just the job that failed last en-core-web-sm resolves to a direct GitHub *release* URL rather than a package index, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own retries all land inside the same ~10 seconds and fail together, so a job dies on a dependency unrelated to the change under test. Tonight that cost four otherwise-green runs across #1515, #1517 and #1518 — and the first fix only covered the Tests job, so the next failure simply moved to Smoke (Linux), which syncs separately. The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh backs off between whole attempts (15s, 45s, 90s) and every workflow that syncs now goes through it — ci.yml (tests + the platform matrix), release.yml, security.yml, evals.yml. It still fails loudly after four attempts, so a genuinely broken lockfile is not disguised as a flake. The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke matrix has always set, which is part of why it was the one that kept dying; it has them now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ci): pin the Intel-Mac contract by intent, not by command spelling test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted the literal line `run: uv sync --extra pockettts`, so routing every sync through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The contract it exists to protect is that the pockettts extra installs ONLY on backend_supported legs — which the regex now pins, while leaving how the sync is invoked free to change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: keep every uv run out of the resolver, and bound the retry budget CodeRabbit, #1517: - `uv run` re-resolves before running, so the smoke suite, the worker-artifact tests, the release test run and the eval run were each a fresh chance to hit the flaky direct-URL fetch outside the retry loop. All of them pass --no-sync now; the environment is already synced by the step that owns the retries. security.yml's `uv run --with pip-audit` is deliberately left alone — it layers an ephemeral package rather than running the project's own tests. - The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a 120 s timeout on the smoke matrix). Three attempts and 60 s of total backoff outlast the refusals actually observed while staying well inside the jobs' timeout-minutes. - The Intel-Mac contract test pinned the smoke command literally too, so --no-sync tripped it exactly like the sync line did. Same fix: assert the contract (smoke runs only on backend_supported legs), not its spelling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4fc07c21fd |
fix(appimage): stop shipping a dangling .DirIcon, and prove it in CI (#1518)
* fix(appimage): stop shipping a dangling .DirIcon, and prove it in CI
The Linux icon is blank because the AppImage's .DirIcon is an absolute
symlink into the machine that built it. From the published v0.4.2:
.DirIcon -> /home/runner/work/OmniVoice-Studio/OmniVoice-Studio/frontend/
src-tauri/target/x86_64-unknown-linux-gnu/release/bundle/
appimage/OmniVoice Studio.AppDir/OmniVoice Studio.png
That path exists on nobody's computer. The link dangles the moment the
AppImage leaves CI, so file managers have no icon for the file, and the
integration tools that read .DirIcon install nothing. A dangling symlink is
not a build error — the bundle packs, runs, and passes every check we had —
which is how it shipped for a whole release without anyone noticing.
Locally built AppDirs are worse: both .DirIcon AND the root .desktop symlink
come out absolute, so a from-source bundle has no readable desktop entry
either, which is why the icon is missing in the menu and the dock too.
- `.DirIcon` is now a real file, copied in through `appimage.files` — the
same seam that already places the WebKitGTK marker.
- `bundle.category` is set, so the generated desktop entry stops emitting an
empty `Categories=`. That is not the same as omitting the key:
desktop-file-validate rejects the entry and menu builders skip it.
- verify-apprun-bundle.sh — already run against the extracted AppImage in the
release job — now fails when .DirIcon is missing or resolves outside the
bundle, when the .desktop entry does not resolve inside it, when Icon=
names a file that is not at the AppImage root, or when Categories= is
present but empty. Its unit test covers each of those, including the exact
shape v0.4.2 shipped.
The `.DirIcon` copy cannot be verified without a full release build, so the
guard is the load-bearing part: the next release either passes it or fails
loudly. It can no longer ship blank in silence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(changelog): stamp the AppImage icon entries with their PR ref
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(changelog): fold the AppImage icon fix into the existing Fixed section
CodeRabbit (#1518): the Unreleased block must carry one `### Fixed`
section of one-line entries. Merge the two entries in and drop the
narrative and the version reference.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
64973d1829 | fix(release): validate wrapped AppImage launcher | ||
|
|
037a5689de | fix(worker): address legacy transport review findings | ||
|
|
b0a6fdfdc2 |
Merge remote-tracking branch 'upstream/main'
# Conflicts: # CHANGELOG.md # frontend/src/components/settings/ModelStoreTab.jsx # frontend/src/i18n/locales/ar.json # frontend/src/i18n/locales/de.json # frontend/src/i18n/locales/en.json # frontend/src/i18n/locales/es.json # frontend/src/i18n/locales/fr.json # frontend/src/i18n/locales/hi.json # frontend/src/i18n/locales/id.json # frontend/src/i18n/locales/it.json # frontend/src/i18n/locales/ja.json # frontend/src/i18n/locales/ko.json # frontend/src/i18n/locales/nl.json # frontend/src/i18n/locales/pl.json # frontend/src/i18n/locales/pt.json # frontend/src/i18n/locales/ru.json # frontend/src/i18n/locales/sv.json # frontend/src/i18n/locales/th.json # frontend/src/i18n/locales/tr.json # frontend/src/i18n/locales/uk.json # frontend/src/i18n/locales/vi.json # frontend/src/i18n/locales/zh-CN.json # frontend/src/i18n/locales/zh-TW.json |
||
|
|
40569d0657 |
Merge remote-tracking branch 'upstream/main'
# Conflicts: # CHANGELOG.md |
||
|
|
5832a81bb6 | feat(onboarding): refresh the bundled demo voice | ||
|
|
673e544812 |
Merge origin/main into feat/worker-protocol-v1
Three conflicts, all additive on both sides — resolved by keeping both
rather than choosing, since either side's entries were real shipped work:
* CHANGELOG.md — remote-GPU entries against branding, IndexTTS 2.5 and
the recording-input work
* setup/download.py — the per-target progress reset against main's
active-install tracking; both belong in the same finally block
* docs/features.yaml — the remote-worker and model docs against
docs/branding.md
Backend 5349 passed, frontend 1871 passed. `bun install --frozen-lockfile`
reports no changes, so the Docker build sees the same tree CI does.
|
||
|
|
c0e1753248 | fix: close AppImage termination races | ||
|
|
c16db04f86 | fix: make process cleanup explicit | ||
|
|
670d9bc37a | fix: stop extracted AppImage before prod reset | ||
|
|
aa1d739843 |
feat(workers): dubbing goes remote, and the protocol stops lying to old workers
The remote-GPU line, verified on hardware rather than asserted. **Dubbing renders on the worker.** dub_generate.py dispatches the coarse `dub_segments` operation through the gateway, following the audiobook pattern: per-unit local fallback after consecutive remote failures, one aggregated notice rather than one per segment. A 40-minute dub that loses its worker at segment 200 degrades instead of producing 200 error rows. **An out-of-date worker is now refused by name.** This was the worst defect in the plan and it was silent: an un-upgraded worker registered cleanly, then ignored `inputs` and rendered a clone with NO reference audio — returned as success. A plausible wrong result with nothing anywhere to surface it. Workers now declare features, and one missing them is turned away with the features named and `no task was run`. Verified live: a worker one commit behind was correctly refused. **"Offline" and "cannot run this" are different facts.** Asking a live worker for an engine it lacks answered "is offline or cannot be reached. Wake the selected worker" — while that worker reported ready, one free slot and 3.6 ms latency. The user was sent to wake a machine that was already awake. The scheduler now distinguishes absent from present-but- incapable, and names the engine rather than the operation, because the engine is the thing a user can install. **An engine with no catalog entry is no longer hidden.** A `repo_ids` non-emptiness check had been implemented as a runtime filter, so a worker silently refused to advertise any engine lacking a models.yaml entry — which is four registered engines, including CosyVoice. Users with those already installed would have lost remote support with only a log line. Empty `repo_ids` now means "not downloadable here", never "not runnable". **And a script so this stops being done by hand.** scripts/verify-remote-worker.sh runs the per-phase acceptance checks against a live worker, non-destructively. Its preconditions are the mistakes that cost the most time: exactly one listener on the control port (two instances silently shared it), and never detecting the worker with a pgrep pattern that matches the ssh shell running it. Its first real run found the dubbing picker claiming remote placement. That turned out to be the CHECK being stale, not the picker — the port had landed since it was written. It now asserts self-consistency instead: the picker may claim remote only for an operation the control plane actually advertises as remotely producible, which cannot rot the next time an op is ported. Backend 5291 passed, frontend 1812 passed. Acceptance script: no automated failures across Phases 4-8 on an RTX 4090. Four checks remain MANUAL by design — true airplane mode, concurrent downloads, killing a worker mid-audiobook, and the model-list UI — and are reported as unverified rather than passed. |
||
|
|
04410a458d |
Release VoiceStudio 5.0.0 (#1487)
Complete the VoiceStudio identity, release documentation, assets, version mirrors, and safe cross-platform development startup. |
||
|
|
c643706d07 |
feat(workers): make a remote GPU actually run a task, end to end
Selecting a remote worker repainted a badge and nothing else. The cause was not subtle: `scheduler.submit` had no production caller, and `routing.decide()` was read only by the status endpoint that paints the header. Remote execution was a complete, tested pipeline with no producer at its head. This adds the producer and fixes the defects that made the pipeline unable to carry a real job: - Nothing routed to the scheduler. Adds `POST /workers/tasks` (loopback-gated, **development-only** until the gateway lands) and `Scheduler.wait`, backed by per-task futures rather than the unregisterable `on_change` listener list. - Every task over two minutes died. No worker ever sent `TaskProgress`, so the 120s progress lease expired mid-render — including during the cold model load, which happens after `TaskStarted`. Workers now report progress and emit a keepalive, bounded by the phase's absolute budget so it renews the lease without deleting the only enforced bound in the system. - The executor rebuilt its engine per task (`return cls()`), so every job paid a cold load. Engines now share one instance cache with the router, resolved by the assignment's engine — never `get_active_tts_backend()`, which returns the worker machine's own Settings preference and would silently run the wrong engine. - One lease expiry took a worker offline permanently: parked slots were never reclaimed. Parks now expire on a TTL, and are deliberately NOT reconciled against the worker's own load report — at a ceiling of one the only task such a worker can report is the wedged one, so "busy" would drop the park and the next idle heartbeat would hand out a slot with a live GPU thread (#730/#1190). - A worker that dropped and reconnected mid-render had every liveness frame discarded: task frames were fenced on the live session epoch, which bumps on every reconnect, while the worker echoes the ref stamped at dispatch. The control plane then expired a task whose GPU was still rendering, and swallowed the failure report when it went wrong. Fenced per attempt instead. - A result from one worker could commit another's task, after which the owner's real delivery arrived as a duplicate and its audio was discarded. "Unknown attempt" and "another worker's attempt" are no longer the same answer. - An oversized result was a poison pill, re-sent identically on every reconnect and permanently disconnecting the worker. It is now a terminal `RESULT_TOO_LARGE`, which is also classified — it was falling through to TRANSIENT and retrying a re-render that could never fit. - `_store_inline` joined the artifact directory with worker-supplied ids, and `os.path.join` discards its prefix on an absolute component. Paths are now minted control-plane-side and resolved through `core.path_security`. - Remote synthesis bypassed `mark_synthetic`, and the guard that exists to catch exactly that walked only `backend/api` and `backend/services` — so it stayed green while a fourth unmarked producer shipped. Marking moved to the worker's tensor stage; the guard now walks `backend/worker` too. Also adds pre-rendered voice previews (`services/gallery.py`), so browsing the gallery no longer needs a GPU or a downloaded model. The manifest is verified against the updater's release key already baked into the binary; a fresh install hears voices without downloading 2.4GB first, and everything falls back to local rendering when the gallery is unreachable. Verified on hardware, not just in CI: 1728 characters submitted to an RTX 4090 returned 105.94s of 24kHz audio in 23.9s, committed and served from the artifact store. Not yet done, and deliberately not claimed: the keepalive fix cannot be exercised end-to-end on fast hardware, because any job long enough to reach the 120s lease produces audio past the 8MiB inline cap. Chunked `UploadResult` has to land first. Pinning to the worker the user chose is also still absent, so "Remote" reaches a remote GPU but not necessarily the one on the badge. |
||
|
|
7924b35f8d | Merge remote-tracking branch 'origin/main' into feat/worker-protocol-v1 | ||
|
|
43de1c794c |
feat(workers): remote GPU workers over a versioned gRPC protocol
Send individual jobs to GPUs on your other machines while everything else stays local. Opt-in, off by default: with the toggle off there is no listening socket, no certificate and no background loop. Design follows remote/goal_v2.md, the council-revised goal doc. The decisions that shaped the code, and why: * A disconnect is an unknown outcome, not a failure. The original design reassigned on disconnect while also describing the case where the worker had already finished — following both guarantees duplicate execution. An attempt now holds a grace window; a worker returning inside it commits its result and no second attempt is ever made. * At-least-once execution, exactly-once result commit. The result is persisted BEFORE it is acknowledged, so a crash between the two cannot silently lose a finished render. * Deadlines are phased (accept -> model load -> execute -> deliver) and liveness is a progress lease. The old fixed 30s execution budget was two orders of magnitude below what this product actually does; silence is the failure signal, not slowness. * Capacity is derived from free VRAM, never configured: a static value corrupts output under torch.compile thread affinity (#315) and aborts the process on small cards (#567). * A circuit breaker replaces the reliability-score/quarantine machinery, which had no recovery path (no probation workload exists in a TTS product) and penalised consumer networks for existing. * Identity is a keypair the worker generates and never sends. A server-assigned id is a name, not an authenticator, so revocation of one would be theatre. Enrollment tokens are single-use and carry the control plane's certificate fingerprint for pin-on-first-use. Adds the domain core, scheduler, durable task store, gRPC transport, worker agent, management API, Settings panel, and docs. Protobuf reserves the tenant/trace/usage fields a hosted control plane would need, since adding them later means upgrading a whole fleet. Includes tests for the failure paths that matter: duplicate delivery, stale-session fencing, reconnect reconciliation, grace expiry, breaker attribution, and a real end-to-end TLS round trip. |
||
|
|
fb78d87ad9 | test(appimage): verify packaged WebKit marker | ||
|
|
5be4a26903 | fix(appimage): ship the compatibility launcher | ||
|
|
93025d9a81 |
fix(dictation): stop the widget stranding an empty square, repair the swept data dirs (#1398)
The dictation hotkey could leave a blank dark square stuck on the desktop with no way to dismiss it. Three defects compounded: the tray listener's effect depended on [state], so it detached across an await on every state change and a press landing in that gap was lost; an idle pill renders null, so the window Rust had already shown was empty; and the opaque chrome background made that empty window a hard-edged square. Nothing could hide it — dismiss() is only reachable from the X button, Esc, or a post-session timer, none of which exist for a session that never started. Fixed at the invariant rather than the call sites: the listener subscribes once for the component's lifetime, the widget window's chrome background is transparent, and an idle-but-visible window reconciles itself to hidden. The reconcile is polled (a dropped press changes no React state, so there is nothing to key an effect off) and aborts if its effect is torn down mid-check, so it can never hide a dictation that has just started. Also in scope: - The rename sweep had repointed three data-dir literals at a brand-named directory that does not exist, so smoke-test.sh verified a directory the backend never writes and desktop-prod.sh silently stopped clearing backend state on Windows. Both invisible on macOS, where they are usually run. A guard test now pins the assignments specifically. - The dictation model picker's download sizes were wrong for all seven models, in both directions — Parakeet TDT v3 (the recommended default) understated 180 MB against an actual 670 MB, while the low-RAM fallbacks were overstated threefold, discouraging exactly the choice that would have helped. Measured from the published repos and pinned by a test. - The 0.6B Parakeet models now decode on more threads, capped by host cores and still overridable. - uninstall.ps1 gained a UTF-8 BOM (Windows PowerShell 5.1 mis-decodes its non-ASCII output without one), and sponsor.yml lost its last OmniVoice references. |
||
|
|
5cab8e0149 |
feat: rename the product to VoiceStudio (previously OmniVoice-Studio)
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
|
||
|
|
a23e69d014 |
chore: point every repo reference at github.com/debpalash/VoiceStudio (#1394)
The repository was renamed. 724 references across 59 files now point at the new URL — README badges, docs, install guides, the updater's releases API call, CONTRIBUTING, the Colab link and the probe harness. GitHub redirects the old URLs, so nothing was broken in the meantime. Deliberately NOT renamed, because each breaks something on a user's machine: the Tauri bundle identifier (the path to every existing user's data), /usr/lib/omnivoice-studio and the compose container names, and the published Docker image paths. The image path needed a code change to STAY still: docker.yml derived it from github.repository, so the next build would have published to ghcr.io/debpalash/voicestudio while Docker Hub, a hardcoded literal, stayed put — everyone pulling the documented GHCR path would have kept receiving the last pre-rename image forever. It is now pinned, with a test that fails if it ever derives from the repo name again. Also makes the probe's repo-name assertion shape-based: it hardcoded the old name and failed on every PR after the rename while the code it tests worked perfectly. |
||
|
|
9877e7c218 |
fix(ci): bind the preview manifest to its run, not to sibling upload times (#1387)
The nightly preview build had been refusing to publish its own healthy manifest since 2026-08-05 — all four matrix legs green, but the macOS bundles uploaded a few minutes ahead of the slowest versioned artifact, and the freshness check compared the version-less darwin tarballs against their siblings with two minutes of slack. Legs finishing minutes apart is normal, so the comparison itself was wrong, and Preview-channel users quietly stopped getting builds. The tarballs are now tied to the run that produced them: anything uploaded after this run's first job began executing belongs to it. A concurrency group serializes preview runs so that holds, and the anchor is the earliest job start rather than the run's created_at (which is stamped while a run is still queued, and would let a queued run claim the previous run's uploads). The preview-notes job also gains the actions: read scope its run-metadata lookup needs, with a warning-and-degrade path so a permissions regression cannot take the channel down again. Regression tests cover the 2026-08-05 shape, the genuinely stale case, clock skew at the boundary, the no-timestamp fallback, per-ref concurrency scoping, and the required permission. |
||
|
|
eeffe6c2d1 |
fix(gguf): a source-built runtime must actually run (#1348) (#1384)
A meticulous report from an LXC/CPU-only source install surfaced three real defects: the build script deleted the libggml shared libraries a dynamically-linked build needs (first spawn died with exit 127), the hardcoded 120s per-spawn kill switch reaped legitimate CPU-only renders, and OMNIVOICE_ALLOWED_ORIGINS — the only fix for cross-origin browser access — was documented nowhere. All platform branches of scripts/build-omnivoice-tts.sh now copy the shared libs next to the binary, the CI artifact glob uploads them, and the backend puts bin/ on the loader path for every spawn of the engine binary. The timeout defaults to 600s (above the pool guard's well-diagnosed 300s deadline), is tunable via OMNIVOICE_GGUF_GENERATE_TIMEOUT_S with non-finite values rejected, and the timeout error names the knob. CORS documented in api-auth.md with a pointer from remote-gpu.md. Regression tests pin the spawn-env rule, the per-branch copy rule, the artifact glob, and the timeout behavior. |
||
|
|
c117bca09e |
fix(release): rebuild + cryptographically verify the preview updater manifest (#1327) (#1362)
* fix(release): rebuild + cryptographically verify the preview updater manifest Since ~2026-07-13 every nightly matrix leg logs 'Signature not found for the updater JSON. Skipping upload...' - tauri-action uploads the bundles and .sig companions but never refreshes latest.json. Combined with the 'Clear this arch's stale preview updater bundle' step (which deletes and replaces the version-less macOS tar.gz every night), the preview manifest's darwin signatures no longer match the published files: macOS Preview users hit 'The signature verification failed' on every update (latest.json frozen at 2026-07-13, tar.gz replaced nightly). Two changes, both in the single post-matrix preview-notes job (no per-leg race): 1. Rebuild latest.json from the release's real assets and their .sig companions, then clobber-upload. The manifest can no longer drift from the files it describes, regardless of what tauri-action's own updater-JSON path does or skips. 2. Extend the existing manifest verification with a cryptographic check: every signature in latest.json must verify (minisign file sig + trusted-comment sig) against the artifact it points at, using the updater pubkey from tauri.conf.json. Parity and version format both passed for 2+ weeks while every darwin entry was unverifiable - this is the check that was missing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(release): refuse a preview manifest built from two different runs Bot review findings on this branch, all fixed here: - The AppImage and MSI were picked independently by highest run number and the larger N became *the* version, so a matrix where one leg failed or was re-run published a manifest advertising X.Y.Z-5 while handing Windows users the -4 MSI. That is the same manifest/artifact drift this job exists to end, reintroduced by the fix for it. Require both legs to come from one run and fail loudly otherwise: leaving the previous manifest in place is a visible, already-understood state; shipping a mismatched one is not. The darwin tarballs carry no run number, so the signature check in the following step is what pins those to the published bytes. - persist-credentials: false on the checkout — nothing here pushes to git. - Floor-pin the cryptography install; this step decides whether a signed manifest is trustworthy, so it is the one dependency worth a bound. tests/test_release_preview_manifest_rebuild.py runs the step body extracted from release.yml against stubbed gh, so it cannot drift from the workflow. Fails before / passes after on the mismatch case. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(changelog): note the preview updater manifest fix (#1327) Co-Authored-By: Pinkers01 <pinky.bouw@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci(release): verify the preview manifest before publishing it, not after Two more review findings on this branch, both valid, both about the manifest being wrong in a way the existing checks structurally cannot see. greptile P1 — verification ran AFTER the clobber-upload. A manifest that failed the check was already live and stayed served; the job merely went red, and every macOS Preview user stayed broken until someone noticed. Verification now runs against the file about to be published, and the upload is the last thing in the step. A refusal leaves the previous manifest in place, which is a visible, already-understood state. CodeRabbit — the darwin entries were not tied to this run. The version comes from the AppImage name; the macOS tarballs were only checked for existence. Signature verification cannot help there, because a stale tarball and its stale .sig match each other perfectly — so a run whose macOS legs never uploaded would advertise this version while serving Mac users the previous build, and since those clients keep reporting the old version the updater would re-offer it forever. They are now bound by upload time, with two minutes of slack for legs that finish apart. The selection rules move out of the YAML heredoc into scripts/build_preview_manifest.py. Three findings in a row have been about WHICH artifacts may be described together, and a heredoc can only be tested by extracting it and stubbing a shell — which is what the previous test file did, asserting against gh stubs rather than against the rules. build_manifest is pure: assets in, manifest out, ManifestRefused on anything it will not describe. 14 tests, including both new refusals and two that pin the workflow still calls the module and still uploads last — an inline copy would pass every other test and ship the original bug. Co-Authored-By: Pinkers01 <pinky.bouw@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Pinkers01 <pinky.bouw@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8f9f778307 |
fix(scripts): desktop-prod:run wiped the data it was documented to preserve (#1333) (#1339)
* fix(scripts): desktop-prod:run wiped the data it was documented to preserve (#1333) `scripts/desktop-prod.sh` emulates a first install, so wiping is its default: it removes the app data dir, `~/.omnivoice` (the SQLite database, every voice profile, all outputs), the Tauri logs and the WebKit profile. `--keep-data` is the only thing that suppresses that block. `--skip-build` is an independent flag that only skips the cargo compile, and `desktop-prod:run` passed it alone — while the script's own header calls that command "re-launch last build (skip compile)" and its closing banner tells you to use it that way. So "just start it again without recompiling" silently deleted the developer's voice profiles and project database, every time. The fix is in the package scripts rather than the flag parsing: making --skip-build imply --keep-data would remove a legitimate combination (fresh data without paying for a recompile). The two stay independent, and the help text now says so. desktop-fresh:run is deliberately untouched — that script is a stricter new-user emulation, so wiping is the point of its name. Tests pin all three rules, and were confirmed fail-before by reverting the desktop-prod:run line. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(scripts): kill the live instance before every launch, not only before a wipe Greptile P1 on this PR, and it was a regression I introduced. The app registers tauri_plugin_single_instance, and that callback ignores the incoming argv — it just refocuses the window the RUNNING process already owns. So starting a second copy over a live one does nothing visible. That was previously masked: kill_running_instances sat inside the `KEEP_DATA = false` branch, so every run happened to kill first *because* every run wiped. Adding --keep-data to the re-launch aliases removed the wipe and would have taken the kill with it — `desktop-prod:run:pill` would have left the user in studio mode with --pill silently discarded, and plain `desktop-prod:run` would have refocused the OLD build instead of the one just compiled, which is the entire point of that command. The kill is now unconditional, before the wipe branch. Its two reasons are independent — zombie-backend-after-wipe, and single-instance-swallows-argv — and only the first was ever about wiping. Adjusted its closing line, which said "safe to wipe" and now also runs when nothing is being wiped. New test asserts the call is not nested inside the KEEP_DATA branch; confirmed fail-before by moving it back. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(scripts): scope the kill to this checkout, warn about an installed app Making kill_running_instances unconditional (so --keep-data re-launches still get past single-instance) widened the blast radius of its pgrep: "OmniVoice Studio.app" also matches an installed /Applications copy, so desktop-prod:run would kill the shipped app a developer was using and take their unsaved work with it. That was previously masked — the kill only ran on wipe runs, where a clean slate had been asked for explicitly. Scope the pattern to ${TAURI_DIR}/target/debug/, which covers both launch shapes and nothing else. An installed instance still gets named rather than ignored: single-instance keys on the bundle id, so it swallows this launch too, and silence would just trade one confusing failure for another. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a3ac1be12d |
fix(ui): report action can't fail silently; correct the Linux data dir (#1177)
CodeRabbit findings 2 and 3 (finding 1, the LAST_FAILURE test race, was already
fixed in
|
||
|
|
aab138ea57 |
fix(backend): surface the shell's start-failure diagnosis instead of "can't reach the backend" (#1177)
The reporter's string is apiFetch's LAST fallback, reached only when no crash
marker exists AND the shell's lifecycle stage is 'failed' or 'unknown'. The
'failed' half was the bug: `BootstrapStage::Failed { message }` carries the
whole diagnosis — exit code plus a ~30-line stderr tail, or the precise reason
`ensure_venv_ready` refused (Intel Mac, a failed `uv sync`, a blocked GitHub) —
and `backendLifecycleStage()` returned only the stage tag, throwing the message
away. Every backend-start failure mode collapsed into one generic, evidence-free
sentence that was also factually wrong: it is not starting, and it will not
recover on its own.
- backendLifecycleStage() returns `{ stage, message }`; a `failed` stage gets
its own branch in apiFetch that surfaces the shell's diagnosis, scrubbed.
- BackendStartFailureNotice renders it after the splash is gone, reusing the
splash's `detectHints` matcher (shared, so the two can't drift) and the
existing bug-report affordance. No Retry advice for unrecoverable failures.
- Rust retains the last `Failed { message }` past a later stage transition
(Retry sets Checking, the supervisor sets StartingBackend) so a respawn can't
erase the first diagnosis; exposed as `last_bootstrap_failure`.
- `bun desktop` prints the exit code and where to look instead of exiting
silently — the from-source twin of the same class ("builds but won't launch").
- Scrub primitives extracted to utils/scrub.js so the transport layer can scrub
without a bugReport -> client import cycle.
Non-Tauri deployments are untouched: there is no shell to fail this way, so the
stage stays 'unknown' and #1164's deployment-specific message still stands.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
939b4d4b78 |
test+dev: production-bundle black-screen gate, and stop dev instances colliding
Two independent causes of a blank app window. Never SHIP a blank (the prod hole). v0.3.22 shipped a black screen: a minifier temporal-dead-zone reorder threw before React mounted, leaving an empty #root (#1178). It reached users because every existing check — vitest, node:test, and the whole Playwright e2e suite — runs the UN-MINIFIED dev server, so a bug living only in the minified bundle passes them all. playwright.prod.config.ts + e2e-prod/ close that hole: build the real bundle, serve dist/ via vite preview, and assert the app actually mounts (#root has children, renders visible text, no pageerror). The core assertion is deliberately structural — "did anything mount?" — because that is what a pre-render crash always breaks, whatever its cause. retries: 0, so a blank screen can never be flaky-passed away. Never DISPLAY a blank in dev (the collision). Running `bun desktop` while one is already up does not fail politely, it cascades into a blank window. Reproduced deterministically and measured over CDP: the healthy main window has #root childElementCount 1; after a second launch it is 0. The new launch's port grab makes the running instance's dev:api exit, and `concurrently --kill-others-on-fail` then tears down that instance's whole stack including its Vite server — leaving its window open, pointed at a dev URL that no longer answers. desktop-dev.mjs now clears a leftover dev app first, loudly. The safety boundary for that cleanup is `isDevAppProcess` in desktop-common.mjs: it matches the cargo dev binary (`omnivoice-studio`) ONLY, never the installed release app (`OmniVoice Studio`) — killing a user's real app would be far worse than the bug being fixed. Unit-tested both ways. Also makes the gate runnable off Linux: the dev e2e config hardcodes /usr/bin/chromium, which doesn't exist on Windows/macOS. The new config falls back to Playwright's own browser so a contributor can run the gate before a release. The ci.yml step that runs this gate is NOT in this commit — pushing workflow changes needs a token scope this session lacks. It is provided separately for the maintainer to apply. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
175989241d |
fix(setup): force UTF-8 output so setup.py doesn't crash on piped stdout (Windows)
`scripts/setup.py` prints ✓/⚙ status glyphs. When its stdout is a pipe rather than an interactive console — which is exactly the case under `bun run setup:api` and in CI — Windows Python encodes with the cp1252 codepage, which can't represent those characters, so the script dies with UnicodeEncodeError mid-setup and takes `bun desktop` down at the setup:api step. It only "works" interactively by luck of the console encoding. Reconfigure sys.stdout/sys.stderr to UTF-8 (errors="replace") at startup so the output is identical whether run interactively or piped. No-op where the streams already speak UTF-8 (macOS/Linux, modern Windows Terminal) or can't be reconfigured. Verified: `bun desktop` from a stale terminal now completes setup:api with piped stdout and launches the app. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f4fe0844ba |
fix(dev): self-heal cargo PATH so bun desktop works from a stale terminal (#1180)
* fix(dev): self-heal cargo PATH so `bun desktop` works from a stale terminal `tauri dev` shells out to cargo, so `bun desktop` died with "failed to run 'cargo metadata' ... program not found" on any terminal opened before rustup was installed — the shell holds a stale PATH snapshot without ~/.cargo/bin even though cargo is installed and on the persisted User PATH (a new terminal finds it). That's a confusing first-run-from-source papercut, hit repeatedly on Windows. The frontend `desktop` script now runs through scripts/desktop-dev.mjs, which prepends ~/.cargo/bin when cargo isn't already resolvable, then launches `tauri dev` with that healed env. Cross-platform (~/.cargo/bin everywhere), a no-op when cargo is already on PATH, and it passes an explicit env with the correct-case Path key (Bun doesn't propagate process.env mutations to children, and Windows uses "Path" not "PATH"). If Rust isn't installed at all, it prints an actionable install hint instead of the cryptic cargo error. Verified E2E: from a cargo-less PATH, `tauri dev` now compiles instead of failing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): note the bun desktop cargo-PATH self-heal (#1180) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: debpalash <tapudattaht@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
08e19e0c1c |
feat(uninstall): opt-in, content-free app_uninstalled ping in the uninstall scripts
Before deleting anything, uninstall.sh / uninstall.ps1 now send ONE
best-effort `app_uninstalled` event — but only when the user opted in:
consent is read from the same prefs.json the app writes, AND the ping needs
the backend-written analytics_info.json (present only while analytics is
enabled — consent + build token — and removed on opt-out), because the
generic scripts ship no token of their own. Payload is content-free: app
version, OS name, random per-install id. 2-second timeout, silent failure,
one honest console line ('Sending anonymous uninstall ping (you opted in to
analytics).'); not opted in => nothing sent, nothing printed, and the
dry-run never sends either way.
Tested by exercising the bash script for real (fake $HOME + a curl shim
recording argv: consented/not-consented/dry-run/missing-info flows, plus
data-still-deleted-after-ping) and by static contract checks on the ps1
(consent gate, TimeoutSec 2, try/catch, no baked token). docs/install/
uninstall.md documents the behavior (docs-sync).
|
||
|
|
e6b1179001 |
feat(dev): loud backend exit banner for bun run dev + docs + changelog (#1164)
In dev there is no supervisor: concurrently's --kill-others-on-fail tears the whole stack down the moment uvicorn exits, the cause scrolls away with the terminal, and the browser tab just says it can't reach the backend — which is exactly how #1164 arrived with zero diagnostics. - scripts/dev-backend.mjs: dev:api now runs uvicorn through a wrapper (command args byte-identical, stdio inherited). On a non-Ctrl+C, non-zero exit it prints a boxed banner: exit code/signal, the last 20 lines of omnivoice.log (data dir resolved exactly like backend/core/config.py), an OOM hint (SIGKILL/137 + the Linux journalctl -k check), and a pointer to the crash notice the run sentinel raises on the next backend start. Exits with the child's own code so --kill-others-on-fail still works. Verified live: started the dev backend, SIGKILLed it, banner printed with the real log tail and exit code 137. - docs-sync: troubleshooting.md gains §14c (browser/dev/Docker crash forensics: the mode-aware error, the dev banner, run_sentinel.json / last_run_crash.json / GET /system/last-run-crash, cap+ack+version-gate semantics) and §14's crash-notice blockquote no longer implies the notice is desktop-only; CONTRIBUTING.md documents the dev:api wrapper. - CHANGELOG.md: [Unreleased] entry for the #1164 class fix. Tests: tests/frontend/devBackend.test.mjs (5) — the uvicorn args are pinned byte-identical, data-dir resolution mirrors config.py, tail/banner content incl. the OOM shapes. |
||
|
|
bec916c348 |
perf(bench): a memory-safe profiler for the pipeline — so "make it faster" stops being a guess (#1129)
Every performance question this week ("can we batch by cores?", "why is dubbing
slow?") was answerable only by measuring, and twice the intuitive answer was wrong:
* Concurrency on Apple Silicon buys NOTHING. Measured, 4 segments:
1 worker 19.3s | 2 workers 20.7s (0.93x) | 3 workers 19.2s (1.00x)
One GPU, already saturated — extra workers interleave. Scaling the GPU pool by
free RAM (the "intelligent batching" that sounds obviously right) would have
added OOM risk on a 16 GB box for zero throughput. _pick_gpu_workers()'s
hardcoded `MPS -> 1` is correct, and now provably so.
* The clone-prompt cache misses on every segment (a dub writes one reference per
segment: 166 distinct keys, cache can never hit). That looked like the dub's
hidden cost. It is 0.40s/segment — ~2% — and it is not even waste: each
reference is genuinely different audio, and encoding it is the *feature*
(per-line prosody). Dropping to per-speaker refs would save ~65s/dub and cost
quality. Not a free win; not taken.
What actually dominates is TTS itself, which scales with text length (3.2s for a
short line, 8.7s for a 2.5x longer one) and is GPU-bound on a GPU that one
inference already fills.
The profiler is deliberately gentle with memory, because a profiler that OOMs the
machine reproduces the very bug class it exists to fix (#1119): stages run one at a
time, models are unloaded between them, a stage is SKIPPED if free RAM is under the
floor rather than starting a load the OS would kill, and each measurement is a fixed
small number of passes — no looping to convergence.
Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
4e5d795832 |
fix(uninstall,storage): remove the saved-env leftover; count sidecar engines in disk usage (#1108)
Two recon findings from the reset work, fixed properly (whole class + tests +
docs), plus the destructive reset path is now exercised end-to-end.
1. ~/.config/omnivoice/env survived every uninstall. The app persists the
model-cache location (and a possible HF_TOKEN) there via
backend/core/user_env.py, but the in-app "Remove all data" (uninstall.rs),
uninstall.sh, and uninstall.ps1 all walked past it — so a reinstall silently
inherited the old file and redirected downloads to a maybe-deleted location.
All three now remove it. It's the same expanduser("~/.config/omnivoice/env")
path on every OS, so the Windows script uses %USERPROFILE%\.config\omnivoice.
is_recognizably_ours accepts it (contains "omnivoice"); docs tables updated.
2. Disk usage measured the wrong engines dir. storage_report.default_engines_dir()
returned backend/engines (built-in engine *modules*, no venvs), while sidecar
installs live in DATA_DIR/engines/<id>. So a multi-GB IndexTTS-2 install was
invisible in the engine-venv category and rolled into data/"other". Now points
at DATA_DIR/engines and sizes the WHOLE install (venv + checkout + weights),
with the data category claiming that subtree so it isn't double-counted.
Reset hardening: extracted purge_scopes() as a pure fs function (no AppHandle),
so the actual delete loop runs in tests against a real on-disk install tree —
"everything" wipes the install but spares the venv/foreign temp/sibling folders,
a settings reset keeps content+config+models, and a poisoned data_dir="$HOME"
deletes NOTHING. This is the live drive-through of the destructive path, minus
the GUI.
Also: gitignore the node_modules symlink form (the directory rule node_modules/
never matched a worktree symlink, so it kept slipping into commits).
Tests: Rust 78 (6 new), storage_report 20 (2 new incl. once-not-twice count +
default-dir guard), frontend 1207, i18n probe green, format+lint clean.
Co-authored-by: mergetest <nizam4103@gmail.com>
|
||
|
|
0421be966e |
feat(settings): in-app uninstall — Settings → Storage → Remove all data (#1089) (#1099)
* feat(settings): in-app uninstall — Settings → Storage → "Remove all data" The v0.3.19 uninstaller was a SCRIPT, which never reaches the people who need it: anyone who installed the .dmg / .msi / AppImage has no repo to run scripts/uninstall.sh from — exactly the reporter in #1089, an AppImage user. "Where is uninstall in the app?" had no answer. Now it does. New Tauri commands (uninstall.rs): - uninstall_scan — every folder this install owns, with real sizes, resolved through the same setup.rs helpers the app itself uses, so custom + portable locations are cleaned instead of the defaults being assumed. - uninstall_purge — stops the backend (marking the kill intentional so the #567 supervisor doesn't respawn one into the directories being deleted), removes the folders, and lets the UI quit the app: the Python env it runs on is gone, so there is nothing to return to. This lives in the Rust shell, not the backend, because the biggest thing to remove is the managed Python environment and the backend is RUNNING FROM IT — a process can't delete its own interpreter (and Windows locks the files). Safety: every path must pass is_recognizably_ours() before any remove_dir_all — absolute, not `/` or $HOME, and carrying an OmniVoice-owned component (unit tested both ways). The shared Hugging Face cache is reported separately and is OPT-IN behind its own checkbox with the caveat spelled out: it's the standard HF cache other ML tools share, so sweeping it up silently would delete models this app never downloaded. Deleting voices/projects is irreversible, so the confirm requires TYPING the word, not just a click. Also fixes a real bug in what shipped in v0.3.19: the scripts and docs missed where the BACKEND writes its logs — ~/.local/state/OmniVoice on Linux and %LOCALAPPDATA%\OmniVoice\Logs on Windows (backend_log_path(), backend.rs) — so every Linux/Windows uninstall left a stray log dir behind. Covered now in the scripts, the docs, and the in-app scan. And the scripts now ship as release assets, so cleanup is possible without launching the app at all. Rust: 2 new guard tests. Frontend: 6 new tests (the size on the confirm button must equal what actually gets deleted); suite 1182 passed. Docs synced. Refs #1089 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(settings): drop token border utilities from UninstallPanel (design guard) tests/test_no_literal_borders.py::test_no_token_border_utilities_in_jsx is a backend guard that scans JSX — so a frontend-only test run misses it. It forbids `border-[var(--chrome-border)]` structural utilities: the app-wide border removal converted every panel/row frame away from them, and they render a stray hairline the moment the token doesn't resolve transparent. Row dividers → spacing + an alternating `--chrome-hover-bg` tint; the opt-in checkbox card → a background tint; the confirm input → the sanctioned arbitrary `[border:1px_solid_var(--chrome-border)]` property form the other settings inputs already use (explicitly not flagged by the guard). Guard green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b1a7ddc374 |
feat(install): clean uninstaller + a straight answer to "where is my data?" (#1097)
A Linux AppImage user asked which folders to delete to remove OmniVoice and whether an uninstaller exists (#1089) — they had to guess. They shouldn't have to: the app is fully local, so uninstalling IS just deleting the folders it wrote, and we never documented them. - scripts/uninstall.sh (macOS/Linux) + scripts/uninstall.ps1 (Windows): find every OmniVoice folder — app data, the multi-GB managed Python env, config, logs — plus, listed SEPARATELY because it is a shared cache, the Hugging Face model cache. Print each with its size as a DRY RUN and stop; delete only on --yes (--models / -Models to include the shared cache). They honor the same env overrides the app reads (OMNIVOICE_DATA_DIR, OMNIVOICE_CACHE_DIR, HF_HOME, HF_HUB_CACHE), and never touch the app binary or anything outside the paths they list. - docs/install/uninstall.md: the complete per-platform path table (what each folder holds and how big it is), the shared-HF-cache caveat, custom/portable locations, per-platform steps to remove the app itself, and what to keep if you plan to reinstall. - Linked from the README FAQ, SUPPORT.md, and install troubleshooting. Paths mirror backend/core/config.py + frontend/src-tauri/src/setup.rs. Verified on macOS: dry-run lists the real dirs; sandboxed HOME runs confirm --yes removes app folders while KEEPING the shared cache, --models removes it, and the env overrides retarget correctly. Closes #1089 Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bd85bab624 |
chore: retire finished planning archives from the repo root (#1095)
Removes ~110 files of process noise (all preserved in git history): .planning/ (GSD-era phases/quick-plans/issue-clusters; workflow retired 2026-07-08), specs/ (spec-kit specs for shipped features 001-007), design/ (pre-React ASCII mockups), research/ (legacy Gradio archive), and .agents/ (rules for a third-party agent tool no longer in use). The four load-bearing decision docs move to docs/adr/ with an archival note; every live pointer follows (gguf engine module docs + quant_map, inject-apprun.sh, pyproject/test comments, fixture README + its seed script — kept byte-identical). The CJK allowlist drops the deleted legacy_gradio entries; STRUCTURE.md and ROADMAP.md document the removal instead of linking into it. Backend suite: 2891 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d3e88c0f3e |
fix(scripts): desktop-fresh kill guard referenced the wrong dry-run flag (#1078)
The kill-before-wipe block used DRY_RUN; the script's flag is dryRun — any run with a live instance crashed with ReferenceError before wiping. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fb0508f4f6 |
fix(shell): deep health probe before attaching to a running backend; scripts kill before wiping (#1077)
A backend that keeps running while its install is deleted or replaced underneath it still answers /health and /system/info from memory — the launcher's version check passed and the UI attached to a process that 500s every DB-touching route (raw errors without CORS headers, so the webview reports access-control failures). The attach path now requires a DB-touching probe (/profiles) to return an actual 200 status line, and replaces the squatter otherwise — the status line is parsed explicitly because the raw HTTP helper previously returned 500 bodies as Ok. desktop-prod/desktop-fresh now terminate our own running processes (bundle, dev binary, app-scoped port-3900 listener) before wiping, which is how the zombie was produced. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
34c8a33628 |
fix(scripts): desktop-prod builds clean, desktop-fresh emulates a brand-new machine (#1070)
desktop-prod fixes:
- `tauri build --debug` used to produce every bundle and THEN exit 1 at the
updater-artifact signing step (no TAURI_SIGNING_PRIVATE_KEY on dev
machines); the script papered over it with a blanket "non-fatal bundle
error" grep that also swallowed real bundling failures. Local emulation
builds now pass `--config '{"bundle":{"createUpdaterArtifacts":false}}'`
and only build the bundle the script launches (--bundles app / appimage,
--no-bundle on Windows), so the build exits 0. Any nonzero exit now FAILS
the script — the sole tolerated case is a specifically-detected
linuxdeploy/FUSE failure on Linux when the raw debug binary was produced.
- The HF cache wipe ran `rm -rf ~/.cache/huggingface` on macOS/Linux — the
SHARED global cache (backend/core/config.py only relocates it on Windows),
deleting models unrelated to OmniVoice. Non-app-scoped cache paths are now
kept with a "models will be reused" notice; FRESH_NUKE_HF=1 opts in.
- Honest clean marks (removed ✓ / already-clean ○ instead of ✗ for success),
`open -n` always (plain `open` focused a stale running instance instead of
launching the freshly built one), stale-AppImage removal on Linux.
New `bun desktop-fresh` (+ desktop-fresh:run), macOS-only with explicit
refusal elsewhere: true new-user emulation.
- Blank slate: everything desktop-prod cleans PLUS the traces that survive a
reinstall + data wipe — ~/Library/WebKit (webview localStorage), Caches,
HTTPStorages*, Preferences plist (+ defaults delete), Saved Application
State. Per-path found/removed/absent status with sizes; --dry-run prints
the full plan without touching anything.
- Dev-machine camouflage: launches by direct exec of the bundle's Mach-O
(which inherits env — `open` hands off to launchd and drops it) with PATH
stripped of /opt/homebrew/{bin,sbin} + /usr/local/bin and HF_TOKEN /
HUGGING_FACE_HUB_TOKEN / HF_HOME / HF_HUB_CACHE / HF_ENDPOINT /
OMNIVOICE_* unset, and prints a banner of what is hidden.
Shared pure helpers live in scripts/desktop-common.mjs, covered by 9 node
tests (tests/frontend/desktopScripts.test.mjs): every cleanable path is
app-scoped and under $HOME, the PATH/env sanitizers strip exactly the
intended entries, and the build args carry the updater-artifacts-off config.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
1808a373a1 |
fix(linux): AppRun workaround detection reads the BUNDLED WebKitGTK version, not the host's (#961 follow-up) (#1024)
The launcher decided whether to export WEBKIT_DISABLE_COMPOSITING_MODE by asking the host's pkg-config — but LD_LIBRARY_PATH makes the BUNDLED libwebkit2gtk the one that actually runs, so on any machine where the two diverge the detection read the wrong number. This was the second bug identified during #961's investigation (the reporter built from source, so their dev packages answered pkg-config with a healthy version while the shipped bundle ran an older lib) and was explicitly deferred in #1007 as not-safely-fixable at runtime. The fix makes it knowable by construction instead: inject-apprun.sh runs at bundle time ON the build host whose libwebkit2gtk gets bundled, so it stamps that version into .bundled-webkitgtk-version inside the AppDir. AppRun reads the stamp first and only falls back to host pkg-config for bundles predating it. Empty/unreadable stamp fails safe (workaround on), same philosophy as the missing-pkg-config path. Tests: 3 new cases in AppRun.test.sh — marker-beats-host in both directions (broken-marker/healthy-host and the #961 inversion, healthy-marker/broken-host) plus empty-marker fail-safe. Also wires AppRun.test.sh into pytest (tests/test_apprun_launcher.py) — it was previously run by NO CI job, so the launcher could regress silently. Also documents Windows install-to-another-drive behavior in docs/install/windows.md (#938): local drives work via the wizard's directory picker, mapped network drives are a Windows Installer limitation, and the data directory moves independently of the app. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4aa9abe22a |
docs+scripts: install fixes — desktop-prod tauri resolution, Ubuntu white-screen guidance, honest GPU/prereq docs (#960 #961 #962) (#964)
* docs+scripts: install fixes — desktop-prod tauri resolution, Ubuntu white-screen guidance, honest GPU/prereq docs (#960 #961 #962) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add the install-fixes batch under [Unreleased] (#964) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
72d137e1f3 |
fix(bootstrap): port cuDNN 8 (NVIDIA CUDA GPU) + VC++ redist to packaged installs (#869)
* fix(bootstrap): port cuDNN 8 (NVIDIA CUDA GPU) + VC++ redist install into ensure_venv_ready() * fix(bootstrap): address #869 review — drop dead VC++ half, cache negative CUDA probe, gate on ROCm, sync docs Per maintainer review on #869: 1. Drop the VC++ Redistributable half: LoadLibraryA("vcruntime140.dll") from the running Tauri exe is a tautology (the exe itself links the MSVC CRT, so the process wouldn't be running without it), and torch's real failure mode is msvcp140.dll inside the venv python process. Dead code removed; a comment records why for future readers. 2. Stop taxing every non-CUDA launch: a negative torch probe (CPU / Intel / AMD — most installs) is now cached in a .venv/.cudnn8_probe_negative marker, so the synchronous `import torch` runs at most once per venv lifetime. Invalidated on every path that can change the torch build (drift sync #307, repair sync, first-run sync, ROCm reinstall) and implicitly by a venv rebuild. A probe that fails to run cleanly is skipped WITHOUT caching so a transient error can't wedge a real CUDA machine. 3. Rewrite docs/install/troubleshooting.md §10 to the actual root cause: packaged installs never had the cudnn8_compat libs (so reinstalling never restored them); the bootstrap now installs them automatically on CUDA machines, with the manual uv pip command as the offline fallback and PyTorch Whisper as the sidestep. 4. Gate the ~700 MB nvidia-cudnn-cu12 download on the venv torch being a real CUDA build: the probe now reports 'hip' before checking cuda.is_available() (which HIP spoofs), so opt-in ROCm installs (#124) never fetch the CUDA wheel. Also reflow the CHANGELOG entry to house style (bold one-line lead, 1-3 lines of why, (#827, #869) refs) and extend the bootstrap unit tests: classify_cuda_probe verdict mapping and the marker write/invalidate round-trip (6 cuDNN tests total, 43 lib tests green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c44bae0d0d |
feat(ui): migrate waveform-* to utilities, trim index.css (P4) (#816)
Move the waveform-* global class family off index.css onto Tailwind v4 utilities on WaveformTimeline.jsx, then delete the now-dead rules. Migrated to utilities (rules deleted): waveform-timeline (mb), waveform-controls + -left/-right (flex/items/justify/gap), waveform-btn + :hover/:disabled and waveform-btn-play + :hover (shared WF_BTN/WF_BTN_PLAY consts; UA <button> padding/font preserved since the app ships no preflight), waveform-time (text/border/bg/mono/tabular-nums), waveform-zoom-slider (important w/h/mt). States -> hover:/disabled: variants; no-preflight borders -> explicit [border:1px_solid_...]; exact px via arbitrary values. Deleted as dead (zero usages anywhere): waveform-video-preview, waveform-track-bg (+ nth-child + the 800px media-query track rows). Kept (irreducible): .waveform-container and its .waveform-container [data-id^="wavesurfer-region"] descendant rules (+ the 800px container/region media query) — those style WaveSurfer-generated DOM we don't render in JSX, so they can't be utilities. The class stays as a hook. index.css net -55 lines (+7/-62). Cascade-correctness verified live (Playwright getComputedStyle, both stylesheets loaded): new utilities reproduce the pre-migration computed styles exactly. Caught two subtleties: (1) controls margin-top is 3px (unlayered wfm-controls already wins over the old 4px), so no mt utility is added; (2) referencing var(--chrome-font-mono) in a class string tripped the global [class*="chrome-font-mono"] selector (adds slashed-zero + ss02) — switched the time font to var(--font-mono) (identical stack) to avoid the substring match. Screenshot pixel-diff old vs new = 0 (AE). Updated record_promo.js's fallback selector (.waveform-controls -> [aria-label="Playback controls"]). Gates: oxlint 0, oxfmt clean, vite build, vitest 641 pass, test:visual 48 pass, bun install --frozen-lockfile no change. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
4eac1b50d5 |
chore(desktop-prod): add --keep-models for fast fresh-app runs (#650)
`bun desktop-prod` (clean) wipes everything including the HF model cache, so every fresh-install emulation re-downloads multi-GB weights — slow and bandwidth -heavy, and the exact pain users on flaky networks hit. --keep-models wipes app/backend data, logs, and webview state for an honest first-run, but KEEPS the model cache so the weights aren't re-pulled. Ignored under --keep-data (which keeps everything). Adds the `desktop-prod:keep-models` convenience script. Scripts-only package.json change — no deps, bun.lock unaffected. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
de80856cd9 |
test: backend route-inventory + webUI feature-coverage guards (#609)
* test: backend route-inventory snapshot + webUI feature-coverage guards A reusable testing system that verifies every feature surface is present: - tests/test_api_route_inventory.py: boots the app, diffs all 213 routes vs a committed snapshot (tests/fixtures/api_routes.txt), guards a critical-endpoint set, and floors the route count — any endpoint drift fails CI. - scripts/dump_api_routes.py: regenerates the snapshot. - frontend featureCoverage.test.js: every AppMode has a render branch, every lazy-imported page file exists, every feature has an i18n namespace. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): note the feature-coverage test system * test(api-inventory): isolate via subprocess + exclude env-dependent mounts CI surfaced two flaws in the first cut: - the in-process app import + sys.modules purge polluted later DB-touching tests (a cascade of 404s in test_dub_subtitles_309 etc.); - the snapshot included StaticFiles mounts (/demo_audio) and a conditional GET / root that register based on filesystem state, so a macOS-generated snapshot didn't match a fresh Linux CI runner. Compute routes in an isolated subprocess (scripts/dump_api_routes.py --print) and cover only the deterministic router surface (drop Mounts + root). 209 routes; inventory + previously-polluted tests now pass together. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e862f0faf0 |
feat(asr): crash-isolated faster-whisper subprocess backend (Wave 4.2) (#393)
* feat(asr): crash-isolated faster-whisper subprocess backend (Wave 4.2) Native ASR engines (faster-whisper / CTranslate2) can segfault on GPU teardown — a process-level crash that kills the whole backend. Running the engine in a child process turns that into a failed job: the sidecar dies, the parent raises a decorated error (engine id + device), and the next request respawns a fresh sidecar. - services/subprocess_asr.py: SubprocessASRBackend reuses SubprocessBackend's wire protocol + lifecycle — including respawn-on-dead-process (_spawn relaunches when the child isn't alive) and GPU-slot acquire/release — adding a 'transcribe' op (the TTS 'generate' surface is stubbed). IsolatedFasterWhisperBackend wraps faster-whisper using the PARENT venv (already a dep — only the process boundary is new); opt-in via OMNIVOICE_ASR_BACKEND=faster-whisper-isolated. - engines/_asr_sidecar/main.py: the faster-whisper runner (stdlib wire protocol; torch/CT2 import lazily so the ready handshake fits the timeout). - engines/_echo/main.py: a 'transcribe' echo op so the round-trip + crash recovery are testable without a real engine. - asr_backend._REGISTRY is now a lazy dict (mirrors the TTS registry) so the isolated backend lists/resolves without importing the subprocess stack unless selected. Tests (echo sidecar, stdlib-only): round-trip, single long-lived sidecar across calls, crash-mid-transcribe → decorated error + backend healthy + next call respawns, registry exposure, generate-not-supported. Spec 7 / parity program Wave 4.2. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(asr): deterministic crash test + drift marker for lazy ASR registry (Wave 4.2 CI) CI surfaced two issues: - The echo crash test relied on the crash-AFTER-reply hook, whose reply may still reach the parent (timing-dependent) — and a leaked OMNIVOICE_ECHO_CRASH from a sibling subprocess test poisoned the non-crash tests. Fix: a deterministic OMNIVOICE_ECHO_CRASH_NO_REPLY hook that exits BEFORE replying (guaranteed dead pipe → decorated error), and the asr fixture clears both crash envs so the round-trip/two-call tests can't inherit a leak. - check-docs-drift's _ASR_MARKER didn't match the new lazy registry line (_LazyASRRegistry({); updated the marker + the self-test fixture. Verified the no-reply crash hook by driving the sidecar directly (reply=None, exit 1); drift self-test + real-repo check green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(asr): allowlist the 'segments' op so transcribe replies aren't dropped (Wave 4.2 CI) The parent's PARENT_INBOUND_OPS frozenset gated inbound sidecar frames but never included 'segments' — the ASR transcribe reply op. _recv() dropped the frame as disallowed, tail-recursed, hit EOF, and returned None, so every transcribe surfaced as a bogus 'sidecar crashed mid-transcription'. TTS ('audio') was allowlisted; ASR ('segments') was missed. Add it (and list 'transcribe' in the informational SIDECAR_INBOUND_OPS), update the exact-shape allowlist test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
11c498eeb5 |
ci(docs): daily docs-drift job — canonical inventory vs README/docs/registries (Wave 0.1) (#353)
docs/features.yaml is the curated single source of truth (12 features, 11 TTS + 7 ASR engine ids, required install docs). scripts/check-docs-drift.py diffs it against README.md, docs/, and the engine registries — parsing registry keys from source so the CI runner never imports torch. The daily workflow updates ONE rolling 'docs-drift' issue in place and auto-closes it when clean (pattern adapted from Patter, MIT). Self-test includes a real-repo-is-clean gate, so any PR that changes engines/features without updating the inventory fails CI too. Spec: docs/competitive-analysis.md Spec 9a / parity program Wave 0.1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
948bc76543 |
macOS: ad-hoc sign so users open without Terminal + signing/notarization verification (#290)
* chore(release): add macOS signing/Gatekeeper/notarization verification Codify and enforce the macOS build-signing requirements. The release pipeline built bundles and had opt-in Apple signing, but never verified codesign/spctl/notarization — unsigned or broken bundles could ship silently. - scripts/verify-macos-signing.sh: runs codesign --verify --deep --strict, spctl Gatekeeper assessment, per-nested-Mach-O signature check, stapler validate, and (opt-in) notarytool history. Report-only by default (unsigned dev/preview is expected); --require-signed fails on any unsigned/un-notarized component so a broken release stops instead of publishing an unsigned artifact. - scripts/macos-dev-unquarantine.sh: local-dev-only quarantine stripper, with a loud "never a substitute for notarization" warning. - release.yml: new "Verify macOS signing" step on the macOS leg — report-only on unsigned paths, STRICT on the opt-in signed stable path (same condition as "Configure Apple signing"), so signing/notarization failures fail the job. - docs/macos-signing-verification.md: the canonical 10-point requirements + how-to-verify checklist, cross-linked to docs/install/macos.md and DESKTOP_RELEASE.md. Verified locally: report-only PASS (exit 0) and --require-signed FAIL (exit 1) against the real unsigned debug .app; release.yml parses as valid YAML. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(macos): ad-hoc sign bundle so users open it without Terminal (no Apple ID) The "app is damaged and can't be opened" error is caused by a broken/incomplete code-signature seal (codesign --verify failed: "code has no resources but signature indicates they must be present") on the quarantined download — there is no GUI bypass for that variant on modern macOS, forcing users to run `xattr`. Give the bundle a VALID ad-hoc signature at build time (free, no Apple Developer account) via tauri.conf.json bundle.macOS.signingIdentity = "-". Verified through a real `tauri build`: the produced .app is now flags=adhoc,runtime and passes codesign --verify --deep --strict. A valid seal flips the Gatekeeper prompt from the un-bypassable "damaged" to the GUI-bypassable "unidentified developer", which users clear with right-click → Open / Settings → "Open Anyway" — no Terminal. Still not notarized (that needs the paid Apple ID), so there's a one-time confirmation rather than a clean double-click. The opt-in Developer-ID path is unchanged: APPLE_SIGNING_IDENTITY (env) overrides the "-" default on the signed stable release. - tauri.conf.json: signingIdentity "-" (ad-hoc default). - verify-macos-signing.sh: detect ad-hoc tier; report the no-Terminal GUI path in report-only, still FAIL it under --require-signed (production must notarize). - docs/install/macos.md: lead the Gatekeeper section with right-click → Open; keep xattr as fallback for the harsher "damaged"/corrupted-download case. - docs/macos-signing-verification.md: signing-tiers table + ad-hoc default note. - release.yml: comment the ad-hoc default + env override on the signed path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ea26893bfc |
fix(scripts): desktop-prod works from cmd/PowerShell via cross-platform launcher (#282) (#333)
`bun run desktop-prod` (and its :run/:upgrade/:pill/:run:pill variants) invoked `bash scripts/desktop-prod.sh` directly. On Windows, cmd and PowerShell have no `bash` on PATH unless Git Bash happens to be there, so the documented from-source install path died with a cryptic spawn failure before printing anything — the exact first step in issue #282's repro. Add scripts/desktop-prod.mjs, a tiny launcher (runs under bun or node): - macOS/Linux: execs the bash script unchanged — zero behavior change. - Windows: locates Git Bash via `where.exe bash`, well-known Git for Windows install paths, or derived from git.exe's location; explicitly skips C:\Windows\System32\bash.exe (the WSL launcher, which would run the script inside Linux and wipe/launch the wrong paths). - No usable bash: prints an actionable error (install Git for Windows, use `bun run desktop`, or use the installer) instead of a spawn error. All flags are forwarded untouched and the child's exit code is propagated. scripts/desktop-prod.sh itself is unchanged, and docs/install/windows.md now lists Git for Windows as a prerequisite for from-source installs. Refs #282 Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |