fc2b0d4cb5ceea4ff13f56a96df007f20dbf79fb
147
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f3c2d745c0 |
Merge #1212: PIN + API-key auth guide for the local API
# Conflicts: # docs/api-auth.md |
||
|
|
84d34658fe | Merge #1213: admin routes require the API key in SERVER_MODE (trusted-network privilege escalation) | ||
|
|
204eff2be1 |
fix(security): #1213 review — the share PIN must not gate RCE-class admin
CodeRabbit: the 6-digit share PIN is brute-forceable (10^6, no lockout), so letting it unlock admin over the network was still weak. Admin now requires the API key (a long operator secret) or loopback; the PIN is consumption-only and never gates /system/* or /api/settings/*. A PIN-only deployment keeps admin loopback-only. Docs (api-auth.md, remote-gpu.md) aligned with the conditional gate (no credential -> admin open; API key -> admin) and the PIN exclusion. Test inverted: presenting the PIN over the network is now 403 on admin. |
||
|
|
3b879f298e |
fix(auth): keep admin gate independent of trusted-network trust under server mode (#1213)
OMNIVOICE_SERVER_MODE=1 made require_loopback an unconditional no-op, so with OMNIVOICE_TRUSTED_NETWORKS also set a trusted-CIDR client — a consumption-only exemption that bypasses the PIN/API-key middleware via is_local_host — could reach the RCE-class admin surface (/system/set-env, /api/settings/*) with no credential. That collapsed the documented two-tier privilege model (consumption trust != admin trust) in exactly the "lock the backend with a key, exempt a LAN proxy for TTS" configuration. Server mode still can't require true loopback (Docker NAT, #261), but it now applies the credential rule to admin routes: open only when NO credential is configured; otherwise the request must present the API key or share PIN. Trusted-network membership alone never satisfies it. Loopback, credential holders, and no-credential Docker deployments are unchanged; consumption routes (require_local / middleware) keep exempting trusted networks. Regression tests cover the server-mode x trusted-network x credential matrix (fail-before/pass-after). Docs: new docs/api-auth.md two-tier model + quick reference; docs/remote-gpu.md corrected (previously documented the hole as accepted behavior). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a12a7e7ee1 |
feat(audiobook): expressive maturity — overrides, emotion, cache opt-out, discoverability (#1210)
Audiobook renders were locked to the model's most deterministic preset (32
steps / 2.0 guidance / model-default temps) with no way to change it, which is
why books sounded flatter than the same voice on the Voice page. Open that up
without changing any default byte-for-byte.
- Production Overrides in the Audiobook tab: position_temperature,
class_temperature, num_step, guidance_scale, postprocess_output (+ seed),
reusing the Voice page's panel. Unset reproduces today exactly.
- IndexTTS2 graded emotion (emo_vector / emo_text / emo_alpha) reaches the
longform path via a typed engine-options object; engines that don't
understand an option ignore it (no crash across the ~14 backends).
- Cache opt-out ("vary repeated lines") so identical lines can get distinct
takes; default off keeps the content-addressed replay.
- Markup reference now lists the reaction tags that already work in audiobooks;
docs/expressive-speech.md corrected so no recipe it names is unreachable.
- Fix AudiobookGenerateBody dropping `language`, so audiobook language
selection actually reaches the backend.
Every new param is folded into BOTH cache layers (chapter + segment) and the
per-chapter preview, so changing a knob re-renders instead of replaying stale
audio — and an all-default request keeps its old cache key, so existing books
don't re-render. Regression tests: tests/test_audiobook_expressive.py (backward
compat, cache-signature loop, preview/render parity, engine-ignores-unknown,
cache opt-out, emotion reaches engine) + audiobookOverrides.test.jsx.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
668962133a |
docs(api): PIN + API-key auth guide for the local API (#1210)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
415ae6d351 |
Merge remote-tracking branch 'origin/main' into fix/1177-surface-backend-diagnosis
# Conflicts: # CHANGELOG.md |
||
|
|
433d4bc659 |
Merge remote-tracking branch 'origin/main' into fix/1177-surface-backend-diagnosis
# Conflicts: # CHANGELOG.md |
||
|
|
865be7510f |
Merge remote-tracking branch 'origin/main' into fix/1191-tts-stranded-on-cpu
# Conflicts: # CHANGELOG.md |
||
|
|
fc7fbf1227 |
fix(gpu-pool): bound execution, not queue wait (#1190, #1202)
A job queued behind a busy 1-worker pool burned its whole 300s budget without executing an instruction, then reported "too heavy for the available compute". The clock now starts when a worker picks the job up; queue wait has its own generous bound and surfaces as a retryable saturation error. Also: reset() no longer cancels innocent queued peers; the timeout message stops claiming capacity was restored (the abandoned job keeps the device until it drains); every GPU dispatch uses the shared length-scaled budget; watermark embeds move off the GPU pool; /v1/audio/speech gets 429/503 + Retry-After; a timed-out batch segment fails the job instead of shipping a silent gap. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
aab138ea57 |
fix(backend): surface the shell's start-failure diagnosis instead of "can't reach the backend" (#1177)
The reporter's string is apiFetch's LAST fallback, reached only when no crash
marker exists AND the shell's lifecycle stage is 'failed' or 'unknown'. The
'failed' half was the bug: `BootstrapStage::Failed { message }` carries the
whole diagnosis — exit code plus a ~30-line stderr tail, or the precise reason
`ensure_venv_ready` refused (Intel Mac, a failed `uv sync`, a blocked GitHub) —
and `backendLifecycleStage()` returned only the stage tag, throwing the message
away. Every backend-start failure mode collapsed into one generic, evidence-free
sentence that was also factually wrong: it is not starting, and it will not
recover on its own.
- backendLifecycleStage() returns `{ stage, message }`; a `failed` stage gets
its own branch in apiFetch that surfaces the shell's diagnosis, scrubbed.
- BackendStartFailureNotice renders it after the splash is gone, reusing the
splash's `detectHints` matcher (shared, so the two can't drift) and the
existing bug-report affordance. No Retry advice for unrecoverable failures.
- Rust retains the last `Failed { message }` past a later stage transition
(Retry sets Checking, the supervisor sets StartingBackend) so a respawn can't
erase the first diagnosis; exposed as `last_bootstrap_failure`.
- `bun desktop` prints the exit code and where to look instead of exiting
silently — the from-source twin of the same class ("builds but won't launch").
- Scrub primitives extracted to utils/scrub.js so the transport layer can scrub
without a bugReport -> client import cycle.
Non-Tauri deployments are untouched: there is no shell to fail this way, so the
stage stays 'unknown' and #1164's deployment-specific message still stands.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
0f70744105 |
fix(tts): stop stranding the TTS model on CPU after a dub abort (#1191)
`offload_tts_for_asr()` moves the TTS model to CPU to make VRAM room for WhisperX, but its partner `restore_tts_after_asr()` was only reachable on the dub-transcribe success path. Any abort, terminal error, or client disconnect skipped it, and `get_model()` never re-checked placement — so EVERY subsequent /generate ran on CPU (10-50x slower, CPU pegged) until the ~15-minute idle unload happened to fire. Reported as "speed varies by time of day"; it is fully deterministic. Two independent guarantees: - Balance the pair at the call site: gen()'s `finally` now pays the restore debt on every exit path, chained off the ASR unload so the two never contend for VRAM (and fire-and-forget, since the finally also runs under GeneratorExit where awaiting is illegal). - Self-heal placement (the class fix): `get_model()` verifies the model is on the resolved target device and moves it back if not, so a future unbalanced offload path cannot strand it either. Cheapest-first probe — one parameter check on the hot path; unified memory is exempt (its offload releases the model rather than moving it). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f3286c5e6e |
feat(dub): paste a translation from an external source onto existing segments
After transcription the user can paste a translation produced elsewhere (ChatGPT, DeepL, a human translator) and have it map onto the segments that already exist — no re-transcription, no timing loss. Three input shapes are auto-detected: a timestamped .srt/.vtt (cues matched to segments by time overlap, greedy one-to-one so one long cue can't be copied onto several rows), numbered lines (`1.` / `2)` / `[3]`, mapped by number and falling back to order when a model renumbers mid-answer), and plain lines (positional, blank lines treated as separators rather than empty translations). Nothing is applied until the preview dialog has shown every row as before→after with unmatched rows flagged. Applying goes through `pasteTranslations` in useSegmentEditing, which mirrors `segmentEditField`'s duties across rows in ONE undo step: write `text` and `translations[dubLangCode]` in lock-step and clear the stale machine-translation badges. It never writes `text_original` (the translate source `handleTranslateAll` reads — overwriting it would poison every later re-translate) and never touches a language other than the active one. Changing `text` alone marks those rows stale via the existing per-language fingerprints, so no new flag is needed. The new `POST /dub/parse-subtitle-text` is a stateless wrapper over the existing `services.srt_parser.parse_srt`, so the lenient cue parsing stays single-sourced instead of being reimplemented in JavaScript. Also fixes a ReDoS in that parser, reachable today via /dub/import-srt: `_TIMING_RE` used `^\s*` under re.MULTILINE, so at every line start the engine consumed all remaining blank lines before failing on the first digit — quadratic. 20k blank lines already took 1.7s and a 2 MB blank-line file never returned, pinning the request thread. Horizontal-whitespace-only classes make the scan linear. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
83f943bead |
fix: bot-review harvest (16 findings) + deterministic style/locale CI + reviewer configs
Harvested and verified every CodeRabbit/Greptile finding from PRs #1175, #1189, #1192, #1195: 16 real ones fixed (fallback ASR preflight bypass, VRAM release on stream exit, typed 409 parity, uv env independence, path-privacy in errors, MCP clone_voice hardening, CaptureWidget WS guard, test hygiene), 4 refuted with evidence, rest documented as deliberate design or deferred. Deterministic CI replaces hand-enforcement: tests/test_changelog_style.py (quiet one-liner format) and tests/test_locale_parity.py (21-locale key/placeholder lockstep with a ratchet baseline) — the latter surfaced and fixes 151 already-broken locale strings. CodeRabbit/Greptile carry the house rules via .coderabbit.yaml + greptile.json; CLAUDE.md gains the harvest-before-merge and never-accept-as-is rules. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9b49a3ba4f |
feat(mcp): clone_voice tool — clone a new voice from reference audio (#1194)
AI agents driving OmniVoice via MCP could use and list voices but couldn't create one. Add a clone_voice MCP tool that takes a base64-encoded reference audio sample (consistent with transcribe's audio_base64 pattern), decodes it, and POSTs it as a multipart ref_audio to POST /profiles (kind=clone). Returns the new profile_id so the agent can immediately use it with generate_speech. Update test_mcp_mount.py to include clone_voice in the asserted tool surface. CHANGELOG entry. |
||
|
|
933743e336 |
fix: resolve open issue batch — #1172 #1173 #1174 #1185 #1186 #1188
- validate managed binaries before exec; 0-byte GGUF placeholders fail actionably instead of "Exec format error" (#1172) - KittenTTS: tokenizer-measured chunking to the ONNX 512-token cap; clear 400 for unspeakable input (#1173) - clean SIGTERM during weight load: shutdown-aware loader, benign cancelled-load classification, lifespan hardening, scoped log silencers (transformers load + alembic fileConfig) (#1174) - broken ASR deep-imports (lightning_fabric) mark the engine unavailable with a repair hint and fall through (#1185) - uv cache + managed Python follow the chosen install drive on Windows (cherry-picked cross-drive class fix + spaces/D: tests) (#1186) - adaptive silence-removal ladder for quiet clone references; localized actionable error for truly silent clips, all 21 locales (#1188) - CHANGELOG: consolidated Unreleased into the quiet one-liner style Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
32f63469d1 |
Merge main into curated-ASR branch (resolve CHANGELOG)
# Conflicts: # CHANGELOG.md |
||
|
|
f669687ed1 |
feat(backend): trust a local network/proxy via OMNIVOICE_TRUSTED_NETWORKS (#1170)
Self-hosting behind a reverse proxy or on a LAN used to force a blunt choice: OMNIVOICE_SERVER_MODE (trust every non-loopback source) or the API-key/PIN gates — which a proxy that strips the Authorization header breaks for browser clients entirely. Add OMNIVOICE_TRUSTED_NETWORKS (comma-separated CIDRs) whose addresses are treated as trusted by the CONSUMPTION gates (PIN/API-key middleware, dictation WebSocket) via is_local_host — a LAN/proxy client is exempted from consumption auth. Admin routes (require_loopback → /system/set-env, /api/settings/*) stay true-loopback-only (is_loopback, not is_local_host) to preserve the two-tier privilege model: consumption trust ≠ admin trust (RCE-class surface). Opt-in, default empty → zero behavior change. The granular companion to server-mode (#261). Tests: is_loopback / is_local_host / require_loopback contract for trusted CIDRs, adjacent subnets, malformed entries, the two-tier split (trusted-network rejected by the admin gate), and the default (no-trust) case. Docs + CHANGELOG. |
||
|
|
63fd497caf |
feat: TTS-only first run, platform-curated ASR, guided OS permissions, parakeet-mlx
Only the TTS model (~2.4 GB) is required on first run; ASR models are per-platform curated picks (curated_on in models.yaml) installed on demand. Every transcription surface returns a typed asr_model_missing error with a one-click download CTA instead of silently pulling multi-GB Whisper weights. Settings -> Models is a grouped, platform-aware catalog. New guided permissions UX (wizard System Check + Settings -> Permissions + mic pre-flight) with native mic-state checks and OS settings deep-links. New parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated capture preference so multilingual dictation never regresses). Docs: expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ. Hardening: preflight fails open for custom model pins, ROCm curation no longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged consent key, CaptureWidget setup race fixed, offline-cache CI simulation fixes so empty-cache runners stay green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
08e19e0c1c |
feat(uninstall): opt-in, content-free app_uninstalled ping in the uninstall scripts
Before deleting anything, uninstall.sh / uninstall.ps1 now send ONE
best-effort `app_uninstalled` event — but only when the user opted in:
consent is read from the same prefs.json the app writes, AND the ping needs
the backend-written analytics_info.json (present only while analytics is
enabled — consent + build token — and removed on opt-out), because the
generic scripts ship no token of their own. Payload is content-free: app
version, OS name, random per-install id. 2-second timeout, silent failure,
one honest console line ('Sending anonymous uninstall ping (you opted in to
analytics).'); not opted in => nothing sent, nothing printed, and the
dry-run never sends either way.
Tested by exercising the bash script for real (fake $HOME + a curl shim
recording argv: consented/not-consented/dry-run/missing-info flows, plus
data-still-deleted-after-ping) and by static contract checks on the ps1
(consent gate, TimeoutSec 2, try/catch, no baked token). docs/install/
uninstall.md documents the behavior (docs-sync).
|
||
|
|
e6b1179001 |
feat(dev): loud backend exit banner for bun run dev + docs + changelog (#1164)
In dev there is no supervisor: concurrently's --kill-others-on-fail tears the whole stack down the moment uvicorn exits, the cause scrolls away with the terminal, and the browser tab just says it can't reach the backend — which is exactly how #1164 arrived with zero diagnostics. - scripts/dev-backend.mjs: dev:api now runs uvicorn through a wrapper (command args byte-identical, stdio inherited). On a non-Ctrl+C, non-zero exit it prints a boxed banner: exit code/signal, the last 20 lines of omnivoice.log (data dir resolved exactly like backend/core/config.py), an OOM hint (SIGKILL/137 + the Linux journalctl -k check), and a pointer to the crash notice the run sentinel raises on the next backend start. Exits with the child's own code so --kill-others-on-fail still works. Verified live: started the dev backend, SIGKILLed it, banner printed with the real log tail and exit code 137. - docs-sync: troubleshooting.md gains §14c (browser/dev/Docker crash forensics: the mode-aware error, the dev banner, run_sentinel.json / last_run_crash.json / GET /system/last-run-crash, cap+ack+version-gate semantics) and §14's crash-notice blockquote no longer implies the notice is desktop-only; CONTRIBUTING.md documents the dev:api wrapper. - CHANGELOG.md: [Unreleased] entry for the #1164 class fix. Tests: tests/frontend/devBackend.test.mjs (5) — the uvicorn args are pinned byte-identical, data-dir resolution mirrors config.py, tail/banner content incl. the OOM shapes. |
||
|
|
d4ee0e3b00 |
docs(specs): dictation flow program — local WhisperFlow-class dictation on Parakeet
Six-phase plan: VAD + true-streaming Parakeet, personal dictionary + hotwords, app-aware/agent-prompting modes, insertion reliability + Wayland chain, local command mode, docs/evals. |
||
|
|
b17df5c7a3 |
docs(docker): refresh the Docker Hub overview and docker guide
- Add a what-you-need line (RAM/disk/GPU from the README requirements table, compressed pull sizes measured from the registry) so homelab users can size the deployment before pulling. - Update stale version examples (:0.3.6 / :0.3.17 -> :0.3.22). - Fix the 'main is always one patch ahead' claim — with AUTO_VERSION_BUMP off, main can equal the released version; say 'at or ahead of the last release', which is true in both modes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f609f57fbc |
chore(release): codify all deployment channels as release rules; preview always builds from main
A release now has an explicit channel checklist (docs/RELEASING.md §5b): GH Release + stable updater manifest, preview updater channel, GHCR + Docker Hub in both CUDA and ROCm flavors, and the Docker Hub overview sync (whose continue-on-error step must be verified by step log — it 403s silently on tokens without description-edit scope). Preview/RC policy is now enforced, not just documented: release.yml's preview-gate fails publish_preview dispatches from any branch but main, since the preview manifest and rolling Docker tags all track main. Also fixes docs/RELEASING.md §4-5, which still described the pre-2026-06 versioning scheme (tauri.conf.json + Cargo.toml as sources, 'Tauri ignores package.json') — the exact opposite of the current single-source rule — and docs/update-channels.md, which invited previews off feature branches. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
99e01610bb |
feat(docker): publish ROCm/AMD GPU image variant (#1165) (#1166)
The Docker image was CUDA-only, so AMD GPUs (e.g. RX 7900 XTX under Podman) silently ran on CPU. Every preview and release now also ships a ROCm variant built from the same Dockerfile: - deploy/Dockerfile: parameterize the runtime base with a BASE_IMAGE build-arg (default unchanged: pytorch/pytorch 2.8.0 CUDA). Add PIP/UV_BREAK_SYSTEM_PACKAGES for the ROCm base's PEP-668-marked Ubuntu 24.04 Python (no-op on the conda CUDA base), and a build-time GPU_FLAVOR guard asserting the dependency install did not clobber the base image's GPU torch/torchaudio — a future dep bump that forces a torch reinstall now fails the build instead of shipping a CPU-only "ROCm" image. - .github/workflows/docker.yml: new build-and-push-rocm job (separate job for runner disk — the ROCm base is ~25 GB unpacked, so it frees the preinstalled toolchains first). Tags mirror the CUDA semantics with a -rocm suffix (:rocm rolling preview, :stable-rocm, :X.Y.Z-rocm, :X.Y-rocm, :sha-xxxx-rocm) on both GHCR and Docker Hub, same secret gating. flavor latest=false so release tags can't clobber :latest. No cache-to: the ROCm layers would blow the 10 GB GHA cache budget. - deploy/docker-compose.yml: new opt-in 'rocm' profile passing the GPU through via /dev/kfd + /dev/dri, with HSA_OVERRIDE_GFX_VERSION=11.0.0 documented (user-set, not baked in — backend auto-sets it for known consumer GFX IDs). - Docs-sync: docker.md (ROCm quick start incl. Podman/Quadlet, tag table, troubleshooting), dockerhub-overview.md, README AMD note, linux.md ROCm section cross-link, CHANGELOG [Unreleased]. Base image: rocm/pytorch:rocm7.2.4_ubuntu24.04_py3.12_pytorch_release_2.8.0 — torch 2.8.0 exactly matches the CUDA image (identical resolution, so uv keeps it), py3.12 satisfies requires-python >=3.11 (the ubuntu22.04 variants are py3.10 and do not). Closes #1165 Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
55c852c6f2 |
fix(remote-auth): show an API-key gate (not PIN) for API-key 401s in remote-backend mode (#1154)
* fix(remote-auth): route API-key 401 to an API-key gate, not the PIN form
When OMNIVOICE_API_KEY is set (remote-backend mode), a non-loopback browser
gets 401 "API key required" from BearerKeyMiddleware. But client.ts fired
`ov:pin-required` on every 401, surfacing the PIN gate — whose payload
(sessionStorage ov_pin / X-OmniVoice-Pin) can never satisfy the API-key
middleware. A remote user was stuck on a PIN form they could not pass.
Read the 401 `detail` and dispatch a single `ov:auth-required` CustomEvent
carrying the mode; RemoteAuthGate renders the matching PIN or API-key form.
Adds a `?api_key=` deep-link bootstrap (one-shot — scrubbed from the URL so a
reload can't re-clobber a corrected key) and a guarded saveApiKey helper.
Backend is unchanged — the two 401s are distinguishable by their `detail`
body ("API key required" vs "PIN required"). Docs: remote-gpu.md gains a
"From a browser" subsection for the new ?api_key= deep link.
* fix(remote-auth): preserve URL hash when scrubbing credentials
The replaceState that scrubs ?api_key=/?pin= rebuilt the URL from pathname
(+ optional query) and dropped url.hash, nuking any deep-link fragment
(e.g. #settings). Rebuild with pathname + (?query) + hash.
Addresses greptile + coderabbit review feedback on #1154.
* fix(remote-auth): guard 401 routing against a non-string/malformed detail
String(detail) can itself throw on a 401 detail whose toString is broken
(e.g. { toString: null }), aborting the auth-event dispatch. Match only real
strings with typeof; anything else falls back to PIN mode.
Addresses coderabbit's 17:03 re-review finding on #1154.
* fix(remote-auth): read the deep-link API key from the URL fragment (#api_key=)
Move the remote-backend deep link from ?api_key= (query) to #api_key=
(fragment): fragments are never sent to the server, so the durable key stays
out of the GPU box's and any reverse proxy's request logs on the page load
(greptile P1). ?pin= stays on the query (QR flow, session PIN).
The bootstrap is extracted into a pure, unit-tested _parseDeepLinkCredentials
helper (pin from the query, api_key from the fragment, one-shot scrub of both,
plus a legacy ?api_key= scrubbed-without-reading so a stray query key never
lingers). Docs document #api_key= with encoding guidance for keys containing
+ / & / # / =.
|
||
|
|
9ecb810946 |
fix(shell): version-gate crash markers, pin the WebView repair contract, run the Rust suite in CI (#1145)
* fix(shell): version-gate crash markers, pin the WebView repair contract, run the Rust suite in CI Two deferred items from the recurrence audit, plus the CI gap that made them possible: - crash.rs: a persisted "backend crashed" marker now only surfaces for the release that wrote it. After an upgrade, markers from the previous version (quite possibly the build whose crash the upgrade fixed) are ignored and pruned on read instead of resurfacing unacknowledged as if the new build had crashed. backend_version gains #[serde(default)] so legacy version-less markers still deserialize — as "", which the gate treats as stale by design. Preview stamps (X.Y.Z-N) count as their release. - commands.rs: the #879 WebView2 cache repair's filesystem half is extracted into clear_webview_cache_at() (paths + retry policy as parameters, zero behavior change) and its contract is pinned by tests: no marker → nothing touched; marker consumed first, unconditionally (one-shot — a failing repair can never loop across launches); missing cache is success; a locked cache is retried then abandoned with a log, never bricking startup. - ci.yml: the Tauri shell check only ran `cargo check`, which neither compiles nor runs #[cfg(test)] code — so the shell's ~90 unit tests (crash.rs, reset.rs, bootstrap.rs, …) never executed anywhere in CI. `cargo test --lib` now runs them natively on all three OSes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(shell): crash-notice read path is strictly read-only — a prune-save there could destroy a fresh marker Greptile's P1 is real, and hotter than stated: get_last_backend_crash is not just a startup check — streamDropError (#1119) polls it every second for 8 s after a stream drops, which is exactly when the death watcher is inside record_crash's load→push→save. The previous commit's read path did load→prune→save when stale-version markers existed (the post-upgrade state), so a poll could load the pre-crash snapshot, lose the race, and save over the freshly recorded marker — silently deleting the only evidence of the crash it was being polled to find. Smallest fix: reads never write. The read path (extracted as read_notice_from(path, version) so the contract is testable) filters stale-version markers in memory only; disk pruning stays on the write paths (record_crash, acknowledge_backend_crash), where load-modify-save already existed pre-PR and is paced by a crash or a user click rather than a 1 Hz poll. Regression test pins the file as byte-identical across reads, stale markers filtered and current ones surfacing as before. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
58c6f37252 |
perf(dub): single-use per-segment refs no longer evict the prompts a dub reuses; add docs/performance.md (#1132)
* perf(dub): single-use per-segment refs no longer evict the prompts a dub reuses; add docs/performance.md The scan-resistance fix: A dub cuts a distinct reference clip per segment (Wave 3.2 / #486 — each line clones its own source delivery) and falls back to the per-speaker clone for segments under 3 s. Both paths flow through the voice-clone prompt cache — an LRU of 8. Streaming hundreds of one-shot per-segment clips through that LRU evicts the per-speaker and locked-profile prompts that every fallback segment reuses, so the speaker ref was re-encoded (~0.4 s each, measured with scripts/bench_pipeline.py) again and again across the render. Note what this deliberately does NOT do: the bench's "166 misses vs 2 speakers" framing suggested keying refs per speaker — but per-segment refs are the intentional prosody-matching feature, and the re-transcription behind them is the #1004 correctness fix. Their encode cost is the price of the feature, not waste. The waste was only the eviction side-effect, and that's what this removes: _get_clone_prompt(store=False) still reads the cache (a hit is free) but never inserts, and the dub loop marks exactly the segment-scoped refs (auto-seg: bindings and auto: bindings resolved to a segment clip) as single-use. Per-speaker, locked-profile, and preview refs cache as before. cache_ref is popped in generate_with_cached_ref before the model call — the model's generate() has an explicit signature and would TypeError — and unknown engines ignore it (**kw adapters). The doc: docs/performance.md is the first performance documentation in the repo — none of the ~15 perf env vars appeared anywhere in docs/, the Performance panel's only control is Windows-only, and slowness reports (#1032) arrived as mysteries instead of settings checks. Covers the three classic causes of "it got slow", where generation/dub time goes, every knob with defaults and warnings (raising OMNIVOICE_GPU_WORKERS on a small GPU is the #567 crash, not a speedup), platform notes, and how to run the bench so reports carry numbers. Linked from README's install section. Tests: store=False semantics (encodes, never inserts, still reads), the flood scenario end to end (a speaker prompt stays warm through 3x the cache cap of one-shots), and the pop contract (cache_ref never reaches the model). Full suite: 2974 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs,dub: review round — qualify the per-file cache claim; note the OOM-retry tradeoff - CodeRabbit: docs/performance.md's "the reference encode is cached per file" now carves out the dub's per-line clips (single-use by design — nothing for a cache to save). - Greptile P2 (OOM retry re-encodes a single-use ref): acknowledged in a code comment as deliberate — caching the retry's ref would reintroduce the eviction this flag prevents, to optimize a path that only runs after an OOM already cost seconds. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(performance): probe-based torch.compile wording; honest accelerator + cache claims (review) Greptile's repeated OOM-retry finding is deliberately skipped: retaining the prompt across the retry would require passing prompt objects through the adapter protocol (backend.generate takes paths), to save 0.4s on a path that only runs after an OOM already cost seconds — the tradeoff is documented at the call site. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4e5d795832 |
fix(uninstall,storage): remove the saved-env leftover; count sidecar engines in disk usage (#1108)
Two recon findings from the reset work, fixed properly (whole class + tests +
docs), plus the destructive reset path is now exercised end-to-end.
1. ~/.config/omnivoice/env survived every uninstall. The app persists the
model-cache location (and a possible HF_TOKEN) there via
backend/core/user_env.py, but the in-app "Remove all data" (uninstall.rs),
uninstall.sh, and uninstall.ps1 all walked past it — so a reinstall silently
inherited the old file and redirected downloads to a maybe-deleted location.
All three now remove it. It's the same expanduser("~/.config/omnivoice/env")
path on every OS, so the Windows script uses %USERPROFILE%\.config\omnivoice.
is_recognizably_ours accepts it (contains "omnivoice"); docs tables updated.
2. Disk usage measured the wrong engines dir. storage_report.default_engines_dir()
returned backend/engines (built-in engine *modules*, no venvs), while sidecar
installs live in DATA_DIR/engines/<id>. So a multi-GB IndexTTS-2 install was
invisible in the engine-venv category and rolled into data/"other". Now points
at DATA_DIR/engines and sizes the WHOLE install (venv + checkout + weights),
with the data category claiming that subtree so it isn't double-counted.
Reset hardening: extracted purge_scopes() as a pure fs function (no AppHandle),
so the actual delete loop runs in tests against a real on-disk install tree —
"everything" wipes the install but spares the venv/foreign temp/sibling folders,
a settings reset keeps content+config+models, and a poisoned data_dir="$HOME"
deletes NOTHING. This is the live drive-through of the destructive path, minus
the GUI.
Also: gitignore the node_modules symlink form (the directory rule node_modules/
never matched a worktree symlink, so it kept slipping into commits).
Tests: Rust 78 (6 new), storage_report 20 (2 new incl. once-not-twice count +
default-dir guard), frontend 1207, i18n probe green, format+lint clean.
Co-authored-by: mergetest <nizam4103@gmail.com>
|
||
|
|
94093605eb |
feat(settings): factory reset gets scopes — preferences, settings, assets, everything (#1100)
* feat(settings): factory reset gets scopes — preferences, settings, assets, everything Factory reset did exactly one thing: clear localStorage. The only other option was "Remove all data", which deletes the Python env and quits. Between "forget my theme" and "wipe the machine" sat every reset a user actually needs — drop a corrupt model download, remove a wedged sidecar engine, put the settings back without losing a single voice — and none of them existed. Settings → Storage → "Reset & remove" now offers four tiers (UI preferences / all settings / downloaded assets & models / everything OmniVoice did) plus a per-scope checklist. Every scope shows its real on-disk size, and the number on the confirm button is exactly what gets freed. Why the shell and not the backend: a loaded model memory-maps its weights out of the HF cache (locked on Windows while mapped), and ensure_dirs() runs at import, so a backend cannot delete voices/ or outputs/ and still write to them. reset.rs stops the backend, deletes, and starts it again — and that restart is also the repair: the fresh process re-runs ensure_dirs() and alembic, so a removed database comes back empty rather than missing. retry_bootstrap's respawn path is extracted to bootstrap::respawn_backend so both callers share one implementation. Deliberate scope choices: - "Everything" stops short of the managed Python env, so a reset hands back a working app on the first-run screen. The env is the uninstaller's business. - A settings reset keeps the storage locations (config.json, the user env file). Clearing the model-cache pointer would strand gigabytes at a path the app no longer looks in — install shape is not a preference. - content deletes the DB with the media: rows without files is how you get a library full of broken entries. - The shared HF cache is flagged as shared only when it IS — computed, so Windows and portable installs (app-private cache) get no caveat they don't need. Safety: nothing is removed unless it sits inside a validated root — one carrying an OmniVoice-owned path component OR holding an OmniVoice signature file, which is what lets a custom data dir on an external volume be cleared while a mis-set data_dir: "/" is refused. Voices/projects/audio need the word typed. 9 Rust tests (guard, scope composition, shared-cache computation) + 14 frontend (planning purity, typed confirm, disk-vs-frontend split, shared warning). Border utilities follow the design guard (tests/test_no_literal_borders.py). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(settings): give the Storage panels a design — proportional sizes, live totals, real tokens "Remove all data" listed four folders as a flat run of text: a 7.5 GB model cache and a 391-byte config file rendered at identical visual weight, so the one thing worth seeing — where the space actually went — was the one thing you couldn't. And the 391 B folder said "0 KB", which reads as "nothing here". - New shared StorageTargetRow, used by BOTH destructive panels so they read as one system: icon, label, dimmed path (truncated, full text on hover), size, and a proportional bar showing that row's share of what will be freed. Unticked rows claim none of the bar — the bars must sum to what the button promises. - The shared HF cache moves OUT of the confirm dialog into its own "Optional" row with the checkbox and the caveat in the list. Ticking it now moves the running total in front of the user, instead of springing a different number on them at the point of no return. The dialog lists exactly what is going. - One byte formatter for both panels (settings/bytes.js). models/format.fmtBytes floors at kilobytes, hence "0 KB"; it stays where it is for the model store. Real fix underneath: three of the tokens these panels styled with DO NOT EXIST (--chrome-fg-subtle, --chrome-bg-raised, --color-warning). An undefined var() makes the declaration invalid, the browser drops it, and the element silently inherits — which is why the paths that were meant to recede rendered at full body weight. That is a whole class of bug that fails invisibly, so it gets a guard: src/test/cssTokens.test.js fails on any var(--token) in JSX not defined in a stylesheet, with runtime-injected tokens (Radix, inline-style hues) allowlisted by reason. Six pre-existing offenders elsewhere in the app are recorded as known-broken and ratcheted so the list can only shrink — they are real bugs, but each is a visual change that wants its own review. Frontend suite 1196 → 1205 (6 UninstallPanel component tests incl. the live total and the bar proportions; 3 token-guard tests, verified fail-before). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: green CI + finish the token sweep + snapshot the panels Three things on top of the redesign: 1. CI was red on tests/probe/test_probe_i18n.py — removing the eight dead `factory_reset_*` keys from en.json orphaned them in all 20 other locales (the probe forbids a non-en key absent from en). Removed them everywhere. This guard scans locales at pytest time; a frontend-only run never sees it. 2. Finished the undefined-token sweep instead of grandfathering it. Six bare `var(--token)` references resolved to nothing; the only genuinely undefined, fallback-less one in shipping panels was `--chrome-input-bg` (input fields AND progress-bar tracks AND skeletons across StoragePanel, StorageUsagePanel, HistoryRetentionPanel, ModelStoreTab — tracks were rendering with no background at all). Repointed to --chrome-hover-bg. The rest (--chrome-menu-bg, --chrome-bg-inset, --border, --input-bg, --muted) already carry `var(--x, fallback)`, which is valid CSS. So cssTokens.test.js now checks only the BARE form and ships with zero exceptions — no known-broken ratchet, because there is nothing left broken. 3. Registered both Storage panels in the visual-regression harness (a Tauri `invoke` stub added to providers.jsx alongside the existing fetch stub) and committed baselines across all three themes. This is how I actually looked at the redesign: the bars render proportional (the 720 KB voices row fills, the 391 B row is a sliver), the shared-cache row sits in its own Optional group, and every token now resolves in default/midnight/catppuccin. `_forceAdvanced` on ResetPanel opens the checklist for the snapshot; no effect on the toggle. Full backend suite 2897 passed (incl. the i18n probe). Frontend 1205. * style: oxfmt the new panels and specs Format-check is a CI gate (oxfmt --check); the new files weren't run through oxfmt --write. No behavior change. * chore: stop tracking the node_modules symlink A worktree-local symlink slipped past .gitignore (which lists node_modules/ — the directory form — so it never matched the symlink file). Removed from the index; the symlink stays on disk for local test runs. --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0421be966e |
feat(settings): in-app uninstall — Settings → Storage → Remove all data (#1089) (#1099)
* feat(settings): in-app uninstall — Settings → Storage → "Remove all data" The v0.3.19 uninstaller was a SCRIPT, which never reaches the people who need it: anyone who installed the .dmg / .msi / AppImage has no repo to run scripts/uninstall.sh from — exactly the reporter in #1089, an AppImage user. "Where is uninstall in the app?" had no answer. Now it does. New Tauri commands (uninstall.rs): - uninstall_scan — every folder this install owns, with real sizes, resolved through the same setup.rs helpers the app itself uses, so custom + portable locations are cleaned instead of the defaults being assumed. - uninstall_purge — stops the backend (marking the kill intentional so the #567 supervisor doesn't respawn one into the directories being deleted), removes the folders, and lets the UI quit the app: the Python env it runs on is gone, so there is nothing to return to. This lives in the Rust shell, not the backend, because the biggest thing to remove is the managed Python environment and the backend is RUNNING FROM IT — a process can't delete its own interpreter (and Windows locks the files). Safety: every path must pass is_recognizably_ours() before any remove_dir_all — absolute, not `/` or $HOME, and carrying an OmniVoice-owned component (unit tested both ways). The shared Hugging Face cache is reported separately and is OPT-IN behind its own checkbox with the caveat spelled out: it's the standard HF cache other ML tools share, so sweeping it up silently would delete models this app never downloaded. Deleting voices/projects is irreversible, so the confirm requires TYPING the word, not just a click. Also fixes a real bug in what shipped in v0.3.19: the scripts and docs missed where the BACKEND writes its logs — ~/.local/state/OmniVoice on Linux and %LOCALAPPDATA%\OmniVoice\Logs on Windows (backend_log_path(), backend.rs) — so every Linux/Windows uninstall left a stray log dir behind. Covered now in the scripts, the docs, and the in-app scan. And the scripts now ship as release assets, so cleanup is possible without launching the app at all. Rust: 2 new guard tests. Frontend: 6 new tests (the size on the confirm button must equal what actually gets deleted); suite 1182 passed. Docs synced. Refs #1089 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(settings): drop token border utilities from UninstallPanel (design guard) tests/test_no_literal_borders.py::test_no_token_border_utilities_in_jsx is a backend guard that scans JSX — so a frontend-only test run misses it. It forbids `border-[var(--chrome-border)]` structural utilities: the app-wide border removal converted every panel/row frame away from them, and they render a stray hairline the moment the token doesn't resolve transparent. Row dividers → spacing + an alternating `--chrome-hover-bg` tint; the opt-in checkbox card → a background tint; the confirm input → the sanctioned arbitrary `[border:1px_solid_var(--chrome-border)]` property form the other settings inputs already use (explicitly not flagged by the guard). Guard green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b1a7ddc374 |
feat(install): clean uninstaller + a straight answer to "where is my data?" (#1097)
A Linux AppImage user asked which folders to delete to remove OmniVoice and whether an uninstaller exists (#1089) — they had to guess. They shouldn't have to: the app is fully local, so uninstalling IS just deleting the folders it wrote, and we never documented them. - scripts/uninstall.sh (macOS/Linux) + scripts/uninstall.ps1 (Windows): find every OmniVoice folder — app data, the multi-GB managed Python env, config, logs — plus, listed SEPARATELY because it is a shared cache, the Hugging Face model cache. Print each with its size as a DRY RUN and stop; delete only on --yes (--models / -Models to include the shared cache). They honor the same env overrides the app reads (OMNIVOICE_DATA_DIR, OMNIVOICE_CACHE_DIR, HF_HOME, HF_HUB_CACHE), and never touch the app binary or anything outside the paths they list. - docs/install/uninstall.md: the complete per-platform path table (what each folder holds and how big it is), the shared-HF-cache caveat, custom/portable locations, per-platform steps to remove the app itself, and what to keep if you plan to reinstall. - Linked from the README FAQ, SUPPORT.md, and install troubleshooting. Paths mirror backend/core/config.py + frontend/src-tauri/src/setup.rs. Verified on macOS: dry-run lists the real dirs; sandboxed HOME runs confirm --yes removes app folders while KEEPING the shared cache, --models removes it, and the env overrides retarget correctly. Closes #1089 Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bd85bab624 |
chore: retire finished planning archives from the repo root (#1095)
Removes ~110 files of process noise (all preserved in git history): .planning/ (GSD-era phases/quick-plans/issue-clusters; workflow retired 2026-07-08), specs/ (spec-kit specs for shipped features 001-007), design/ (pre-React ASCII mockups), research/ (legacy Gradio archive), and .agents/ (rules for a third-party agent tool no longer in use). The four load-bearing decision docs move to docs/adr/ with an archival note; every live pointer follows (gguf engine module docs + quant_map, inject-apprun.sh, pyproject/test comments, fixture README + its seed script — kept byte-identical). The CJK allowlist drops the deleted legacy_gradio entries; STRUCTURE.md and ROADMAP.md document the removal instead of linking into it. Backend suite: 2891 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
23367cccaf |
fix: first-run wizard version + mirror-unreachable rescue + lifecycle-aware backend reachability (#1094)
Three fixes from the same first-run session report:
- SetupWizard shows v{APP_VERSION} in its masthead (same identity mark as
the install splash footer), so setup screenshots identify the build.
- A dead configured HF mirror no longer strands the wizard: the
install_error SSE now carries docs_topic (core.failure.classify), and
WizardLibrary renders the MirrorRescue quick-pick (extracted from
SetupWizard, now including the official preset) next to the failed row,
retrying it the moment a new endpoint is applied. PUT /hf-mirror clears
the install cooldowns (no 429 on the immediate retry) and clearing to
official also drops the legacy hf_endpoint pref that silently kept the
dead mirror in effect. The hint's false "applied when the app starts"
claim is corrected: downloads resolve the endpoint per call, retry
first, restart only if it still fails.
- "Can't reach the local OmniVoice backend" stops firing during real
start/restart windows: a respawn takes 10-20+s (venv spawn + torch
import) but the transport cascade gave up at ~2.9s. apiFetch now asks
the shell (bootstrap_status via utils/backendLifecycle) whether a
start/restart is in progress and keeps retrying while it is (capped at
120s, matching the supervisor's respawn budget); the new
BackendRestartBanner finally implements the reconnecting banner the
#567 supervisor has emitted events for all along. Truly dead backends
(or non-Tauri deploys) still error promptly.
Regression tests for all three layers; docs synced
(downloading-models.md, troubleshooting.md §14b); CHANGELOG [Unreleased].
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
3d8799d9d2 |
feat(asr): complete activation flow for the OpenAI-compatible remote ASR engine (#1087)
The openai-compat-asr backend (#877) shipped with settings routes but no discoverable activation path: the config panel hid in Settings → Models, its hint text claimed "there's no in-app engine picker for ASR yet" (stale — the matrix has one), and there was no way to check a server actually answers before pointing a dub/dictation run at it. Configure → test → activate now live on one screen, Settings → Engines: - The ASR family tab mounts the config panel (URL / model / optional API key) below the engine matrix; saving refetches the matrix via a new reloadToken prop so the engine row flips unavailable → available and its "Use" button appears without a manual refresh. - New "Test connection" button + loopback-gated POST /api/settings/asr-openai-compat/test: saves first (same stale-config contract as /llm-providers/{id}/test), then probes GET {base_url}/models — no audio leaves the machine. The structured verdict maps to localized, actionable messages: latency + whether the configured model is listed on success; classified auth_failed / http_error / timeout / unreachable / ok_no_models failures. detail is core.scrub-ed; the key is never logged or echoed. - Engine reads persisted config fresh per transcribe (regression test) — config changes need no backend restart. Never default-active: ASR auto-detect only picks local engines. - i18n for every new string (en.json); no hardcoded CJK; identical behavior on macOS/Windows/Linux (pure HTTP + React). - Docs-sync: docs/engines/openai-compatible-asr.md rewritten around the one-screen flow with LM Studio / llama.cpp / Groq / OpenAI examples and the privacy note; README engine table cell updated. Verified end-to-end against a fake OpenAI-compatible server: UI drive (configure → test → row flip → Use) plus a real transcription through the backend's /v1/audio/transcriptions immediately after a config change, no restart. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5ca9f54e47 |
docs: migration guide for stranded Real-Time-Voice-Cloning users + sharper local-first promise (#1085)
RTVC (CorentinJ's 50k-star SV2TTS repo) is archived; its users need a maintained home. New docs/migration/real-time-voice-cloning.md maps every RTVC concept to its OmniVoice equivalent (encoder+utterance → reference clip, toolbox → app, vocoder choice → Settings → Engines, demo_cli.py → REST API/CLI/MCP), is honest about what RTVC did that we don't (research toolbox, three-stage training, MIT license, smaller footprint), and walks the first clone with verified UI labels only. Wired into docs/features.yaml's existence-checked docs list and linked from the README Quickstart. README tagline now states the local-first promise verbatim at the very top: "No accounts. No API keys. No cloud." — everything else on the front page is unchanged. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ff56865cf7 |
feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines (#1083)
* feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines
IndexTTS-2 required four manual terminal steps (git clone, uv venv,
uv pip install -e ., export OMNIVOICE_INDEXTTS_DIR). This turns that into
a guided in-app install:
- backend/services/sidecar_install.py — parametrized sidecar provisioner
(SidecarSpec/SPECS so future sidecar engines are one entry, not another
installer). Resumable background job with step-by-step status: disk-space
preflight (needs-X/have-Y message), source fetch (git clone --depth 1
primary, GitHub tarball fallback when git is absent/fails), dedicated
venv via uv (OMNIVOICE_BUNDLED_UV → PATH resolution; transformers<5
isolation preserved — the parent env is never touched), import-probe
verification, IndexTeam/IndexTTS-2 weights into <checkout>/checkpoints
(where the sidecar actually loads from) via snapshot_download with the
auto-selected/configured HF endpoint + token — no hardcoded
huggingface.co — and persistence of OMNIVOICE_INDEXTTS_DIR (os.environ
for immediate use, prefs.json env.* for the next launch). Idempotent:
partial installs repair, downloads resume, healthy installs (incl. a
user's own clone) report already_installed and are never touched.
- API: POST /engines/{id}/install starts the job, GET
/engines/{id}/install/status polls it, DELETE /engines/{id}/install
removes an app-managed install (loopback-gated; refuses user-managed
clones). list_backends() gains one_click_install.
- Frontend: Settings → Engines shows an Install button on the IndexTTS2
row with per-step progress, live log tail, weight-download %, and
error+remediation; the manual setup snippet is demoted to a collapsed
"Manual install" fallback. All strings via i18n (en.json).
- OMNIVOICE_INDEXTTS_DIR joins the Settings env-var allowlist
(single-sourced from the installer SPECS).
- Docs: docs/engines/indextts.md leads with the one-click flow; manual
steps become the fallback section. CHANGELOG Unreleased entry added.
- Tests: tests/test_sidecar_install.py (24 cases — happy path, disk-space
fail, git-absent/git-failing tarball fallback, partial-install repair,
already-installed/running gating, uninstall safety, spec↔bootstrap
contract, router wiring) + 6 new EngineCompatibilityMatrix RTL cases.
API route snapshot regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): harden the sidecar installer — review findings
- Route namespace: /engines/sidecar/{id}/install — a dynamic
/engines/{id}/install would shadow the literal
POST /engines/sonitranslate/install (engines router registers first);
regression-guarded by test_sidecar_routes_never_shadow_literal_engine_routes.
- Weights completion marker: a killed-mid-download multi-shard weights dir
(config.yaml + plausible shards) no longer passes for healthy; the marker
is written only after snapshot_download returns, so re-runs resume.
- _run_logged: drain thread + proc.wait(timeout) + POSIX process-group kill
— a grandchild holding the stdout pipe can no longer hang the step past
its timeout.
- Job log lock: the status poll's list(deque) copy no longer races the
worker's appends (RuntimeError under active logging).
- Self-heal: a healthy managed install whose env var was lost (prefs wiped)
is re-pointed by start_install instead of reported already_installed
while the engine stays unavailable; legacy bootstrap installs (Probe-2
venv) are trusted via the engine's own probe.
- Single-sourced uv/venv-layout resolution: engines.indextts.bootstrap now
delegates _locate_uv/_venv_python_path to services.sidecar_install.
- Frontend: stable poll interval (keyed on the running-id set, not the
status map), reload on a job that finishes before the first poll,
re-attach to an in-flight job on remount, i18n'd Install aria-label,
manual-install <details> auto-opens on failure, snippet block hoisted
out of the JSX IIFE.
- list_backends: sidecar-installable set hoisted out of the per-engine
loop; exhaustive-shape registry test updated for one_click_install.
- Tests rebind the live services.sidecar_install module per test (other
suites purge sys.modules["services"], which made router tests
order-dependent).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): fill in the PR ref (#1083)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): validated tarball fallback + scanner-clean installer
- The pre-filter= extractall fallback (Python < 3.11.4) now extracts
member-by-member behind the same guards extractall(filter="data")
enforces — regular files/dirs only, no absolute paths, no ../ escapes,
resolved-path containment. Kills the new CodeQL py/tarslip (high) and
Bandit B202 (error) alerts; regression-tested with a malicious tarball
(test_safe_extract_members_blocks_tar_slip).
- snapshot_download tracks the weights repo's default branch on purpose
(same policy as every other model download; artifacts are
checksum-verified by hf_hub) — documented + B615 waived at the call.
- Explanatory comments on the intentional empty-except blocks
(CodeQL py/empty-except notes).
Verified locally: bandit -ll -ii on the module reports 0 MEDIUM+ findings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): address Greptile review — Windows tree kill, prefs write race, poll robustness
- _kill_tree: Windows now uses taskkill /F /T so a git/uv helper spawned by
the timed-out child can't keep writing into the checkout (POSIX already
killed the process group). Unit-tested with os.name patched to nt.
- core/prefs: mutations (set_/delete) are serialized behind a module lock —
the installer worker persisting its env.* key concurrently with a Settings
write could previously drop whichever key saved first (whole-class fix:
every threaded prefs writer, not just the installer). Fail-before/
pass-after: tests/test_prefs_thread_safety.py.
- Matrix polling: at most one in-flight status request per engine (an old
'running' response can no longer land after a newer 'succeeded' and
restart the poller), and four consecutive poll failures drop the stale
snapshot instead of showing "Installing…" and hammering a dead backend
forever.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9c81e3389d |
feat(network): automatic Hugging Face endpoint selection — probe, pick, remember (#1082)
* feat(network): automatic Hugging Face endpoint selection — probe, pick, remember Restricted-network first-runs (the #984 class: huggingface.co unreachable, user dead-ends before discovering the mirror setting) now self-heal by default, while explicit endpoint choices are never second-guessed. - New backend/services/endpoint_race.py: parallel HTTPS reachability + latency probes of huggingface.co and the hf-mirror.com community mirror (3s timeouts). Probes are the only signal — no geo-IP, no third-party calls. Reachable beats unreachable; with both reachable the official endpoint wins unless the mirror is decisively faster (anti-flap hysteresis). The pick is cached in prefs and re-raced only on first run, a network-classified download failure, staleness (>7 days), or an explicit "Test again". - Manual mode is sacred: HF_ENDPOINT env, an hf_endpoint pref, or any explicit Settings pick disables auto-switching entirely; OMNIVOICE_HF_ENDPOINT_MODE=manual is a hard opt-out. - Wiring: the wizard preflight races endpoints when nothing is configured (honest copy when the mirror wins; warn-not-block when nothing is reachable); Model Store installs and the model-cache auto-repair resolve their per-call endpoint= through the cached decision, and a network-classified failure re-races once per repo per process and retries on the new winner (same guard pattern as the cache-recovery ladder). - Settings → Models → Hugging Face mirror gains "Auto (recommended)": shows the current pick, measured latency, last-checked time, and a "Test again" button (POST /api/settings/hf-mirror/test). Existing explicit configs surface as the matching manual mode. Panel notes that hf_hub checksums every download regardless of endpoint. - Tests: policy/cache/failover matrices in tests/test_endpoint_race.py, preflight + settings + repair-failover integration with mocked probers, HFMirrorPanel mode tests, and a suite-wide conftest guard that pins the probers so no test can hit the real network. - Docs: downloading-models.md and install/troubleshooting.md describe the automatic default and both opt-outs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): Unreleased entry for automatic HF endpoint selection Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint-probe pin uses an isolated MonkeyPatch and clears the decision cache; dtype guard tolerates stubbed torch The autouse probe pin requested the shared monkeypatch fixture, hoisting its setup earlier for every test and reordering teardown against the fp16 guard — which then ran torch.get_default_dtype() on test_torch_compile_gate's SimpleNamespace stub. The pin now uses its own MonkeyPatch context and also clears the prefs-cached endpoint decision per test (one test's auto pick leaked into other tests' preflight labels on CI ordering). The dtype guard additionally skips non-module torch stubs outright. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint env vars can no longer leak out of the mirror-settings suite set_hf_mirror writes os.environ[HF_ENDPOINT] during the test, and monkeypatch.delenv(raising=False) on an absent var records nothing to undo — so the write leaked process-wide and flipped later suites' preflight network checks into the explicit-endpoint branch (the CI-order failures). Guaranteed save/restore autouse fixture at the source, plus defensive env shedding in the preflight suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
254f071b45 |
fix(docker): ship alembic.ini in the image, add an image-level HEALTHCHECK (#1080)
Migrations in Docker fell back to the additive-column self-heal because alembic.ini was never copied; the real migration chain now runs. The HEALTHCHECK covers plain docker-run (compose files keep their own), with a start period sized for first-boot schema creation. Docs example tag freshened. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5562aa16a7 |
feat(setup): media tools become invisible — bundled by default, controllable in Settings (#1071)
Most users should never learn what ffmpeg is. The Setup Wizard's SYSTEM
PREFLIGHT stops listing FFmpeg / FFprobe / yt-dlp as user-installed
requirements ("brew install ffmpeg…"): they are internal dependencies the
app provisions for itself. Genuine user facts (OS, RAM, disk, GPU,
network, Python) are untouched.
Backend
- New services/media_tools.py: per-tool status {version, path, origin:
sidecar|bundled|system|custom}; background acquisition of a pinned,
SHA-256-verified static ffmpeg+ffprobe build (immutable-commit fetch
from the same upstream the static-ffmpeg pip package uses — that
package itself was audited and rejected: mutable raw/main URL, no
checksums, writes into site-packages); binaries are `-version`-probed
via the existing _binary_runs before being trusted, installed under
DATA_DIR (update-surviving, frozen-build-safe), zero new Python deps.
- ffmpeg_utils resolution chain gains the acquired-bundled tier — and
ffprobe finally has a bundled tier at all (imageio-ffmpeg ships none),
closing the source-install gap.
- New /media-tools router (loopback-gated, same contract as
/system/set-env): status, acquire, {tool}/custom-path | use-system |
restore, ytdlp/update | restore. Overrides persist via the existing
env.FFMPEG_PATH / env.FFPROBE_PATH prefs convention — one store, no
competing controls.
- yt-dlp updates: audited in-venv pip/uv upgrade and rejected (venv is
uv-managed with no pip; yt-dlp is a locked dep, so the updater's
--inexact drift sync would revert it). Instead the newest wheel —
verified against PyPI's own sha256 — lands in a DATA_DIR overlay
prepended to sys.path at startup: survives app updates, works in
frozen builds, and "Restore tested version" is just deleting the
overlay. Gallery now runs yt-dlp via `python -m yt_dlp` (module, not
PATH) so the CLI can never be a user-install task either.
- /setup/preflight drops the three tool rows, carries a media_tools
verdict, and self-heals: kicks the bundled download in the background
when no tier resolves (never re-fires after a failure — the wizard's
card owns Retry). diagnose + the ffmpeg-missing notification now point
at Settings → Audio tools instead of package managers.
Frontend
- Wizard: new MediaEngineCard — renders NOTHING when the engine is ready,
a one-line progress while acquiring, and only on failure an actionable
card (Retry / Use a system copy / Choose file…).
- Settings → Audio tools (new category, System group): FFmpeg + FFprobe
rows with version, path, origin badge, Use system copy / Choose file… /
Restore bundled, header-level "Update bundled build"; yt-dlp row with
Update + Restore tested version (+ restart affordance). Package-manager
commands appear only as copyable prose, never executed.
- The FFmpeg-path override moved out of Settings → Network (pointer row
deep-links to Audio tools; no second writer of env.FFMPEG_PATH).
Notifications gain a settings-tab action type.
- All strings i18n (en + defaultValue), a11y labels on every control.
Tests: 29 new backend (origin classification, checksum/size/probe
rejection, override persistence, overlay update/restore, router gating +
route-shadowing) + preflight contract tests (tool rows gone, verdict
present, auto-acquire fires once); 14 new frontend (wizard hide/progress/
failure-card, Audio tools rows/badges/actions). Route snapshot
regenerated. Docs (macos/linux install, troubleshooting §7b) describe the
new reality in the same commit.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
fd083749c6 |
fix(settings): network & privacy panels — clearable proxy after reload, HF-mirror panel never vanishes, guarded remote-backend save, honest privacy claims (#1063)
Settings → Network / Models / Sharing / Privacy / OpenAPI fixes:
- NetworkTab: a proxy persisted in a previous session can now be cleared —
the Clear button and "Set" badge derive from the backend-persisted value
(sysInfo.proxy_url), not only from a save in the current session. Proxy row
copy now matches its real semantics ("Applies now" badge; desc/toast no
longer claim a restart is needed or leak yt-dlp jargon — reworded in all
21 locales). FFmpeg path placeholder is platform-appropriate instead of
Windows-only on every OS.
- HFMirrorPanel: the panel no longer disappears when the initial GET fails —
the section shell always renders, with a loading state and an error +
Retry affordance. Saving now toasts, the active preset is marked
(aria-pressed), and the custom-URL row is labelled "Custom mirror URL"
instead of raw HF_ENDPOINT jargon (env var moved to the row note).
- RemoteBackendPanel: full i18n (was 100% hardcoded English); Save & reload
now validates the URL (http/https, parseable) and asks for confirmation
before saving a URL that hasn't passed a connection test — a typo'd base
no longer bricks every API call after reload. Dropped the contradictory
"Restart required" badge (saving reloads the app itself; description says
so). docs/remote-gpu.md updated to match (docs-sync).
- PrivacyTab: the "Network calls" row no longer shows the green "Offline
translator" assurance when the backend is down or reports 'unknown' —
green is reserved for confirmed-offline providers (nllb/argos/
libretranslate), everything unconfirmed shows a neutral "Unknown" badge.
The online-translator warning now deep-links to Translation settings.
- OpenApiPanel: a failed clipboard copy toasts an error instead of silence.
- a11y: all five text inputs across these panels now carry accessible names
(aria-label), previously announced only by their vanishing placeholders.
Tests: new colocated suites for NetworkTab, HFMirrorPanel,
RemoteBackendPanel, PrivacyTab; OpenApiPanel suite extended with copy
success/failure. Frontend suite 140 files / 1061 tests green; i18n parity
probes green (new keys en-only with defaultValue, reworded keys updated in
every locale).
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
5ce9d0e51d |
feat(dub): predict segment fit before synthesis — tight/impossible badges + opt-in shorter rewrites (#1051)
New pure planning layer (services/duration_planner.py) runs after translation, before TTS: estimates each translated line's natural speech duration (self- calibrating from the job's already-synthesized segments, static per-language rates as cold-start fallback) and classifies it fits/tight/impossible against slot + capped gap borrow, with thresholds derived from fit_planner's own caps so "impossible" means "would be trimmed". Verdicts ride the /dub/translate response and badge the segment table; an opt-in (default OFF) LLM pass attaches one-click shorter-rewrite suggestions for impossible lines. Never blocks generation — informs before GPU time is burned. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7dbb95fa15 |
feat(dub): LLM translations keep terms consistent and sound spoken — auto-glossary brief + reflect pass (#1050)
One up-front LLM pass over the full transcript extracts a theme summary + terminology map, merges it under the user's manual glossary (user entries always win), caches it on the dub job per target language (job_data blob, no schema change), and injects the brief into every per-segment prompt. A new reflect pass then critiques each segment's direct translation for wordiness / stiff register and rewrites it as natural spoken dialogue — any failure or divergence silently keeps the direct translation. Both stages have Dub-tab toggles (default ON for the LLM engine, persisted; MT engines unaffected), with i18n strings across all 21 locales and docs updated. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
df870e9ed3 |
docs(agents): add verified Tesla T4 (16GB) inference notes (#1014)
* docs: add AGENTS.md with verified Tesla T4 (16GB) inference notes
Documents two things found while verifying inference on a real T4:
1. Cold-cache first /v1/audio/speech call can hit the 300s
OMNIVOICE_GENERATE_TIMEOUT_S because the checkpoint download happens
inside that budget — workaround via existing POST /models/install or
raising the timeout, no code change needed.
2. The OpenAI-compatible endpoint silently ignores num_step/guidance_scale
(schema doesn't declare them) — use native /generate for those.
Also documents the T4 acceleration checklist (dtype/attention/int8/CUDA
graphs) and measured VRAM (peak 2.05GB). No code changes.
* fix(docs): make /models/install workaround command actually executable
Addresses Greptile review: the instruction omitted the required
repo_id body field (InstallModelRequest rejects an empty body).
* fix(docs): correct port in /models/install example (3900, not 8000)
The app serves on port 3900 (confirmed: /health returns 200 there,
connection refused on 8000). Verified the exact corrected curl command
returns 200 {"status":"install_started",...}.
* move T4 notes to docs/hardware-notes-tesla-t4.md — AGENTS.md is the auto-loaded agent-instructions filename
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
80f10289fe |
feat(asr): ASR engines get the same Settings picker TTS has (env var still wins) (#1026)
Settings → Engines now stacks one pinned Engine Compatibility Matrix per family (TTS, ASR, LLM) instead of a single TTS-titled table with the other families tucked behind a low-discoverability tab. The backend select/prefs path (family="asr" → prefs.asr_backend, env > prefs > auto-detect) already worked but was unexercised and undocumented — it's now locked by API and resolution-order tests, and README + the openai-compat-asr doc stop promising a picker that didn't exist / denying one that now does. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1808a373a1 |
fix(linux): AppRun workaround detection reads the BUNDLED WebKitGTK version, not the host's (#961 follow-up) (#1024)
The launcher decided whether to export WEBKIT_DISABLE_COMPOSITING_MODE by asking the host's pkg-config — but LD_LIBRARY_PATH makes the BUNDLED libwebkit2gtk the one that actually runs, so on any machine where the two diverge the detection read the wrong number. This was the second bug identified during #961's investigation (the reporter built from source, so their dev packages answered pkg-config with a healthy version while the shipped bundle ran an older lib) and was explicitly deferred in #1007 as not-safely-fixable at runtime. The fix makes it knowable by construction instead: inject-apprun.sh runs at bundle time ON the build host whose libwebkit2gtk gets bundled, so it stamps that version into .bundled-webkitgtk-version inside the AppDir. AppRun reads the stamp first and only falls back to host pkg-config for bundles predating it. Empty/unreadable stamp fails safe (workaround on), same philosophy as the missing-pkg-config path. Tests: 3 new cases in AppRun.test.sh — marker-beats-host in both directions (broken-marker/healthy-host and the #961 inversion, healthy-marker/broken-host) plus empty-marker fail-safe. Also wires AppRun.test.sh into pytest (tests/test_apprun_launcher.py) — it was previously run by NO CI job, so the launcher could regress silently. Also documents Windows install-to-another-drive behavior in docs/install/windows.md (#938): local drives work via the wizard's directory picker, mapped network drives are a Windows Installer limitation, and the data directory moves independently of the app. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
da4bef8e42 |
docs: correct troubleshooting §16 — mic bug was a missing entitlement, not an upstream limitation; changelog for #1016/#1020/#1021 (#1022)
troubleshooting.md §16 claimed the macOS microphone-permission bug was an unresolved upstream Tauri/wry limitation with no available fix. That was wrong: @MahdiHedhli read the wry/tauri sources more carefully and found the real cause — Tauri's Hardened Runtime default blocks mic hardware access without com.apple.security.device.audio-input in the bundle's entitlements, which also explains why TCC never listed the app. Their fix (#1016) is merged; §16 now documents the real mechanism, credits the correction, and keeps the record-elsewhere workaround for users on ≤0.3.12 builds. Also brings CHANGELOG [Unreleased] current for the three merges that lacked entries: #1016 (mic fix), #1020 (shutdown wait 3s→20s), #1021 (CI flaky-trio root cause + guard). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7f8a42ce51 |
fix(tts): mlx-audio CSM cloning drops ref_text, breaking every clone attempt (#1012, #1013) (#1017)
MLXAudioBackend.generate() reads voice/ref_audio/language/speed from its kwargs but never extracted ref_text — it was built, then silently never passed through to self._model.generate(). CSM (sesame.py) only builds its cloning context when BOTH ref_audio AND ref_text are present; with ref_text missing, the context list stays empty and indexing into it raises "IndexError: list index out of range" deep inside mlx-audio, instead of the clone ever being attempted. Voice cloning on the CSM engine could never have worked as shipped. generation.py already threads ref_text all the way through — even auto-transcribing it via the GPU pool when the caller supplies ref_audio without one (~line 780) — so the value was always available in kwargs; it just never survived the crossing into this specific backend. Reported with the precise root cause and a working fix (community member independently diagnosed and patched it locally, confirmed working on MPS/0.3.12). Two-line fix: extract ref_text and pass it through when both ref_audio and ref_text are present (guards against passing an orphaned ref_text with no accompanying audio to engines that don't expect it). Tests: tests/test_engines.py — ref_text is passed through when paired with ref_audio, omitted when ref_audio is absent. Also documents the second bug from the same report (#1013): macOS microphone permission never prompts, so OmniVoice never appears in System Settings to grant access. Root-caused to an unresolved upstream Tauri/WebKit limitation (WKWebView's requestMediaCapturePermissionFor delegate — wry#1195, tauri#11951, fix wry#1196 still open/unmerged, no released version to bump to) — not something fixable here without an unverified native Rust/WKWebView hack this session has no way to test. Documented in docs/install/troubleshooting.md with the confirmed workaround (record elsewhere, upload the file). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5a7d9cc05c |
feat(asr): generic OpenAI-compatible transcription backend (#877) (#1003)
First slice of the community's two-track proposal for #877: a generic OpenAI-compatible ASR backend that works TODAY, without waiting on transformers to ship a direct Qwen3-ASR integration (tracked separately, still blocked upstream). Points OmniVoice's transcription at any server exposing POST /v1/audio/transcriptions — a self-hosted Qwen3-ASR/ FunASR/SenseVoice server, or OpenAI's own API. - New OpenAICompatASRBackend (backend/services/asr_backend.py): a pure network client, no local model, no install. Prefers response_format=verbose_json for real per-segment timestamps, degrades to plain text (matching MoonshineASRBackend's shape) when a minimal server rejects that format. Never leaks a raw SDK/httpx exception to the caller (#977 convention) — wraps network/auth failures in a clean, actionable RuntimeError naming the server. - Settings persist via the same encrypted-secret convention as services/llm_providers.py (settings_store.set_secret for the API key — Fernet-encrypted, never a .env row, never echoed back; get_text/ set_text for base_url/model). New GET/PUT /api/settings/ asr-openai-compat, loopback-gated like every other settings route. - Frontend: a small settings panel (Settings → Models) mirroring HFMirrorPanel's exact structure. No ASR engine picker exists yet for ANY ASR backend (only TTS has one) — activating this engine still needs OMNIVOICE_ASR_BACKEND=openai-compat-asr; documented plainly rather than pretending otherwise. - README's ASR Engines table (9 → 10 engines) and docs/features.yaml's drift-checker inventory updated; the '9 engines, all fully local' claim corrected since this one genuinely isn't. - docs/engines/openai-compatible-asr.md: setup steps + an explicit privacy note (unlike every other ASR engine, audio leaves the machine to whatever server is configured). Regression tests: tests/test_asr_openai_compat_877.py (12 tests) — is_available() gating, verbose_json + plain-text response adaptation, network-failure error hygiene, SDK retry disabling, and the settings endpoints' persist/mask/clear-vs-unchanged semantics. Fixed two real full-suite-only failures found during verification (not brushed aside): the API route inventory snapshot needed regenerating for the two new routes, and this file's own tests had a module- staleness bug — a collection-time settings_store import went stale relative to a test-time-fresh fixture when another test elsewhere in the ~2400-test suite reimports the module — fixed by making settings_store itself a fixture resolved at test-run time, same lesson already applied to tests/test_mm2_lifecycle.py earlier this session. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
93aab6dadc |
docs(linux): mention yt-dlp as an optional prerequisite (#973) (#997)
The preflight system check already warns in-app when yt-dlp is missing (Voice Gallery/Dub YouTube downloads fail without it), but the install docs never mentioned it — a user has to hit the in-app warning first instead of seeing it up front alongside the other optional prereqs. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |